Adenosine deaminase capable of acting on DNA and application of adenosine deaminase

CN121152875APending Publication Date: 2025-12-16BEIJING QI BIODESIGN BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480030901.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-09
Filing Date
2024-05-09
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

The existing adenosine deaminase is mainly based on E. coli TadA. It has limited application scenarios, low editing efficiency, and lacks other adenosine deaminases that can act on DNA.

Method used

A new adenosine deaminase, called QB34, was developed to screen and optimize its amino acid sequence through bioinformatics technology, enhance its editing ability on DNA, and combine nucleic acid-targeting domains to form a base editing system.

Benefits of technology

It significantly improves the editing efficiency and accuracy of adenine base editing, expands the application scenarios of adenine base editor, and provides more efficient gene editing tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121152875A_ABST
    Figure CN121152875A_ABST
Patent Text Reader

Abstract

Belongs to the field of gene engineering. In particular to adenosine deaminase capable of acting on DNA and application of adenosine deaminase. More specifically, the invention provides adenosine deaminase capable of acting on DNA, a gene editing system based on the adenosine deaminase and application of the adenosine deaminase. The application scene and selection of the adenine base editor are enriched, and the adenine base editor has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Adenosine deaminase capable of acting on DNA and its application

[0001] Priority and related applications

[0002] This application claims priority to Chinese patent application No. 202310513300.9, filed on May 9, 2023, entitled “Adenosine deaminase capable of acting on DNA and its application”, the entire contents of which, including the appendix, are incorporated herein by reference. Technical Field

[0003] The present invention belongs to the field of genetic engineering. Specifically, the present invention relates to an adenosine deaminase that can act on DNA and its application. More specifically, the present invention provides an adenosine deaminase that can act on DNA adenine deoxyribonucleotides, an adenine base editing system based on the adenosine deaminase and its application. A method for base editing a target sequence in an organism's genome using the base editing system, as well as a genetically modified organism and its offspring produced by the method. Background Art

[0004] Base editing refers to the process of replacing nucleotides at specific DNA sites through genetic engineering. Base editing is a type of genome editing technology. In plants, many excellent agronomic traits are produced by single nucleotide variations (SNVs); in humans and animals, many genetic diseases are caused by point mutations in functional genes. Base editing technology can be used to modify and transform the genomic DNA of animals, plants, and microorganisms to create new genotypes and obtain target traits that are beneficial to production applications. It can also be used to correct some serious congenital genetic variations in the treatment of genetic diseases. Therefore, base editing has important application prospects in the creation of excellent germplasm and disease treatment.

[0005] Adenine base editors (ABEs) are one of the most important base editing systems. In 2017, the ABE base editing system was first reported and successfully demonstrated efficient base editing in mammalian cells (Gaudelli, NM et al. Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage. Nature 551, 464-471 (2017).). The system consists of three components: 1) a nucleic acid targeting domain, such as the Cas9 nickase (nCas9, containing the D10A amino acid mutation), 2) an adenosine deaminase, and 3) sgRNA. The principle of ABE base editing is that, under the action of adenosine deaminase, adenine deoxynucleotide (A) is deaminated and converted into the intermediate inosine deoxynucleotide (I), which is ultimately converted to guanine deoxynucleotide (G) during cell repair and DNA replication, ultimately converting AT base pairs on double-stranded DNA to GC base pairs, thereby achieving base editing.

[0006] The key to ABE base editing is the use of adenosine deaminase that can act on DNA. However, the substrate of natural adenosine deaminase is RNA rather than DNA, so it cannot deaminate DNA. To solve this problem, the researchers used the tRNA adenosine deaminase TadA from Escherichia coli as the chassis protein and carried out multiple rounds of protein evolution. They eventually obtained the adenosine deaminase TadA*7.10 that can act on DNA and achieved the conversion of AT base pairs to GC base pairs. Although the researchers introduced some new amino acid mutations based on TadA*7.10 and obtained some new effective mutants, such as TadA-8e, TadA-8s, etc., these adenosine deaminases are still tRNA deaminases derived from Escherichia coli.

[0007] Since the ABE system was first reported in 2017, the adenosine deaminases used in the field have all been based on Escherichia coli TadA, and no other DNA-based adenosine deaminase chassis have been reported. Therefore, finding a DNA-based adenosine deaminase chassis is extremely important for expanding existing adenosine base editing systems and developing gene editing tools that can precisely manipulate target DNA sequences.

[0008] Summary of the Invention

[0009] To address the current problems of adenosine deaminases, which are limited in application scenarios and low editing efficiency, the inventors of this application, after extensive experimentation and exploration, unexpectedly discovered a new type of adenosine deaminase that can act on adenosine in DNA. Based on this discovery, the inventors developed a new type of adenine base editor, which enriches the application scenarios and options of adenine base editors.

[0010] Solutions for solving problems

[0011] In a first aspect of the present invention, an adenosine deaminase is provided, wherein the adenosine deaminase is capable of deaminating adenine bases in DNA, and the adenosine deaminase is selected from one or more of the following (i) to (iii):

[0012] (i) comprising an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 4, and which retains the deamination activity of the amino acid sequence shown in SEQ ID NO: 4;

[0013] (ii) an amino acid sequence comprising consecutive amino acids added to the C-terminus of the amino acid sequence shown in SEQ ID NO: 4, and which retains the deamination activity of the amino acid sequence shown in SEQ ID NO: 4; or

[0014] (iii) A multimer comprising any two or more amino acid sequences as described in (i) or (ii).

[0015] In some embodiments, in (ii), the adenosine deaminase comprises an amino acid sequence having at least 8 consecutive amino acids added to the C-terminus of SEQ ID NO: 4.

[0016] In some embodiments, in (ii), the adenosine deaminase comprises the amino acid sequence shown in SEQ ID NO:5.

[0017] In some embodiments, the adenosine deaminase comprises an amino acid sequence in which one or more amino acid residues are added, substituted, deleted, or inserted into the amino acid sequence shown in SEQ ID NO:4 or SEQ ID NO:5, and retains the deamination activity of the amino acid sequence shown in SEQ ID NO:4.

[0018] In some embodiments, the adenosine deaminase comprises an amino acid sequence having a W71X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except W (tryptophan), preferably, X is Y (tyrosine).

[0019] In some embodiments, the adenosine deaminase comprises an amino acid sequence having an A91X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except A (alanine). In some specific embodiments, X is T (threonine).

[0020] In some embodiments, the adenosine deaminase comprises an amino acid sequence having a V80X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except V (valine). In some specific embodiments, X is T (threonine).

[0021] In some embodiments, the adenosine deaminase comprises an amino acid sequence having an S122X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except S (serine). In some specific embodiments, X is G (glycine).

[0022] In some embodiments, the adenosine deaminase comprises an amino acid sequence having one or more mutations selected from the group consisting of a W71Y mutation, an A91T mutation, a V80T mutation, and an S122T mutation in SEQ ID NO: 4 or SEQ ID NO: 5.

[0023] In some embodiments, the adenosine deaminase comprises the amino acid sequence of SEQ ID NO: 4 or SEQ ID NO: 5 having a W71Y mutation and another adenine deaminase mutation.

[0024] In some specific embodiments, the adenosine deaminase comprises an amino acid sequence having W71Y and A91T mutations in SEQ ID NO:4 or SEQ ID NO:5.

[0025] In some specific embodiments, the adenosine deaminase comprises an amino acid sequence having W71Y and V80T mutations in SEQ ID NO:4 or SEQ ID NO:5.

[0026] In some specific embodiments, the adenosine deaminase comprises an amino acid sequence having W71Y and S122T mutations in SEQ ID NO:4 or SEQ ID NO:5.

[0027] In some embodiments, in (iii), the multimer comprises a dimer consisting of any two amino acid sequences as described in (i) and (ii), and the any two amino acid sequences as described in (i) and (ii) are connected by a linker.

[0028] In some specific embodiments, the linker comprises the amino acid sequence shown in SEQ ID NO:30.

[0029] In a second aspect of the present invention, provided is the use of the adenosine deaminase according to the first aspect of the present invention for gene editing in an organism or an organism cell, or for preparing a reagent for gene editing in an organism or an organism cell.

[0030] In a third aspect of the present invention, a fusion protein is provided, comprising:

[0031] (a) a nucleic acid targeting domain; and

[0032] (b) an adenosine deamination domain, wherein the adenosine deamination domain comprises at least one polypeptide of the adenosine deaminase according to the first aspect of the present invention.

[0033] In some embodiments, the nucleic acid targeting domain is a TALE, ZFP, or CRISPR effector protein domain.

[0034] In some embodiments, the CRISPR effector protein is selected from Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9, Cas10, Cast10d, Cas12, Cas13, Cas14, CasX, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE), Cse4 (or CasC), Csc1, Csc2, Csa5, Csn1, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Cpf1, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1, Csx15, Csf1, Csf2, Csf3, Csf4, Cu1966, TraC, and functional variants or any combination thereof.

[0035] In some embodiments, the nucleic acid targeting domain and the adenosine deamination domain are fused via a linker.

[0036] In some embodiments, the fusion protein further comprises a nuclear localization sequence (NLS).

[0037] In a fourth aspect of the present invention, a base editing system for modifying a target region of a nucleic acid molecule is provided, comprising:

[0038] i) the adenosine deaminase according to the first aspect of the present invention or the fusion protein according to the third aspect of the present invention, and / or an expression construct comprising a nucleotide sequence encoding the adenosine deaminase or the fusion protein.

[0039] In some embodiments, the base editing system further comprises:

[0040] ii) at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding said at least one guide RNA; and / or

[0041] iii) Nuclear localization sequence (NLS).

[0042] In some embodiments, the at least one guide RNA is bound to the nucleic acid targeting domain of the fusion protein, and the guide RNA is directed against at least one target sequence within the target region of the nucleic acid molecule.

[0043] In some embodiments, the guide RNA is 15-100 nucleotides in length and comprises a sequence of at least 10, at least 15, or at least 20 contiguous nucleotides that are complementary to the target sequence.

[0044] In some embodiments, the guide RNA comprises a 15 to 40 contiguous nucleotide sequence that is complementary to the target sequence.

[0045] In some embodiments, the guide RNA is 15-50 nucleotides in length.

[0046] In some embodiments, the nucleic acid molecule is DNA.

[0047] In some embodiments, the nucleic acid molecule is in the genome of an organism.

[0048] In some embodiments, the organism is a prokaryotic organism such as a bacterium; a eukaryotic organism such as a plant, a fungus, or a vertebrate.

[0049] In some embodiments, the vertebrate is a mammal such as a human, mouse, rat, monkey, dog, pig, sheep, cow, or cat.

[0050] In some embodiments, the plant is a crop plant, such as wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, kiwifruit, lettuce, tobacco, cassava, or potato.

[0051] In a fifth aspect of the present invention, a base editing method is provided, wherein the base editing method comprises contacting the base editing system described in the fourth aspect of the present invention with a nucleic acid molecule target sequence.

[0052] In some embodiments, the nucleic acid molecule is a double-stranded DNA molecule or a single-stranded DNA molecule.

[0053] In some embodiments, the nucleic acid molecule target sequence comprises a sequence associated with a plant trait or expression.

[0054] In some embodiments, the nucleic acid molecule target sequence comprises a sequence or point mutation associated with a disease or disorder.

[0055] In some embodiments, the base editing system contacts a target sequence of a nucleic acid molecule to exert a deamination effect, and the deamination effect causes one or more nucleotides in the target sequence to be replaced.

[0056] In some embodiments, the target sequence comprises the DNA sequence 5'-MAN-3', wherein M is A, T, C, or G; N is A, T, C, or G; and wherein the A in the middle of the 5'-MAN-3' sequence is deaminated.

[0057] In some embodiments, the deamination results in the introduction or removal of a splice site.

[0058] In some embodiments, the deamination results in the introduction of a mutation in a gene promoter that results in increased or decreased transcription of a gene operably linked to the gene promoter.

[0059] In some embodiments, the deamination results in the introduction of a mutation in the gene suppressor that results in increased or decreased transcription of a gene operably linked to the gene suppressor.

[0060] In some embodiments, the contacting is performed in vivo or wherein the contacting is performed in vitro.

[0061] In the fifth aspect of the present invention, a method for producing at least one genetically modified cell is provided, comprising introducing the base editing system described in the fourth aspect of the present invention into at least one of the cells, thereby causing one or more nucleotides in the target nucleic acid region in the at least one cell to be replaced, wherein the one or more nucleotide replacements are A to G replacements.

[0062] In some embodiments, the method further comprises the step of screening the at least one cell for cells having the desired one or more nucleotide substitutions.

[0063] In some embodiments, the base editing system is introduced into cells by a method selected from the following: calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus or other virus), N-acetylgalactosamine (GalNAc) mediation, gene gun method, PEG-mediated protoplast transformation, soil Agrobacterium-mediated transformation.

[0064] In some embodiments, the cells are from mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, cats; poultry such as chickens, ducks, geese; plants, preferably crop plants, such as wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, kiwifruit, lettuce, tobacco, cassava and potatoes.

[0065] Effects of the Invention

[0066] The present invention provides an adenosine deaminase that can act on DNA, and a gene editing system based on this deaminase and its applications. These systems enrich the application scenarios and options of adenine base editors and have broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] FIG1 shows the amino acid sequence similarity comparison results of QB34, wild-type TadA, and TadA-8e.

[0068] Figure 2 shows a schematic diagram of the structure of the QB34-based base editor expression construct.

[0069] Figure 3 shows the editing efficiency of the QB34-based adenine base editor.

[0070] Figure 4 shows the editing window of the QB34-based adenine base editor at the site 1 target site.

[0071] Figure 5 shows the editing window of the QB34-based adenine base editor at the site 3 target site.

[0072] Figure 6 shows a comparison of the editing efficiency of TadA- and QB34-based adenine base editors.

[0073] Figure 7 shows the editing window of the TadA- and QB34-based adenine base editor at the site 1 target site.

[0074] Figure 8 shows the editing window of the TadA- and QB34-based adenine base editor at the site 3 target site.

[0075] Figure 9 shows the editing efficiency of the adenine base editor based on the optimized QB34.

[0076] Figure 10 shows the editing window of the optimized QB34-based adenine base editor at the site 3 target site.

[0077] Figure 11 shows the editing window of the adenine base editor based on the QB34C mutant at the site 4 target site.

[0078] Figure 12 shows the editing window of the adenine base editor based on the QB34C mutant at the site 10 target site.

[0079] Figure 13 shows the editing efficiency of adenine base editors based on different dimer forms of the QB34C mutant.

[0080] Figure 14 shows the editing window of the adenine base editor at the site 1 target site based on different dimer forms of the QB34C mutant.

[0081] Figure 15 shows the editing window of the adenine base editor at the site 3 target site based on different dimer forms of the QB34C mutant.

[0082] Figure 16 shows the editing window of the adenine base editor at the site 10 target site based on different dimer forms of the QB34C mutant. DETAILED DESCRIPTION

[0083] 1. Definition

[0084] In the present invention, unless otherwise indicated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. In addition, the terms and laboratory procedures related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology used herein are terms and routine procedures widely used in the corresponding fields. For example, the standard recombinant DNA and molecular cloning techniques used in the present invention are well known to those skilled in the art and are more fully described in the following literature: Sambrook, J., Fritsch, EF and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter referred to as "Sambrook"). At the same time, in order to better understand the present invention, definitions and explanations of relevant terms are provided below.

[0085] As used herein, the term "and / or" encompasses all combinations of items connected by the term, and should be treated as if each combination had been individually listed herein. For example, "A and / or B" encompasses "A," "A and B," and "B." For example, "A, B, and / or C" encompasses "A," "B," "C," "A and B," "A and C," "B and C," and "A and B and C."

[0086] When the word "comprising" is used herein to describe a sequence of a protein or nucleic acid, the protein or nucleic acid may be composed of the sequence, or may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, but still have the activity described in the present invention. In addition, it is clear to those skilled in the art that the methionine encoded by the start codon at the N-terminus of the polypeptide may be retained in certain practical situations (for example, when expressed in a specific expression system), but it does not substantially affect the function of the polypeptide. Therefore, when describing a specific polypeptide amino acid sequence in the specification and claims of this application, although it may not contain a methionine encoded by a start codon at the N-terminus, a sequence containing the methionine is also covered, and accordingly, its encoding nucleotide sequence may also contain a start codon; and vice versa.

[0087] "Genome" as used herein encompasses not only the chromosomal DNA present in the nucleus of a cell, but also the organelle DNA present in subcellular components of the cell (eg, mitochondria, plastids).

[0088] As used herein, "organism" includes any organism suitable for genome editing, preferably a eukaryotic organism. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants, including monocots and dicots, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, and Arabidopsis.

[0089] "Genetically modified organism" or "genetically modified cell" refers to an organism or cell that contains an exogenous polynucleotide or a modified gene or expression control sequence within its genome. For example, the exogenous polynucleotide is capable of stably integrating into the genome of the organism or cell and being inherited through successive generations. The exogenous polynucleotide can be integrated into the genome alone or as part of a recombinant DNA construct. A modified gene or expression control sequence is one that contains single or multiple deoxynucleotide substitutions, deletions, and additions within the genome of the organism or cell.

[0090] "Exogenous" with respect to a sequence refers to a sequence that is from a foreign species, or, if from the same species, has been significantly altered from its native form in composition and / or locus through deliberate human intervention.

[0091] "Polynucleotide," "nucleic acid sequence," "nucleotide sequence," or "nucleic acid fragment" are used interchangeably and are single-stranded or double-stranded RNA or DNA polymers that optionally contain synthetic, non-natural, or altered nucleotide bases. Nucleotides are referred to by their single-letter names as follows: "A" for adenosine or deoxyadenosine (RNA or DNA, respectively), "C" for cytidine or deoxycytidine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, "R" for purine (A or G), "Y" for pyrimidine (C or T), "K" for G or T, "H" for A or C or T, "I" for inosine, and "N" for any nucleotide. The terms "adenine deoxyribonucleotide" and "adenosine" are used interchangeably herein.

[0092] "Polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. The terms apply to amino acid polymers in which one or more amino acid residues is an artificial chemical analog of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" may also include modified forms including, but not limited to, glycosylation, lipid attachment, sulfation, gamma-carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation.

[0093] According to the present invention, the amino acid three-letter code and the one-letter code used are as described in J.biol.chem, 243, p3558 (1968). The amino acids and their abbreviations and English abbreviations in the present invention are as follows: histidine (His, H); serine (Ser, S); glutamic acid (Glu, E); glutamine (Gln, Q); glycine (Gly, G); threonine (Thr, T); phenylalanine (Phe, F); aspartic acid (Asp, D); tyrosine (Tyr, Y); leucine (Leu, L); isoleucine (Ile, I); arginine (Arg, R); alanine (Ala, A); valine (Val, V); tryptophan (Trp, W); methionine (Met, M); asparagine (Asn, N); cysteine ​​(Cys, C); lysine (Lys, K); proline (Pro, P).

[0094] Sequence "identity" has a meaning recognized in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using published techniques. Sequence identity can be measured along the entire length of a polynucleotide or polypeptide or along a region of the molecule. (See, for example: Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While there are many methods to measure the identity between two polynucleotides or polypeptides, the term "identity" is well known to those of skill in the art (Carrillo, H. & Lipman, D., SIAM J Applied Math 48: 1073 (1988)).

[0095] In peptides or proteins, suitable conservative amino acid substitutions are known to those skilled in the art and can generally be made without altering the biological activity of the resulting molecule. Generally, those skilled in the art recognize that single amino acid substitutions in non-essential regions of a polypeptide do not substantially alter biological activity (see, e.g., Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub.co., p. 224).

[0096] As used herein, "expression construct" or "construct" refers to a vector, such as a recombinant vector, suitable for expressing a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, expression of a nucleotide sequence can refer to the transcription of the nucleotide sequence (e.g., transcription to produce mRNA or functional RNA) and / or translation of RNA into a precursor or mature protein.

[0097] The "expression construct" of the present invention can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, can be an RNA (such as mRNA) that can be translated.

[0098] An "expression construct" of the present invention may comprise regulatory sequences and a nucleotide sequence of interest from different sources, or regulatory sequences and a nucleotide sequence of interest from the same source but arranged in a manner different from that normally found in nature.

[0099] "Regulatory sequence" and "regulatory element" are used interchangeably to refer to nucleotide sequences located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding sequence and that influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences may include, but are not limited to, promoters, translation leader sequences, introns, and polyadenylation recognition sequences.

[0100] "Promoter" refers to a nucleic acid fragment that is capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the present invention, a promoter is a promoter that is capable of controlling the transcription of a gene in a cell, whether or not it is derived from the cell. A promoter can be a constitutive promoter, a tissue-specific promoter, a developmentally regulated promoter, or an inducible promoter.

[0101] "Constitutive promoter" refers to a promoter that will generally cause a gene to be expressed in most cell types under most circumstances. "Tissue-specific promoter" and "tissue-preferred promoter" are used interchangeably and refer to a promoter that is expressed primarily, but not necessarily exclusively, in one tissue or organ, and may also be expressed in one specific cell or cell type. "Developmentally regulated promoter" refers to a promoter whose activity is determined by developmental events. "Inducible promoter" selectively expresses an operably linked DNA sequence in response to endogenous or exogenous stimuli (environmental, hormonal, chemical signals, etc.).

[0102] As used herein, the term "operably linked" refers to the connection of a regulatory element (e.g., but not limited to, a promoter sequence, a transcription termination sequence, etc.) to a nucleic acid sequence (e.g., a coding sequence or an open reading frame) such that transcription of the nucleotide sequence is controlled and regulated by the transcriptional regulatory element. Techniques for operably linking regulatory element regions to nucleic acid molecules are known in the art.

[0103] "Introducing" a nucleic acid molecule (e.g., a plasmid, a linear nucleic acid fragment, RNA, etc.) or a protein into an organism refers to transforming an organism cell with the nucleic acid or protein so that the nucleic acid or protein can function in the cell. "Transformation" as used herein includes stable transformation and transient transformation.

[0104] "Stable transformation" refers to the introduction of an exogenous nucleotide sequence into a genome, resulting in the stable inheritance of the exogenous nucleotide sequence. Once stably transformed, the exogenous nucleic acid sequence is stably integrated into the genome of the organism and any successive generations thereof.

[0105] "Transient transformation" refers to the introduction of a nucleic acid molecule or protein into a cell where it functions without the exogenous nucleotide sequence being stably inherited. In transient transformation, the exogenous nucleic acid sequence does not integrate into the genome.

[0106] "Trait" refers to a physiological, morphological, biochemical, or physical characteristic of a cell or organism.

[0107] "Agronomic traits" specifically refer to measurable indicator parameters of crop plants, including but not limited to: leaf greenness, grain yield, growth rate, total biomass or accumulation rate, fresh weight at maturity, dry weight at maturity, fruit yield, seed yield, plant total nitrogen content, fruit nitrogen content, seed nitrogen content, plant vegetative tissue nitrogen content, plant total free amino acid content, fruit free amino acid content, seed free amino acid content, plant vegetative tissue free amino acid content, plant total protein content, fruit protein content, seed protein content, plant vegetative tissue protein content, herbicide resistance and drought resistance, nitrogen absorption, root lodging, harvest index, stem lodging, plant height, ear height, ear length, disease resistance, cold resistance, salt resistance and tiller number, etc.

[0108] 2. Adenosine deaminase

[0109] The term "deaminase" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminase provided herein (e.g., engineered adenosine deaminase, evolved adenosine deaminase) can be from any organism, such as bacteria. In some embodiments, the deaminase is a variant of a naturally occurring deaminase from an organism. In some embodiments, the deaminase does not exist in nature.

[0110] The term "dimer" refers to substances of the same or identical species that appear in pairs, potentially possessing properties or functions not present in the single form. In some embodiments, a dimer refers to two identical or different adenosine deaminases. In some embodiments, the dimers are connected by a linker. In some embodiments, the linker is a polypeptide.

[0111] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% identical to SEQ ID NO:4, and retains the deamination activity of the amino acid sequence shown in SEQ ID NO:4.

[0112] In some embodiments, the adenosine deaminase comprises an amino acid sequence that adds or subtracts consecutive amino acids at the C-terminus or N-terminus of SEQ ID NO: 4 without changing the function of the adenosine deaminase to deaminize adenine from deoxyadenosine in DNA (e.g., retaining the deamination activity of the amino acid sequence shown in SEQ ID NO: 4).

[0113] In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 8 consecutive amino acids added to the C-terminus of SEQ ID NO:4.

[0114] In some exemplary embodiments, the adenosine deaminase comprises the amino acid sequence shown in SEQ ID NO:5.

[0115] In some embodiments, the adenosine deaminase comprises an amino acid sequence having a W71X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except W (tryptophan). In some embodiments, X is Y (tyrosine). In some exemplary embodiments, the adenosine deaminase comprises an amino acid sequence as shown in SEQ ID NO: 6.

[0116] In some embodiments, the adenosine deaminase comprises an amino acid sequence having an A91X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except A (alanine). In some embodiments, X is T (threonine).

[0117] In some embodiments, the adenosine deaminase comprises an amino acid sequence having a V80X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except V (valine). In some embodiments, X is T (threonine).

[0118] In some embodiments, the adenosine deaminase comprises an amino acid sequence having an S122X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except S (serine). In some embodiments, X is G (glycine).

[0119] In some embodiments, the adenosine deaminase comprises an amino acid sequence having one or more mutations selected from the group consisting of a W71Y mutation, an A91T mutation, a V80T mutation, and an S122T mutation in SEQ ID NO: 4 or SEQ ID NO: 5.

[0120] In some embodiments, the adenosine deaminase comprises the amino acid sequence of SEQ ID NO: 4 or SEQ ID NO: 5 having a W71Y mutation and another adenine deaminase mutation.

[0121] In some embodiments, the adenosine deaminase comprises an amino acid sequence having W71Y and A91T mutations in SEQ ID NO: 4 or SEQ ID NO: 5. In some exemplary embodiments, the adenosine deaminase comprises an amino acid sequence as shown in SEQ ID NO: 7.

[0122] In some embodiments, the adenosine deaminase comprises an amino acid sequence having W71Y and V80T mutations in SEQ ID NO: 4 or SEQ ID NO: 5. In some exemplary embodiments, the adenosine deaminase comprises an amino acid sequence as shown in SEQ ID NO: 8.

[0123] In some embodiments, the adenosine deaminase comprises an amino acid sequence having W71Y and S122T mutations in SEQ ID NO: 4 or SEQ ID NO: 5. In some exemplary embodiments, the adenosine deaminase comprises an amino acid sequence as shown in SEQ ID NO: 9.

[0124] In some embodiments, the adenosine deaminase comprises a multimer consisting of any two or more of the amino acid sequences of the adenosine deaminase.

[0125] In some embodiments, the adenosine deaminase comprises a dimer composed of any two adenosine deaminases (ie, any two amino acid sequences of the adenosine deaminase).

[0126] In some embodiments, the dimer is formed by connecting any two adenosine deaminases (i.e., any two amino acid sequences of the adenosine deaminases) via a linker. In some embodiments, the linker is a 32 amino acid (32aa) polypeptide chain comprising the amino acid sequence shown in SEQ ID NO: 30.

[0127] In some exemplary embodiments, the adenosine deaminase comprises an amino acid sequence as shown in any one of SEQ ID NOs: 10-13.

[0128] 3. Base editing fusion proteins containing adenosine deaminase

[0129] In one aspect, the present application relates to the use of the adenosine deaminase of the present invention for gene editing in an organism or an organism cell, or the use of the adenosine deaminase of the present invention in the preparation of a reagent for gene editing in an organism or an organism cell. In some embodiments, the use may be in disease diagnosis or treatment methods. In other embodiments, the use may be in non-disease diagnosis or treatment methods.

[0130] In another aspect, the present invention provides a fusion protein comprising: (a) a nucleic acid targeting domain; and (b) an adenosine deamination domain, wherein the adenosine deamination domain comprises at least one adenosine deaminase polypeptide of the present invention.

[0131] In the embodiments herein, "fusion protein", "base editing fusion protein" and "base editor" are used interchangeably and refer to a reagent comprising a polypeptide that can modify a base (e.g., A, T, C, G, or U). In some embodiments, a base editor can deaminize a base within a nucleic acid. In some embodiments, a base editor can deaminize a base within a DNA molecule. In some embodiments, a base editor can deaminize adenine (A) in DNA.

[0132] As used herein, a "nucleic acid targeting domain" refers to a domain that can mediate the attachment of the base editing fusion protein to a specific target sequence in the genome in a sequence-specific manner (e.g., by a guide RNA). In some embodiments, the nucleic acid targeting domain can include one or more zinc finger protein domains (ZFPs) or transcription factor effector domains (TALEs) for a specific target sequence. In some embodiments, the nucleic acid targeting domain comprises at least one (e.g., one) CRISPR effector protein (CRISPR effector) polypeptide.

[0133] A "zinc finger protein domain (ZFP)" typically contains 3-6 individual zinc finger repeats, each of which can recognize a unique sequence of, for example, 3 bp. By combining different zinc finger repeats, different genomic sequences can be targeted.

[0134] A "transcription activator-like effector domain" is the DNA binding domain of a transcription activator-like effector (TALE). TALEs can be engineered to bind to virtually any desired DNA sequence.

[0135] As used herein, the term "CRISPR effector protein" generally refers to a nuclease (CRISPR nuclease) present in a naturally occurring CRISPR system or a functional variant thereof. The term encompasses any effector protein based on the CRISPR system that is capable of achieving sequence-specific targeting within a cell.

[0136] As used herein, a "functional variant" with respect to a CRISPR nuclease means that it at least retains the sequence-specific targeting ability mediated by the guide RNA. Preferably, the functional variant is a variant in which the nuclease is inactivated, i.e., it lacks double-stranded nucleic acid cleavage activity. However, CRISPR nucleases lacking double-stranded nucleic acid cleavage activity also encompass nickases, which form a nick in a double-stranded nucleic acid molecule but do not completely cut the double-stranded nucleic acid. In some preferred embodiments of the present invention, the CRISPR effector protein of the present invention has nickase activity. In some embodiments, the functional variant recognizes a different PAM (protospacer adjacent motif) sequence relative to the wild-type nuclease.

[0137] "CRISPR effector protein" can be derived from a Cas9 nuclease, including a Cas9 nuclease or a functional variant thereof. The Cas9 nuclease can be a Cas9 nuclease from a different species, such as spCas9 from Streptococcus pyogenes (S.pyogenes) or SaCas9 derived from Staphylococcus aureus (S.aureus). "Cas9 nuclease" and "Cas9" are used interchangeably herein and refer to an RNA-guided nuclease comprising a Cas9 protein or a fragment thereof (e.g., a protein comprising an active DNA cleavage domain of Cas9 and / or a gRNA binding domain of Cas9). Cas9 is a component of the CRISPR / Cas (clustered regularly interspaced short palindromic repeats and related systems) genome editing system that can target and cut a DNA target sequence to form a DNA double-strand break (DSB) under the guidance of a guide RNA. An exemplary amino acid sequence of wild-type SpCas9 is shown in SEQ ID NO: 16.

[0138] Useful "CRISPR effector proteins" can also be derived from Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9, Cas10, Cast10d, Cas12, Cas13, Cas14, CasX, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE), Nucleases such as Cse4 (or CasC), Csc1, Csc2, Csa5, Csn1, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Cpf1, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1, Csx15, Csf1, Csf2, Csf3, Csf4, Cu1966, and TraC, for example, include these nucleases or functional variants thereof. The TraC nuclease is disclosed in PCT / CN2023 / 097783 (Publication No. WO / 2023 / 232109), which is incorporated herein by reference.

[0139] The "CRISPR effector protein" can also be derived from a Cpf1 nuclease, including a Cpf1 nuclease or a functional variant thereof. The Cpf1 nuclease can be a Cpf1 nuclease from a different species, such as Francisella novicida U112, Acidaminococcus sp. BV3L6, and Lachnospiraceae bacterium ND2006.

[0140] In some embodiments, the CRISPR effector protein can be a nuclease-inactivated Cas9. The DNA cleavage domain of the Cas9 nuclease is known to comprise two subdomains: the HNH nuclease subdomain and the RuvC subdomain. The HNH subdomain cleaves the strand complementary to the guide RNA, while the RuvC subdomain cleaves the non-complementary strand. Mutations in these subdomains can inactivate the nuclease activity of Cas9, resulting in a "nuclease-inactivated Cas9." This nuclease-inactivated Cas9 still retains the ability to bind DNA under the guidance of the guide RNA.

[0141] The nuclease-inactivated Cas9 of the present invention can be derived from Cas9 of different species, for example, derived from Streptococcus pyogenes (S. pyogenes) Cas9 (SpCas9), or derived from Staphylococcus aureus (S. aureus) Cas9 (SaCas9). Simultaneously mutating the HNH nuclease subdomain and the RuvC subdomain of Cas9 (for example, comprising mutations D10A and H840A) inactivates the nuclease of Cas9, resulting in nuclease-inactivated Cas9 (dCas9). Mutating and inactivating one of the subdomains can impart nickase activity to Cas9, i.e., obtaining a Cas9 nickase (nCas9), for example, nCas9 having only the D10A mutation.

[0142] In some embodiments of the present invention, the adenosine deamination domain in the fusion protein is capable of converting the adenosine in the single-stranded DNA generated during the formation of the fusion protein-guide RNA-DNA complex into inosine I, and then achieving A to G base substitution through base mismatch repair.

[0143] In some embodiments of the invention, the nucleic acid targeting domain and the adenosine deamination domain are fused via a linker.

[0144] As used herein, a "linker" can be a non-functional amino acid sequence of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 20-25, 25-50) or more amino acids in length, without secondary or higher structure. For example, the linker can be a flexible linker.

[0145] In some embodiments of the present invention, the fusion protein of the present invention may further include a nuclear localization sequence (NLS). Generally speaking, the one or more NLS in the fusion protein should have sufficient strength to drive the fusion protein in the nucleus of the cell to accumulate in an amount that can realize its base editing function. Generally speaking, the intensity of nuclear localization activity is determined by the number, position, one or more specific NLS used in the fusion protein or a combination of these factors.

[0146] In some embodiments of the present invention, the NLS of the fusion protein of the present invention can be located at the N-terminus and / or the C-terminus. In some embodiments of the present invention, the NLS of the fusion protein of the present invention can be located between the adenine deamination domain and the nucleic acid targeting domain. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the N-terminus. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the C-terminus. In some embodiments, the polypeptide comprises a combination of these, such as comprising one or more NLSs at the N-terminus and one or more NLSs at the C-terminus. When more than one NLS is present, each can be selected to be independent of the other NLSs.

[0147] Generally speaking, NLS consists of one or more short sequences of positively charged lysine or arginine amino acids exposed on the surface of the protein, but other types of NLS are also known. Non-limiting examples of NLS include: KKRKV (SEQ ID NO: 33), PKKKRKV (SEQ ID NO: 34), KRPAATKKAGQAKKKK (SEQ ID NO: 35), or MKRTADGSEFEPKKKRKV (SEQ ID NO: 36).

[0148] In addition, depending on the DNA location to be edited, the fusion protein of the present invention may also include other localization sequences, such as cytoplasmic localization sequences, chloroplast localization sequences, mitochondrial localization sequences, etc.

[0149] 4. Genome Editing System

[0150] In another aspect, the present invention provides a base editing system comprising: i) an adenosine deaminase or base editing fusion protein of the present invention, and / or an expression construct containing a nucleotide sequence encoding the adenosine deaminase or base editing fusion protein.

[0151] In some embodiments, the base editing system is used to modify a target region of a nucleic acid.

[0152] In some embodiments, the base editing system further comprises ii) at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding the at least one guide RNA. However, one skilled in the art will appreciate that if the base editing fusion protein is not based on a CRISPR effector protein, the system may not require a guide RNA or an expression construct encoding the same.

[0153] In some embodiments, the at least one guide RNA can bind to the nucleic acid targeting domain of the fusion protein. In some embodiments, the guide RNA is directed against at least one target sequence within the nucleic acid target region.

[0154] As used herein, a "base editing system" refers to a combination of components required for base editing a nucleic acid sequence, such as a genomic sequence in a cell or organism. The individual components of the system, such as adenosine deaminase, a base editing fusion protein, and one or more guide RNAs, can exist independently or in any combination as a composite.

[0155] In some embodiments, it comprises an adenosine deaminase of the present invention or a fusion protein of the present invention and a guide RNA that can bind to a nucleic acid-targeting binding protein.

[0156] As used herein, "guide RNA" and "gRNA" are used interchangeably and refer to an RNA molecule that can form a complex with a CRISPR effector protein and can target the complex to a target sequence due to a certain homology with the target sequence. The guide RNA targets the target sequence by base pairing with the complementary strands of the target sequence. For example, the gRNA used by the Cas9 nuclease or its functional variants is generally composed of crRNA and tracrRNA molecules that are partially complementary to form a complex, wherein the crRNA comprises a guide sequence (also known as a seed sequence) that has sufficient homology to the target sequence so as to hybridize with the complementary strand of the target sequence and guide the CRISPR complex (Cas9+crRNA+tracrRNA) to specifically bind to the target sequence sequence. However, it is known in the art that single guide RNA (sgRNA) can be designed, which includes the features of crRNA and tracrRNA at the same time. The gRNA used by the Cpf1 nuclease or its functional variants is generally composed of only mature crRNA molecules, which can also be referred to as gRNA. It is within the capabilities of those skilled in the art to design suitable gRNA based on the other CRISPR nucleases used and the target sequence to be edited.

[0157] In some embodiments, the guide RNA is 15-100 nucleotides in length and comprises a sequence of at least 10, at least 15, or at least 20 consecutive nucleotides that are complementary to the target sequence.

[0158] In some embodiments, the guide RNA comprises a 15 to 40 contiguous nucleotide sequence that is complementary to the target sequence.

[0159] In some embodiments, the guide RNA is 15-50 nucleotides in length.

[0160] In some embodiments, the target sequence is a DNA sequence.

[0161] In some embodiments, wherein the target sequence is in the genome of an organism. In some embodiments, wherein the organism is a prokaryotic organism. In some embodiments, wherein the prokaryotic organism is a bacterium. In some embodiments, wherein the organism is a eukaryotic organism. In some embodiments, wherein the organism is a plant or a fungus. In some embodiments, wherein the organism is a vertebrate. In some embodiments, wherein the vertebrate is a mammal. In some embodiments, wherein the mammal is a mouse, a rat, or a human. In some embodiments, wherein the organism is a cell. In some embodiments, wherein the cell is a mouse cell, a rat cell, or a human cell. In some embodiments, wherein the cell is a HEK-293T cell.

[0162] In some embodiments, after the base editing system of the present invention is introduced into the cell, the base editing fusion protein and the guide RNA are able to form a complex, and the complex specifically targets the target sequence under the mediation of the guide RNA, and causes one or more A in the target sequence to be replaced by G.

[0163] In some embodiments, the at least one guide RNA can be directed against a target sequence on a sense strand (e.g., a protein-coding strand) and / or an antisense strand located within a genomic target nucleic acid region. When the guide RNA targets the sense strand (e.g., a protein-coding strand), the base editing composition of the present invention can cause one or more A's within the target sequence on the sense strand (e.g., a protein-coding strand) to be replaced by G's. When the guide RNA targets the antisense strand, the base editing composition of the present invention can cause one or more T's within the target sequence on the sense strand (e.g., a protein-coding strand) to be replaced by C's.

[0164] In order to obtain efficient expression in cells, in some embodiments of the present invention, the nucleotide sequence encoding the adenosine deaminase or base editing fusion protein is codon-optimized for the organism whose genome is to be modified.

[0165] Codon optimization refers to a process of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon of the native sequence (e.g., about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) with codons that are more frequently or most frequently used in the genes of the host cell, while maintaining the native amino acid sequence. Different species exhibit specific preferences for certain codons for specific amino acids. Codon bias (differences in codon usage between organisms) is often correlated with the efficiency of translation of messenger RNA (mRNA), which is believed to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell generally reflects the codons that are most frequently used for peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the Codon Usage Database available at www.kazusa.orjp / codon / , and these tables can be adapted in various ways. See, Nakamura et al., 2002, 2010, 2011, 2012, 2013, 2014, 2015, 2016, 2017, 2018, 2019, 2020, 2021, 2022, 2023, 2024, 2025, 2026, 2027, 2028, 2029, 2030, 2031, 2032, 2033, 2034, 2035, 2036, 2037, 2038, 2039, 2040, 2041, 2042, 2043, 2044 Y. et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).

[0166] Organisms that can be genomic modified using the base editing system of the present invention include any organism suitable for base editing, preferably eukaryotic organisms. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants, including monocots and dicots. For example, the plants are crop plants, including but not limited to wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, kiwis, lettuce, tobacco, cassava, and potatoes.

[0167] 5. Base Editing Methods

[0168] In another aspect, the present invention provides a base editing method, comprising contacting the base editing system of the present invention with a nucleic acid molecule target sequence.

[0169] In some embodiments, the nucleic acid molecule is a DNA molecule. In some preferred embodiments, the nucleic acid molecule is a double-stranded DNA molecule or a single-stranded DNA molecule.

[0170] In some embodiments, the nucleic acid molecule target sequence comprises a sequence associated with a plant trait or expression.

[0171] In some embodiments, the nucleic acid molecule target sequence comprises a sequence or point mutation associated with a disease or disorder.

[0172] In some embodiments, the base editing system contacts a target sequence of a nucleic acid molecule and exerts a deamination effect, wherein the deamination effect causes one or more nucleotides in the target sequence to be replaced.

[0173] In some embodiments, the target sequence comprises the DNA sequence 5'-MAN-3', wherein M is A, T, C, or G; N is A, T, C, or G; and wherein the A in the middle of the 5'-MAN-3' sequence is deaminated.

[0174] In some embodiments, the deamination results in the introduction or removal of a splice site.

[0175] In some embodiments, the deamination results in the introduction of a mutation in a gene promoter, the mutation resulting in increased or decreased transcription of a gene operably linked to the gene promoter.

[0176] In some embodiments, the deamination results in the introduction of a mutation in the gene suppressor that results in increased or decreased transcription of a gene operably linked to the gene suppressor.

[0177] In some embodiments, the contacting occurs in vivo.

[0178] In some embodiments, the contacting is performed in vitro.

[0179] VI. Methods for Producing Genetically Modified Cells

[0180] In another aspect, the present invention also provides a method for producing at least one genetically modified cell, comprising introducing a base editing system of the present invention into at least one of the cells, thereby causing one or more nucleotide substitutions within a target nucleic acid region in the at least one cell. In some embodiments, the one or more nucleotide substitutions are A to G substitutions.

[0181] In some embodiments, the method further comprises the step of screening the at least one cell for cells having the desired one or more nucleotide substitutions.

[0182] In some embodiments, the methods of the present invention are performed in vitro. For example, the cells are isolated cells, or cells in isolated tissues or organs.

[0183] In another aspect, the present invention also provides a genetically modified organism comprising a genetically modified cell or progeny thereof produced by the method of the present invention. Preferably, the genetically modified cell or progeny thereof has a desired one or more nucleotide substitutions.

[0184] In the present invention, the target nucleic acid region to be modified can be located at any position in the genome, for example, in a functional gene such as a protein-coding gene, or, for example, in a gene expression regulatory region such as a promoter region or an enhancer region, thereby achieving modification of the gene function or modification of gene expression. In some embodiments, the desired nucleotide substitution results in a desired gene function modification or gene expression modification.

[0185] In some embodiments, the target nucleic acid region is related to the proterties of the cell or organism. In some embodiments, the mutation in the target nucleic acid region causes a change in the proterties of the cell or organism. In some embodiments, the target nucleic acid region is located in the coding region of an albumen. In some embodiments, the function-related motif or domain of the target nucleic acid region encoding protein. In some preferred embodiments, one or more nucleotide substitutions in the target nucleic acid region cause an amino acid substitution in the amino acid sequence of the albumen. In some embodiments, one or more nucleotide substitutions cause a change in the function of the albumen.

[0186] In the method of the present invention, the base editing system can be introduced into cells by various methods well known to those skilled in the art.

[0187] Methods that can be used to introduce the base editing system of the present invention into cells include, but are not limited to, calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus and other viruses), N-acetylgalactosamine (GalNAc)-mediated, gene gun method, PEG-mediated protoplast transformation, and Agrobacterium-mediated transformation.

[0188] Cells that can be base edited by the methods of the present invention can be from, for example, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, cats; poultry such as chickens, ducks, geese; plants, including monocots and dicots, preferably crop plants, including but not limited to wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, kiwis, lettuce, tobacco, cassava and potatoes.

[0189] 7. Application in plants

[0190] The base editing fusion proteins, base editing systems, and methods for generating genetically modified cells of the present invention are particularly suitable for genetically modifying plants. Preferably, the plants are crop plants, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, kiwifruit, lettuce, tobacco, cassava, and potato.

[0191] In another aspect, the present invention provides a method for producing a genetically modified plant, comprising introducing the base editing system of the present invention into at least one of the plants, thereby causing one or more nucleotide substitutions within a target nucleic acid region in the genome of the at least one plant.

[0192] In some embodiments, the method further comprises screening the at least one plant for plants having the desired one or more nucleotide substitutions.

[0193] In the method of the present invention, the base editing composition can be introduced into the plant by various methods well known to those skilled in the art. Methods that can be used to introduce the base editing system of the present invention into plants include, but are not limited to, gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube channel method, and ovary injection method. Preferably, the base editing composition is introduced into the plant by transient transformation.

[0194] In the method of the present invention, modification of the target sequence can be achieved by simply introducing or producing the base editing fusion protein and guide RNA in plant cells, and the modification can be stably inherited without the need to stably transform the plant with exogenous polynucleotides encoding the components of the base editing system. This avoids the potential off-target effects of the stably existing (continuously produced) base editing composition and also avoids the integration of exogenous nucleotide sequences into the plant genome, thereby having higher biosafety.

[0195] In some preferred embodiments, the introduction is performed in the absence of selection pressure, thereby avoiding integration of the exogenous nucleotide sequence into the plant genome.

[0196] In some embodiments, the introduction comprises transforming the base editing system of the present invention into isolated plant cells or tissues, and then regenerating the transformed plant cells or tissues into complete plants. Preferably, the regeneration is performed in the absence of selection pressure, that is, no selection agent for the selection gene carried on the expression vector is used during the tissue culture process. Not using a selection agent can improve the regeneration efficiency of the plant and obtain a modified plant that does not contain exogenous nucleotide sequences.

[0197] In other embodiments, the base editing system of the present invention can be transformed into specific parts of intact plants, such as leaves, stem tips, pollen tubes, young ears, or hypocotyls. This is particularly suitable for the transformation of plants that are difficult to regenerate through tissue culture.

[0198] In some embodiments of the present invention, in vitro expressed proteins and / or in vitro transcribed RNA molecules (e.g., the expression construct is an in vitro transcribed RNA molecule) are directly transformed into the plant. The protein and / or RNA molecule can be base edited in plant cells and then degraded by the cells, thereby avoiding the integration of exogenous nucleotide sequences in the plant genome.

[0199] Therefore, in some embodiments, genetic modification and breeding of plants using the methods of the present invention can obtain plants whose genomes have no exogenous polynucleotide integration, ie, non-transgenic (transgene-free) modified plants.

[0200] In some embodiments of the invention, the modified target nucleic acid region is associated with a plant trait such as an agronomic trait, whereby the one or more nucleotide substitutions result in the plant having an altered (preferably improved) trait, such as an agronomic trait, relative to a wild-type plant.

[0201] In some embodiments, the method further comprises the step of screening plants for desired one or more nucleotide substitutions and / or desired traits, such as agronomic traits.

[0202] In some embodiments of the present invention, the method further comprises obtaining offspring of the genetically modified plant. Preferably, the genetically modified plant or its offspring has a desired one or more nucleotide substitutions and / or a desired trait such as an agronomic trait.

[0203] In another aspect, the present invention also provides a genetically modified plant, or a progeny thereof, or a part thereof, wherein the plant is obtained by the method of the present invention described above. In some embodiments, the genetically modified plant, or a progeny thereof, or a part thereof is non-transgenic. Preferably, the genetically modified plant, or a progeny thereof, has a desired genetic modification and / or a desired trait, such as an agronomic trait.

[0204] In another aspect, the present invention also provides a plant breeding method, comprising crossing a first genetically modified plant obtained by the above-described method of the present invention, comprising one or more nucleotide substitutions in a target nucleic acid region, with a second plant that does not contain the one or more nucleotide substitutions, thereby introducing the one or more nucleotide substitutions into the second plant. Preferably, the genetically modified first plant has a desired trait, such as an agronomic trait.

[0205] 8. Therapeutic Applications

[0206] The present invention also covers the use of the base editing system of the present invention in the treatment of diseases or the preparation of drugs for the treatment of diseases.

[0207] By modifying the disease-related genes through the base editing system of the present invention, it is possible to achieve upregulation, downregulation, inactivation, activation or mutation correction of the disease-related genes, thereby achieving the prevention and / or treatment of the disease. For example, the target nucleic acid region described in the present invention can be located in the protein coding region of the disease-related gene, or, for example, can be located in a gene expression regulatory region such as a promoter region or an enhancer region, so as to achieve functional modification of the disease-related gene or modification of the expression of the disease-related gene. Therefore, the modified disease-related genes described herein include modification of the disease-related gene itself (e.g., protein coding region), and also include modification of its expression regulatory region (e.g., promoter, enhancer, intron, etc.).

[0208] "Disease-associated" gene refers to any gene that produces a transcription or translation product at an abnormal level or in an abnormal form in cells derived from tissues affected by the disease, compared to tissues or cells of non-disease controls. In the case where the altered expression is related to the appearance and / or progression of the disease, it can be a gene expressed at an abnormally high level; it can be a gene expressed at an abnormally low level. Disease-associated genes also refer to genes with one or more mutations or genetic variations that are directly responsible for or disequilibrium with one or more genes responsible for the etiology of the disease. The mutation or genetic variation is, for example, a single nucleotide variation (SNV). The transcribed or translated product can be known or unknown and can be at normal or abnormal levels.

[0209] Therefore, the present invention also provides a method for treating a disease in a subject in need thereof, comprising delivering an effective amount of a base editing system of the present invention to the subject to modify a gene associated with the disease (e.g., deaminating mitochondrial DNA by a fusion protein or multiple fusion proteins). The present invention also provides the use of a base editing system in the preparation of a pharmaceutical composition for treating a disease in a subject in need thereof, wherein the base editing system is used to modify a gene associated with the disease. The present invention also provides a pharmaceutical composition for treating a disease in a subject in need thereof, comprising a base editing system of the present invention, and an optional pharmaceutically acceptable carrier, wherein the base editing system is used to modify a gene associated with the disease.

[0210] In some embodiments, the fusion protein or base editing system described in the present invention is used to introduce point mutations into nucleic acids by deaminating target nuclear bases (e.g., A residues). In some embodiments, the deamination of target nuclear bases leads to the correction of genetic defects, such as in point mutations that cause loss of function in gene products during correction. In some embodiments, genetic defects are associated with diseases or conditions (e.g., lysosomal storage diseases or metabolic diseases, such as, for example, type I diabetes). In some embodiments, the methods provided herein can be used to introduce inactivating point mutations into genes or alleles encoding gene products associated with diseases or conditions.

[0211] In some embodiments, the purpose of the scheme described in the present invention is to restore the function of dysfunctional genes via genome editing. The nucleobase editing proteins provided herein are for use in human cell in vitro gene editing, such as correcting disease-related mutations in human cell cultures. The nucleobase editing proteins provided herein, for example, fusion proteins containing nucleic acid editable DNA proteins (e.g., CRISPR effector proteins Cas9) and adenosine deaminase domains can be used to correct any single-point G to A or C to T mutations. In the first case, the A of the mutant is corrected for mutation by deamination, while in the latter case, the A paired with the mutant T is corrected for mutation by deamination and subsequent rounds of replication.

[0212] In some embodiments, the purpose of the scheme described in the present invention is to treat diseases associated with or caused by point mutations, which can be corrected by the DNA base editing fusion proteins provided herein. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a neoplastic disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a lysosomal storage disease.

[0213] In some embodiments, the methods described herein are intended to be used to treat mitochondrial diseases or disorders. As used herein, "mitochondrial diseases" refer to diseases caused by abnormal mitochondria, such as mutations in mitochondrial genes, enzyme pathways, etc. Examples of diseases include, but are not limited to, neurological diseases, loss of motor control, muscle weakness and pain, gastrointestinal diseases and difficulty swallowing, poor growth, heart disease, liver disease, diabetes, respiratory complications, epilepsy, vision / hearing problems, lactic acidosis, developmental delays, and susceptibility to infection.

[0214] Examples of diseases described herein include, but are not limited to, genetic diseases, circulatory system diseases, muscle diseases, brain, central nervous and immune system diseases, Alzheimer's disease, secretase disorders, amyotrophic lateral sclerosis (ALS), autism, trinucleotide repeat expansion disorders, hearing diseases, gene-targeted therapy of non-dividing cells (neurons, muscles), liver and kidney diseases, epithelial cell and lung diseases, cancer, Usher syndrome or retinitis pigmentosa-39, cystic fibrosis, HIV and AIDS, beta thalassemia, sickle cell disease, herpes simplex virus, autism, drug addiction, age-related macular degeneration, schizophrenia. Other diseases that can be treated by correcting point mutations or introducing inactivating mutations into disease-related genes are known to those skilled in the art, and therefore the present disclosure is not limited in this regard. In addition to the diseases exemplarily described herein, other related diseases can also be treated using the strategies and fusion proteins provided by the present invention, and this application will be apparent to those skilled in the art. The diseases or targets to which the present invention can be applied refer to WO2015089465A1 (PCT / US2014 / 070135), WO2016205711A1 (PCT / US2016 / 038181), WO2018141835A1 (PCT / EP2018 / 052491), WO2020191234A1 (PCT / US2020 / 023713), WO2020191233A1 (PCT / US2020 / 023712), WO2019079347A1 (PCT / US2018 / 056146), and WO2021155065A1 (PCT / US2021 / 015580) for the related diseases to which the base editing systems are applicable.

[0215] The administration of the base editing system or pharmaceutical composition of the present invention can be adjusted according to the weight and species of the patient or subject. The frequency of administration is within the range permitted by medical or veterinary medicine. It depends on conventional factors including the age, sex, general health, other conditions of the patient or subject, and the specific condition or symptom being addressed.

[0216] IX. Test Kit

[0217] The present invention also includes a kit for use in the methods of the present invention, comprising a genome editing system of the present invention, and instructions for use. The kit generally includes a label indicating the intended use and / or instructions for use of the contents of the kit. The term label includes any written or recorded material provided on or with the kit or otherwise provided with the kit.

[0218] Example

[0219] Example 1: Acquisition of a new adenosine deaminase

[0220] The inventors used bioinformatics technology to utilize all functional annotations related to TAD (tRNA adenine deaminase) in public databases. Through sequence filtering, AlphaFold2 algorithm protein structure analysis, and TMalign structure alignment, 91 candidate proteins were initially screened. Among them was a potential adenosine deaminase QB34 (SEQ ID NO: 4). Blast results showed that the new deaminase protein had low amino acid sequence similarity with the reported wild-type TadA (SEQ ID NO: 1) and TadA-8e (SEQ ID NO: 2) with DNA deamination function. The sequence similarity alignment results are shown in Figure 1. The amino acid sequence identity of QB34 and TadA is 40.0%, and the amino acid sequence identity of QB34 and TadA-8e is 35.2%.

[0221] Example 2: DNA editing function of QB34 adenosine deaminase

[0222] To verify whether the above-mentioned adenosine deaminase can produce an editing effect on DNA in cells. In this example, an adenine base editor based on 91 candidate proteins including QB34 adenosine deaminase was constructed (the amino acid sequence of QB34 adenine base editor is shown in SEQ ID NO: 20). The schematic diagram of the expression construct structure of the QB34 adenine base editor is shown in Figure 2. The sgRNA was designed for the target human embryonic kidney cell 293T (HEK293T) genome, targeting the site3 (SEQ ID NO: 15) site of the genomic target DNA. The adenine base editor based on the candidate protein was transfected into HEK293T cells, and DNA was extracted 72 hours after transfection; the mock group was used as a blank control without transfection of the adenine editor plasmid, and the other treatments were the same as above. The base editing efficiency of the target was detected using second-generation sequencing technology. The results showed that among the 91 candidate proteins verified, QB34 showed significant editing efficiency compared with the blank control group (Table 1).

[0223] Table 1 Base editing efficiency of adenine base editors of candidate proteins at site 3

[0224] Subsequently, the inventors designed sgRNA targeting the genome of human embryonic kidney 293T cells (HEK293T) to target site 1 (SEQ ID NO: 14) of the genomic target DNA. HEK293T cells were transfected with an adenine base editor based on QB34 adenosine deaminase, and DNA was extracted 72 hours after transfection. A mock control group was treated as a blank control without transfection of the adenine editor plasmid. The results showed that the editing efficiency of the QB34 adenosine deaminase at the target site 1 reached 0.75%, and the editing efficiency at the target site 3 reached 1.44% (Figure 3).

[0225] Further analysis of the editing window showed that QB34 adenosine deaminase edits sites including A2, A4, A5, A6, A9, A11, and A12 at site 1; and edits sites including A2, A3, A5, A7, A8, A9, A12, A14, and A16 at site 3 ( Figures 4 and 5 ), indicating a wide editing window.

[0226] In addition, the inventors compared QB34 adenosine deaminase with wild-type TadA adenosine deaminase (Figures 6-8). The results showed that wild-type TadA, as a tRNA adenosine deaminase, has an editing efficiency close to zero in DNA, while QB34 has a significantly higher editing efficiency than TadA (Figure 6). Comparing the editing windows of the two adenosine deaminases at site 1 and site 3 (Figures 7-8), TadA's editing efficiency at every editing site that QB34 has editing ability is less than 0.1%, and even close to zero in most cases.

[0227] This shows that QB34 has the function of DNA adenosine deaminase and can serve as a potential deaminase element of adenine base editor.

[0228] Example 3: Optimizing the DNA editing function of QB34 adenosine deaminase

[0229] Screening in the E. coli evolution system revealed that a variant, QB34C (SEQ ID NO: 5), obtained by extending the C-terminus of QB34 by 8 amino acids, could improve the DNA editing efficiency of QB34 deaminase. These 8 amino acids were derived from the original coding frame of QB34. On this basis, key site mutations such as W71Y, A91T, V80T, and S122G were performed on QB34C, and the mutations were combined to test whether the DNA editing efficiency of QB34 deaminase could be further improved. The expression construct structure and fusion protein sequence of the base editor are shown in Table 2, and the preparation method is shown in Example 2.

[0230] The same method as in Example 2 was used to construct an ABE system based on nCas9 nickase (Cas9 (D10A)). The base editing efficiency of QB34C and QB34C mutant adenine base editors on HEK293T cells at site1, site3, site4 (SEQ ID NO: 18), and site10 (SEQ ID NO: 19) targets was detected. The results of the second-generation sequencing technology showed that the editing efficiency of QB34C at the site3 target site was 1.83%, and the editing efficiency of the single mutant QB34C-W71Y based on QB34C was also slightly improved compared to QB34C. It can be seen that mutating the key sites of QB34C can improve the editing efficiency (Figure 9).

[0231] The inventors combined single mutation sites and found that the QB34C double mutant had comparable or significantly improved editing efficiency compared to QB34C. For example, the double mutant QB34C-W71Y-S122G mutant had an editing efficiency of up to 21.82% at the site 3 target site. In addition, the QB34C-W71Y-S122G mutant had an editing efficiency of 5.61% at the site 1 target site, 20.53% at the site 4 target site, and 18.1% at the site 10 target site (Figure 9). This shows that extending the C-terminus of the QB34 adenosine deaminase and single or multiple mutants containing W71Y, A91T, V80T, and S122G significantly improves the editing efficiency of DNA and has higher editing activity.

[0232] In addition, QB34, QB34C and their mutants have a lower indel (random insertion / deletion) frequency than existing adenosine deaminases, and their indel frequencies are all below 0.2% (Table 3). The results show that when the QB34 series adenosine deaminases are used as the deaminase domain of the base editor to perform deamination editing on specific bases, the frequency of random base insertion / deletion is low, that is, no or less other types of non-target editing are introduced. In summary, the QB34 series adenosine deaminases have the characteristics of high specificity and high editing safety compared to the adenosine deaminases in the prior art.

[0233] The inventors further analyzed the editing windows of QB34 functional variants. Figure 10 shows that for the site 3 target, the editing sites of QB34 and QB34C include A2, A3, A5, A7, and A8. In comparison, the base editors corresponding to the QB34C-W71Y single mutant and the QB34C-W71Y-S122G, QB34C-W71Y-A91T, and QB34C-W71Y-V80T double mutants only include A2, A3, A5, and A7 in the site 3 target window, indicating that the QB34C mutant has a narrower editing window, which is conducive to improving the precision of base editing.

[0234] Figures 11-12 show the editing status of target sites 4 and 10, respectively. The study showed that the QB34C-W71Y-S122G mutant edited sites A5 and A9 at target site 4 and A5 at target site 10 (Figures 11-12). In particular, only one site at site 10 was edited, and the editing efficiency was high. These results indicate that the QB34C-W71Y-S122G double mutant has a narrower editing window and higher editing activity, demonstrating its potential for precise gene editing.

[0235] Table 2 QB34-based and optimized QB34 adenosine deaminase editor fusion protein framework structure and amino acid sequence

[0236] Table 3 Indel efficiency of QB34-based and optimized QB34 adenosine deaminase editors

[0237] Example 4: DNA editing function of QB34C adenosine deaminase dimer

[0238] In this example, QB34C adenosine deaminase and its preferred mutants constitute a deaminase dimer, and the dimer is used as the adenosine deaminase domain to be linked to the adenine base editor. The framework structure and amino acid sequence of the QB34C dimer base editor fusion protein are shown in Table 2, and the schematic diagram of the expression construct structure is shown in Figure 2.

[0239] Results showed that QB34C adenosine deaminase dimers could further enhance DNA editing efficiency. The (QB34C-W71Y-S122G)-(QB34C-W71Y-S122G) double mutant dimer exhibited the highest editing efficiency, achieving 16.46% editing efficiency at target site 1, 42.33% editing efficiency at target site 3, and 31.59% editing efficiency at target site 10 (Figure 13). This efficiency was nearly threefold higher than that of the (QB34C-W71Y)-(QB34C-W71Y) single mutant dimer. Furthermore, both the (QB34C-W71Y-S122G)-(QB34C-W71Y-S122G) double mutant dimer and the single mutant dimer introduced low levels of indels (Table 3), demonstrating the high specificity and safety of the deaminase dimer. Further analysis of its editing window revealed that the QB34C adenosine deaminase dimer contained editing sites A2, A4, A5, A6, and A9 at the site 1 target site (Figure 14), the editing window at site 3 contained editing sites A2, A3, A5, and A7 (Figure 15), and the editing window at site 10 contained A5 (Figure 16). This shows that the adenosine deaminase dimer also has the same editing window as its adenosine deaminase.

[0240] In summary, the QB34C adenosine deaminase dimer, when used as the deaminase domain of a base editor, has improved editing activity while maintaining its precision and safety.

[0241] Some of the sequences involved in this disclosure:

[0242] >SEQ ID NO:1 TadA

[0243] >SEQ ID NO:2 TadA-8e

[0244] >SEQ ID NO:3 TadA*7.10

[0245] >SEQ ID NO:4 QB34

[0246] >SEQ ID NO:5 QB34C

[0247] >SEQ ID NO:6 QB34C-W71Y

[0248] >SEQ ID NO:7 QB34C-W71Y-A91T

[0249] >SEQ ID NO:8 QB34C-W71Y-V80T

[0250] >SEQ ID NO:9 QB34C-W71Y-S122G

[0251] >SEQ ID NO:10(QB34C)-(QB34C-W71Y)

[0252] >SEQ ID NO:11(QB34C-W71Y)-(QB34C)

[0253] >SEQ ID NO:12(QB34C-W71Y)-(QB34C-W71Y)

[0254] >SEQ ID NO:13(QB34C-W71Y-S122G)-(QB34C-W71Y-S122G)

[0255] >SEQ ID NO:14 site1(HEK293T)

[0256] >SEQ ID NO:15 site3(HEK293T)

[0257] >SEQ ID NO:16 SpCas9

[0258] >SEQ ID NO:17 nCas9(D10A)

[0259] >SEQ ID NO:18 site4(HEK293T)

[0260] >SEQ ID NO:19 site10(HEK293T)

[0261] >SEQ ID NO:20 QB34-ABE

[0262] >SEQ ID NO:21 QB34C-ABE

[0263] >SEQ ID NO:22 QB34C-W71Y-ABE

[0264] >SEQ ID NO:23 QB34C-A91T-ABE

[0265] >SEQ ID NO:24 QB34C-V80T-ABE

[0266] >SEQ ID NO:25 QB34C-S122G-ABE

[0267] >SEQ ID NO:26(QB34C-W71Y-S122G)-(QB34C-W71Y-S122G)-ABE

[0268] >SEQ ID NO:27(QB34C-W71Y)-(QB34C-W71Y)-ABE

[0269] >SEQ ID NO:28(QB34C)-(QB34C-W71Y)-ABE

[0270] >SEQ ID NO:29(QB34C-W71Y)-(QB34C)-ABE

[0271] >SEQ ID NO:30 32aa

[0272] >SEQ ID NO:31 NLS

[0273] >SEQ ID NO:32 TraC

[0274] >SEQ ID NO:33 NLS

[0275] >SEQ ID NO:34 NLS

[0276] >SEQ ID NO:35 NLS

[0277] >SEQ ID NO:36 NLS

Claims

1. An adenosine deaminase capable of deaminating adenine bases in DNA, wherein the adenosine deaminase is selected from one or more of the following (i) to (iii): (i) comprising an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence shown in SEQ ID NO:4, and which retains the deamination activity of the amino acid sequence shown in SEQ ID NO:4; (ii) an amino acid sequence comprising consecutive amino acids added to the C-terminus of the amino acid sequence shown in SEQ ID NO: 4, and which retains the deamination activity of the amino acid sequence shown in SEQ ID NO: 4; or (iii) A multimer comprising any two or more amino acid sequences as described in (i) or (ii).

2. The adenosine deaminase according to claim 1, wherein In (ii), the adenosine deaminase comprises an amino acid sequence having at least 8 consecutive amino acids added to the C-terminus of SEQ ID NO:

4.

3. The adenosine deaminase according to claim 2, wherein In (ii), the adenosine deaminase comprises the amino acid sequence shown in SEQ ID NO:

5.

4. The adenosine deaminase according to any one of claims 1 to 3, wherein The adenosine deaminase comprises an amino acid sequence in which one or more amino acid residues are added, substituted, deleted or inserted into the amino acid sequence shown in SEQ ID NO:4 or SEQ ID NO:5, and retains the deamination activity of the amino acid sequence shown in SEQ ID NO:

4.

5. The adenosine deaminase according to claim 4, wherein The adenosine deaminase comprises an amino acid sequence having a W71X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except W (tryptophan), preferably, X is Y (tyrosine).

6. The adenosine deaminase according to any one of claims 1 to 5, wherein The adenosine deaminase comprises an amino acid sequence having an A91X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except A (alanine).

7. The adenosine deaminase according to claim 6, wherein X is T (threonine).

8. The adenosine deaminase according to any one of claims 1 to 7, wherein The adenosine deaminase comprises an amino acid sequence having a V80X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except V (valine).

9. The adenosine deaminase according to claim 8, wherein X is T (threonine).

10. The adenosine deaminase according to any one of claims 1 to 9, wherein The adenosine deaminase comprises an amino acid sequence having an S122X mutation in SEQ ID NO: 4 or SEQ ID NO: 5, wherein X is any amino acid except S (serine).

11. The adenosine deaminase according to claim 10, wherein X is G (glycine).

12. The adenosine deaminase according to any one of claims 1 to 11, wherein The adenosine deaminase comprises an amino acid sequence having one or more mutations selected from the group consisting of W71Y mutation, A91T mutation, V80T mutation and S122T mutation in SEQ ID NO: 4 or SEQ ID NO:

5.

13. The adenosine deaminase according to any one of claims 1 to 12, wherein The adenosine deaminase comprises an amino acid sequence having a W71Y mutation in SEQ ID NO: 4 or SEQ ID NO: 5 and another mutation of an adenine deaminase.

14. The adenosine deaminase according to any one of claims 1 to 13, wherein The adenosine deaminase comprises an amino acid sequence having W71Y and A91T mutations in SEQ ID NO:4 or SEQ ID NO:

5.

15. The adenosine deaminase according to any one of claims 1 to 13, wherein The adenosine deaminase comprises an amino acid sequence having W71Y and V80T mutations in SEQ ID NO:4 or SEQ ID NO:

5.

16. The adenosine deaminase according to any one of claims 1 to 13, wherein The adenosine deaminase comprises an amino acid sequence having W71Y and S122T mutations in SEQ ID NO:4 or SEQ ID NO:

5.

17. The adenosine deaminase according to any one of claims 1 to 16, wherein In (iii), the multimer comprises a dimer consisting of any two amino acid sequences as described in (i) and (ii), and the any two amino acid sequences as described in (i) and (ii) are connected by a linker.

18. The adenosine deaminase according to claim 17, wherein The linker comprises the amino acid sequence shown in SEQ ID NO:

30.

19. Use of the adenosine deaminase according to any one of claims 1 to 18 for gene editing in an organism or an organism cell, or use of the adenosine deaminase according to any one of claims 1 to 18 for preparing a reagent for gene editing in an organism or an organism cell.

20. A fusion protein comprising: (a) a nucleic acid targeting domain; and (b) an adenosine deamination domain, wherein the adenosine deamination domain comprises at least one polypeptide of the adenosine deaminase according to any one of claims 1 to 18.

21. The fusion protein according to claim 20, wherein The nucleic acid targeting domain is a TALE, ZFP or CRISPR effector protein domain.

22. The fusion protein according to claim 21, wherein The CRISPR effector protein is selected from Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9, Cas10, Cast10d, Cas12, Cas13, Cas14, CasX, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE ), Cse4 (or CasC), Csc1, Csc2, Csa5, Csn1, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Cpf1, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1, Csx15, Csf1, Csf2, Csf3, Csf4, Cu1966, TraC, and their functional variants or any combination thereof.

23. The fusion protein according to any one of claims 20 to 22, wherein The nucleic acid targeting domain and the adenosine deamination domain are fused via a linker.

24. The fusion protein according to any one of claims 20 to 22, wherein The fusion protein also comprises a nuclear localization sequence (NLS).

25. A base editing system for modifying a target region of a nucleic acid molecule, comprising: i) the adenosine deaminase according to any one of claims 1 to 18 or the fusion protein according to any one of claims 20 to 24, and / or an expression construct containing a nucleotide sequence encoding the adenosine deaminase or the fusion protein.

26. The base editing system of claim 25, wherein: The base editing system further comprises: ii) at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding the at least one guide RNA; and / or iii) Nuclear localization sequence (NLS).

27. The base editing system of claim 26, wherein: The at least one guide RNA is bound to the nucleic acid targeting domain of the fusion protein, and the guide RNA is directed against at least one target sequence within the target region of the nucleic acid molecule.

28. The base editing system of claim 27, wherein: The guide RNA is 15-100 nucleotides in length and comprises a sequence of at least 10, at least 15, or at least 20 consecutive nucleotides that are complementary to the target sequence.

29. The base editing system of claim 28, wherein: The guide RNA comprises a 15 to 40 consecutive nucleotide sequence complementary to the target sequence.

30. The base editing system according to any one of claims 26-29, wherein: The guide RNA has a length of 15-50 nucleotides.

31. The base editing system according to any one of claims 26 to 30, wherein: The nucleic acid molecule is DNA.

32. The base editing system of any one of claims 26-31, wherein: The nucleic acid molecule is in the genome of an organism.

33. The base editing system of claim 32, wherein: The organism is a prokaryotic organism such as a bacterium; a eukaryotic organism such as a plant, a fungus or a vertebrate.

34. The base editing system of claim 33, wherein: The vertebrate is a mammal such as a human, mouse, rat, monkey, dog, pig, sheep, cow, or cat.

35. The base editing system of claim 33, wherein: The plant is a crop plant, for example wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, kiwifruit, lettuce, tobacco, cassava or potato.

36. A base editing method, wherein: The base editing method comprises contacting the base editing system of any one of claims 25-35 with a nucleic acid molecule target sequence.

37. The base editing method according to claim 36, wherein: The nucleic acid molecule is a double-stranded DNA molecule or a single-stranded DNA molecule.

38. The base editing method according to any one of claims 36-37, wherein: The nucleic acid molecule target sequence comprises a sequence associated with a plant trait or expression.

39. The base editing method according to any one of claims 36-37, wherein: The nucleic acid molecule target sequence comprises a sequence or point mutation associated with a disease or disorder.

40. The base editing method according to any one of claims 36 to 39, wherein: The base editing system contacts a target sequence of a nucleic acid molecule and exerts a deamination effect, which causes one or more nucleotides in the target sequence to be replaced.

41. The base editing method according to any one of claims 36 to 40, wherein: The target sequence comprises a DNA sequence 5'-MAN-3', wherein M is A, T, C or G; N is A, T, C or G; and A in the middle of the 5'-MAN-3' sequence is deaminated.

42. The base editing method according to any one of claims 36 to 41, wherein: The deamination results in the introduction or removal of a splice site.

43. The base editing method according to any one of claims 36 to 42, wherein: The deamination results in the introduction of mutations in the gene promoter, which result in increased or decreased transcription of a gene operably linked to the gene promoter.

44. The base editing method according to any one of claims 36 to 43, wherein: The deamination results in the introduction of mutations in the gene suppressor that result in increased or decreased transcription of a gene operably linked to the gene suppressor.

45. The base editing method according to any one of claims 36 to 44, wherein: The contacting is performed in vivo or wherein the contacting is performed in vitro.

46. ​​A method for producing at least one genetically modified cell, comprising introducing the base editing system of any one of claims 25-35 into at least one of the cells, thereby causing one or more nucleotides in the target nucleic acid region in the at least one cell to be replaced, wherein the one or more nucleotide replacements are A to G replacements.

47. The method of claim 46, wherein: The method further comprises the step of screening the at least one cell for cells having the desired one or more nucleotide substitutions.

48. The method according to claim 46 or 47, wherein: The base editing system is introduced into cells by a method selected from the following: calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus or other virus), N-acetylgalactosamine (GalNAc) mediation, gene gun method, PEG-mediated protoplast transformation, soil Agrobacterium-mediated transformation.

49. The method according to any one of claims 46 to 48, wherein: The cells are from mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, cats; poultry such as chickens, ducks, geese; plants, preferably crop plants, such as wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, kiwifruit, lettuce, tobacco, cassava and potatoes.