Novel adenine deaminase mutant and base editing method using the same

Adenine deaminase mutants and fusion proteins with specific amino acid substitutions address the off-target editing issues in ABEs, providing precise A-to-G base editing in DNA with reduced side effects, improving the reliability and applicability of base editing compositions.

JP2025532113APending Publication Date: 2025-09-29INST FOR BASIC SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025517270
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2023-07-04
Publication Date
2025-09-29

AI Technical Summary

Technical Problem

Existing adenine base editors (ABEs) face challenges in delivering guide RNA to organelles for targeted editing of mitochondrial DNA, leading to off-target base editing due to residual deaminase activity on RNA substrates, resulting in undesired modifications.

Method used

Development of adenine deaminase mutants and fusion proteins comprising DNA-binding proteins to reduce off-target editing effects, utilizing specific amino acid substitutions in TadA8e to minimize unintended base modifications.

Benefits of technology

The adenine deaminase mutants and fusion proteins significantly reduce off-target editing, enabling precise A-to-G base editing in DNA with minimal side effects, enhancing the reliability and applicability of base editing compositions for therapeutic and scientific applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025532113000001_ABST
    Figure 2025532113000001_ABST
Patent Text Reader

Abstract

The present invention relates to a novel adenine deaminase mutant, a fusion protein comprising the adenine deaminase mutant, a base editing composition comprising the fusion protein for A-to-G base editing in DNA, and a method for A-to-G base editing in DNA, the method comprising delivering the base editing composition to a cell containing target DNA. The novel adenine deaminase mutant can significantly reduce off-target effects, including unintended base modifications in DNA and / or RNA, and can induce base editing only at a single nucleotide residue in target DNA without any unintended off-target editing.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention provides an adenine deaminase mutant that can reduce off-target editing; a fusion protein comprising a DNA-binding protein and the adenine deaminase mutant; a base editing composition for A-to-G base editing in DNA comprising the fusion protein; and a method for A-to-G base editing in DNA, comprising the step of delivering the base editing composition to a cell containing target DNA. [Background technology]

[0002] Targeted base editing of mammalian mitochondrial DNA (mtDNA) is a powerful and versatile technique that can be used to model mitochondrial genetic diseases in cell lines and animals and to develop new therapeutic approaches to correct pathogenic mutations in patients. Programmable deaminase enzymes, consisting of custom-made DNA-binding proteins and nucleobase deamidates, enable targeted editing of mtDNA.

[0003] Unlike cytosine base editors, adenine base editors (ABEs), also known as TALEDs, contain TadA8e, a deoxyadenine deaminase engineered from the tRNA-specific TadA protein derived from Escherichia coli (E. coli). TadA8e and related variants of TadA are key components of CRISPR RNA-guided adenine base editors, which are widely used for A-to-G base editing of nuclear DNA. Nevertheless, editing organelle DNA using RNA-guided ABEs is challenging due to the difficulty of delivering guide RNA to organelles. Furthermore, TadA8e present in ABEs maintains residual deaminase activity on RNA substrates, resulting in unintended, transcriptome-wide off-target base editing. Off-target effects occur when each editing apparatus mistakenly acts on unintended genomic regions, resulting in undesired modifications.

[0004] Despite widespread interest in programmable deaminase-mediated base editing, there is currently a lack of strategies developed for adenine base editors with improved specificity and methods that can reduce the occurrence of off-target editing effects. Developing these base editors that minimize such side effects and implementing methods for reducing them is of considerable importance, enabling broader applicability with greater reliability for scientific and therapeutic applications. Summary of the Invention [Problem to be solved by the invention]

[0005] The present invention relates to base editing compositions and methods for their application in gene therapy and genome engineering. More specifically, the compositions and methods of the present invention comprise base editor variants that, when used in a clinical setting, exhibit significantly reduced side effects, such as off-target genome cleavage and RNA off-target toxicity. By introducing a novel utilization approach involving an innovative adenine base editor, the present invention provides a new avenue for targeted base editing with broad applications in the fields of medicine and biotechnology. [Means for solving the problem]

[0006] In one embodiment of the present invention, adenine deaminase mutants are provided herein.

[0007] In another embodiment of the present invention, provided herein is a fusion protein comprising a DNA binding protein and an adenine deaminase mutant.

[0008] In another embodiment of the present invention, provided herein are polynucleotides encoding adenine deaminase mutants or fusion proteins and expression vectors comprising said polynucleotides.

[0009] In one embodiment of the present invention, the present specification describes a base editing composition for A-to-G base editing in DNA, comprising a fusion protein, a polynucleotide encoding the fusion protein, or an expression vector comprising the polynucleotide.

[0010] In another embodiment of the invention, provided herein is a method for A-to-G base editing in DNA, comprising delivering a base editing composition to a cell containing target DNA.

[0011] In another embodiment of the present invention, the present specification describes a method for reducing off-target editing effects, comprising delivering a base editing composition to a cell containing target DNA.

[0012] Other embodiments of the invention provide uses of adenine deaminase variants or fusion proteins in A-to-G base editing of DNA and / or reducing off-target effects of A-to-G base editing of DNA, or in manufacturing compositions for A-to-G base editing of DNA, and / or reducing off-target effects of A-to-G base editing of DNA. [Brief explanation of the drawings]

[0013] [Figure 1a] Figure 1 shows the structures of the base editors used in the present invention: AD (TadA8e adenine deaminase); AD* (TadA8e adenine deaminase mutant); MTS (mitochondrial targeting sequence); UGI (uracil glycosylase inhibitor). [Figure 1b] Graph showing the number of RNA edits and editing frequency for Cox3.1-specific sTALEDs. [Figure 1c] Graphs showing the number of A-to-G edits and C-to-T edits for Cox3.1-specific sTALEDs (left), and the targeting activity of Cox3.1-specific sTALED mutants relative to wild-type Cox3.1-specific sTALED (right). [Figure 1d] Graph showing the number of RNA edits and editing frequency for ND1-specific sTALEDs. [Figure 1e] Graphs showing the number of A-to-G edits and C-to-T edits for ND1-specific sTALEDs (left), and the targeting activity of ND1-specific sTALED mutants relative to wild-type ND1-specific sTALED (right). [Figure 2(A)] FIG. 1 is a structural representation of the TadA portion of ABE8e (Protein Data Bank (PDB) accession number 6VPC). [Figure 2(B)]Heatmap showing the DNA on-target activity (left), RNA off-target activity (middle), and the relative ratio of DNA on-target to RNA off-target editing frequency (right) of 101 sTALED mutants that maintain mitochondrial DNA on-target activity out of a total of 209 sTALED mutants. The relative ratios were normalized to the ratio of the original sTALED, which has a value of 1. [Figure 3] Graph showing DNA on-target base editing frequencies induced by 209 sTALED mutants. [Figure 4(A)] 1 is a graph showing the editing frequency at six RNA off-target sites as measured by targeted RNA sequencing. [Figure 4(B)] This graph shows the frequency of RNA off-target editing induced by sTALED and sTALED variants identified by transcriptome-wide sequencing at six selected RNA off-target sites. [Figure 4(C)] Graph showing the frequency of RNA off-target editing induced by sTALED and sTALED mutants analyzed by targeted RNA amplicon sequencing. [Figure 5a] Graph showing the number of RNA edits and editing frequency for Cox3.1-specific sTALEDs. [Figure 5b] Graph showing A-to-G editing and C-to-T editing for Cox3.1-specific sTALEDs. [Figure 5c] Graph showing the number of RNA edits and editing frequency for ND1-specific sTALEDs. [Figure 5d] Graph showing A-to-G editing and C-to-T editing for ND1-specific sTALEDs. [Figure 5e] Graph showing the number of RNA edits and editing frequency for ND6-specific sTALEDs. [Figure 5f]Graph showing A-to-G editing and C-to-T editing for ND6-specific sTALEDs. [Figure 6(A)] Graph showing the on-target activity of ND1-specific sTALED mutants relative to wild-type ND1-specific sTALED. [Figure 6(B)] Graph showing the on-target activity of ND6-specific sTALED mutants relative to wild-type ND6-specific sTALED. [Figure 6(C)] Heatmap showing RNA off-target activity of ND1-specific sTALED and sTALED mutants at six representative sites. [Figure 6(D)] Heatmap showing RNA off-target activity of ND6-specific sTALED and sTALED mutants at six representative sites. [Figure 6(E)] 1 is a bar graph showing the frequency of RNA off-target editing induced by ND1-specific sTALED and sTALED mutants at six representative sites. The editing efficiency of each replicate is represented by a point on the graph. Error bars are sem for n = 2 (replicates) × 6 (RNA off-target sites) biologically independent samples. [Figure 6(F)] 1 is a bar graph showing the frequency of RNA off-target editing induced by ND6-specific sTALED and sTALED mutants at six sites. The editing efficiency of each replicate is represented by a point on the graph. Error bars are sem for n = 2 (replicates) × 6 (RNA off-target sites) biologically independent samples. [Figure 6(G)] A bar graph showing the ratio of DNA on-target editing frequency relative to RNA off-target editing frequency induced by ND1-specific sTALED and sTALED mutants. [Figure 6(H)] 1 is a bar graph showing the ratio of DNA on-target editing frequency relative to RNA off-target editing frequency induced by ND6-specific sTALED and sTALED mutants, normalized to the ratio for sTALED, which has a value of 1. [Figure 6(I)] Graph showing the average relative ratio values ​​for sTALEDs and sTALED mutants targeting Cox3.1, ND1 (Figure 6(G)), and ND6 (Figure 6(H)). [Figure 7a] Heatmap showing A-to-G transitions generated by sTALED or sTALED mutants targeting the Cox3.1 site. [Figure 7b] Heatmap showing A-to-G transitions generated by sTALED or sTALED mutants targeting the ND1 site. [Figure 7c] Heatmap showing A-to-G transitions generated by sTALED or sTALED mutants targeting the ND6 site. [Figures 7d-7f]

[0033] Figure 7a-c show analyses of Cox3.1, ND1, and ND6 alleles summarized in Figures 7a-7c, respectively. Spacer sequences are shown on the left, and bar graphs displaying the frequency of each allele are shown on the right. Reference sequences are all written in uppercase, and lowercase letters indicate positions where base editing occurred. [Figure 8(A)] Figure 1 shows a plot of the locations of on-target and off-target edits across the mitochondrial genome at day 4 post-transfection. Black and gray dots represent off-target edits and naturally occurring single-nucleotide variations (SNVs), respectively, and arrows indicate on-target (and bystander) edits. Nucleotide positions in the human mitochondrial genome are represented on the x-axis. [Figure 8(B)] Figure 1 shows the average frequency of genome-wide off-target editing induced by wild-type TALED and TALED mutants. [Figure 9(A)]Figure 1 shows a plot of the locations of on-target and off-target edits across the mitochondrial genome at day 2 post-transfection. Black and gray dots represent off-target edits and naturally occurring single-nucleotide variations (SNVs), respectively, and arrows indicate on-target (and bystander) edits. Nucleotide positions in the human mitochondrial genome are represented on the x-axis. [Figure 9(B)] Graph showing the average frequency of genome-wide off-target edits induced by wild-type TALED and TALED mutants. Error bars are sem for n=2 biologically independent samples. [Figures 10(A)-10(C)] Line graphs showing the frequency of on-target base edits induced by (A) Cox3.1-, (B) ND1-, and (C) ND6-specific sTALEDs and sTALED mutants (V106W, V28R, R111S) over time. Error bars are sem for n=2 biologically independent samples. [Figures 10(D)-10(F)] Line graphs showing the frequency of RNA off-target base edits induced by (D) Cox3.1-, (E) ND1-, and (F) ND6-specific sTALEDs and sTALED mutants (V106W, V28R, R111S) over time at six representative sites. Error bars are sem for n=2 biologically independent samples. [Figure 10(G)] 1 is a line graph showing the average RNA off-target editing frequency over time for Cox3.1-, ND1-, and ND6-specific sTALEDs and sTALED mutants. Error bars are sem for n=2 biologically independent samples. [Figure 11(A)] FIG. 1 illustrates an experimental design for one embodiment of the present invention. [Figure 11(B)-11(C)]Bar graphs showing the viability of cells transfected with plasmids expressing sTALED, sTALED-V106W, sTALED-V28R, and sTALED-R111S, as determined by the color change due to formazan formation in an MTS assay at 2 (B) and 4 (C) days post-transfection. Absorbance values ​​were normalized to the absorbance value of cells transfected with pEGFP as a control. Error bars are s.e.m. for n=2 biologically independent samples. [Figure 12(A)] Figure 1 shows the structures of ABE8e and ABE8e mutant constructs: AD (TadA8e adenine deaminase); AD* (TadA8e adenine deaminase mutant); NLS (nuclear localization sequence). [Figure 12(B)] 1 is a graph showing the on-target activity of ABE8e and ABE8e mutants (ABE8e-V106W, ABE8e-V28R, ABE8e-R111S) at the nuclear TYRO3 site. [Figure 12(C)] 1 is a heat map showing the frequency of A-to-G transversions generated by ABE8e and ABE8e mutants (ABE8e-V106W, ABE8e-V28R, ABE8e-R111S) at the nuclear TYRO3 locus. [Figure 12(D)] Graph showing RNA off-target activity of TYRO3-targeted ABE8e and ABE8e mutants (ABE8e-V106W, ABE8e-V28R, ABE8e-R111S) at six representative sites. [Figure 12(E)-12(F)] Graph showing the total number of RNA edits discovered in HEK293T cells expressing ABE8e or ABE8e mutants targeted to the nuclear TYRO3 site as assessed through transcriptome-wide sequencing. [Figure 13(a)] This is a graph showing the DNA on-target activity of Cox3-specific TALEDs, including dimeric TALEDs (dTALEDs), semi-monomeric TALEDs (dTALED-ADs), monomeric TALEDs (mTALEDs), and untreated samples. [Figure 13(b)] This is a graph showing the RNA off-target activity of Cox3-specific TALEDs, including dimeric TALEDs (dTALEDs), semi-monomeric TALEDs (dTALED-ADs), monomeric TALEDs (mTALEDs), and untreated samples, at six representative sites. [Figure 13(c)] A graph showing the specificity ratio of RNA off-target editing to on-target editing induced by Cox3-specific TALEDs. DETAILED DESCRIPTION OF THE INVENTION

[0014] The following definitions supplement those in the art and are relevant to this application and should not be attributed to, for example, any commonly owned patent or application, whether related or unrelated. Although any methods and materials similar or equivalent to those described herein can be used in the practice for testing herein, the preferred materials and methods are described herein. Therefore, the terminology used herein is for the purpose of describing particular examples only and is not intended to be limiting. Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs.

[0015] As used herein, the use of the singular includes the plural unless expressly stated otherwise. It should be noted that as used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. As used herein, the use of "or" means "and / or" and is understood to be inclusive unless expressly stated otherwise. Additionally, the use of the term "including" as well as other forms such as "include" and "included" is not limiting.

[0016] As used herein, terms such as "about" or "approximately" mean within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which may vary in part depending on the manner in which the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, according to practice in the art. Alternatively, "about" can mean within a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Or, particularly in relation to biological systems or processes, the term can mean within a range of magnitude, such as within 5-fold or within 2-fold of a value. When specific values ​​are described in the specification and claims, unless otherwise specified, the term "about" should be assumed to mean within an acceptable error range for the particular value.

[0017] As used herein, the term "corresponding" refers to an amino acid residue at a position listed in a polypeptide, or an amino acid residue similar, identical, or homologous to one listed in a polypeptide. Identifying an amino acid at a corresponding position can refer to determining a specific amino acid in a sequence that references a particular sequence. As used herein, "corresponding region" generally refers to a similar or corresponding position in a related or reference protein. For example, any amino acid sequence can be aligned to SEQ ID NO: 3, and based on this, each amino acid residue in the amino acid sequence can be numbered with reference to the amino acid residue in SEQ ID NO: 3 and the numeric position of the corresponding amino acid residue. For example, the sequence alignment algorithms described herein can determine amino acid positions or positions where modifications such as substitutions, insertions, or deletions occur through comparison with positions in a query sequence (also referred to as a "reference sequence").

[0018] As used herein, the term "alignment" refers to mapping a sequence read to a reference genome and then aligning bases that share identical positions in the genome. Therefore, any computer program can be used as long as it can align sequence reads in the same manner as described above. The program may be one already known in the art or may be selected from programs tailored to the purpose. In one embodiment of the present invention, the alignment is performed using ISAAC, but is not limited thereto.

[0019] In this specification, references to "various embodiments," "some embodiments," "embodiments," "one embodiment," or "other embodiments" mean that the particular features, structures, or characteristics described in connection with the embodiments include at least some embodiments herein, but not necessarily all embodiments.

[0020] The term "host cell" (or "recombinant host cell"), as used herein, is intended to refer to a cell that has been genetically modified or is capable of being genetically modified by the introduction of an exogenous polynucleotide molecule, such as a recombinant plasmid or vector. Such terms should be understood to refer not only to the particular subject cell but also to the progeny of such a cell. Because specific modifications may occur in successive generations due to mutation or environmental influences, such progeny may not actually be identical to the parent cell, but are nevertheless included within the scope of the term "host cell" as used herein.

[0021] As used herein, the term "base editor (BE)" refers to an agent that binds to a polynucleotide and has nucleobase-modifying activity. In one embodiment of the present invention, the base editor comprises a nucleobase-modifying polypeptide (e.g., a deaminase) and a nucleic acid-programmable nucleotide-binding domain. In another embodiment of the present invention, the base editor comprises a nucleic acid-programmable nucleotide-binding domain together with a nucleobase-modifying polypeptide (e.g., a deaminase) and a guide polynucleotide (e.g., a guide RNA). However, in other embodiments of the present invention, the agent is a biomolecular complex comprising a protein domain with base-editing activity, i.e., a domain that can modify a base (e.g., A, T, C, G, I, or U) in a nucleic acid molecule (e.g., DNA). In some embodiments of the present invention, a polynucleotide-programmable DNA-binding domain is fused or linked to a deaminase domain. In one embodiment of the present invention, the agent is a fusion protein comprising a domain with base-editing activity. In other embodiments of the present invention, the protein domain with base-editing activity is linked to a guide RNA (e.g., via an RNA-binding motif on the guide RNA and an RNA-binding domain fused to a deaminase). In some embodiments of the present invention, the domain having base editing activity is capable of deaminating a base in a nucleic acid molecule. In some embodiments of the present invention, the base editor is capable of deaminating one or more bases in a DNA molecule. In some embodiments of the present invention, the base editor is capable of deaminating adenine (A) in DNA. In some embodiments of the present invention, the base editor is an adenine base editor (ABE).

[0022] As used herein, "administration" means providing one or more compositions described herein to a patient or subject. By way of example and without limitation, administration of a composition, e.g., injection, can be intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection, and can be accomplished using one or more of these routes. Parenteral administration can be, for example, a bolus injection or a gradual infusion over time. In some embodiments of the present invention, parenteral administration includes intravascular, intravenous, intramuscular, intraarterial, intraspinal, intratumoral, intradermal, intraperitoneal, intraorgan, subcutaneous, subcuticular, intraarticular, subcapsular, subarachnoid, and intrasternal infusion or injection. Oral administration can be alternated or simultaneous.

[0023] The expression "other amino acids" as used herein may be intended to refer to amino acids selected from alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, asparagine, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartic acid, glutamic acid, arginine, histidine, lysine, and all known variants thereof, excluding amino acids that maintain the wild-type protein at the original substitution position.

[0024] As used herein, the term "off-target site" may refer to a site where an adenine base editor exhibits activity, but is not an on-target site. That is, the off-target site may refer to a site other than the on-target site where base editing occurs. In one embodiment of the present invention, the term "off-target site" may be used to encompass not only a site that is not an on-target site for an adenine base editor, but also a site that may be an off-target site.

[0025] As used herein, the term "whole genome sequencing (WGS)" refers to a method of reading a genome at various multiples, such as 10X, 20X, and 40X formats, for whole genome sequencing through next-generation sequencing. The term "next-generation sequencing" refers to a technology that fragments the whole genome or target regions of the genome in chip-based and PCR-based paired-end formats and sequences the fragments in a chemical reaction (hybridization)-based high-throughput manner.

[0026] As used herein, the term "nucleic acid" refers to DNA or RNA. "Nucleic acid sequence" or "polynucleotide sequence" refers to a single- or double-stranded polymer of deoxyribonucleotide or ribonucleotide bases reading from the 5' to the 3' end. This term includes all autonomously replicating plasmids, infectious polymers of DNA or RNA, and non-functional DNA or RNA.

[0027] The phrase "nucleic acid molecule encoding" refers to a nucleic acid molecule that directs the expression of a specific protein or peptide. The nucleic acid sequence includes both the DNA strand sequence that is transcribed into RNA and the RNA sequence that is translated into the protein or peptide. The nucleic acid molecule includes not only full-length nucleic acid sequences but also non-full-length sequences derived from full-length proteins. It should also be understood that the sequence includes the native sequence or degenerate codons of each sequence that may be introduced to provide codon preference in a particular host cell.

[0028] The term "vector" refers to viral expression systems, autonomous self-replicating original DNA (plasmids), and includes both expression and non-expression plasmids. When a recombinant microorganism or cell is described as carrying an "expression vector," this includes both extrachromosomal original DNA and DNA integrated into a host chromosome. When a vector is maintained by a host cell, the vector can either be an autonomous structure, stably replicated by the cell during cell division, or integrated into the host's genome.

[0029] The term "plasmid" refers to an autonomous, prototypic DNA molecule capable of replicating within a cell, and includes both expression and non-expression types. When a recombinant microorganism or cell is described as carrying an "expression plasmid," this includes latent viral DNA integrated into the host chromosome. When a plasmid is maintained by a host cell, it is either stably replicated by the cell during cell division as an autonomous structure, or integrated into the host's genome.

[0030] As used herein, "percent amino acid sequence identity" or "percent amino acid sequence identity" refers to the sequence identity between a first amino acid sequence and a second amino acid sequence. This can be calculated by dividing the number of amino acid residues in the first amino acid sequence that are identical to the amino acid residues at the corresponding positions in the second amino acid sequence by the total number of amino acid residues in the first amino acid sequence, and then multiplying this result by 100%. Each deletion, insertion, substitution, or addition of an amino acid residue in the second amino acid sequence is considered a difference at a single amino acid residue (position) compared to the first amino acid sequence, i.e., an "amino acid difference" as defined herein. Alternatively, the degree of sequence identity between two amino acid sequences can be recalculated using standard settings using known computer algorithms, such as those mentioned above for determining the degree of sequence identity of nucleotide sequences.

[0031] In various embodiments of the present invention, a process for selecting TadA8e mutants with reduced RNA or DNA off-target effects (including bystander off-target effects) was described to address the problem of undesired off-target effects associated with TadA8e. In one embodiment of the present invention, this goal was achieved by substituting amino acid residues at specific positions within TadA8e that interact with nucleotides. Furthermore, to assess the impact of these mutants on RNA or DNA off-target effects, RNA sequencing was performed in addition to whole mitochondrial genome sequencing. Through this comprehensive analysis, the entire spectrum of RNA off-target sites or a representative selection of six significant RNA off-target sites was identified and confirmed. Furthermore, recognizing the temporal nature of RNA, the dynamics of RNA off-target effects over time were measured. These measurements were performed at various time points to assess how the RNA off-target profile changed during expression.

[0032] In one embodiment of the present invention, there is provided an adenine deaminase comprising the amino acid sequence set forth in SEQ ID NO:1 or an amino acid sequence having at least 80% sequence identity to the amino acid sequence set forth in SEQ ID NO:1, wherein at least one amino acid residue selected from residues 28, 30, 46, 48, 49, 82, 84, 106, 108, 110, and 111 of SEQ ID NO:1 or a corresponding amino acid residue in an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity (or homology) to the amino acid sequence set forth in SEQ ID NO:1 is substituted with another amino acid.

[0033] The term "adenine deaminase" refers to a polypeptide or fragment capable of catalyzing the hydrolytic deamination of adenine or adenosine. In one embodiment of the present invention, the deaminase or deaminase domain represents an adenine deaminase that facilitates the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In another embodiment of the present invention, the adenine deaminase performs the hydrolytic deamination of adenine or adenosine in DNA (deoxyribonucleic acid). The adenosine deaminases, such as engineered or evolved adenosine deaminases described herein, can be derived from any organism, including bacteria.

[0034] In some embodiments of the invention, the adenosine deaminase is TadA deaminase. In some embodiments of the invention, the TadA deaminase is a TadA mutant. In some embodiments of the invention, the TadA mutant is TadA8e. In some embodiments of the invention, the deaminase or deaminase domain is a mutant of a naturally occurring deaminase from an organism such as bacteria, archaea, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments of the invention, the deaminase or deaminase domain does not occur in nature. For example, in some embodiments of the invention, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a naturally occurring deaminase.

[0035] The TadA8e adenosine deaminase has the following sequence:

[0036] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 1)

[0037] In some embodiments of the present invention, the amino acid substitution may be at least one selected from the group consisting of V28Q, or V28R, A48W, F84M, V106A, K110S, K110T, or K110V, and R111F, R111Q, R111S, R111T, or R111Y in the amino acid sequence represented by SEQ ID NO: 1.

[0038] In one embodiment of the present invention, the amino acid substitution with the lowest RNA or DNA off-target editing efficiency (e.g., including bystander off-target effects) may be V28Q of the amino acid sequence represented by SEQ ID NO: 1, or at least one selected from the group consisting of V28R, A48W, and R111S.

[0039] In various embodiments of the present invention, the adenine deaminase mutant can exhibit significantly reduced off-target effects, including unintended base modifications in DNA and / or RNA. In other embodiments of the present invention, the adenine deaminase mutant can narrow its range of activity while simultaneously reducing undesired bystander effects. Alternatively, the adenine deaminase mutant can induce base editing only at a single nucleotide residue without any intended off-target editing in target DNA at a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0040] In one embodiment of the present invention, a fusion protein is provided, comprising a DNA-binding protein and the adenine deaminase mutant. The amino acid substitution in the adenine deaminase mutant may be at least one selected from the group consisting of V28Q, V28R, A48W, F84M, V106A, K110S, K110T, or K110V ​​in the amino acid sequence represented by SEQ ID NO: 1, and R111F, R111Q, R111S, R111T, or R111Y. In one embodiment of the present invention, the amino acid substitution in the adenine deaminase mutant having the lowest RNA or DNA off-target editing effect (e.g., including bystander off-target effect) may be at least one selected from the group consisting of V28Q, V28R, A48W, and R111S in the amino acid sequence represented by SEQ ID NO: 1.

[0041] In various embodiments of the present invention, the fusion protein can exhibit significantly reduced off-target effects, including unintended base modifications in DNA and / or RNA. In other embodiments of the present invention, the fusion protein can narrow its range of activity while simultaneously reducing undesired bystander effects. Alternatively, the fusion protein can induce base editing at only a single nucleotide residue without any intended off-target editing in the target DNA at a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0042] The DNA-binding protein may be, for example, but is not limited to, 1) a zinc finger protein, 2) a transcriptional activator-like effector (TALE) protein, or 3) a CRISPR-associated nuclease. The nuclease may be type II and / or type V, such as a Cas protein (e.g., Cas9 protein (CRISPR (Clustered regularly interspaced short palindromic repeats)-associated protein 9)) or a Cpf1 protein (CRISPR from Prevotella and Francisella 1). Nucleases (e.g., endonucleases) associated with the CRISPR system or similar nucleases may be used. Specifically, in one embodiment of the present invention, the nuclease may be a Cas protein such as Cas3, Cas9, Cpf1, Cas6, or C2c2, specifically a CRISPR / Cas II type Cas protein, more specifically a Cas9 protein derived from Streptococcus pyogenes.

[0043] As used herein, the term "transcription activator-like effector" (TALE) refers to a DNA-binding protein containing a TALE repeat sequence, which contains a number of highly conserved 33-34 amino acid sequences containing two highly variable amino acid motifs (Repeat Variable Diresidues, RVDs). The RVD motifs determine binding specificity for nucleic acid sequences and can be designed to specifically bind to desired DNA sequences by methods well known to those skilled in the art. Due to the simple relationship between amino acid sequence and DNA recognition, a combination of repeat segments containing appropriate RVDs can be selected to design a specific DNA-binding domain.

[0044] In one embodiment of the present invention, the DNA binding protein may be a TALE. In another embodiment of the present invention, the TALE may be a dual TALE module composed of a first TALE module and a second TALE module. In some embodiments of the present invention, each of the first and second TALE modules may be linked to various deaminase enzymes. For example, the first TALE module may be linked to DddA in its full-length form. tox and the second TALE module may be linked to a cytosine deaminase such as adenine deaminase mutant.

[0045] The Cas9 protein is the main protein component of the CRISPR / Cas system and can function as an activated endonuclease or nickase.

[0046] The Cas9 protein or its gene information can be obtained from known databases such as GenBank at the National Center for Biotechnology Information (NCBI, USA). For example, the Cas9 protein may be at least one selected from the group consisting of, but not limited to, the following:

[0047] Cas9 proteins derived from Streptococcus sp., such as Streptococcus pyogenes (e.g., SwissProt accession number Q99ZW2 (NP_269215.1) (encoding gene: SEQ ID NO: 229);

[0048] Cas9 proteins derived from Campylobacter sp., e.g., Campylobacter jejuni;

[0049] a Cas9 protein derived from Streptococcus sp., such as Streptococcus thermophiles or Streptococcus aureus;

[0050] Cas9 protein derived from Neisseria meningitidis;

[0051] A Cas9 protein derived from Pasteurella sp., e.g., Pasteurella multocida; and

[0052] A Cas9 protein derived from a Francisella sp., e.g., Francisella novicida.

[0053] The Cpf1 protein, a novel endonuclease in the CRISPR system that is distinct from the CRISPR / Cas system, is smaller than Cas9, does not require tracrRNA, and can function as a single guide RNA. Cpf1 also recognizes thymidine-rich protospacer-adjacent motif (PAM) sequences and generates cohesive double-strand breaks (cohesive ends).

[0054] For example, the Cpf1 protein can be an endonuclease derived from Candidatus spp., Lachnospira spp., Butyrivibrio spp., Peregrinibacteria, Acidaminococcus spp., Porphyromonas spp., Prevotella spp., Francisella spp., Candidatus Methanoplasma, or Eubacterium spp.Examples of microorganisms from which the Cpf1 protein can be derived include Parcubacteria bacterium (GWC2011_GWC2_44_17), Lachnospiraceae bacterium (MC2017), Butyrivibrio proteoclasticus, Peregrinibacteria bacterium (GW2011_GWA_33_10), Acidaminococcus sp. (BV3L6), Porphyromonas macacae, Lachnospiraceae bacterium (ND2006), and Porphyromonas creviolicanis. crevioricanis, Prevotella disiens, Moraxella bovoculi (237), Smithella sp. (SC_KO8D17), Leptospira inadai, Lachnospiraceae bacterium (MA2020), Francisella novicida (U112), Candidatus Methanoplasma termitum, Candidatus Paceibacter, and Eubacterium eligens.

[0055] In one embodiment of the present invention, when the DNA-binding protein is a Cas9 protein, the Cas9 protein may be at least one selected from the group consisting of a modified Cas9 protein (e.g., SwissProt accession number Q99ZW2 (NP_269215.1)) that loses endonuclease activity but maintains nickase activity as a result of introducing a mutation (e.g., substitution with another amino acid) into D10 of the Streptococcus pyogenes-derived Cas9 protein, and a modified Cas9 protein (e.g., Streptococcus pyogenes-derived Cas9 protein) that loses both endonuclease activity and nickase activity as a result of introducing a mutation (e.g., substitution with another amino acid) into both D10 and H840 of the Streptococcus pyogenes-derived Cas9 protein. For example, in the Cas9 protein, the mutation at D10 may be a D10A mutation (substitution of amino acid D with A at the 10th position of the Cas9 protein), and the mutation at H840 may be an H840A mutation.

[0056] If the nuclease has nickase activity, nicks can be introduced simultaneously with the deaminase-mediated base modification (e.g., cytidine to uridine conversion) or sequentially in either the strand where the base modification occurred or the opposite strand (e.g., the strand opposite the strand where the base conversion occurred) (e.g., a nick can be introduced between the third and fourth nucleotide positions toward the 5' end of the PAM sequence on the strand opposite to the strand where the PAM was located). Nuclease mutations (e.g., amino acid substitutions, etc.) can occur in the catalytically active domain of the nuclease (e.g., the RuvC catalytic domain in the case of Cas9).

[0057] In one embodiment of the present invention, in the case of the Cas9 protein derived from Streptococcus pyogenes, the mutation may be a substitution of at least one amino acid selected from the group consisting of the catalytic aspartic acid at position 10 (D10), glutamic acid at position 762 (E762), histidine at position 840 (H840), asparagine at position 854 (N854), asparagine at position 863 (N863), and aspartic acid at position 986 (D986) with another amino acid. Specifically, the mutation may include a variant in which one or more amino acids selected from the group consisting of H839, H840, and N863 of Cas9 are substituted with another amino acid. Specifically, the mutation may include a variant in which amino acids N863, H840-N863, or H839-H840-N863 of Cas9 are substituted with another amino acid. In addition to H840A, D10A SpCas9 nickase, which is produced by removing part of the catalytic domain, can also be used.

[0058] In some embodiments of the present invention, the fusion protein may further comprise a guide RNA. The guide RNA may be, for example, at least one selected from the group consisting of CRISPR RNA (crRNA), trans-activating crRNA (tracrRNA), and single guide RNA (sgRNA). Specifically, the guide RNA may be a double-stranded crRNA:tracrRNA complex in which the crRNA and tracrRNA are bound to each other, or a single-stranded guide RNA (sgRNA) in which the crRNA or a portion thereof and the tracrRNA or a portion thereof are linked by an oligonucleotide linker.

[0059] The adenine deaminase mutant and the DNA-binding protein may be used in the form of a fusion protein fused to each other directly or via a peptide linker (e.g., the adenine deaminase mutant-DNA-binding protein order from N- to C-terminus (i.e., the DNA-binding protein fused to the C-terminus of the adenine deaminase mutant) or the DNA-binding protein-adenine deaminase mutant order from N- to C-terminus (i.e., the adenine deaminase mutant fused to the C-terminus of the DNA-binding protein)); a mixture of the adenine deaminase mutant or mRNA encoding it and the DNA-binding protein or mRNA encoding it; or a plasmid carrying both the adenine deaminase mutant-encoding gene and the DNA-binding protein-encoding gene (e.g., two genes arranged to encode the fusion protein described above, a mixture of an adenine deaminase mutant expression plasmid and a DNA-binding protein expression plasmid, or plasmids carrying the adenine deaminase mutant-encoding gene and the DNA-binding protein-encoding gene, respectively).

[0060] In one embodiment of the present invention, the fusion protein may further comprise a cytosine deaminase. The cytosine deaminase refers to any enzyme having the activity of converting cytosine found in nucleotides (e.g., cytosine present in double-stranded DNA or RNA) to uracil (CU converting activity or CU editing activity). The cytosine deaminase converts cytosine located in a strand containing a PAM sequence linked to a target sequence to uracil. In one embodiment of the present invention, the cytosine deaminase may be derived from bacteria, archaea, or mammals, including primates such as humans and monkeys, and rodents such as rats and mice, but is not limited thereto. For example, the cytosine deaminase may be derived from PmCDA1 (Petromyzon marinus cytosine deaminase 1) from sea lamprey or DddA from Burkholderia cenocepacia. toxand APOBEC (apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like), but is not limited thereto.

[0061] In some embodiments of the invention, the cytosine deaminase is wild-type sea lamprey CDA1 (pmCDA1) or a catalytic domain thereof, and in some embodiments of the invention, the cytosine deaminase comprises one or more mutations in the pmCDA1 sequence that alter the editing efficiency and / or substrate preference of pmCDA1 according to specific needs.

[0062] pmCDA1 has the following amino acid sequence:

[0063] MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMIQVKILHTTKSPAVSRGSG (SEQ ID NO: 2)

[0064] In some embodiments of the present invention, an exemplary deaminase is DddA. tox Since DddA is cytotoxic, it is split into two inactive halves to avoid toxicity in host cells, and each is fused to a DNA-binding protein, a DddA-derived cytosine base editor (DdCBE). A functional deaminase is reassembled at the target DNA site when the two inactive halves are joined by the DNA-binding protein. Full-length DddA tox has the following amino acid sequence:

[0065] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO: 3), which corresponds to Burkholderia cenocepacia PDB accession number 6U08_A, and can include fragments or variants thereof comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to DddA of 6U08_A.

[0066] In another embodiment of the present invention, the APOBEC may be at least one selected from the following group, but is not limited thereto:

[0067] APOBEC1: Homo sapiens APOBEC1 (protein: GenBank accession numbers NP_001291495.1, NP_001635.2, NP_005880.2, etc.; gene (mRNA or cDNA; described in the order of the corresponding proteins listed above): GenBank accession numbers NM_001304566.1, NM_001644.4, NM_005889.3, etc.), Mus musculus APOBEC1 (protein: GenBank accession numbers NP_001127863.1, NP_112436.1, etc.; gene: GenBank accession numbers NM_001134391.1, NM_031159.3, etc.);

[0068] APOBEC2: Homo sapiens APOBEC2 (protein: GenBank accession number NP_006780.1, etc.; gene: GenBank accession number NM_006789.3, etc.), mouse APOBEC2 (protein: GenBank accession number NP_033824.1, etc.; gene: GenBank accession number NM_009694.3, etc.);

[0069] APOBEC3B: Homo sapiens APOBEC3B (protein: GenBank accession numbers NP_001257340.1, NP_004891.4, etc.; gene: GenBank accession numbers NM_001270411.1, NM_004900.4, etc.), Mus musculus APOBEC3B (protein: GenBank accession numbers NP_001153887.1, NP_001333970.1, NP_084531.1, etc.; gene: GenBank accession numbers NM_001160415.1, NM_001347041.1, NM_030255.3, etc.);

[0070] APOBEC3C: Homo sapiens APOBEC3C (protein: GenBank accession number NP_055323.2, etc.; gene: GenBank accession number NM_014508.2, etc.);

[0071] APOBEC3D (including APOBEC3E): Homo sapiens APOBEC3D (protein: GenBank accession number NP_689639.2, etc.; gene: GenBank accession number NM_152426.3, etc.);

[0072] APOBEC3F: Homo sapiens APOBEC3F (protein: GenBank accession numbers NP_660341.2, NP_001006667.1, etc.; gene: GenBank accession numbers NM_145298.5, NM_001006666.1, etc.);

[0073] APOBEC3G: Homo sapiens APOBEC3G (protein: GenBank accession numbers NP_068594.1, NP_001336365.1, NP_001336366.1, NP_001336367.1, etc.; gene: GenBank accession numbers NM_021822.3, NM_001349436.1, NM_001349437.1, NM_001349438.1, etc.);

[0074] APOBEC3H: Homo sapiens APOBEC3H (protein: GenBank accession numbers NP_001159474.2, NP_001159475.2, NP_001159476.2, NP_861438.3, etc.; gene: GenBank accession numbers NM_001166002.2, NM_001166003.2, NM_001166004.2, NM_181773.4, etc.);

[0075] APOBEC4 (including APOBEC3E): Homo sapiens APOBEC4 (protein: GenBank accession number NP_982279.1, etc.; gene: GenBank accession number NM_203454.2, etc.); mouse APOBEC4 (protein: GenBank accession number NP_001074666.1, etc.; gene: GenBank accession number NM_001081197.1, etc.); and

[0076] Activation-induced cytidine deaminase (AICDA or AID): Homo sapiens AID (protein: GenBank accession numbers NP_001317272.1, NP_065712.1, etc.; gene: GenBank accession numbers NM_001330343.1, NM_020661.3, etc.); mouse AID (protein: GenBank accession number NP_033775.1, etc.; gene: GenBank accession number NM_009645.2, etc.), and others.

[0077] The cytosine deaminase can be a non-toxic full-length deaminase (i.e., a monomeric cytosine deaminase) or can be two split forms comprising separated first and second domains (i.e., a dimeric deaminase), each of which can be characterized by the absence of deaminase activity.

[0078] In another embodiment of the present invention, the adenine deaminase mutant can be bound to the N- or C-terminus of a DNA-binding protein, or a cytosine deaminase or a mutant thereof. For example, the DNA-binding protein is a ZFP, the adenine deaminase mutant is a TadA8e mutant, and the cytosine deaminase or a mutant thereof is a DddA. tox These may include, but are not limited to, the following order: ZFP-TadA8e mutant-DddA tox , ZFP-DddA tox -TadA8e mutant, TadA-DddA tox -ZFP or DddA tox -TadA8e mutant-ZFP.

[0079] In some embodiments of the present invention, when the cytosine deaminase is included in a split form and the DNA-binding protein is a zinc finger protein, the C-terminus of the first domain of the cytosine deaminase is bound to the N-terminus of the zinc finger protein (ZF-Left), the N-terminus of the second domain of the cytosine deaminase is bound to the C-terminus of the zinc finger protein (ZF-Right) (NC configuration), and the adenine deaminase mutant can be bound to the C-terminus of the zinc finger protein (ZF-Left), the N-terminus or C-terminus of the first domain of the cytosine deaminase, the N-terminus of the zinc finger protein (ZF-Right), or the N-terminus or C-terminus of the second domain of the cytosine deaminase.

[0080] In various embodiments of the present invention, the adenine deaminase mutant can be bound to:

[0081] the C-terminus of a zinc finger protein linked to the N-terminus of the first domain of cytosine deaminase (ZF-Left) and the C-terminus of a zinc finger protein linked to the N-terminus of the second domain of cytosine deaminase (ZF-Right) (CC configuration);

[0082] the C-terminus of a zinc finger protein (ZF-Left) linked to the N-terminus of the first domain of cytosine deaminase, and the N-terminus of a zinc finger protein (ZF-Right) linked to the C-terminus of the second domain of cytosine deaminase (CN configuration); or

[0083] The N-terminus of the zinc finger protein is linked to the C-terminus of the first domain of cytosine deaminase (ZF-Left), and the N-terminus of the zinc finger protein is linked to the C-terminus of the second domain of cytosine deaminase (ZF-Right) (NN configuration).

[0084] Thus, in some embodiments of the present invention, the adenine deaminase mutant can bind to the C-terminus of a zinc finger protein (ZF-Left), the N-terminus or C-terminus of the first domain of a cytosine deaminase, the zinc finger protein (ZF-Right), or the N-terminus or C-terminus of the second domain of a cytosine deaminase.

[0085] In some embodiments of the present invention, when the cytosine deaminase is in a split form and the DNA-binding protein is a TALE, the first domain of the cytosine deaminase binds to a first TALE, and the second domain of the cytosine deaminase binds to a second TALE, which have the structures N'-TALE-first domain DDDA-C' and N'-TALE-second domain DDDA-C', respectively. The adenine deaminase mutant can bind to the N-terminus or C-terminus of the first domain of the cytosine deaminase, or the N-terminus or C-terminus of the second domain of the cytosine deaminase.

[0086] In one embodiment of the present invention, when the cytosine deaminase is included in its full-length form and the DNA-binding protein is a TALE, it includes a single TALE module and a single TALE module including cytosine deaminase in the N- or C-direction, wherein the adenine deaminase mutant can be bound to the C-terminus of the single TALE module or to the N- or C-terminus of the cytosine deaminase.

[0087] In another embodiment of the present invention, when the cytosine deaminase is included in its full-length form and the DNA-binding protein is a TALE, a dual TALE module may be included. The first TALE module and cytosine deaminase are included in the N-C direction, and a second domain including an adenine deaminase mutant and a second TALE may be further included. In the structures N'-TALE-cytosine deaminase-C' and N'-TALE-adenine deaminase mutant-C', the adenine deaminase mutant can be bound to the N-terminus or C-terminus of the TALE.

[0088] In one embodiment of the present invention, the fusion protein may further comprise a uracil glycosylase inhibitor (UGI), which can increase the efficiency of base editing by inhibiting the activity of uracil DNA glycosylase (UDG), a mutation-repairing enzyme that promotes the removal of U from DNA.

[0089] In other embodiments of the present invention, the DNA may be nuclear DNA or organelle DNA.

[0090] In one embodiment of the present invention, the fusion protein may further comprise a nuclear localization signal (NLS). The nuclear localization signal protein may be derived from, for example, but not limited to, Simian Virus 40 large tumor antigen (SV40 large T antigen). The nuclear localization signal protein may comprise, for example, but not limited to, the following amino acid sequence:

[0091] PKKKRKV (SEQ ID NO: 4)

[0092] In another embodiment of the present invention, the fusion protein may further comprise a mitochondrial targeting sequence (MTS) or a chloroplast transit peptide (CTP). The mitochondrial targeting sequence protein may be, for example, SOD2-MTS or COX8A-MTS, and may comprise, but is not limited to, the following amino acid sequence:

[0093] SOD2-MTS: LSRAVCGTSRQLAPVLGYLGSRQKHSLPD (SEQ ID NO: 5)

[0094] COX8A-MTS:SVLTPLLLRGLTGSARRLPVPRAKIHSL (SEQ ID NO: 6).

[0095] The chloroplast transit peptide protein may be derived from, for example, but is not limited to, Arabidopsis RECA1. The chloroplast transit peptide protein may include, for example, but is not limited to, the following amino acid sequence:

[0096] MDSQLVLSLKLNPSFTPLSPLFPFTPCSSFSPSLRFSSCYSRRLYSPVTVYAAK (SEQ ID NO: 7)

[0097] In another embodiment of the present invention, the fusion protein may further comprise a nuclear export signal (NES). The nuclear export signal protein may be derived from, for example, but not limited to, MVM (Minute virus of mice). The nuclear export signal protein may comprise, for example, but not limited to, the following amino acid sequence:

[0098] VDEMTKKFGTLTIHDTEK (SEQ ID NO: 8)

[0099] In one embodiment of the present invention, when the signal peptide is attached to a fusion protein, its structure is as follows: signal peptide-DNA binding protein-deaminase. In another embodiment of the present invention, the structure may be signal peptide-deaminase-DNA binding protein. In one embodiment of the present invention, the nuclear export signal protein, CTP (chloroplast transit peptide), or a polynucleotide encoding the same, may be attached to the N-terminus of the DNA binding protein, cytosine deaminase (DdCBE), or a polynucleotide encoding the same.

[0100] In various embodiments of the present invention, the fusion protein may further comprise a nickase, such as, but not limited to, MutH, a MutH mutant, or Nt.BspD6I(C). MutH is a weak endonuclease that is activated upon binding to MutL. It nicks the unmethylated strand of unmethylated and semi-methylated DNA, but not fully methylated DNA. On the other hand, the nicking endonuclease Nt.BspD6I (Nt.BspD6I) is the large subunit of the heterodimeric restriction endonuclease R.BspD6I. It recognizes the short, specific DNA sequence 5′-GAGTC and cleaves only the top strand of dsDNA at a distance of four nucleotides downstream of the recognition site toward the 3′-end. The resulting strand-specific nick induces the generation of transient single-stranded DNA. The TadA mutant is a nucleobase deaminase that specifically targets single-stranded DNA, and thus has the ability to induce A-to-G editing by being in close proximity to the nick site. In another embodiment of the present invention, the fusion protein further comprising a nickase may be in a dimeric form comprising a first fusion protein and a second protein. The first fusion protein may comprise a DNA-binding protein and an adenine deaminase mutant, and the second protein may comprise another DNA-binding protein and a nickase.

[0101] In one embodiment of the present invention, a polynucleotide encoding the adenine deaminase or the fusion protein is provided. The term "polynucleotide" is used interchangeably with "nucleic acid," "oligonucleotide," "nucleotide," and "nucleotide sequence." It can include a polymeric form of nucleotides of any length, deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. Polynucleotides can include one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. Modifications to the nucleotide structure can be made before or after assembly of the polymer.

[0102] The nucleic acid can be an RNA sequence, in particular an mRNA sequence, a DNA sequence, or a combination thereof (combined RNA-DNA sequence).

[0103] The nucleic acid can be delivered using other viral vectors, such as adeno-associated viral vectors (AAV), adenoviral vectors (AdV), lentiviral vectors (LV), retroviral vectors (RV), or episomal vectors containing Simian virus 40 (SV40) origin, bovine papilloma virus (BPV) origin, or Epstein-Barr nuclear antigen (EBV). In various embodiments of the present invention, the delivery can be achieved using non-viral vectors or through plasmid or mRNA delivery.

[0104] The vectors can be delivered in vivo or intracellularly by local injection methods (e.g., direct injection into a lesion or target site), electroporation, lipofection, viral vectors, nanoparticles, PTD (protein translocation domain) fusion protein methods, or similar methods.

[0105] As a means for expressing the protein, known expression vectors such as plasmid vectors, cosmid vectors, bacteriophage vectors, etc. Such vectors can be easily produced by those skilled in the art using any known method using DNA recombination technology.

[0106] Recombinant expression vectors are designed to deliver nucleic acids in a format that facilitates expression in a host cell. The nucleic acid sequence desired for expression is operably linked to the recombinant expression vector, which contains one or more regulatory elements that can be selected for a particular host cell. In the context of a recombinant expression vector, "operably linked" means that the nucleotide sequence of interest is linked to regulatory elements in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).

[0107] In one embodiment of the present invention, a base editing composition for A-to-G base editing of DNA is provided, the composition comprising the fusion protein, a polynucleotide encoding the fusion protein, or an expression vector comprising the polynucleotide. The fusion protein has been described in detail above.

[0108] In one embodiment of the present invention, the fusion protein of the base editing composition may further comprise cytosine deaminase. The cytosine deaminase has been described above. In some embodiments of the present invention, the cytosine deaminase may exist in the form of two segments, and the fusion protein may comprise a first fusion protein comprising a first segment of the cytosine deaminase and a second fusion protein comprising a second segment of the cytosine deaminase. The first segment may comprise the amino acid sequence of SEQ ID NO: 9 or 10, and the second segment may comprise the amino acid sequence of SEQ ID NO: 11 or 12, but is not limited thereto. In one embodiment of the present invention, for example, two TALEDs may be N-terminally DddA at G1397. tox Left or right TALE (L-1397N or R-1397N, respectively) fused to the half split, and C-terminal DddA at G1397 and TadA8e tox It may consist of a right or left sided TALE (R-1397C-AD or L-1397C-AD) fused to a half-body.

[0109] [Table 1]

[0110] In various embodiments of the present invention, the base editing composition can exhibit significantly reduced off-target effects associated with unintended base modifications in DNA and / or RNA. In other embodiments of the present invention, the base editing composition can narrow its range of activity while simultaneously reducing undesired bystander effects. Alternatively, the base editing composition can induce base editing only at a single nucleotide residue without any intended off-target editing in target DNA at a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0111] In another embodiment of the present invention, a method for A-to-G base editing of DNA is provided, comprising delivering a base editing composition to a cell containing target DNA.

[0112] The cells may be, but are not limited to, eukaryotic cells (e.g., fungi such as yeast, eukaryotic animal and / or eukaryotic plant-derived cells (e.g., germ cells, stem cells, somatic cells, germ cells, etc.), eukaryotic animals (e.g., humans, monkeys, primates, dogs, pigs, cows, sheep, goats, mice, rats, etc.), or eukaryotic plants (e.g., algae such as green algae, corn, soybeans, wheat, rice, etc.).

[0113] Delivery of the base editing composition to cells containing target DNA can be performed in vitro or in vivo.

[0114] In some embodiments of the present invention, the target DNA may be nuclear DNA, organelle DNA, or mitochondrial DNA of a human subject having a genetic disease.

[0115] The term "genetic disease" as used herein refers to a pathological condition caused by a deleterious mutation in a gene or chromosome. Examples of such genetic diseases include, but are not limited to, MELAS (mitochondrial encephalopathy, lactic acidosis, and stroke-like episodes syndrome), DEAF, Leber hereditary optic neuropathy (LHON), Leigh syndrome, myopathy, and chronic progressive external ophthalmoplegia (CPEO).

[0116] The method may exhibit reduced off-target effects compared to using a base editor comprising an adenine deaminase having the amino acid sequence of SEQ ID NO: 1, where the off-target editing is characterized by unintended base modifications in DNA and / or RNA. In other embodiments of the invention, the method may narrow the scope of activity and simultaneously reduce undesired bystander effects compared to using a base editor comprising an adenine deaminase having an amino acid represented by a mutation in SEQ ID NO: 1. Alternatively, the method may induce base editing only at a single nucleotide residue without any intended off-target editing in the target DNA at a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0117] In another embodiment of the invention, the method may exhibit reduced off-target effects compared to using a base editor comprising an adenine deaminase with a V106W mutation in the amino acid sequence represented by SEQ ID NO: 1, where the off-target editing is characterized by unintended base modifications in DNA and / or RNA. In another embodiment of the invention, the method may narrow the scope of action and simultaneously reduce undesired bystander effects compared to using a base editor comprising an adenine deaminase with a V106W mutation in the amino acid sequence represented by SEQ ID NO: 1. Alternatively, the method may induce base editing only at a single nucleotide residue without any intended off-target editing in the target DNA at a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0118] In one embodiment of the present invention, a method for reducing undesired bystander effects while narrowing off-target effects and / or range of action is provided, the method comprising a step of delivering the base editing composition to a cell containing target DNA.

[0119] The cells may be, but are not limited to, eukaryotic cells (e.g., fungi such as yeast, eukaryotic animal and / or eukaryotic plant-derived cells (e.g., germ cells, stem cells, somatic cells, germ cells, etc.), eukaryotic animals (e.g., humans, monkeys, primates, dogs, pigs, cows, sheep, goats, mice, rats, etc.), or eukaryotic plants (e.g., algae such as green algae, corn, soybeans, wheat, rice, etc.).

[0120] In some embodiments of the present invention, the target DNA may be nuclear DNA, organelle DNA, or mitochondrial DNA of a human subject with a genetic disease. The genetic disease is defined above. The method may exhibit reduced off-target effects compared to using a base editor comprising an adenine deaminase having the amino acid sequence of SEQ ID NO: 1, where the off-target editing is characterized by unintended base modifications in DNA and / or RNA. In other embodiments of the present invention, the method may narrow the scope of action compared to using a base editor comprising an adenine deaminase having the amino acid sequence of SEQ ID NO: 1, while simultaneously reducing undesired bystander effects. Alternatively, the method may induce base editing only at a single nucleotide residue without any intended off-target editing in the target DNA at a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0121] In another embodiment of the invention, the method may exhibit reduced off-target effects compared to using a base editor comprising an adenine deaminase with a V106W mutation in the amino acid sequence represented by SEQ ID NO: 1, where the off-target editing is characterized by unintended base modifications in DNA and / or RNA. In another embodiment of the invention, the method may narrow the scope of action and simultaneously reduce undesired bystander effects compared to using a base editor comprising an adenine deaminase with a V106W mutation in the amino acid sequence represented by SEQ ID NO: 1. Alternatively, the method may induce base editing only at a single nucleotide residue without any intended off-target editing in the target DNA at a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0122] However, in another embodiment of the present invention, there is provided a use of the adenine deaminase mutant or the fusion protein for A-to-G base editing of DNA and / or reducing off-target effects in A-to-G base editing of DNA, or for manufacturing a composition for A-to-G base editing of DNA and / or reducing off-target effects in A-to-G base editing of DNA. The adenine deaminase mutant or the fusion protein have been described in detail above.

[0123] In one embodiment of the present invention, the fusion protein may further comprise cytosine deaminase. The cytosine deaminase is described above. In some embodiments of the present invention, the cytosine deaminase may exist in the form of two segments, and the fusion protein comprises a first fusion protein comprising a first segment of the cytosine deaminase and a second fusion protein comprising a second segment of the cytosine deaminase. The first segment may comprise the amino acid sequence of SEQ ID NO: 9 or 10, and the second segment may comprise the amino acid sequence of SEQ ID NO: 11 or 12, but is not limited thereto.

[0124] In various embodiments of the present invention, the compositions for A-to-G base editing in DNA and / or for reducing off-target effects in A-to-G base editing in DNA can exhibit significantly reduced off-target effects associated with unintended base modifications in DNA and / or RNA. In one embodiment of the present invention, the compositions can narrow the range of activity while simultaneously reducing undesired bystander effects. Alternatively, the compositions can induce base editing only at a single nucleotide residue without any intended off-target editing in the target DNA at a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0125] The present invention will now be described in more detail with reference to examples. Those skilled in the art will appreciate that these examples are merely for the purpose of illustrating the invention and should not be construed as limiting the scope of the invention.

[0126] Reference Example

[0127] 1. Cell Line Preparation

[0128] HEK 293T cells were purchased from the American Type Culture Collection (ATCC) (CRL-11268). NIH3T3 and B16F10 cells were purchased from the American Type Culture Collection (ATCC) (CRL-1658, CRL-6475). HEK 293T cells were cultured in Dulbecco's Modified Eagle Medium (DMEM; Welgene) supplemented with 10% (v / v) fetal bovine serum (Welgene) and 1% (v / v) antibiotic-antimycotic solution (Welgene). NIH3T3 and B16F10 cells were cultured in DMEM supplemented with 10% (v / v) fetal bovine serum (Gibco) without antibiotics (NIH3T3 cells) and in DMEM supplemented with 10% (v / v) fetal bovine serum (Gibco) without antibiotics (B16F10 cells). The cells were cultured at 5% CO2 and 37°C. All cell lines were passaged before reaching 90% confluency.

[0129] 2.PyMOL analysis

[0130] The ABE8e (PDB accession number 6VPC) structure was downloaded from the PDB and visualized using PyMOL v.2.5.4. Some elements, including Cas9, single guide RNA, and duplex DNA, were excluded from the PDB file. The transition-state analog of the adenosine deamination reaction, 8AZ, and the TadA monomer were retained. Eleven residues near 8AZ (V28, V30, N46, A48, I49, V82, F84, V106, N108, K110, and R111), including residues known to contact DNA, were selected.

[0131] 3. Plasmid Construction

[0132] To construct plasmids encoding site-specific targeting DdCBEs, DddA-split TALEDs (sTALEDs), and mTALEDs, we created plasmids containing a "stuffer," a sequence between two restriction enzyme sites that aids in fragment separation during gel electrophoresis. The construction of these plasmids is as follows: sTALED plasmid (p3s-stuffer-DddA) tox Half (1397C)-AD and p3s-Stuffer-DddA tox half (1397N); mTALED plasmid (p3s-stuffer-E1347A DddA tox (Total - AD). The plasmids were purchased from Addgene (DdCBE, #187168, #187171, #187173, #178174; sTALED, #187167, #187169, #187170, #187172; mTALED, #187163, #187166). Plasmids encoding DdCBE, sTALED, and mTALED targeting specific sites were constructed by inserting custom-designed TALE array sequences as shown in Table 2 below. To remove the stuffer sequence, each plasmid was digested with BsaI or BsmbI (NEB), and an insert containing the custom-designed TALE array sequence synthesized by IDT was inserted into the digested vector using the HiFi DNA Assembly Kit (NEB). Alternatively, the desired TALED construct was generated by digesting the master vector with BsaI or BsmbI, cleaving the site with stuffer, and then assembling six TALE arrays at that location using the Golden Gate method.

[0133] [Table 2] TIFF2025532113000004.tif89170

[0134] 4. Cell Culture and Transfection

[0135] HEK 293T cells were cultured in DMEM (Welgene) supplemented with 10% fetal bovine serum (Welgene) and 1% antibiotic-antimycotic solution (Welgene). NIH3T3 (CRL-1658, ATCC) and B16F10 (CRL-6475, ATCC) cells were cultured in DMEM supplemented with 10% (v / v) calf serum (Gibco) for NIH3T3 cells and 10% fetal bovine serum (Gibco) for B16F10 cells without antibiotics. Cell lines were maintained at 37°C in 5% CO2 and passaged before reaching 90% confluency, depending on the doubling period of the particular cell line.

[0136] For HEK 293T cells, 7.5 × 10 cells were cultured per well in a 48-well plate (Corning) before transfection. 4 The cells were then transfected at a density of 1 x 10 cells per well. After 24 hours, the cells were transfected with plasmids (1 µg total) using 1.5 µL of Lipofectamine 2000 (Invitrogen). When transfecting CRISPR-base editors with sTALED pairs, DdCBE pairs, and sgRNAs, the total amount of plasmid was 1 µg (500 ng each). When a single construct was used, the amount of transfected plasmid was 500 ng. After 96 hours, the transfected cells were harvested. For NIH3T3 and B16F10 cells, 1 x 10 cells were transfected per well into a 24-well cell culture plate (SPL, Seoul, Korea) 18 to 24 hours before transfection. 5 Cells were transfected at a density of 1000 ng per well. Lipofection using Lipofectamine 3000 (Invitrogen) was performed with 500 ng of each sTALED-encoding plasmid, constituting 1000 ng of total plasmid DNA. For mTALED, 500 ng of plasmid was used. Cells were harvested 3 days post-transfection.

[0137] 5. Transcriptome Sequencing

[0138] Total RNA was isolated 48 or 96 hours post-transfection using the NucleoSpin RNA Kit (MN #Macherey-Nagel) according to the manufacturer's instructions. RNA libraries were prepared using the TruSeq Stranded Total RNA Library Prep Gold Kit (Illumina). RNA library quality was assessed using a 2200 TapeStation with a D1000 ScreenTape system (Agilent). Total RNA sequencing was performed using a Macrogen NovaSeq 6000 sequencer (Illumina) with paired-end sequencing (2x100bp).

[0139] 6.RNA Mutation Calling

[0140] To analyze the NGS data from RNA sequencing, we created an RNA mutation calling pipeline based on a previously used one for RNA off-target analysis of CRISPR DNA base editors. Briefly, fastq sequencing reads were aligned to the hg38 human reference genome (GRCh38, release v105) using STAR aligner (v.2.7.10a). The resulting BAM files were processed with GATK (v.4.2.4.1) MarkDuplicates, BaseRecalibrator, and ApplyBQSR. RNA base editing mutations were called using GATK HaplotypeCaller. RNA mutation loci were compared with those in control samples and filtered according to the following criteria: (1) loci with a minimum read depth of 10 were retained; (2) loci with a minimum number of mutations of 2 were retained; (3) loci present in control samples were removed; and (4) loci that could not be determined due to insufficient sequencing depth in the control samples were excluded. For replicate set-1, untreated replicate set-2 was used as a control for filtering, and for replicate set-2, untreated replicate set-1 was used as a control. For A-to-G editing, the number of RNA variant loci with A-to-G edits on the positive strand or T-to-C edits on the negative strand was counted. For C-to-T editing, the number of RNA variant loci with C-to-T edits on the positive strand or G-to-A edits on the negative strand was counted.

[0141] 7. Cell Lysis for Genomic DNA Analysis

[0142] After removing the growth medium, HEK 293T cells were treated with 100 μL of cell lysis buffer (50 mM Tris-HCl; pH 8.0, 1 mM EDTA, 0.005% sodium dodecyl sulfate) supplemented with 5 μL of Proteinase K (Qiagen). The cells were incubated at 55°C for 1 hour and then lysed at 95°C for 10 minutes. The genomic DNA mixture was then subjected to targeted deep sequencing.

[0143] 8. Targeted Deep Sequencing

[0144] NGS libraries for targeted deep sequencing were generated using overlapping PCR. The target region was first amplified by PCR using PrimeSTAR® GXL polymerase (Takara). The amplicons were then amplified again by PCR using primers containing TruSeq DNA-RNA CD indexes. Each fragment was labeled with an adapter and index sequence to construct an NGS library. The PCR primers are listed in Tables 3 to 6 below. The final PCR products were purified using a PCR purification kit (MGmed) and sequenced using a MiniSeq sequencer (Illumina). Base editing frequencies in the targeted deep sequencing data were measured using the source code (https: / / github.com / ibs-cge / maund).

[0145] [Table 3]

[0146] [Table 4]

[0147] [Table 5]

[0148] [Table 6]

[0149] 9. RNA Purification and Targeted RNA Sequencing in Cultured Mammalian Cells

[0150] RNA was extracted from the cultured cells using the NucleoSpin RNA Plus Kit (Macherey-Nagel) according to the manufacturer's instructions. The RNA was then reverse transcribed to generate cDNA using the SuperScript IV Reverse Transcriptase Kit (Thermo Fisher) according to the manufacturer's instructions. Regions of interest were then amplified by PCR using the primers listed in Tables 7 and 8 below. The amplified regions were sequenced using the targeted deep sequencing procedure described above.

[0151] [Table 7]

[0152] [Table 8]

[0153] 10. Relative ratio of RNA off-target editing frequency to DNA on-target editing frequency

[0154] To analyze the DNA on-target editing frequency versus RNA off-target editing frequency of mutants at six representative sites, the DNA on-target activity and RNA off-target activity of the mutants were normalized to the sTALED value. The normalized DNA on-target value was then divided by the normalized RNA off-target value as shown below:

[0155] TIFF2025532113000011.tif16169

[0156] For wild-type sTALED, this value is 1, as the normalized DNA on-target and RNA off-target values ​​are both 1. A higher relative ratio indicates lower off-target activity compared to on-target activity.

[0157] 11. Whole Mitochondrial Genome Sequencing

[0158] For whole mitochondrial genome sequencing, three steps were required: PCR amplification, NGS library generation, and NGS. First, after removing the growth medium, cells were treated with 100 μL of cell lysis buffer (50 mM Tris-HCl; pH 8.0, 1 mM EDTA, 0.005% sodium dodecyl sulfate) supplemented with 5 μL of Proteinase K (Qiagen). The cells were incubated at 55°C for 1 hour and then lysed by incubation at 95°C for 10 minutes. Mitochondrial DNA was amplified by PCR using PrimeSTAR® GXL polymerase (Takara). To reduce primer bias, PCR was performed using two sets of slightly overlapping primers listed in Tables 3 to 6. Each primer pair amplified approximately 50% of the mitochondrial DNA. The PCR product was then purified using a PCR purification kit (MGmed). Finally, NGS libraries were generated from the purified PCR products using an Illumina DNA Prep kit with Nextera DNA CD index (Illumina), which were then pooled and transferred to a MiniSeq sequencer (Illumina).

[0159] 12. Mitochondrial Genome-wide Off-target Editing Analysis

[0160] To analyze NGS data from whole mitochondrial genome sequencing, we used a method commonly used to investigate off-target effects in mitochondrial genomes. First, Fastq sequences were aligned to the GRCh38 (release v102) reference genome using BWA (v.0.7.17). Then, SAMtools (v.1.9) was used to correct read pairing information and flags and generate BAM files. Subsequently, the REDItoolDenovo.py script in REDItools (v.1.2.1) was used to identify all thymines and adenines in the mitochondrial genome with a conversion rate >0.1%. Positions with a conversion rate ≥10% in both treated and untreated samples were identified as SNVs in the cell lines and removed. On-target sites in the construct were excluded. Remaining sites were considered off-target sites, and the number of edited A / T nucleotides with an editing frequency >0.1% was counted. The average A / T to G / C editing frequency was calculated for all bases in the mitochondrial genome by averaging the conversion rates at each base position within the off-target site, as shown below.

[0161] TIFF2025532113000012.tif16169

[0162] A mitochondrial genome-wide graph was constructed by plotting the conversion rates of on-target and off-target sites with editing frequencies ≥ 1% across the entire mitochondrial genome.

[0163] 13. Cell Viability Assay

[0164] Cell viability assays were performed using CellTiter 96® Aqueous One Solution (Promega) on days 2 or 4 after plasmid transfection. The MTS assay measured the number of viable cells colorimetrically. Cells were treated with CellTiter 96® Aqueous One Solution, and quantification of the biologically reduced products was measured by recording the absorbance at 490 nm according to the manufacturer's guidelines.

[0165] Example 1: Mitochondrial DNA-targeted TALEDs induce transcriptome-wide off-target editing.

[0166] To investigate whether DddA-split TALEDs (sTALEDs) targeting the COX3.1 or ND1 site can induce undesired off-target RNA editing in human embryonic kidney 293T (HEK 293T) cells, transcriptome-wide sequencing of total RNA isolated from cells 2 days post-transfection was performed as described in the reference example. Transcriptome-wide sequencing of two independent biological replicates revealed an N-terminal DddA at G1397. tox Left or right TALE (L-1397N or R-1397N, respectively) fused to the half split, and C-terminal DddA at G1397 and TadA8e tox We demonstrated that two sTALEDs composed of either a right- or left-sided TALE (R-1397C-AD or L-1397C-AD) fused to a half-body induced transversions at >50,000 sites with a frequency of at least 7% (Figure 1). Most of these transversions (>99.8%) were A-to-G edits (in reverse-transcribed cDNA) rather than C-to-T edits, indicating that adenine deaminase, not cytosine deaminase, is responsible for these single-nucleotide modifications. In contrast, DdCBEs targeting the same sites or a single sTALED subunit (L-1397N or R-1397N), used as negative controls, did not induce such off-target RNA editing. These results suggest that other subunits, including TadA8e (L-1397C-AD or R-1397C-AD), are responsible for the A-to-G transversions observed in RNA via A-to-I modifications.

[0167] To avoid or minimize undesired transcriptome-wide off-target A-to-I conversion induced by TALEDs, we introduced site-specific mutations into TadA8e, including V106W, V106G, K20A / R21A (double mutation), or F148A, which are known to reduce off-target RNA editing when integrated into CRISPR RNA-guided adenine base editors (ABEs). Transcriptome-wide sequencing showed that sTALED mutants incorporating these mutations in TadA8e maintained DNA on-target editing efficiency while significantly, if not completely, reducing the number of off-target A-to-G edits (Figure 1). Thus, the Cox3.1-specific sTALED mutants with these site-specific mutations induced RNA off-target edits at sites 44,627 (F148A) to 107,304 (K20A / R21A), reducing the number of such sites from 18.7% (=(132,064-107,304) / 132,064) to 66.2% (=(132,064-44,627) / 132,064) compared to the parent sTALED (132,064 AG edits) (Figure 1b and Figure 1c). Similarly, the ND1-specific sTALED mutants reduced the number of RNA off-target edits by a minimum of 46.2% (=(61,189-32,910) / 61,189) (K20A / R21A) and a maximum of 84.2% (=(61,189-9,674) / 61,189) (V106G) (Figure 1d and Figure 1e). These results indicate that the same TadA8e mutant used to reduce the RNA off-target editing effects of CRISPR RNA-guided ABEs can also reduce collateral damage caused by sTALEDs. However, a minimum of 9,600 and a maximum of 107,000 off-target A-to-G edits were maintained even with the best-performing TadA8e mutant.

[0168] Example 2: Protein engineering of TALEDs to avoid RNA off-target editing

[0169] To further minimize transcriptome-wide off-target editing, we engineered TALEDs by mutating amino acid residues, including V106, in the substrate-binding site of TadA8e. Based on the 3D cryo-EM structure of ABE8e bound to DNA (Figure 2(A)), we selected 11 amino acid residues in the substrate-binding site and replaced each of these residues with the other 19 amino acid residues of the Cox3.1-specific sTALED to generate a total of 209 (=11 × 19) sTALED mutants. HEK 293T cells were transfected with plasmids encoding each of the generated sTALED pairs, as described in Reference Example 3. The on-target editing frequency of the Cox3.1 site was then measured 4 days post-transfection. Of the 209 sTALED mutants, 101 sTALED pairs maintained high activity in targeted mitochondrial DNA editing, with at least 50% efficiency compared to the wild-type sTALED pair (Figure 3). We then used targeted RNA amplicon sequencing to measure the off-target editing frequencies of 101 sTALEDs at six representative sites identified by transcriptome-wide sequencing analysis (Figure 4(A)). These representative sites were mutated at high frequencies ranging from 60% to 80% by all Cox3.1- and ND1-specific TALEDs (Figure 2(B)). The RNA off-target editing frequencies measured by targeted deep sequencing were in good agreement with those estimated by transcriptome-wide sequencing (Figure 4(B) and Figure 4(C)). A total of 12 TadA8e mutants were selected to minimize RNA off-target editing efficiency at the six representative sites while maintaining mtDNA on-target editing efficiency (Figure 2(B)).

[0170] We performed transcriptome-wide sequencing to investigate whether the 12 sTALED mutants could avoid off-target editing at sites other than the six representative sites (Figure 5). These mutants all significantly reduced the number of off-target edits. For example, the sTALED mutants with V28R and R111S induced off-target edits at only 852 and 829 sites, respectively, whereas the original sTALED and sTALED-V106W (a sTALED with a V106W mutation in TadA8e) induced off-target edits at 96,559 and 81,156 sites, respectively (Figure 5a and Figure 5b). Thus, sTALED-V28R and -R111S avoided >99% of RNA off-target editing. Considering that 316 A-to-G edits were found in the untreated DNA sample used as a negative control, which were likely caused by high-throughput sequencing errors, these results indicated that these sTALED mutants almost completely avoided RNA off-target editing.

[0171] We further investigated whether these TadA8e mutations could reduce RNA off-target editing when integrated into sTALEDs targeting other mitochondrial DNA target sites (Figures 5c-f and 6). Targeted RNA amplicon sequencing demonstrated that most ND1- and ND6-specific sTALED mutants significantly reduced RNA off-target editing frequencies at six representative sites (Figures 6(C)-6(F)). Furthermore, we measured mitochondrial DNA on-target editing frequencies through targeted deep sequencing (Figures 6(A) and 6(B)), and then obtained the ratios of DNA on-target editing frequencies to RNA off-target editing frequencies (Figures 6(G) and 6(H)). Based on these results, we selected four TadA8e mutants, V28Q, V28R, A48W, and R111S, which exhibited higher mitochondrial DNA on-target activity and lower RNA off-target activity compared to the original sTALED and sTALED-V106W, which targets the Cox3.1, ND1, and ND6 sites (Figure 6(I)). Transcriptome-wide sequencing showed that the ND1- and ND6-specific sTALED mutants harboring these four mutations, respectively, significantly reduced the number of RNA off-target edits (Figure 5), consistent with the results for the Cox3.1-specific sTALED mutant. In other words, the ND1- and ND6-specific sTALED-V28R or -R111S mutants differentially induced the least amount of RNA off-target edits (Figures 5c-f).

[0172] Example 3: Engineered TALEDs reduce bystander and off-target editing.

[0173] sTALED mutants with site-specific mutations in the TadA8e substrate binding site not only reduce activity on RNA substrates, but also potentially fine-tune adenine deaminase activity on DNA substrates, potentially reducing bystander editing at the target site and off-target editing of the mitochondrial genome. Nevertheless, we investigated the frequency of base editing at each nucleotide position. The Cox3.1, ND1, and ND6 site-specific sTALED-V28R and -R111S mutants induced A-to-I editing in a narrower region than wild-type sTALED and sTALED-V106W mutants (Figure 7). For example, wild-type sTALED and sTALED-V106W targeting the Cox3.1 site induced A-to-I editing at various positions, not only in the spacer region between the two TALE binding sites but also at the TALE binding site, with a frequency of >1.1%. In contrast, Cox3.1-specific sTALED-V28R induced A-to-I editing at a single position in the spacer region (Fig. 7a).

[0174] Next, we compared the frequencies of edited alleles induced by these sTALEDs targeting the three mtDNA sites. Most mutant alleles induced by the original sTALEDs or the sTALED-V106W mutant contained multiple, but not single, base edits. In stark contrast, sTALED-V28R and -R111S induced mutant alleles with single-base substitutions much more frequently than multiple-base substitutions. Thus, the most abundant allele with a single A-to-I edit in the middle of the spacer region was induced by the original sTALEDs at low frequencies of 3.06% (Cox3.1), 1.90% (ND1), and 1.70% (ND6), whereas the same allele was induced at high frequencies of 10.1% (Cox3.1), 17.0% (ND1), and 10.5% (ND6) by the sTALED-V28R mutant (Figures 7d-f). These results have important implications for the use of TALEDs in disease modeling as well as therapeutic applications. Because most pathogenic mitochondrial DNA mutations responsible for mitochondrial genetic diseases are single-nucleotide variants rather than multi-nucleotide variants, TALEDs that induce single-base substitutions with little or no bystander editing are preferred.

[0175] We also performed whole-mitochondrial genome sequencing to assess and compare the off-target effects of wild-type sTALEDs and sTALED mutants on day 4 post-transfection (Figure 8). Wild-type sTALEDs targeting three mitochondrial DNA sites induced off-target mutations with an average frequency of mitochondrial genome-wide off-target editing, ranging from 0.0024% to 0.0066%, 4- to 10-fold higher than that observed in the untreated control group (0.0006%). In contrast, sTALED-V28R and R111S induced off-target editing with an average frequency of <0.0006%, similar to the baseline frequency observed in the negative control group (Figure 8(B)). Furthermore, the number of off-target edits induced in the mitochondrial genome was also reduced by the sTALED mutants. Thus, the ND1-specific wild-type sTALED induced A-to-I off-target editing at 108 sites in human mitochondrial DNA at a frequency of >0.1%, while sTALED-V28R and -R111S induced off-target editing at 14 and 17 sites, respectively, which was similar to the baseline number (i.e., 24) observed in untreated samples (Figure 8(A)). Whole mitochondrial genome sequencing was also performed on day 2 post-transfection (Figure 9), which revealed the highest level of RNA off-target editing. The original sTALEDs targeting three sites induced off-target editing at an average frequency ranging from 0.0071% to 0.0090%, which was 11- to 14-fold higher than that observed in the untreated control (0.0006%). However, the V28R and R111S mutants did not induce off-target editing compared to the negative control (Figure 9(B)).

[0176] Example 4: Measuring DNA on-target and RNA off-target editing frequencies over time

[0177] In Figure 10, we investigated whether on-target mutations induced by various forms of sTALEDs are stably maintained and how long RNA off-target mutations persist over time. Using targeted deep sequencing, we measured the frequency of DNA on-target editing at three mitochondrial DNA sites (Figures 10(A)-10(C)) and the frequency of RNA off-target editing at six representative sites co-edited by the three sTALEDs (Figures 10(D)-10(G)) at various time points. RNA editing was significantly induced by sTALEDs and the sTALED-V106W mutant on days 1 and 2 post-transfection but was almost completely eliminated by day 8 post-transfection. The novel sTALED mutant also reduced the frequency of RNA off-target editing several-fold compared to sTALEDs or the sTALED-V106W mutant on days 1 and 2 post-transfection (Figure 10(G)). The mitochondrial DNA on-target editing induced by the sTALEDs and the sTALED-V106W mutant was not stably maintained. Thus, the frequency of mitochondrial DNA on-target editing induced by sTALEDs and the sTALED-V106W mutant significantly decreased over time. For example, mitochondrial DNA on-target editing induced by the ND6-specific sTALED and sTALED-V106W showed high frequencies of up to 47% and 40%, respectively, at 1 and 2 days post-transfection, but showed low frequencies of <10% at 8 days or later (Figure 10(C)). In contrast, the frequency of on-target editing induced by sTALED-V28R and -R111S increased or even remained stable over time. Consequently, the editing frequency observed with the new sTALED mutants was low at 1 and 2 days post-transfection, but was higher at 8 days or later than that induced by previous versions of sTALEDs (Figures 10(A) to 10(C)).

[0178] Based on these results, sTALEDs and the sTALED-V106W mutant frequently induced excessive RNA off-target editing, resulting in cytotoxicity, and mtDNA-edited cells either failed to divide or died over time. However, the sTALED mutants of the present invention were acceptable because they avoided RNA off-target editing or did not induce excessive bystander editing at the target site or off-target mutations in the mitochondrial genome. Furthermore, an MTS cell proliferation assay (Figure 11(A)) confirmed that sTALEDs and the sTALED-V106W mutant were indeed cytotoxic, significantly reducing cell viability compared to the negative control (pEGFP-transfected) (Figure 11(B)). The sTALED mutants were much better tolerated, with no reduction in cell viability even four days after transfection (Figure 11(C)).

[0179] Example 5: TadA8e-V28R and -R111S integrated into CRISPR RNA-guided ABEs

[0180] In Figure 12, we investigated whether the V28R and R111S mutations in TadA8e could reduce bystander editing and RNA off-target editing induced by CRISPR RNA-guided ABEs, which are widely used for nuclear DNA editing. First, we measured on-target and bystander editing frequencies at the TYRO3 site. ABE8e-V28R and -R111S were as efficient as ABE8e and ABE8e-V106W (ABE8eW) at target sites with editing frequencies >30% (Figures 12(B) and 12(C)). The V28R and R111S mutants showed a narrow editing range with efficient editing (up to 37%) at the fifth and seventh positions of the protospacer region (A5 and A7 in Figure 12(C)), but almost no bystander editing at the tenth position (A10) (0.8% and 0.5%, respectively). In contrast, ABE8e and ABE8eW showed an even broader editing range with maximum editing at A5 (29% and 30%, respectively) and A7 (29% and 30%, respectively), and substantial bystander editing at A10 (9.3% and 3.1%, respectively).

[0181] Next, we evaluated RNA off-target editing activity using targeted RNA amplicon sequencing. ABE8e-V28R and -R111S reduced the average RNA off-target editing frequency measured at a total of six sites by 3.8-fold and 2.5-fold, respectively, compared to ABE8e, and by 3.1-fold and 2.0-fold, respectively, compared to ABE8eW (Figure 12(D)). Furthermore, transcriptome-wide sequencing demonstrated that ABE8e-V28R and -R111S significantly reduced the number of RNA off-target edits compared to ABE8e and ABE8eW (Figures 12(E) and 12(F)). These results indicate that the two novel TadA8e mutants, V28R and R111S, reduce RNA off-target editing and bystander editing, potentially improving ABE8e.

[0182] Example 6: DNA on-target and RNA off-target editing by mTALEDs and sTALEDs

[0183] In Figure 13, in addition to sTALED, we investigated whether dimeric TALEDs (dTALED) and monomeric TALEDs (mTALED) mutants with V28R and R111S can reduce RNA off-target editing. In the dTALED, each TALE unit has a cytosine deaminase (e.g., DddA) on one side. tox On the other hand, the mTALED is characterized in that both cytosine deaminase and adenine deaminase are present in a single TALE unit.

[0184] Figure 13(a) confirms that both dTALEDs and mTALEDs targeting the Cox3 site exhibit on-target activity. Furthermore, it was observed that gene editing did not occur when only adenine deaminase was present in the TALE unit. Figure 13(b) shows the results of editing efficiency at six representative sites. Similar to the sTALEDs, both dTALEDs and mTALEDs with the V28R and R111S mutants exhibited significantly reduced RNA off-target efficiency compared to TadA8e.

[0185] In Figure 13(c), we analyzed the specificity ratio of RNA off-target editing relative to on-target editing induced by Cox3-specific TALEDs. The results show that the TadA mutant with V28R and R111S reduced the RNA off-target efficiency not only in other systems such as sTALED and ABE, but also in dTALEDs and mTALEDs. These results indicate that the RNA off-target effects were primarily induced by TadA adenine deaminase rather than DNA-binding proteins or cytosine deaminase. Furthermore, the TadA mutant significantly reduced these undesirable RNA off-target effects.

[0186] From the above description, those skilled in the art will understand that the present invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. In this regard, it should be understood that the above-described embodiments are illustrative in all respects and not restrictive. The scope of the present invention should be construed as including within the scope of the present invention as defined by the appended claims.

Claims

1. An adenine deaminase comprising an amino acid sequence represented by SEQ ID NO: 1 or an amino acid sequence having at least 80% sequence identity to the amino acid sequence represented by SEQ ID NO: 1, wherein at least one amino acid residue selected from residues 28, 30, 46, 48, 49, 82, 84, 106, 108, 110, and 111 of the amino acid sequence represented by SEQ ID NO: 1, or a corresponding amino acid residue in an amino acid sequence having at least 80% sequence identity to the amino acid sequence represented by SEQ ID NO: 1, is substituted with another amino acid.

2. 2. The adenine deaminase of claim 1, wherein the substitution is one or more selected from the group consisting of: The amino acid sequence represented by SEQ ID NO: 1 V28Q or V28R, A48W, F84M, V106A, K110S, K110T, or K110V, and R111F, R111Q, R111S, R111T, or R111Y.

3. 3. The adenine deaminase of claim 2, wherein the substitution is one or more selected from the group consisting of: The amino acid sequence represented by SEQ ID NO: 1 V28Q or V28R, A48W, and R111S.

4. DNA binding proteins; and The modified adenine deaminase of claim 1; A fusion protein comprising:

5. 5. The fusion protein of claim 4, wherein the substitution in the modified adenine deaminase is one or more selected from the group consisting of: The amino acid sequence represented by SEQ ID NO: 1 V28Q or V28R, A48W, F84M, V106A, K110S, K110T, or K110V, and R111F, R111Q, R111S, R111T, or R111Y.

6. 6. The fusion protein of claim 5, wherein the substitution in the modified adenine deaminase is one or more selected from the group consisting of: The amino acid sequence represented by SEQ ID NO: 1 V28Q or V28R, A48W, and R111S.

7. The fusion protein of claim 4, wherein the DNA-binding protein is a zinc finger protein, a TALE (transcription activator-like effector), or a CRISPR-associated nuclease.

8. The fusion protein of claim 7, wherein the DNA binding protein is a TALE.

9. The fusion protein of claim 8, wherein the TALE is a dual TALE module composed of a first TALE module and a second TALE module.

10. The fusion protein of claim 4, wherein the DNA binding protein is a Cas protein without Cas nickase or endonuclease activity.

11. The fusion protein according to claim 4 , further comprising cytosine deaminase.

12. 12. The fusion protein of claim 11, wherein the cytosine deaminase is in a full-length form or a two-split form.

13. The cytosine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), DddA tox 13. The fusion protein according to claim 11 or 12, which is cytosine deaminase specific to double-stranded DNA (CdD), or a variant thereof.

14. The fusion protein according to any one of claims 4 to 12, further comprising UGI (uracil glycosylase inhibitor).

15. The fusion protein according to any one of claims 4 to 12, wherein the DNA is nuclear DNA or organelle DNA.

16. The fusion protein of claim 15, further comprising an NLS (nuclear localization signal).

17. The fusion protein of claim 15, further comprising a mitochondrial targeting sequence (MTS) or a chloroplast transit peptide (CTP).

18. The fusion protein of claim 15, further comprising a nuclear export signal (NES).

19. The fusion protein of claim 4 further comprising a nickase.

20. 20. The fusion protein of claim 19, comprising a first fusion protein comprising the DNA binding protein and the modified adenine deaminase, and a second fusion protein comprising the DNA binding protein and the nickase.

21. 5. The fusion protein of claim 4, wherein the nickase is MutH, a MutH mutant, or Nt.BspD6I(C).

22. A polynucleotide encoding the modified adenine deaminase of any one of claims 1 to 3, or the fusion protein of any one of claims 4 to 12 and claims 19 to 21.

23. 23. An expression vector comprising the polynucleotide of claim 22.

24. A base editing composition for A-to-G base editing in DNA, comprising the fusion protein according to any one of claims 4 to 10 and claims 19 to 21, a polynucleotide encoding the fusion protein, or an expression vector comprising the polynucleotide.

25. The base editing composition of claim 24, wherein the fusion protein further comprises cytosine deaminase.

26. The base editing composition of claim 25, wherein the cytosine deaminase exists in the form of two split bodies, and the fusion protein comprises a first fusion protein comprising a first split body of cytosine deaminase and a second fusion protein comprising a second split body of cytosine deaminase.

27. The base editing composition of claim 26, wherein the first segment contains an amino acid sequence represented by SEQ ID NO: 9 or 11, and the second segment contains an amino acid sequence represented by SEQ ID NO: 10 or 12.

28. 25. The base editing composition of claim 24, which exhibits reduced off-target effects, wherein the off-target effects are characterized by unintended base alterations in DNA and / or RNA.

29. A method for A-to-G base editing in DNA, comprising a step of transfecting a cell containing target DNA with the base editing composition of claim 24.

30. 30. The method of claim 29, wherein said transferring is performed ex vivo.

31. 30. The method of claim 29, wherein the delivery is performed in vivo.

32. 30. The method of claim 29, wherein the target DNA is nuclear DNA, organelle DNA, or mitochondrial DNA of a human subject with a genetic disease.

33. 30. The method of Claim 29, wherein the method induces base editing at a single nucleotide residue in the target DNA at a frequency of at least 10%.

34. 30. The method of claim 29, wherein the method exhibits reduced off-target editing effects compared to when using a base editor comprising an adenine deaminase comprising the amino acid sequence represented by SEQ ID NO: 1, wherein the off-target editing effects are characterized by unintended base mutations in DNA or RNA.

35. 30. The method of Claim 29, wherein the method exhibits reduced off-target editing effects compared to using a base editor comprising an adenine deaminase comprising the amino acid sequence set forth in SEQ ID NO: 1 with a V106W mutation, wherein the off-target editing effects are characterized by unintended base mutations in DNA or RNA.

36. 30. The method of Claim 29, wherein base editing at a single nucleotide residue in the target gene is induced at a frequency of 10% or greater.

37. A method for reducing off-target effects, comprising a step of delivering the base editing composition of claim 24 to a cell containing target DNA, wherein the method exhibits reduced off-target editing effects compared to when using a base editor comprising an adenine deaminase having an amino acid sequence represented by SEQ ID NO: 1, and the off-target editing effects are characterized by unintended base mutations in DNA or RNA.

Citation Information

Patent Citations

  • Base editor as well as construction method and application thereof

    CN114045277A