Novel adenine deaminase variant and method for base editing using same

By developing a fusion protein containing adenine deaminase variant and DNA binding protein, the off-target editing effect problem of adenine base editor in organelle DNA editing was solved, and a more efficient and specific base editing effect was achieved.

CN119998447APending Publication Date: 2025-05-13INST FOR BASIC SCI
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202380067900.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2023-07-04
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing adenine base editors have off-target editing effects in organelle DNA editing, resulting in unanticipated off-target base editing of the whole transcriptome.

Method used

A fusion protein including adenine deaminase variant and a DNA binding protein was developed to reduce off-target editing effects. The frequency of RNA off-targeting and DNA off-target editing is reduced by introducing specific amino acid substitutions such as V28Q, V28R, A48W and R111S.

Benefits of technology

The off-target effect of unexpected base changes in DNA and RNA is significantly reduced, improving editing specificity and efficiency, and reducing unwanted bystander effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998447A_ABST
    Figure CN119998447A_ABST
Patent Text Reader

Abstract

The present invention relates to a novel adenine deaminase variant, a fusion protein comprising the adenine deaminase variant, a base editing composition for A-to-G base editing in DNA comprising the fusion protein, and a method for A-to-G base editing in DNA comprising delivering the base editing composition to a cell comprising a target DNA. The novel adenine deaminase variants can result in a significant reduction in off-target effects involving unintended base changes in DNA and / or RNA, and induce base editing only at a single nucleotide residue without any unintended off-target editing occurring in the target DNA.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Provided are: an adenine deaminase variant capable of reducing off-target editing; a fusion protein comprising a DNA binding protein and an adenine deaminase variant; a base editing composition for A to G base editing in DNA comprising the fusion protein; and a method for A to G base editing in DNA comprising delivering the base editing composition to a cell containing target DNA. Background Art

[0002] Targeted base editing of mammalian mitochondrial DNA (mtDNA) is a powerful and versatile technology that can be used to model mitochondrial genetic diseases in cell lines and animals and to develop novel therapeutic modalities to correct disease-causing mutations in patients. Programmable deaminases, which consist of a customized DNA-binding protein and a nucleobase deaminase, enable mitochondrial DNA editing in a targeted manner.

[0003] Unlike cytosine base editors, adenine base editors (called TALEDs) incorporate TadA8e, a deoxyadenine deaminase engineered from the tRNA-specific TadA protein from Escherichia coli (E. coli). TadA8e and related variants of TadA are key components in CRISPR RNA-guided adenine base editors, which are widely used for A to G base editing of nuclear DNA. However, editing organellar DNA using RNA-guided ABEs is a challenge due to the difficulty in delivering guide RNA into organelles. In addition, the presence of TadA8e in ABEs retains residual deaminase activity for RNA substrates, leading to unintended off-target base editing of the entire transcriptome. Off-target effects occur when the editing mechanism mistakenly acts on unintended genomic regions, resulting in undesired modifications.

[0004] Despite widespread interest in programmable deaminase-mediated base editing, strategies for developing adenine base editors with enhanced specificity that reduce the occurrence of off-target editing effects are currently lacking. The development of such base editors that minimize these side effects and the implementation of methods to reduce them are of great importance, making them more reliable and applicable to a wider range of scientific and therapeutic applications. Summary of the invention

[0005] The present invention relates to base editing compositions and methods of their use in gene therapy and genome engineering. More specifically, the compositions and methods relate to base editor variants that, when employed in clinical settings, exhibit significant reductions in off-target genomic cleavage and side effects (such as RNA off-target toxicity). By introducing innovative adenine base editors and accompanying novel utilization methods, the present invention provides a new approach for targeted base editing, with broad applications in the fields of medicine and biotechnology.

[0006] In one embodiment, provided herein are adenine deaminase variants.

[0007] In another embodiment, provided herein is a fusion protein comprising a DNA binding protein and an adenine deaminase variant.

[0008] In another embodiment, provided herein are polynucleotides encoding adenine deaminase variants or fusion proteins, and expression vectors comprising the polynucleotides.

[0009] In one embodiment, described herein is a base editing composition for A to G base editing in DNA, which includes a fusion protein, a polynucleotide encoding the fusion protein, or an expression vector including the polynucleotide.

[0010] In another embodiment, provided herein is a method for A to G base editing in DNA, comprising delivering a base editing composition to a cell containing the target DNA.

[0011] In another embodiment, described herein is a method for reducing off-target editing effects comprising delivering a base editing composition to a cell containing target DNA.

[0012] In another embodiment, the use of an adenine deaminase variant or fusion protein in A to G base editing in DNA and / or reducing the off-target effects of A to G base editing in DNA, or in preparing a composition for A to G base editing in DNA and / or reducing the off-target effects of A to G base editing in DNA is provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1(a) illustrates the structure of the base editor used in the present invention. AD (TadA8e adenine deaminase); AD* (TadA8e adenine deaminase variant); MTS (mitochondrial targeting sequence); UGI (uracil glycosylase inhibitor).

[0014] FIG. 1( b ) is a graph showing the amount and frequency of RNA editing by Cox3.1-specific sTALED.

[0015] Figure 1(c) is a graph showing A to G editing and C to T editing of Cox3.1-specific sTALEDs (left), and a graph showing on-target activities of Cox3.1-specific sTALED variants relative to the on-target activities of wild-type Cox3.1-specific sTALEDs (right).

[0016] FIG. 1( d ) is a graph showing the amount and frequency of RNA editing by ND1 -specific sTALED.

[0017] FIG. 1( e ) is a graph showing A to G editing and C to T editing of ND1-specific sTALED (left), and a graph showing on-target activity of ND1-specific sTALED variants relative to the on-target activity of wild-type ND1-specific sTALED (right).

[0018] Figure 2 (a) Structural representation of the TadA portion of ABE8e (Protein Data Bank (PDB) accession number 6VPC).

[0019] Figure 2 (b) is a heat map showing DNA on-target activity (left), RNA off-target activity (middle), and the relative ratio of DNA on-target editing frequency to RNA off-target editing frequency (right) for 101 sTALED variants out of a total of 209 sTALED variants that retain mitochondrial DNA on-target activity. The relative ratios are normalized to the relative ratio of the original sTALED, which has a value of 1.

[0020] Figure 3 is a graph showing the frequency of on-target DNA base editing induced by 209 sTALED variants.

[0021] FIG. 4( a ) is a graph showing editing frequencies at six RNA off-target sites measured by targeted RNA sequencing.

[0022] FIG. 4( b ) is a graph showing the frequencies of RNA off-target editing induced by sTALED and sTALED variants at six selected RNA off-target sites extracted from whole transcriptome sequencing.

[0023] FIG. 4( c ) is a graph showing the frequency of RNA off-target editing induced by sTALED and sTALED variants analyzed by targeted RNA amplicon sequencing.

[0024] FIG. 5( a ) is a graph showing the amount and frequency of RNA editing by Cox3.1-specific sTALED.

[0025] FIG. 5( b ) is a graph showing A to G editing and C to T editing of Cox3.1-specific sTALEDs.

[0026] FIG. 5( c ) is a graph showing the amount and frequency of RNA editing by ND1-specific sTALED.

[0027] FIG. 5( d ) is a graph showing A to G editing and C to T editing of ND1 -specific sTALEDs.

[0028] FIG. 5( e ) is a graph showing the amount and frequency of RNA editing by ND6-specific sTALED.

[0029] FIG. 5( f ) is a graph showing A to G editing and C to T editing of ND6-specific sTALEDs.

[0030] FIG. 6( a ) is a graph showing the on-target activity of ND1 -specific sTALED variants relative to the on-target activity of the wild-type ND1 -specific sTALED.

[0031] FIG. 6( b ) is a graph showing the on-target activity of ND6-specific sTALED variants relative to the on-target activity of wild-type ND6-specific sTALED.

[0032] FIG. 6( c ) is a heat map showing RNA off-target activities of ND1-specific sTALED and sTALED variants at six representative sites.

[0033] FIG6( d ) is a heat map showing RNA off-target activities of ND6-specific sTALED and sTALED variants at six representative sites.

[0034] Figure 6(e) is a bar chart showing the frequency of RNA off-target editing induced by ND1-specific sTALED and sTALED variants at six representative sites. The editing efficiency of each replicate is shown as a point on the chart. The error bars are sem of n = 2 (replicate) × 6 (RNA off-target sites) biologically independent samples.

[0035] Figure 6(f) is a bar chart showing the frequency of RNA off-target editing induced by ND6-specific sTALED and sTALED variants at six sites. The efficiency of each replicate is shown as a point on the chart. The error bars are the sem of n = 2 (replicate) × 6 (RNA off-target sites) biologically independent samples.

[0036] FIG. 6( g ) is a bar graph showing the ratio of DNA on-target editing frequency relative to RNA off-target editing frequency induced by ND1-specific sTALED and sTALED variants.

[0037] Figure 6(h) is a bar chart showing the ratio of DNA on-target editing frequency induced by ND6-specific sTALED and sTALED variants relative to RNA off-target editing frequency. The ratio is normalized to the ratio of sTALED, which has a value of 1.

[0038] FIG. 6( i ) is a graph showing the average relative ratio values ​​of sTALEDs and sTALED variants targeting Cox3.1, ND1 ( FIG. 6( g ) ), and ND6 ( FIG. 6( h ) ).

[0039] FIG. 7( a ) is a heat map depicting A to G transitions caused by sTALED or sTALED variants targeted to the Cox3.1 site.

[0040] FIG. 7( b ) is a heat map depicting A to G transitions caused by sTALED or sTALED variants targeted to the ND1 site.

[0041] FIG. 7( c ) is a heat map depicting A to G transitions caused by sTALED or sTALED variants targeting the ND6 site.

[0042] Figure 7(d)-(f) shows the analysis of Cox3.1, ND1, and ND6 alleles summarized in Figure 7(a)-(c), respectively. The interval sequence is shown on the left, and a bar chart showing the frequency of each allele is shown on the right. The reference sequence is written in all capital letters, while lowercase letters indicate the location where base editing occurs.

[0043] Figure 8(a) shows a graph showing the locations of on-target and off-target edits throughout the mitochondrial genome at day 4 post-transfection. Black and grey dots represent off-target edits and naturally occurring single nucleotide variations (SNVs), respectively, and arrows represent on-target (and bystander) edits. Nucleotide positions in the human mitochondrial genome are indicated on the X-axis.

[0044] Figure 8(b) shows the average frequency of genome-wide off-target editing induced by wild-type TALED and TALED variants.

[0045] Figure 9(a) shows a graph showing the locations of on-target and off-target editing throughout the mitochondrial genome on day 2 post-transfection. Black and gray dots represent off-target editing and naturally occurring single nucleotide variations (SNVs), respectively, and arrows represent on-target (and bystander) editing. Nucleotide positions in the human mitochondrial genome are represented on the X-axis.

[0046] Figure 9(b) is a graph showing the average frequency of genome-wide off-target editing induced by wild-type TALED and TALED variants. Error bars are sem of n=2 biologically independent samples.

[0047] Figure 10 (a)-(c) are line graphs showing the frequency of on-target base editing induced by Cox3.1-(a), ND1-(b) and ND6-specific (c) sTALEDs and sTALED variants (V106W, V28R, R111S) over time. Error bars are sem of n=2 biologically independent samples.

[0048] Figure 10 (d)-(f) is a line graph showing the frequency of RNA off-target base editing induced by Cox3.1-(d), ND1-(e) and ND6-specific (f) sTALEDs and sTALED variants (V106W, V28R, R111S) at six representative sites over time. Error bars are sem of n=2 biologically independent samples.

[0049] Figure 10(g) is a line graph showing the average RNA off-target editing frequency over time for Cox3.1, ND1 and ND6-specific sTALEDs and sTALED variants. Error bars are sem of n=2 biologically independent samples.

[0050] FIG. 11( a ) illustrates an experimental scheme of one embodiment of the present invention.

[0051] Figures 11(b) and (c) are bar graphs showing the viability of cells transfected with plasmids expressing sTALED, sTALED-V106W, sTALED-V28R and sTALED-R111S, which target the indicated sites as determined by observing color changes caused by formazan formation in the MTS assay at day 2 (B) and day 4 (C) post-transfection. The absorbance values ​​were normalized to the absorbance values ​​of cells transfected with pEGFP as a control. Error bars are sem of n=2 biologically independent samples.

[0052] Figure 12(a) illustrates the architecture of ABE8e and ABE8e variant constructs. AD (TadA8e adenine deaminase); AD* (TadA8e adenine deaminase variant); NLS (nuclear localization sequence).

[0053] FIG. 12( b ) is a graph showing the on-target activity of ABE8e and ABE8e variants (ABE8e-V106W, ABE8e-V28R, ABE8e-R111S) at the nuclear TYRO3 site.

[0054] FIG. 12( c ) depicts a heat map showing the frequency of A to G transitions caused by ABE8e and ABE8e variants (ABE8e-V106W, ABE8e-V28R, ABE8e-R111S) at the nuclear TYRO3 site.

[0055] FIG. 12( d ) is a graph showing RNA off-target activities of ABE8e and ABE8e variants (ABE8e-V106W, ABE8e-V28R, ABE8e-R111S) targeting TYRO3 at six representative sites.

[0056] Figures 12(e) and 12(f) are graphs illustrating the total number of RNA edits found in HEK 293T cells expressing ABE8e or ABE8e variants targeted to the nuclear TYRO3 site, assessed by whole transcriptome sequencing.

[0057] Fig.13 (a) is a graph showing the DNA on-target activity of Cox3-specific TALEDs, including dimeric TALEDs (dTALEDs), semi-monomers (dTALED-ADs), monomeric TALEDs (mTALEDs), and untreated samples.

[0058] Fig.13 (b) is a graph showing RNA off-target activities of Cox3-specific TALEDs at six representative sites, including dimeric TALEDs (dTALEDs), semi-monomers (dTALED-ADs), monomeric TALEDs (mTALEDs), and untreated samples.

[0059] Fig.13 (c) is a graph showing the specific ratio of RNA off-target editing relative to on-target editing induced by Cox3-specific TALED. DETAILED DESCRIPTION

[0060] The following definitions supplement those in the art and should not be attributed to any related or unrelated cases, e.g., any commonly owned patents or applications, for the current application. Although any methods and materials similar to or equivalent to those described herein may be used in the practice or testing of the present disclosure, preferred materials and methods are described herein. Therefore, the terms used herein are only used for the purpose of describing specific embodiments and are not intended to be limiting. Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which the present invention belongs.

[0061] In this application, unless otherwise specifically stated, the use of the singular includes the plural. It must be noted that the singular forms "a", "an", and "the" used in this specification include plural referents unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" and should be understood to be inclusive unless otherwise stated. In addition, the use of the term "including" as well as other forms such as "include", "includes", and "included" is not restrictive.

[0062] As used herein, the term "about" or "approximately" means within an acceptable error range for a particular value, as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to the practice in the art, "about" can mean within 1 or more than 1 standard deviation. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude of a value, such as within 5 times or within 2 times of a value. Where specific values ​​are described in the present application and claims, the meaning of the term "about" within an acceptable error range for a particular value should be assumed unless otherwise stated.

[0063] As used herein, the term "corresponding" or "corresponds" refers to an amino acid residue at a listed position in a polypeptide or an amino acid residue that is similar, identical, or homologous to those listed positions in a polypeptide. Identifying an amino acid at a corresponding position can be determining a specific amino acid in a sequence with reference to a specific sequence. As used herein, a "corresponding region" generally refers to a similar or corresponding position in a related protein or a reference protein. For example, an arbitrary amino acid sequence is aligned with SEQ ID NO:3, on which basis, each amino acid residue of the amino acid sequence can be numbered with reference to the numerical position of the amino acid residues of SEQ ID NO:3 and the corresponding amino acid residue. For example, the sequence alignment algorithm described in the present disclosure can determine the position of an amino acid or the position where a modification, such as a substitution, insertion, or deletion, occurs by comparing it to the amino acid position or the position where a modification, such as a substitution, insertion, or deletion occurs in a query sequence (also referred to as a "reference sequence").

[0064] As used herein, the term "alignment" refers to mapping sequence reads to a reference genome, and then aligning bases with identical sites in the genome to fit each site. Therefore, any computer program can be used as long as it can align sequence reads in the same manner as described above. The program can be known in the relevant art, or can be selected from programs customized for the purpose. In one embodiment, ISAAC is used for alignment, but is not limited thereto.

[0065] References in the specification to "various embodiments," "some embodiments," "an embodiment," or "other embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least some embodiments of the present disclosure, but not necessarily all embodiments.

[0066] As used herein, the term "host cell" (or "recombinant host cell") means a cell that has undergone genetic alteration, or a cell that can undergo genetic alteration by the introduction of an exogenous polynucleotide molecule such as a recombinant plasmid or vector. It should be understood that these terms refer not only to the specific subject cell, but also to the progeny of such a cell. Since certain modifications may occur in subsequent generations due to mutations or environmental influences, such progeny may not actually be identical to the parent cell, but are still included within the scope of the term "host cell" as used herein.

[0067] As used herein, the expression "base editor (BE)" refers to an agent that binds to a polynucleotide and has a nuclear base modification activity. In one embodiment, the base editor includes a nuclear base modification polypeptide (e.g., a deaminase) and a nucleic acid programmable nucleotide binding domain. In another embodiment, the base editor includes a nuclear base modification polypeptide (e.g., a deaminase) and a nucleic acid programmable nucleotide binding domain and a guide polynucleotide (e.g., a guide RNA). However, in another embodiment, the agent is a biomolecule complex that includes a protein domain with base editing activity, i.e., a domain that can modify a base (e.g., A, T, C, G, I, or U) in a nucleic acid molecule (e.g., DNA). In some embodiments, a polynucleotide programmable DNA binding domain is fused or connected to a deaminase domain. In one embodiment, the agent is a fusion protein that includes a domain with base editing activity. In another embodiment, a protein domain with base editing activity is connected to a guide RNA (e.g., via an RNA binding motif on a guide RNA and an RNA binding domain fused to a deaminase). In some embodiments, the domain with base editing activity is capable of deaminating a base within a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating one or more bases within a DNA molecule. In some embodiments, the base editor is capable of deaminating adenine (A) within DNA. In some embodiments, the base editor is an adenine base editor (ABE).

[0068] As used herein, "administration" in this article refers to providing one or more compositions described herein to a patient or subject. For example, but not limited to, the administration of the composition, such as injection, can be performed by intravenous (iv) injection, subcutaneous (sc) injection, intradermal (id) injection, intraperitoneal (ip) injection or intramuscular (im) injection. One or more such routes can be used. Parenteral administration can be, for example, by bolus injection or by gradual infusion over time. In some embodiments, parenteral administration includes intravascular, intravenous, intramuscular, intraarterial, intrathecal, intratumoral, intradermal, intraperitoneal, transtracheal, subcutaneous, subcutaneous, intraarticular, subcapsular, subarachnoid and intrasternal infusion or injection. Alternatively, or in parallel, administration can be by oral route.

[0069] As used herein, the expression "another amino acid" may mean an amino acid selected from alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, asparagine, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartic acid, glutamic acid, arginine, histidine, lysine, and all known variants thereof, excluding the amino acid having the wild-type protein retained at the original substitution position.

[0070] As used herein, the term "off-target site" may refer to a site that is not on-target but to which an adenine base editor shows activity. That is, an off-target site may refer to a site at which base editing occurs in addition to the on-target site. In an embodiment, the term "off-target site" may be used to cover not only sites that are not on-target sites of an adenine base editor, but also sites that may become off-target sites thereof.

[0071] As used herein, the term "whole genome sequencing" (WGS) refers to a method of reading a genome at many multiples, such as 10X, 20X, and 40X formats, by next-generation sequencing to perform whole genome sequencing. The term "next-generation sequencing" refers to a technique of fragmenting a whole genome or a target region of a genome in a chip-based and PCR-based double-end format, and sequencing the fragments with high throughput based on chemical reactions (hybridization).

[0072] As used herein, the term "nucleic acid" refers to DNA or RNA. "Nucleic acid sequence" or "polynucleotide sequence" refers to a single-stranded or double-stranded polymer of deoxyribonucleotide or ribonucleotide bases read from the 5' to 3' end. It includes self-replicating plasmids, infectious DNA or RNA polymers, and non-functional DNA or RNA.

[0073] The phrase "nucleic acid molecule encoding" refers to a nucleic acid molecule that directs the expression of a specific protein or peptide. Nucleic acid sequences include DNA chain sequences that are transcribed into RNA and RNA sequences that are translated into proteins or peptides. Nucleic acid molecules include full-length nucleic acid sequences and non-full-length sequences derived from full-length proteins. It is further understood that the sequence includes degenerate codons of the native sequence, or includes degenerate codons that may be introduced to provide codon preference sequences in specific host cells.

[0074] The term "vector" refers to a viral expression system, an autonomous self-replicating circular DNA (plasmid), and includes expression plasmids and non-expression plasmids. When a recombinant microorganism or cell is described as hosting an "expression vector", this includes both extrachromosomal circular DNA and DNA that has been integrated into one or more host chromosomes. When the vector is maintained by a host cell, the vector can be stably replicated by the cell as an autonomous structure during mitosis, or the vector can be integrated into the host's genome.

[0075] The term "plasmid" refers to an autonomous circular DNA molecule that can replicate in a cell, and includes both expressed and non-expressed types. When a recombinant microorganism or cell is described as hosting an "expression plasmid," this includes latent viral DNA that is integrated into one or more host chromosomes. When a plasmid is maintained by a host cell, the plasmid can be stably replicated by the cell as an autonomous structure during mitosis, or the plasmid can be integrated into the host's genome.

[0076] As used herein, "percentage of amino acid sequence homology" or "percentage of amino acid sequence identity" refers to the percentage between a first amino acid sequence and a second amino acid sequence that can be calculated by dividing [the number of amino acid residues in the first amino acid sequence that are identical to the amino acid residues at the corresponding positions in the second amino acid sequence] by [the total number of amino acid residues in the first amino acid sequence] and multiplying by [100%], wherein each deletion, insertion, substitution or addition of an amino acid residue in the second amino acid sequence compared to the first amino acid sequence is regarded as a difference at a single amino acid residue (position), i.e., an "amino acid difference" as defined herein. Alternatively, the degree of sequence identity between two amino acid sequences can be calculated using known computer algorithms, such as those described above for determining the degree of sequence identity of nucleotide sequences, again using standard settings.

[0077] In various embodiments of the present invention, in order to address the problem of adverse off-target effects associated with TadA8e, a process for selecting TadA8e variants with reduced RNA off-target effects or DNA off-target effects (including, for example, bystander off-target effects) is described. In one embodiment, this purpose is achieved by replacing different amino acid residues at specific positions that interact with nucleotides within TadA8e. In addition, in order to evaluate the effects of these variants on RNA or DNA off-target effects, RNA sequencing and whole mitochondrial genome sequencing were performed. This comprehensive analysis allows the identification and confirmation of a representative selection of the entire RNA off-target site range or six prominent RNA off-target sites. In addition, recognizing that RNA has transient properties, the dynamics of RNA off-target effects over time are also measured. Such measurements are performed at different time points to evaluate how RNA off-targets are comprehensively presented and change during expression.

[0078] In one embodiment, an adenine deaminase is provided, comprising an amino acid sequence of SEQ ID NO: 1 or an amino acid sequence having at least 80% sequence homology to the amino acid sequence of SEQ ID NO: 1, wherein at least one amino acid residue selected from residues 28, 30, 46, 48, 49, 82, 84, 106, 108, 110 and 111 of SEQ ID NO: 1 or an amino acid sequence having at least 80% sequence homology to the amino acid sequence of SEQ ID NO: 1 The corresponding amino acid residues of an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% sequence homology (or sequence identity) to NO:1 are replaced by another amino acid.

[0079] The term "adenine deaminase" refers to a polypeptide or fragment that can catalyze the hydrolytic deamination of adenine or adenosine. In one embodiment, the deaminase or deaminase domain represents an adenine deaminase that promotes the hydrolytic deamination of adenosine to inosine or the hydrolytic deamination of deoxyadenosine to deoxyinosine. In addition, in another embodiment, the adenine deaminase performs the hydrolytic deamination of adenine or adenosine in DNA (deoxyribonucleic acid). The adenosine deaminase described herein, such as an engineered adenosine deaminase or an evolved adenosine deaminase, can be derived from any organism, including bacteria.

[0080] In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA8e. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as bacteria, archaea, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a naturally occurring deaminase.

[0081] TadA8e adenosine deaminase has the following sequence:

[0082] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN(SEQ ID NO:1)

[0083] In some embodiments, the amino acid substitution can be at least one selected from the group consisting of V28Q, or V28R, A48W, F84M, V106A, K110S, K110T or K110V ​​and R111F, R111Q, R111S, R111T or R111Y of the amino acid sequence of SEQ ID NO:1.

[0084] In one embodiment, the amino acid substitution with the lowest RNA or DNA off-target editing efficiency (including, for example, bystander off-target effects) can be at least one selected from the group consisting of V28Q, or V28R, A48W and R111S of the amino acid sequence of SEQ ID NO: 1.

[0085] In various embodiments, adenine deaminase variants can exhibit significantly reduced off-target effects involving unexpected base changes in DNA and / or RNA. In another embodiment, adenine deaminase variants can reduce unwanted bystander effects while narrowing the activity window. Alternatively, adenine deaminase variants can induce base editing only at a single nucleotide residue without any intentional off-target editing in the target DNA, with a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0086] In one embodiment, a fusion protein is provided, which includes a DNA binding protein; and an adenine deaminase variant. The amino acid replacement of the adenine deaminase variant can be selected from the group consisting of V28Q, or V28R, A48W, F84M, V106A, K110S, K110T, or K110V ​​and R111F, R111Q, R111S, R111T, or R111Y of the amino acid sequence of SEQ ID NO: 1. At least one of. In one embodiment, the amino acid replacement of the adenine deaminase variant with the lowest RNA or DNA off-target editing effect (including, for example, bystander off-target effects) can be selected from the group consisting of V28Q, or V28R, A48W, and R111S of the amino acid sequence of SEQ ID NO: 1.

[0087] In various embodiments, the fusion protein can exhibit significantly reduced off-target effects involving unexpected base changes in DNA and / or RNA. In another embodiment, the fusion protein can reduce unwanted bystander effects while narrowing the activity window. Alternatively, the fusion protein can induce base editing only at a single nucleotide residue without any intentional off-target editing in the target DNA, with a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0088] The DNA binding protein can be, for example, but not limited to, 1) a zinc finger protein, 2) a transcription activator-like effector (TALE) protein, 3) a CRISPR-associated nuclease. The nuclease is type II and / or type V, such as a Cas protein (e.g., a Cas9 protein (CRISPR (clustered regularly interspaced short palindromic repeats)-associated protein 9)) or a Cpf1 protein (CRISPR from Prevotella and Francisella 1). Nucleases (e.g., endonucleases) associated with a CRISPR system, etc. can be used. Specifically, in one embodiment, the nuclease can be a Cas protein, such as Cas3, Cas9, Cpf1, Cas6, or C2c2, specifically a CRISPR / Cas II type Cas protein, more specifically a Cas9 protein derived from Streptococcus Pyogenes.

[0089] As used herein, the term "transcription activator-like effector" (TALE) refers to a DNA binding protein comprising a TALE repeat array, which includes a plurality of highly conserved 33-34 amino acid sequences comprising a highly variable two-amino acid motif (repeat variable diresidues, RVD). The RVD motif determines the binding specificity to the nucleic acid sequence and can be engineered to specifically bind to the desired DNA sequence according to methods well known to those skilled in the art. The simple relationship between amino acid sequence and DNA recognition allows for engineering of specific DNA binding domains by selecting a combination of repeat fragments containing the appropriate RVD.

[0090] In one embodiment, the DNA binding protein can be a TALE. In another embodiment, the TALE can be a dual TALE module, which consists of a first TALE module and a second TALE module. In some embodiments, each of the first and second TALE modules can be linked to various deaminases. For example, the first TALE module can be linked to a cytosine deaminase such as DddA in full length. tox , while the second TALE module can be attached to the adenine deaminase variant.

[0091] Cas9 protein is the main protein component of the CRISPR / Cas system, which can function as an activated endonuclease or nickase.

[0092] Cas9 protein or its gene information can be obtained from well-known databases such as GenBank of NCBI (National Center for Biotechnology Information). For example, Cas9 protein can be at least one selected from the group consisting of, but not limited to:

[0093] Cas9 protein from Streptococcus sp., such as Streptococcus pyogenes (e.g., SwissProt accession number Q99ZW2 (NP_269215.1) (encoding gene: SEQ ID NO: 229);

[0094] Cas9 protein derived from Campylobacter sp., such as Campylobacter jejuni;

[0095] Cas9 proteins derived from Streptococcus species, such as Streptococcus thermophiles or Staphylococcus aureus;

[0096] Cas9 protein from Neisseria meningitidis;

[0097] Cas9 proteins derived from Pasteurella sp., such as Pasteurella multocida; and

[0098] Cas9 protein derived from Francisella sp., such as Francisella novicida.

[0099] The Cpf1 protein is an endonuclease of a new CRISPR system that is different from the CRISPR / Cas system. Compared with Cas9, it is smaller in size, does not require tracrRNA, and can function with a single guide RNA. In addition, Cpf1 can recognize thymidine-rich PAM (protospacer adjacent motif) sequences and generate sticky double-strand breaks (sticky ends).

[0100] For example, the Cpf1 protein can be an endonuclease derived from a Candidatus spp., Lachnospira spp., Butyrivibrio spp., Peregrinibacteria, Acidominococcus spp., Porphyromonas spp., Prevotella spp., Francisella spp., Candidatus Methanoplasma), or Eubacterium spp. Examples of microorganisms from which the Cpf1 protein can be derived include, but are not limited to, Parcubacteria bacterium (GWC2011_GWC2_44_17), Lachnospiraceae bacterium (MC2017), Butyrivibrioproteoclasiicus, Peregrinibacteria bacterium (GW2011_GWA_33_10), Acidominococcus sp. (BV3L6), Porphyromonas macacae, Lachnospiraceae bacterium (ND2006), Porphyromonas crevioricanis, Prevotella disiens, Moraxella bovoculi)(237), Smiihella sp.(SC_KO8D17), Leptospira inadai, Lachnospiraceae MA2020, Francisella novicida(U112), Candidatus Methanoplasma termitum, Candidatus Paceibacter, and Eubacteria eligens.

[0101] In one embodiment, when the DNA binding protein is a Cas9 protein, the Cas9 protein may be at least one selected from the group consisting of: a modified Cas9 (e.g., SwissProt accession number Q99ZW2 (NP_269215.1)) in which a mutation is introduced into D10 of the Cas9 protein from Streptococcus pyogenes (e.g., substituted with a different amino acid) to lose endonuclease activity but retain nickase activity, and a modified Cas9 protein in which a mutation is introduced into both D10 and H840 of the Cas9 protein from Streptococcus pyogenes (e.g., substituted with a different amino acid) to lose endonuclease activity and nickase activity. For example, in the Cas9 protein, the mutation at D10 may be a D10A mutation (the amino acid D at position 10 in the Cas9 protein is replaced by A), and the mutation at H840 may be an H840A mutation.

[0102] If the nuclease has nickase activity, the nick can be introduced simultaneously with the diaminobenzase-mediated base modification (e.g., conversion of cytidine to uridine), or can be introduced sequentially in any order, either on the strand where the base modification occurs or on the opposite strand (e.g., the strand opposite to the strand where the base conversion occurs) (e.g., a nick is introduced at a position between the third nucleotide and the fourth nucleotide position in the 5' direction of the PAM sequence on the strand opposite to the strand where the PAM is located). Nuclease mutations (e.g., amino acid substitutions, etc.) can occur in the catalytically active domain of the nuclease (e.g., in the case of Cas9, the RuvC catalytic domain).

[0103] In one embodiment, in the case of a Cas9 protein derived from Streptococcus pyogenes, the mutation may be to replace at least one amino acid selected from the group consisting of the following with another amino acid: catalytic aspartic acid (D10) at position 10, glutamic acid (E762) at position 762, histidine (H840) at position 840, asparagine (N854) at position 854, asparagine (N863) at position 863, and aspartic acid (D986) at position 986. Specifically, it may include a variant in which one or more amino acids selected from the group consisting of H839, H840, and N863 of Cas9 are replaced with another amino acid. Specifically, it may include a variant in which an amino acid at N863, H840-N863, or H839-H840-N863 of Cas9 is replaced with another amino acid. In addition to H840A, D10ASpCas9 nickase SpCas9 nickase prepared by removing part of the catalytic domain may also be used.

[0104] In some embodiments, the fusion protein may further include a guide RNA. The guide RNA may be, for example, at least one selected from the group consisting of CRISPR RNA (crRNA), trans-activating crRNA (tracrRNA) and single guide RNA (sgRNA). Specifically, it may be a double-stranded crRNA:tracrRNA complex, in which crRNA and tracrRNA are bound to each other, or a single-stranded guide RNA (sgRNA), in which crRNA or a portion thereof and tracrRNA or a portion thereof are connected by an oligonucleotide linker.

[0105] The adenine deaminase variant and the DNA-binding protein can be used in the form of a fusion protein, wherein the fusion protein is fused to each other directly or via a peptide linker (for example, in the order of adenine deaminase variant-DNA-binding protein in the N-terminal to C-terminal direction (i.e., the DNA-binding protein is fused to the C-terminal end of the adenine deaminase variant) or in the order of DNA-binding protein-adenine deaminase variant in the N-terminal to C-terminal direction (i.e., the adenine deaminase variant is fused to the C-terminal end of the DNA-binding protein), a mixture of an adenine deaminase variant or mRNA encoding the same and a DNA-binding protein or mRNA encoding the same, a plasmid carrying both an adenine deaminase variant encoding gene and a DNA-binding protein encoding gene (for example, two genes arranged to encode the above-mentioned fusion protein, or a mixture of an adenine deaminase variant expression plasmid and a DNA-binding protein expression plasmid, or plasmids carrying an adenine deaminase variant encoding gene and a DNA-binding protein encoding gene, respectively).

[0106] In one embodiment, the fusion protein may further include a cytosine deaminase. A cytosine deaminase refers to any enzyme that has the ability to convert cytosine found in a nucleotide (e.g., cytosine present in double-stranded DNA or RNA) into uracil (C to U conversion activity or C to U editing activity). A cytosine deaminase converts a cytosine located on a chain where a PAM sequence connected to a target sequence is present into uracil. In an embodiment, a cytosine deaminase may be derived from mammals, including bacteria, archaea, primates such as humans and monkeys, rodents such as rats and mice, etc., but is not limited thereto. For example, a cytosine deaminase may be selected from PmCDA1 (sea lamprey cytosine deaminase 1) from sea lamprey (Petromyzon marinus), DddA from Burkholderia cenocepacia (Burkholderia cenocepacia), or a cytosine deaminase from sea lamprey (Petromyzon marinus). tox and APOBEC (Apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like) family, but not limited thereto.

[0107] In some embodiments, the cytosine deaminase is wild-type sea lamprey CDA1 (pmCDA1) or a catalytic domain thereof. In some embodiments, the cytosine deaminase comprises one or more mutations in the pmCDA1 sequence such that the editing efficiency and / or substrate editing preference of pmCDA1 is changed according to specific needs.

[0108] pmCDA1 has the following amino acid sequence:

[0109] MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWY NQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMIQVKILHTTKSPAVSRGSG(SEQ IDNO:2)

[0110] In some embodiments, as an example of a deaminase, DddA tox To avoid toxicity in host cells, DddA tox It is split into two inactive halves, each of which is fused to a DNA-binding protein in the DddA-derived cytosine base editor (DdCBE). When the DNA-binding protein brings the two inactive halves together, a functional deaminase is reassembled at the target DNA site. Full-length DddA tox Has the following amino acid sequence:

[0111] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPT PYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETL LPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO:3), which corresponds to PDB accession number 6U08_A of Burkhol. deria cenocepacia), and can include fragments or variants thereof, including amino acid sequences having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity with DddA of 6U08_A.

[0112] In another embodiment, the APOBEC (Apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like) family can be, for example, at least one selected from the following groups, but is not limited to:

[0113] APOBEC1: Homo sapiens APOBEC1 (protein: GenBank accession number NP_001291495.1, NP_001635.2, NP_005880.2, etc.; gene (mRNA or cDNA; described in the order of the corresponding protein): GenBank accession number NM_001304566.1, NM_001644.4, NM_005889.3, etc.), Mus musculus APOBEC1 (protein: GenBank accession number NP_001127863.1, NP_112436.1, etc.; gene: GenBank accession number NM_001134391.1, NM_031159.3, etc.);

[0114] APOBEC2: Homo sapiens APOBEC2 (protein: GenBank Accession No. NP_006780.1, etc.; gene: GenBank Accession No. NM_006789.3, etc.), mouse APOBEC2 (protein: GenBank Accession No. NP_033824.1, etc.; gene: GenBank Accession No. NM_009694.3, etc.);

[0115] APOBEC3B: Homo sapiens APOBEC3B (protein: GenBank accession number NP_001257340.1, NP_004891.4, etc.; gene: GenBank accession number NM_001270411.1, NM_004900.4, etc.), Mus musculus APOBEC3B (protein: GenBank accession number NP_001153887.1, NP_001333970.1, NP_084531.1, etc.; gene: GenBank accession number NM_001160415.1, NM_001347041.1, NM_030255.3, etc.);

[0116] APOBEC3C: Homo sapiens APOBEC3C (protein: GenBank accession number NP_055323.2, etc.; gene: GenBank accession number NM_014508.2, etc.);

[0117] APOBEC3D (including APOBEC3E): Homo sapiens APOBEC3D (protein: GenBank accession number NP_689639.2, etc.; gene: GenBank accession number NM_152426.3, etc.);

[0118] APOBEC3F: Homo sapiens APOBEC3F (protein: GenBank accession numbers NP_660341.2, NP_001006667.1, etc.; gene: GenBank accession numbers NM_145298.5, NM_001006666.1, etc.);

[0119] APOBEC3G: Homo sapiens APOBEC3G (protein: GenBank Accession Nos. NP_068594.1, NP_001336365.1, NP_001336366.1, NP_001336367.1, etc.; gene: GenBank Accession Nos. NM_021822.3, NM_001349436.1, NM_001349437.1, NM_001349438.1, etc.);

[0120] APOBEC3H: Homo sapiens APOBEC3H (protein: GenBank Accession Nos. NP_001159474.2, NP_001159475.2, NP_001159476.2, NP_861438.3, etc.; gene: GenBank Accession Nos. NM_001166002.2, NM_001166003.2, NM_001166004.2, NM_181773.4, etc.);

[0121] APOBEC4 (including APOBEC3E): Homo sapiens APOBEC4 (protein: GenBank Accession No. NP_982279.1, etc.; gene: GenBank Accession No. NM_203454.2, etc.); mouse APOBEC4 (protein: GenBank Accession No. NP_001074666.1, etc.; gene: GenBank Accession No. NM_001081197.1, etc.); and

[0122] Activation-induced cytidine deaminase (AICDA or AID): Homo sapiens AID (protein: GenBank Accession No. NP_001317272.1, NP_065712.1, etc.; gene: GenBank Accession No. NM_001330343.1, NM_020661.3, etc.); mouse AID (protein: GenBank Accession No. NP_033775.1, etc., gene: GenBank Accession No. NM_009645.2, etc.), etc.

[0123] The cytosine deaminase may be a non-toxic full-length deaminase (ie, a monomeric cytosine deaminase), or a bipartite form comprising separate first and second domains (ie, a dimeric deaminase), each domain being characterized by the absence of deaminase activity.

[0124] In another embodiment, the adenine deaminase variant can be bound to the N-terminus or C-terminus of a DNA binding protein or a cytosine deaminase or its variant. For example, when the DNA binding protein is a ZFP, the adenine deaminase variant is a TadA8e variant, and the cytosine deaminase or its variant is a DddA tox When, they may be included in the following order, but are not limited thereto: ZFP-TadA8e variant-DddA tox , ZFP-DddA tox -TadA8e variant, TadA-DddA tox -ZFP or DddAtox-TadA8e variant-ZFP.

[0125] In some embodiments, when the cytosine deaminase is contained in a split form and the DNA binding protein is a zinc finger protein, the C-terminus of the first domain of the cytosine deaminase binds to the N-terminus of the zinc finger protein (ZF-left), and the N-terminus of the second domain of the cytosine deaminase binds to the C-terminus of the zinc finger protein (ZF-right) (NC configuration), then the adenine deaminase variant can be attached to the C-terminus of the zinc finger protein (ZF-left), the N-terminus or C-terminus of the first domain of the cytosine deaminase, the N-terminus of the zinc finger protein (ZF-right), or the N-terminus or C-terminus of the second domain of the cytosine deaminase. In various embodiments, the adenine deaminase variant can be bound to:

[0126] The C-terminus of the zinc finger protein attached to the N-terminus of the first domain of cytosine deaminase (ZF-left), and the C-terminus of the zinc finger protein attached to the N-terminus of the second domain of cytosine deaminase (ZF-right).

[0127] (CC configuration);

[0128] The C-terminus of a zinc finger protein attached to the N-terminus of the first domain of cytosine deaminase (ZF-left), and the N-terminus of a zinc finger protein attached to the C-terminus of the second domain of cytosine deaminase (ZF-right) (CN configuration); or

[0129] The N-terminus of a zinc finger protein attached to the C-terminus of the first domain of CDase (ZF-left), and the N-terminus of a zinc finger protein attached to the C-terminus of the second domain of CDase (ZF-right). (NN configuration).

[0130] Thus, in some embodiments, the adenine deaminase variant can be bound to the C-terminus of a zinc finger protein (ZF-left), the N-terminus or C-terminus of the first domain of cytosine deaminase, a zinc finger protein (ZF-right), or the N-terminus or C-terminus of the second domain of cytosine deaminase.

[0131] In some embodiments, when the cytosine deaminase is in a split form and the DNA binding protein is a TALE, the first domain of the cytosine deaminase is attached to the first TALE, the second domain of the cytosine deaminase is attached to the second TALE, and each has a structure of N'-TALE-first domain DDDA-C' and N'-TALE-second domain DDDA-C', respectively. The adenine deaminase variant can be bound to the N-terminus or C-terminus of the first domain of the cytosine deaminase, or to the N-terminus or C-terminus of the second domain of the cytosine deaminase.

[0132] In one embodiment, when the cytosine deaminase is included in full-length form and the DNA binding protein is a TALE, it can include a single TALE module, including a single TALE module and a cytosine deaminase in the NC orientation, wherein the adenine deaminase variant can be bound to the C-terminus of the single TALE module, or to the N-terminus or C-terminus of the cytosine deaminase.

[0133] In another embodiment, when the cytosine deaminase is included in full-length form and the DNA binding protein is a TALE, a dual TALE module may be included. A first TALE module and a cytosine deaminase are included in the NC direction, and a second domain may be further included, the second domain including an adenine deaminase variant and a second TALE. The adenine deaminase variant has the structure of N'-TALE-cytosine deaminase-C' and N'-TALE-adenine deaminase variant-C', which can bind to the N-terminus or C-terminus of the TALE.

[0134] In one embodiment, the fusion protein may further include UGI (uracil glycosylase inhibitor). UGI can improve the efficiency of base correction by inhibiting the activity of UDG (uracil DNA glycosylase), an enzyme that repairs mutant DNA and catalyzes the removal of U from DNA.

[0135] In another embodiment, the DNA may be nuclear DNA or organellar DNA.

[0136] In one embodiment, the fusion protein may further include an NLS (nuclear localization signal). The nuclear localization signal protein may be derived from, for example, simian virus 40 large tumor antigen (SV40 large T antigen), but is not limited thereto. The nuclear localization signal protein may contain, for example, the following amino acid sequence, but is not limited thereto:

[0137] PKKKRKV (SEQ ID NO:4)

[0138] In another embodiment, the fusion protein may further include MTS (mitochondrial targeting sequence) or CTP (chloroplast transit peptide). The mitochondrial targeting sequence protein may be, for example, SOD2-MTS or COX8A-MTS, and may contain the following amino acid sequence, but is not limited thereto:

[0139] SOD2-MTS:LSRAVCGTSRQLAPVLGYLGSRQKHSLPD(SEQ ID NO:5)

[0140] COX8A-MTS:SVLTPLLLRGLTGSARRLPVPRAKIHSL (SEQ ID NO: 6).

[0141] The chloroplast transit peptide protein may be derived from Arabidopsis RECA1, for example, but is not limited thereto. The chloroplast transit peptide protein may contain, for example, the following amino acid sequence, but is not limited thereto:

[0142] MDSQLVLSLKLNPSFTPLSPLFPFTPCSSSFSPSLRFSSCYSRRLYSPVT VYAAK(SEQ ID NO:7).

[0143] In another embodiment, the fusion protein may further include NES (nuclear export signal). The nuclear export signal protein may be derived from MVM (Mirute virus of mice), but is not limited thereto. The nuclear export signal protein may contain, for example, the following amino acid sequence, but is not limited thereto:

[0144] VDEMTKKFGTLTIHDTEK(SEQ ID NO:8)

[0145] In one embodiment, when the signal peptide is attached to the fusion protein, the structure may be as follows: signal peptide-DNA binding protein-deaminase. In another embodiment, the structure may be signal peptide-deaminase-DNA binding protein. In one embodiment, the nuclear export signal protein CTP (chloroplast transit peptide) or a polynucleotide encoding the protein may be attached to the N-terminus of the DNA binding protein, cytosine deaminase (DdCBE) or a polynucleotide encoding the protein.

[0146] In various embodiments, the fusion protein may further include a nicking enzyme, such as MutH, a MutH variant, or Nt.BspD6I (C), but is not limited thereto. MutH is a weak endonuclease that is activated once it binds to MutL. It cuts unmethylated DNA and unmethylated strands of hemimethylated DNA, but does not cut fully methylated DNA. On the other hand, the nicking endonuclease Nt.BspD6I (Nt.BspD6I) is the large subunit of the heterodimeric restriction endonuclease R.BspD6I. It recognizes a short specific DNA sequence 5"-GAGTC and cuts the upper chain only in the dsDNA at the 3"-terminal four nucleotides downstream of the recognition site. The resulting chain-specific nicking results in the generation of transient single-stranded DNA. Since TadA variants are nucleobase deaminases that specifically target single-stranded DNA, they are able to induce A to G editing near the nicking site. In another embodiment, the fusion protein further including a nicking enzyme may be in the form of a dimer, including a first fusion protein and a second protein. The first fusion protein may include a DNA binding protein and an adenine deaminase variant, while the second protein may include another DNA binding protein and a nicking enzyme.

[0147] In one embodiment, a polynucleotide encoding adenine deaminase or fusion protein is provided. The term "polynucleotide" is used interchangeably with "nucleic acid", "oligonucleotide", "nucleotide", "nucleotide sequence". It can contain a polymer form of nucleotides, deoxyribonucleotides or ribonucleotides or their analogs of any length. The polynucleotide can have any three-dimensional structure and can perform any known or unknown function. The polynucleotide can include one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. Modification of the nucleotide structure can occur before or after polymer assembly.

[0148] The nucleic acid may be an RNA sequence, in particular an mRNA sequence, a DNA sequence or a combination thereof (RNA-DNA combined sequence).

[0149] Nucleic acids can be delivered using viral vectors, such as adeno-associated viral vectors (AAV), adenoviral vectors (AdV), lentiviral vectors (LV), retroviral vectors (RV), or other viral vectors, such as episomal vectors including simian virus 40 (SV40) replication origin, bovine papilloma virus (BPV) replication origin, or Epstein-Barr nuclear antigen (EBV). In various embodiments, delivery can also be performed using non-viral vectors, or by plasmid or mRNA delivery.

[0150] The vector can be delivered in vivo or into cells by local injection (e.g., directly into the lesion or target site), electroporation, liposome transfection, viral vectors, nanoparticles, PTD (protein transduction domain) fusion protein methods, etc.

[0151] As means for expressing the above-mentioned protein, known expression vectors such as plasmid vectors, cosmid vectors and phage vectors can be used. Those skilled in the art can easily prepare such vectors using DNA recombination technology according to any known method.

[0152] Recombinant expression vectors are designed to carry nucleic acids in a form that is conducive to their expression in host cells. The nucleic acid sequence used for expression is operably connected to the recombinant expression vector, and the vector is equipped with one or more regulatory elements, which can be selected according to the specific host cell. In the recombinant expression vector, "operably connected" refers to the nucleotide sequence of interest being connected to the regulatory element in a manner that allows the nucleotide sequence to be expressed. (For example, in an in vitro transcription / translation system, or in a host cell when the vector is introduced into a host cell).

[0153] In one embodiment, a base editing composition for A to G base editing in DNA is provided, comprising a fusion protein, a polynucleotide encoding the fusion protein, or an expression vector comprising the polynucleotide. The fusion protein has been explained in detail above.

[0154] In one embodiment, the fusion protein of the base editing composition may further include cytosine deaminase. Cytosine deaminase has been explained in detail above. In some embodiments, cytosine deaminase may exist in a double-split form, and the fusion protein may include a first fusion protein and a second fusion protein, wherein the first fusion protein includes a first split body of the cytosine deaminase, and the second fusion protein includes a second split body of the cytosine deaminase. The first split body may include the amino acid sequence of SEQ ID NO: 9 or 10, and the second split body may include the amino acid sequence of SEQ ID NO: 11 or 12, but is not limited thereto. As an embodiment, for example, two TALEDs may be composed of an N-terminal DddA at G1397. tox The left or right TALE (L-1397N or R-1397N, respectively) of the hemizygote fusion and the C-terminal DddA at G1397 and TadA8e tox The hemizygote consists of the right or left TALE (R-1397C-AD or L-1397C-AD) fused together.

[0155]

[0156] In various embodiments, the base editing composition can exhibit a significant reduction in off-target effects involving unexpected base changes in DNA and / or RNA. In another embodiment, the base editing composition can reduce unwanted bystander effects while narrowing the activity window. Alternatively, the base editing composition can induce base editing only at a single nucleotide residue without any intentional off-target editing in the target DNA, with a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0157] In yet another embodiment, a method for A to G base editing in DNA is provided, comprising delivering a base editing composition to a cell containing the target DNA.

[0158] The cell can be a eukaryotic cell (e.g., a fungus such as yeast, a cell of eukaryotic animal and / or eukaryotic plant origin (e.g., embryonic cell, stem cell, somatic cell, gamete, etc.), a eukaryotic animal (e.g., human, monkey, primate, dog, pig, cow, sheep, goat, mouse, rat, etc.) or a eukaryotic plant (e.g., algae such as green algae, corn, soybean, wheat, rice, etc.), but is not limited thereto.

[0159] Delivery of base editing compositions to cells containing target DNA can be performed ex vivo or in vivo.

[0160] In some embodiments, the target DNA can be nuclear DNA, organellar DNA, or mitochondrial DNA of a human subject suffering from a genetic disease.

[0161] As used herein, the term "genetic disease" refers to a pathological condition that occurs due to a deleterious mutation in a gene or chromosome. Examples of genetic diseases include, but are not limited to, mitochondrial encephalopathy, lactic acidosis and stroke-like episodes (MELAS) syndrome, DEAF, Leber hereditary optic neuropathy (LHON), Leigh syndrome, myopathy, and chronic progressive external ophthalmoplegia (CPEO).

[0162] The method can exhibit reduced off-target effects compared to the case where a base editor comprising an adenine deaminase having an amino acid sequence of SEQ ID NO: 1 is used, wherein off-target editing is characterized by unexpected base changes in DNA and / or RNA. In another embodiment, the method can reduce unwanted bystander effects while narrowing the activity window compared to the case where a base editor comprising an adenine deaminase having an amino acid represented by a mutation in SEQ ID: 1 is used. Alternatively, the method can induce base editing only at a single nucleotide residue without any intentional off-target editing occurring in the target DNA, with a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0163] In another embodiment, the method can exhibit reduced off-target effects compared to a base editor using an adenine deaminase comprising an amino acid represented by SEQ ID: 1 with a V106W mutation, wherein the off-target editing is characterized by an unexpected base change in DNA and / or RNA. In another embodiment, the method can reduce unwanted bystander effects while narrowing the activity window compared to a base editor using an adenine deaminase comprising an amino acid represented by SEQ ID: 1 with a V106W mutation. Alternatively, the method can induce base editing only at a single nucleotide residue without any intentional off-target editing occurring in the target DNA, with a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0164] In one embodiment, a method for reducing off-target effects and / or unwanted bystander effects while narrowing the activity window is provided, the method comprising delivering a base editing composition to a cell containing target DNA.

[0165] The cell can be a eukaryotic cell (e.g., a fungus such as yeast, a cell of eukaryotic animal and / or eukaryotic plant origin (e.g., embryonic cell, stem cell, somatic cell, gamete, etc.), a eukaryotic animal (e.g., human, monkey, primate, dog, pig, cow, sheep, goat, mouse, rat, etc.) or a eukaryotic plant (e.g., algae such as green algae, corn, soybean, wheat, rice, etc.), but is not limited thereto.

[0166] In some embodiments, the target DNA can be nuclear DNA, organelle DNA, or mitochondrial DNA of a human subject suffering from a genetic disease. Genetic diseases have been defined above. Compared with the case of using a base editor comprising an adenine deaminase having an amino acid sequence of SEQ ID NO: 1, the method can exhibit reduced off-target effects, wherein off-target editing is characterized by unexpected base changes in DNA and / or RNA. In another embodiment, compared with the use of a base editor comprising an adenine deaminase having an amino acid represented by SEQ ID: 1, the method can reduce unwanted bystander effects while narrowing the activity window. Alternatively, the method can induce base editing only at a single nucleotide residue without any intentional off-target editing in the target DNA, with a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0167] In another embodiment, the method can exhibit reduced off-target effects compared to a base editor using an adenine deaminase comprising an amino acid represented by SEQ ID: 1 with a V106W mutation, wherein the off-target editing is characterized by an unexpected base change in DNA and / or RNA. In another embodiment, the method can reduce unwanted bystander effects while narrowing the activity window compared to a base editor using an adenine deaminase comprising an amino acid represented by SEQ ID: 1 with a V106W mutation. Alternatively, the method can induce base editing only at a single nucleotide residue without any unexpected off-target editing in the target DNA, with a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0168] However, in another embodiment, adenine deaminase variants or fusion proteins are provided for use in A to G base editing in DNA and / or reducing off-target effects of A to G base editing in DNA, or for use in preparing a composition for A to G base editing in DNA and / or reducing off-target effects of A to G base editing in DNA. Adenine deaminase variants and fusion proteins have been explained in detail above.

[0169] In one embodiment, the fusion protein may further include cytosine deaminase. Cytosine deaminase has been explained in detail above. In some embodiments, the cytosine deaminase may exist in a double-split form, and the fusion protein includes a first fusion protein and a second fusion protein, the first fusion protein including a first split body of the cytosine deaminase, and the second fusion protein including a second split body of the cytosine deaminase. The first split body may include the amino acid sequence of SEQ ID NO: 9 or 10 and the second split body may include the amino acid sequence of SEQ ID NO: 11 or 12, but is not limited thereto.

[0170] In various embodiments, compositions for A to G base editing in DNA and / or reducing off-target effects in A to G base editing in DNA can exhibit significantly reduced off-target effects involving unexpected base changes in DNA and / or RNA. In one embodiment, the composition can reduce unwanted bystander effects while narrowing the activity window. Alternatively, the composition can induce base editing only at a single nucleotide residue without any unexpected off-target editing in the target DNA, with a frequency of at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, or at least 15%.

[0171] The present invention will be described in more detail below in conjunction with embodiments. It will be appreciated by those skilled in the art that these embodiments are only used to illustrate the present invention, and the scope of the present invention should not be construed as being limited by these embodiments.

[0172] Reference Examples

[0173] 1. Preparation of Cell Lines

[0174] HEK 293T cells were purchased from the American Type Culture Collection (ATCC) (CRL-11268). NIH3T3 and B16F10 cells were purchased from ATCC (CRL-1658, CRL-6475). HEK 293T cells were cultured in Dulbecco's modified Eagle medium (DMEM; Welgene) supplemented with 10% (v / v) fetal bovine serum (Welgene) and 1% (v / v) antibiotic antimycotic solution (Welgene). NIH3T3 and B16F10 cells were cultured in DMEM, which was supplemented with 10% (v / v) calf serum (Gibco) for NIH3T3 cells or 10% fetal bovine serum (Gibco) for B16F10 cells without any antibiotics. The cells were incubated at 37°C with 5% CO2. All cell lines were passaged before reaching 90% confluence.

[0175] 2. PyMOL Analysis

[0176] The ABE8e (PDB accession number 6VPC) structure was downloaded from the PDB and visualized with PyMOL v.2.5.4. Some elements including Cas9, single guide RNA, and double-stranded DNA were excluded from the PDB file; 8AZ (a transition state analog for adenosine deamination) and TadA monomer were retained. Eleven residues close to 8AZ (V28, V30, N46, A48, I49, V82, F84, V106, N108, K110, and R111) were selected, including residues previously known to contact DNA.

[0177] 3. Plasmid Construction

[0178] To construct plasmids encoding DdCBE, DddA-split TALED (sTALED), and site-specific mTALED, plasmids containing a stuffer fragment, a sequence between two restriction enzyme sites that facilitates separation of the fragments during gel electrophoresis, were prepared. The plasmid structure is as follows: sTALED plasmid (p3s-stuffer-DddA tox Half (1397C)-AD and p3s-filler fragment-DddA tox half (1397N)); mTALED plasmid (p3s-stuffer fragment-E1347A DddA toxAll-AD). These plasmids were obtained from Addgene (DdCBE, #187168, #187171, #187173, #178174; sTALED, #187167, #187169, #187170, #187172; mTALED, #187163, #187166). Plasmids encoding DdCBE, sTALED, and mTALED targeting specific sites were constructed by inserting custom-designed TALE array sequences, as shown in Table 1 below. To remove the stuffer sequence, each plasmid digested with BsaI or BsmbI (NEB) and the insert fragment (including the custom-designed TALE array sequence synthesized by IDT) was inserted into the digested vector using the HiFi DNA Assembly Kit (NEB). Alternatively, the desired TALED construct was generated by digesting the master vector with BsaI or BsmbI to cut a site in the stuffer fragment and then assembling a six TALE array at that location using the Golden Gate method.

[0179] [Table 1]

[0180]

[0181]

[0182]

[0183] 4. Cell Culture and Transfection

[0184] HEK 293T cells were grown in DMEM (Welgene) with 10% fetal bovine serum (Welgene) and 1% antibiotic antimycotic solution (Welgene). NIH3T3 (CRL-1658, ATCC) and B16F10 (CRL-6475, ATCC) cells were grown in DMEM, and the medium was supplemented with 10% (v / v) calf serum (Gibco) for NIH3T3 cells or 10% fetal bovine serum (Gibco) for B16F10 cells without any antibiotics. Cell lines were maintained at 37°C in 5% CO2 and passaged before reaching 90% confluence, depending on the doubling period of the specific cell line.

[0185] For HEK 293T cells, cells were plated at 7.5 × 10 cells per well before transfection. 4The cells were seeded into 48-well plates (Corning) at a density of 1×10 cells / well. After 24 h, the cells were transfected with plasmids (1 ug total) using 1.5 uL of Lipofectamine 2000 (Invitrogen). For sTALED pairs, DdCBE pairs, and CRISPR-base editors with sgRNA, the total amount of plasmid was 1 ug (500 ng each). When a single construct was used, the amount of transfected plasmid was 500 ng. After 96 h, the transfected cells were harvested. For NIH3T3 and B16F10 cells, cells were plated at 1×10 per well 18-24 h before transfection. 5 Cells were seeded at a density of 100 cells / mL into 24-well cell culture plates (SPL, Seoul, South Korea). Lipofectamine 3000 (Invitrogen) was used for lipofectamine transfection with 500 ng of each sTALED encoding plasmid to make up 1000 ng of total plasmid DNA. For mTALED, 500 ng of plasmid was used. Cells were harvested 3 days after transfection.

[0186] 5. Transcriptome Sequencing

[0187] 48h or 96h after transfection, total RNA was isolated using NucleoSpin RNA kit (MN#Macherey-Nagel) according to the manufacturer's instructions. RNA libraries were prepared using TruSeq Stranded Total RNA Library Prep Gold kit (Illumina). RNA library quality was assessed using 2200TapeStation (Agilent) with D1000 ScreenTape system. Total RNA sequencing was performed using Macrogen's NovaSeq 6000 sequencer (Illumina) and double-end sequencing system (2x100bp).

[0188] 6. RNA variant detection

[0189] To analyze NGS data from RNA sequencing, the RNA variant detection process previously used for RNA off-target analysis of CRISPR DNA base editors was referenced. In short, the fastq sequencing reads were aligned to the hg38 human reference genome (GRCh38, version v105) using the STAR aligner (v.2.7.10a). The resulting BAM files were processed using GATK (v.4.2.4.1) MarkDuplicates, BaseRecalibrator, and ApplyBQSR. RNA base editing variants were detected using GATK HaplotypeCaller. RNA variant loci were compared with those of the control sample and filtered based on the following criteria: (1) retain loci with a read depth of at least 10; (2) retain loci with a variant number of at least 2; (3) remove loci that are also present in the control sample; and (4) exclude loci that could not be determined in the control sample due to insufficient sequencing depth. For the replicate 1 experimental group, untreated replicate 2 was used as a filtering control, and for the replicate 2 experimental group, untreated replicate 1 was used as a control. The A to G editing count is the number of RNA variant loci with A to G editing on the plus strand or T to C editing on the minus strand. The C to T editing count is the number of RNA variant loci with C to T editing on the plus strand or G to A editing on the minus strand.

[0190] 7. Cell Lysis for Genomic DNA Analysis

[0191] After removing the growth medium, HEK 293T cells were treated with 100 μL of cell lysis buffer (50 mM Tris-HCl; pH 8.0, 1 mM EDTA, 0.005% sodium dodecyl sulfate) supplemented with 5 μL of proteinase K (Qiagen). Cells were lysed by incubating at 55°C for 1 h and then at 95°C for 10 min. The genomic DNA mixture was subjected to targeted deep sequencing.

[0192] 8. Targeted Deep Sequencing

[0193] Create NGS libraries for targeted deep sequencing using nested PCR. GXL polymerase (Takara) amplified the target region by PCR. The amplicon was amplified again by PCR with primers containing TruSeq DNA-RNACD index to label each fragment with the adapter and index sequence to construct the NGS library. PCR primers are listed in Table 2 below. The final PCR product was purified with a PCR purification kit (MGmed) and sequenced using a MiniSeq sequencer (Illumina). The frequency of base editing from targeted deep sequencing data was measured with the source code (https: / / github.com / ibs-cge / maund).

[0194] [Table 2]

[0195]

[0196] 9. Purify RNA from cultured mammalian cells and perform targeted RNA sequencing

[0197] RNA was extracted from cultured cells using the NucleoSpin RNAPlus kit (Macherey-Nagel) according to the manufacturer's instructions. RNA was then reverse transcribed to create cDNA using the SuperScript IV reverse transcriptase kit (Thermo Fisher), also following the manufacturer's instructions. The region of interest was then amplified by PCR using the primers shown in Table 3 below. The amplified region was sequenced using the above-mentioned targeted deep sequencing program.

[0198] [Table 3]

[0199]

[0200] 10. Relative ratio of DNA on-target editing frequency to RNA off-target editing frequency

[0201] To analyze the DNA on-target editing frequency of the variants versus the RNA off-target editing frequency at six representative sites, the DNA on-target activity and RNA off-target activity of the variants were normalized to the sTALED value. The normalized DNA on-target value was then divided by the normalized RNA off-target value as shown below.

[0202] Relative ratio (DNA on-target / RNA off-target)

[0203]

[0204] For wild-type sTALED, this value is 1, since both the normalized DNA on-target and RNA off-target values ​​are 1. A higher relative ratio indicates lower off-target activity compared to on-target activity.

[0205] 11. Whole mitochondrial genome sequencing

[0206] For whole mitochondrial genome sequencing, three procedures are required: PCR amplification, NGS library creation, and NGS. First, after removing the growth medium, the cells were treated with 100 μL of cell lysis buffer (50 mM Tris-HCl; pH 8.0, 1 mM EDTA, 0.005% sodium dodecyl sulfate) supplemented with 5 μL of proteinase K (Qiagen). The cells were incubated at 55°C for 1 h and then at 95°C for 10 min to lyse the cells. Then, the cells were lysed using Mitochondrial DNA was amplified by PCR using GXL polymerase (Takara). PCR was performed using two sets of slightly overlapping primers shown in Table 2 above to reduce primer bias. Each primer pair amplified nearly 50% of the mitochondrial DNA. The PCR products were then purified using a PCR purification kit (MGmed). Finally, an NGS library was created from the purified PCR products using an Illumina DNAPrep kit (Illumina) with Nextera DNACD Indexes. The libraries were then pooled and transferred to a MiniSeq sequencer (Illumina).

[0207] 12. Analysis of off-target editing in the whole mitochondrial genome

[0208] To analyze NGS data from whole mitochondrial genome sequencing, a method for studying off-target effects in the mitochondrial genome was used. First, Fastq sequences were aligned to the GRCh38 (version v102) reference genome using BWA (v.0.7.17), and then SAMtools (v.1.9) was used to create BAM files by repairing read pairing information and tags. Subsequently, the REDItoolDenovo.py script from REDItools (v.1.2.1) was used to find all thymines and adenines with conversion rates > 0.1% in the mitochondrial genome. Positions with conversion rates ≥ 10% in treated and untreated samples were identified as SNVs in the cell line and removed. The on-target sites of the construct were excluded. The remaining sites were considered off-target sites, and we calculated the number of edited A / T nucleotides with an editing frequency > 0.1%. By averaging the conversion rate at each base position in the off-target site, the average A / T to G / C editing frequency was calculated for all bases in the mitochondrial genome, as shown below.

[0209]

[0210] A graph of the entire mitochondrial genome was constructed by plotting the conversion rates at on-target and off-target sites with editing frequencies ≥1% across the entire mitochondrial genome.

[0211] 13. Cell Viability Assay

[0212] Use CellTiter on day 2 or 4 after plasmid transfection Cell viability was determined using Aqueous One Solution (Promega). The number of live cells was measured using a colorimetric method using the MTS assay. The cells were treated with Aqueous One Solution, and the quantification of bioreduction products was measured by recording the absorbance at 490 nm according to the manufacturer's instructions.

[0213] Example 1: Mitochondrial DNA-targeted TALED induces transcriptome-wide off-target editing

[0214] To investigate whether DddA-split TALEDs (sTALEDs) targeting COX3.1 or ND1 sites would lead to unwanted off-target RNA editing in human embryonic kidney 293T (HEK 293T) cells, whole transcriptome sequencing was performed on total RNA isolated from cells on day 2 after transfection with reference to the above-mentioned reference example. Whole transcriptome sequencing of two independent biological replicates showed that two sTALEDs (consisting of N-terminal DddA at G1397) tox The left or right TALE (L-1397N or R-1397N, respectively) of the hemizygote fusion and the C-terminal DddA at G1397 and TadA8e tox Hemizygote fusions composed of either the right or left TALE (R-1397C-AD or L-1397C-AD) induced base transitions at >50,000 sites with a frequency of at least 7% (Figure 1). The vast majority (>99.8%) of these base transitions were A to G edits (in cDNA reverse transcribed from RNA), rather than C to T edits, indicating that adenine deaminases rather than cytosine deaminases were responsible for these single nucleotide changes. In contrast, DdCBE targeting the same sites or a single sTALED subunit (L-1397N or R-1397N) (used as negative controls) did not induce these off-target RNA edits. These results suggest that the other subunit containing TadA8e (L-1397C-AD or R-1397C-AD) is responsible for the observed A to G transitions, which are caused by A to I changes in the RNA.

[0215] To avoid or minimize unwanted whole-transcriptome off-target A to I conversions induced by TALED, site-specific mutations were introduced into TadA8e, including V106W, V106G, K20A / R21A (double mutation), or F148A, which are known to reduce off-target RNA editing when incorporated into CRISPR RNA-guided adenine base editors (ABEs). Whole-transcriptome sequencing showed that sTALED variants incorporating these mutations into TadA8e significantly (but not completely) reduced the number of off-target A to G editing while retaining DNA on-target editing efficiency (Figure 1). Thus, Cox3.1-specific sTALED variants with these site-specific mutations induced RNA off-target editing at sites ranging in number from 44,627 (F148A) to 107,304 (K20A / R21A), which was reduced by 18.7% (= (132,064-107,304) / 132,064) to 66.2% (= (132,064-44,627) / 132,064) compared to the original sTALED (132,064 A to G edits) ( Figures 1( b ) and 1( c )). Similarly, ND1-specific sTALED variants reduced the number of RNA off-target edits by at least 46.2% (=(61,189-32,910) / 61,189)(K20A / R21A) and up to 84.2% (=(61,189–9,674) / 61,189)(V106G) (Figures 1(d) and 1(e)). These results suggest that the same TadA8e variants used to reduce RNA off-target editing effects of CRISPR RNA-guided ABEs can also reduce collateral damage caused by sTALEDs. However, even with the best-performing TadA8e variants, at least 9,600 and up to 107,000 off-target A to G edits remained.

[0216] Example 2: Protein engineering of TALED to avoid off-target RNA editing

[0217] To further minimize transcriptome-wide off-target editing, TALEDs were engineered by mutating amino acid residues at the substrate binding site in TadA8e, including V106. To this end, based on the 3D cryo-electron microscopy structure of ABE8e bound to DNA ( Figure 2(a)), 11 amino acid residues were selected at the substrate binding site, and each of these residues was replaced with the other 19 amino acid residues in the Cox3.1-specific sTALED, resulting in a total of 209 (=11×19) sTALED variants. As described in Reference Example 3, the plasmid encoding each of the resulting sTALED pairs was transfected into HEK 293T cells. Subsequently, the on-target editing frequency at the Cox3.1 site was measured on day 4 after transfection. Among the 209 sTALED variants, a total of 101 sTALED pairs remained highly active in targeted mitochondrial DNA editing with an efficiency of at least 50% ( ) compared to the wild-type sTALED pair. Figure 3 ). Targeted RNA amplicon sequencing was then used to measure the off-target editing frequencies of 101 sTALEDs at six representative sites revealed by whole transcriptome sequencing analysis (Figure 4(a)). These representative sites were mutated at high frequencies ranging from 60% to 80% by both Cox3.1-specific and ND1-specific TALEDs ( Figure 2 (b)). The frequency of RNA off-target editing measured by targeted deep sequencing was highly consistent with the frequency estimated by whole transcriptome sequencing (Figure 4(b) and Figure 4(c)). A total of 12 TadA8e variants were selected that minimized RNA off-target editing efficiency at six representative sites while retaining mtDNA on-target editing efficiency ( Figure 2 (b)).

[0218] Whole transcriptome sequencing was performed to investigate whether the 12 sTALED variants could avoid off-target editing at sites other than the six representative sites (Figure 5). All of these variants significantly reduced the number of off-target editing. For example, sTALED variants with V28R and R111S induced off-target editing at only 852 and 829 sites, respectively, while the original sTALED and sTALED-V106W (sTALED with the V106W mutation in TadA8e) induced off-target editing at 96,559 and 81,156 sites, respectively (Figures 5(a) and 5(b)). Therefore, sTALED-V28R and -R111S avoided >99% of RNA off-target editing. Considering that 316 A to G edits were found in the untreated DNA sample used as a negative control, which may have been caused by high-throughput sequencing errors, these results indicate that these sTALED variants almost completely avoid RNA off-target editing.

[0219] In addition, we further investigated whether these TadA8e mutations could reduce RNA off-target editing when integrated into sTALEDs targeting other mitochondrial DNA target sites (Figures 5C-F and 6). Targeted RNA amplicon sequencing showed that most ND1-specific and ND6-specific sTALED variants significantly reduced the frequency of RNA off-target editing at six representative sites (Figures 6(c)-(f)). The frequency of mitochondrial DNA on-target editing was also measured via targeted deep sequencing (Figures 6(a) and 6(b)), and then the ratio of DNA on-target editing frequency to RNA off-target editing frequency was obtained (Figures 6(g) and 6(h)). Based on these results, the following four TadA8e variants, V28Q, V28R, A48W, and R111S, were selected, which exhibited higher mitochondrial DNA on-target activity and lower RNA off-target activity compared to the original sTALEDs and sTALED-V106W targeting Cox3.1, ND1, and ND6 sites (Figure 6(i)). Whole transcriptome sequencing showed that ND1-specific and ND6-specific sTALED variants with each of the four variants significantly reduced the number of RNA off-target editing (Figure 5), which is consistent with the results above with Cox3.1-specific sTALED variants. Similarly, ND1-specific and ND6-specific sTALED-V28R or -R111S variants were the most discriminative and induced the least RNA off-target editing (Figure 5(c)-(f)).

[0220] Example 3: Engineered TALEDs to reduce bystander and off-target editing

[0221] sTALED variants with site-specific mutations at the TadA8e substrate binding site may also reduce bystander editing at the target site and off-target editing in the mitochondrial genome, as these mutations may potentially fine-tune adenine deaminase activity on DNA substrates in addition to reducing activity on RNA substrates. That is, the frequency of base editing at each nucleotide position was examined. Both sTALED-V28R and -R111S variants specific for the Cox3.1, ND1, and ND6 sites induced A to I editing in a narrower window compared to the wild-type sTALED and sTALED-V106W variants (Figure 7). For example, wild-type sTALED and sTALED-v106W targeting the Cox3.1 site induced A to I editing at multiple positions with a frequency of >1.1%, not only in the spacer region between the two TALE binding sites, but also in the TALE binding site. In contrast, the Cox3.1-specific sTALED-V28R induced A to I editing at a single position in the spacer region ( Figure 7a ).

[0222] Next, the frequencies of edited alleles induced by these sTALEDs targeting three mitochondrial DNA sites were compared. Most mutant alleles induced by the original sTALED or the sTALED-V106W variant contained multi-base edits, rather than single-base edits. In sharp contrast, the frequencies of mutant alleles induced by single-base substitutions by sTALED-V28R and -R111S were much higher than those by multi-base substitutions. Thus, the original sTALED induced the most abundant allele with a single A to I edit in the middle of the intergenic region at low frequencies of 3.06% (Cox3.1), 1.90% (ND1), and 1.70% (ND6), while the sTALED-V28R variant induced the same allele at high frequencies of 10.1% (Cox3.1), 17.0% (ND1), and 10.5% (ND6) (Figure 7 (d)-(f)). This result has important implications for the use of TALED in disease modeling as well as therapeutic applications. TALEDs that induce single base substitutions with no or few bystander editing are desirable because the vast majority of pathogenic mtDNA mutations that cause mitochondrial genetic disorders are single nucleotide variants rather than multi-nucleotide variants.

[0223] In addition, whole mitochondrial genome sequencing was performed on day 4 after transfection to evaluate and compare the off-target effects of wild-type sTALED and sTALED variants (Figure 8). Wild-type sTALED targeting three mitochondrial DNA sites induced off-target mutations, and the average frequency of off-target editing of the whole mitochondrial genome ranged from 0.0024% to 0.0066%, which was 4 to 10 times higher than the frequency observed in the untreated control (0.0006%). In contrast, the average frequency of off-target editing induced by sTALED-V28R and R111S was <0.0006%, similar to the baseline frequency observed in the negative control (Figure 8B). In addition, the sTALED variants also reduced the number of off-target edits induced in the mitochondrial genome. Thus, the ND1-specific wild-type sTALED caused off-target editing of A to I at 108 sites in human mitochondrial DNA with a frequency of >0.1%, while sTALED-V28R and -R111S induced off-target editing at 14 and 17 sites, respectively, similar to the baseline number seen in untreated samples (i.e., 24) (Figure 8(a)). Whole mitochondrial genome sequencing was also performed on day 2 after transfection (Figure 9), when the level of RNA off-target editing induction was highest. The average frequency of off-target editing induced by the original sTALED targeting these three sites ranged from 0.0071% to 0.0090%, which is 11 to 14 times higher than the frequency observed in the untreated control (0.0006%). However, compared with the negative control, the V28R and R111S variants did not induce off-target editing (Figure 9(b)).

[0224] Example 4: Time course measurement of DNA on-target and RNA off-target editing frequencies

[0225] In Figure 10, it was studied whether the on-target mutations induced by various forms of sTALED were stably maintained and how long the RNA off-target mutations lasted over time. Using targeted deep sequencing, the DNA on-target editing frequencies at three mitochondrial DNA sites (Figure 10 (a)-(c)) and the RNA off-target editing frequencies at six representative sites (Figure 10 (d)-(g)) were measured at different time points, and sTALEDs targeting these three sites were co-edited. On the 1st and 2nd day after transfection, RNA editing was induced in large quantities by sTALED and sTALED-V106W variants, but almost completely disappeared on the 8th day after transfection. Compared with sTALED or sTALED-V106W variants, the new sTALED variants reduced the RNA off-target editing frequency by several times, even on the 1st and 2nd day after transfection (Figure 10 (g)). The mitochondrial DNA on-target editing induced by sTALED and sTALED-V106W variants is not stably maintained. Thus, the frequency of on-target editing of mitochondrial DNA induced by sTALED and sTALED-V106W variants decreased significantly over time. For example, on-target editing of mitochondrial DNA induced by ND6-specific sTALED and sTALED-V106W was performed at high frequencies of up to 47% and 40% on days 1 and 2 after transfection, respectively, but had a low frequency of <10% on day 8 or thereafter (Figure 10(c)). In contrast, the frequency of on-target editing induced by sTALED-V28R and -R111S increased or was maintained more stably over time. As a result, the editing frequencies observed for our new sTALED variants were lower on days 1 and 2 after transfection, but were higher than those induced by previous versions of sTALED on day 8 or thereafter (Figures 10(a)-(c)).

[0226] Based on these results, it can be concluded that sTALED and sTALED-V106W variants are cytotoxic because they induce excessive RNA off-target editing at a high frequency, and mtDNA edited cells cannot divide or die over time. However, the sTALED variants of the present invention can be tolerated, probably because they avoid RNA off-target editing or do not induce excessive bystander editing and off-target mutations in the mitochondrial genome at the target site. In addition, an MTS cell proliferation assay (Figure 11 (a)) was performed to confirm that sTALED and sTALED-V106W variants are indeed cytotoxic, significantly reducing cell viability compared to the negative control (pEGFP transfection) (Figure 11 (b)). The sTALED variants are much better tolerated, so that cell viability is not reduced on the 4th day after transfection (Figure 11 (c)).

[0227] Example 5: Integration of TadA8e-V28R and -R111S into CRISPR RNA-guided ABE

[0228] In Figure 12, it was investigated whether the V28R and R111S mutations in TadA8e can also reduce bystander editing and RNA off-target editing induced by CRISPR RNA-guided ABEs widely used for nuclear DNA editing. First, the on-target and bystander editing frequencies at the TYRO3 site were measured. Both ABE8e-V28R and -R111S were as efficient at the target site as ABE8e and ABE8e-V106W (ABE8eW), with editing frequencies of >30% (Figures 12 (b) and 12 (c)). The V28R and R111S variants exhibited a narrower editing window with editing efficiency (up to 37%) at positions 5 and 7 (A5 and A7 in Figure 12 (c)) in the original spacer region, but there was almost no bystander editing at position 10 (A10) (0.8% and 0.5%, respectively). In contrast, ABE8e and ABE8eW show a wider editing window, with highest editing at A5 (29% and 30%, respectively) and A7 (29% and 30%, respectively), and a large number of bystander edits at A10 (9.3% and 3.1%, respectively).

[0229] Next, target RNA amplicon sequencing was used to assess RNA off-target editing activity. Compared with ABE8e, ABE8e-V28R and -R111S reduced the average RNA off-target editing frequency measured for a total of 6 sites by 3.8-fold and 2.5-fold, respectively, and by 3.1-fold and 2.0-fold compared with ABE8eW (Figure 12(d)). In addition, whole transcriptome sequencing showed that ABE8e-V28R and -R111S significantly reduced the number of RNA off-target editing compared with ABE8e and ABE8eW (Figure 12(e) and Figure 12(f)). These results suggest that the two new TadA8e variants, V28R and R111S, can also improve ABE8e by reducing RNA off-target editing and bystander editing.

[0230] Example 6: On-target DNA editing and off-target RNA editing by mTALED and sTALED

[0231] exist Fig.13 In addition to sTALED, we investigated whether dimeric TALED (dTALED) and monomeric TALED (mTALED) variants with V28R and R111S could reduce RNA off-target editing. In dTALED, each TALE unit contains a cytosine deaminase (e.g., DddA toxvariant (E1347A)) contains an adenine deaminase (eg TadA8e) on the other side. mTALEDs, on the other hand, are characterized in that both cytosine deaminase and adenine deaminase are present in a single TALE unit.

[0232] exist Fig.13 In (a), it was confirmed that both dTALED and mTALED targeting the Cox3 locus exhibited on-target activity. In addition, it was observed that no gene editing occurred when only adenine deaminase was present in the TALE unit. Fig.13 (b) Editing efficiency results at six representative sites. Similar to sTALED, both dTALED and mTALED variants with V28R and R111S had significantly reduced RNA off-target efficiency compared to TadA8e.

[0233] exist Fig.13 In (c), the specific ratio of RNA off-target editing induced by Cox3-specific TALED relative to on-target editing was analyzed. The results showed that TadA variants with V28R and R111S not only reduced the RNA off-target efficiency in sTALED and other systems (such as ABE), but also reduced the RNA off-target efficiency in dTALED and mTALED. These results indicate that the RNA off-target effect is mainly caused by TadA adenine deaminase, rather than by DNA binding proteins or cytosine deaminase. In addition, TadA variants can significantly reduce this undesirable RNA off-target effect.

[0234] From the above description, it will be appreciated by those skilled in the art that the present invention can be embodied in other specific forms without departing from its spirit or essential features. In this respect, it should be understood that the above-mentioned embodiments should be regarded as illustrative rather than restrictive in all respects. The scope of the present invention should be interpreted as being included within the scope of the present invention without departing from the scope of the present invention defined by the appended claims.

[0235] Sequence Listing Free Text

[0236] Attached to the electronic file.

Claims

1. An adenine deaminase comprising the amino acid sequence of SEQ ID NO: 1 or an amino acid sequence having at least 80% sequence homology with the amino acid sequence of SEQ ID NO: 1, wherein: Selected from SEQ At least one of residues 28, 30, 46, 48, 49, 82, 84, 106, 108, 110 and 111 of SEQ ID NO: 1, or the corresponding amino acid residues of an amino acid sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO: 1, is substituted with another amino acid.

2. The adenine deaminase according to claim 1, wherein The substitution is at least one selected from the group consisting of: V28Q or V28R of the amino acid sequence of SEQ ID NO: 1, A48W, F84M, V106A, K110S, K110T, or K110V, and R111F, R111Q, R111S, R111T or R111Y.

3. The adenine deaminase according to claim 2, wherein The substitution is at least one selected from the group consisting of: V28Q or V28R of the amino acid sequence of SEQ ID NO: 1, A48W, and R111S.

4. A fusion protein comprising: DNA binding proteins; as well as The modified adenine deaminase according to claim 1.

5. The fusion protein according to claim 4, wherein The substitution in the modified adenine deaminase is at least one selected from the group consisting of: V28Q or V28R of the amino acid sequence of SEQ ID NO: 1, A48W, F84M, V106A, K110S, K110T, or K110V, and R111F, R111Q, R111S, R111T or R111Y.

6. The fusion protein according to claim 5, wherein The substitution in the modified adenine deaminase is at least one selected from the group consisting of: V28Q or V28R of the amino acid sequence of SEQ ID NO: 1, A48W, and R111S.

7. The fusion protein according to claim 4, wherein The DNA binding protein is a zinc finger protein, a TALE (transcription activator-like effector) or a CRISPR-associated nuclease.

8. The fusion protein according to claim 7, wherein The DNA binding protein is TALE.

9. The fusion protein according to claim 8, wherein The TALE is a dual TALE module consisting of a first TALE module and a second TALE module.

10. The fusion protein according to claim 4, wherein The DNA binding protein is a Cas nickase or a Cas protein lacking endonuclease activity.

11. The fusion protein according to claim 4, wherein The fusion protein further comprises cytosine deaminase.

12. The fusion protein according to claim 11, wherein The cytosine deaminase is in full-length form or in a bipartite form.

13. The fusion protein according to claim 11 or 12, wherein The cytosine deaminase is APOBEC (apolipoprotein B editing complex), AID (activation-induced deaminase), DddAtox (cytosine deaminase specific for double-stranded DNA) or a variant thereof.

14. The fusion protein according to any one of claims 4 to 12, wherein The fusion protein further includes UGI (uracil glycosylase inhibitor).

15. The fusion protein according to any one of claims 4 to 12, wherein The DNA is nuclear DNA or organellar DNA.

16. The fusion protein according to claim 15, wherein The fusion protein further includes an NLS (nuclear localization signal).

17. The fusion protein according to claim 15, wherein The fusion protein further includes MTS (mitochondrial targeting sequence) or CTP (chloroplast transit peptide).

18. The fusion protein according to claim 15, wherein The fusion protein further comprises a NES (nuclear export signal).

19. The fusion protein according to claim 4, wherein The fusion protein further comprises a nickase.

20. The fusion protein according to claim 19, wherein The fusion protein comprises a first fusion protein and a second fusion protein, wherein the first fusion protein comprises the DNA binding protein and the modified adenine deaminase, and the second fusion protein comprises the DNA binding protein and the nickase.

21. The fusion protein according to claim 4, wherein The nicking enzyme is MutH, a MutH variant or Nt.BspD6I (C).

22. A polynucleotide encoding the modified adenine deaminase according to any one of claims 1 to 3, or the fusion protein according to any one of claims 4-12 and 19-21.

23. An expression vector comprising the polynucleotide according to claim 22.

24. A base editing composition for A to G base editing in DNA, comprising a fusion protein according to any one of claims 4-10 and 19-21, a polynucleotide encoding the fusion protein, or an expression vector comprising the polynucleotide.

25. The base editing composition according to claim 24, wherein The fusion protein further comprises cytosine deaminase.

26. The base editing composition according to claim 25, wherein The cytosine deaminase exists in a double-split form, and the fusion protein includes a first fusion protein and a second fusion protein, wherein the first fusion protein includes a first split body of the cytosine deaminase and the second fusion protein includes a second split body of the cytosine deaminase.

27. The base editing composition according to claim 26, wherein The first split body includes the amino acid sequence of SEQ ID NO:9 or 11, and the second split body includes the amino acid sequence of SEQ ID NO:10 or 12.

28. The base editing composition of claim 24, wherein The compositions exhibit reduced off-target effects involving unintended base changes in DNA and / or RNA.

29. A method for A to G base editing in DNA, comprising delivering the base editing composition of claim 24 to a cell containing the target DNA.

30. The method of claim 29, wherein: The delivery is performed ex vivo.

31. The method of claim 29, wherein: The delivery is performed in vivo.

32. The method of claim 29, wherein: The target DNA is nuclear DNA, organellar DNA or mitochondrial DNA of a human subject suffering from a genetic disease.

33. The method of claim 29, wherein: The method induces base editing at a single nucleotide residue in the target DNA at a frequency of at least 10%.

34. The method of claim 29, wherein: The method exhibits reduced off-target editing effects compared to using a base editor comprising an adenine deaminase comprising the amino acid sequence of SEQ ID NO: 1, wherein the off-target editing effects are characterized by unintended base changes in DNA or RNA.

35. The method of claim 29, wherein: The method exhibits reduced off-target editing effects compared to using a base editor comprising an adenine deaminase comprising an amino acid sequence of SEQ IDNO:1 having a V106W mutation, wherein the off-target editing effects are characterized by unintended base changes in DNA or RNA.

36. The method of claim 29, wherein: The method induces base editing at a single nucleotide residue in the target gene with a frequency of at least 10%.

37. A method for reducing off-target effects, the method comprising delivering the base editing composition of claim 24 to a cell containing a target DNA, wherein: The method exhibits reduced off-target editing effects compared to using a base editor comprising an adenine deaminase comprising the amino acid sequence of SEQ ID NO: 1, wherein the off-target editing effects are characterized by unintended base changes in DNA or RNA.

Citation Information

Patent Citations

  • Overhead tractive contact network and matched overhead vehicle current collector

    CN1218849C

  • Press button type switch and its producing method

    CN1409338A

  • Image wafer configuration structure

    CN2518224Y

  • Transistor type high frequency induction heating power supply cabinet

    CN3210875D

  • Decorative stickers (5)

    CN3220306D