Base editor system having non-fused udg
A DNA base correction composition with a DNA binding protein, cytosine deaminase, and adenine deaminase, using UDG independently, addresses the challenge of correcting adenine bases in organelles like mitochondria and chloroplasts, achieving selective adenine-to-guanine conversion while minimizing cytosine-to-thymine errors.
Patent Information
- Application Number
- PCT/KR2025/009620
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-29
- Filing Date
- 2025-07-04
- Publication Date
- 2026-01-08
AI Technical Summary
Conventional genome editing tools, such as CRISPR systems, are ineffective for correcting DNA bases in organelles like mitochondria and chloroplasts due to the inability to deliver guide RNAs, which are essential for activating these systems.
A DNA base correction composition comprising a DNA binding protein, cytosine deaminase, adenine deaminase, and UDG, where UDG exists independently, allowing selective correction of adenine bases to guanine without affecting cytosine bases, particularly useful for organelle DNA.
The composition effectively corrects adenine bases to guanine in organelle DNA, reducing unwanted cytosine-to-thymine corrections, and is particularly effective in plant chloroplasts.
Smart Images

Figure KR2025009620_08012026_PF_FP_ABST
Abstract
Description
Base editor system with non-fused UDG
[0001] The present invention relates to a DNA base correction composition comprising uracil DNA glycosylase (UDG) and a DNA base correction method using the same. Specifically, the present invention relates to a DNA base correction composition comprising a DNA binding protein, cytosine deaminase, adenine deaminase, and UDG, or polynucleotides encoding the proteins, or a DNA base correction method using the same. The present invention is useful for correcting bases in nuclear DNA or organelle DNA, and is particularly useful for correcting bases in organelle DNA such as chloroplasts or mitochondria. In the present invention, UDG exists independently without being fused to a DNA binding protein, cytosine deaminase, and / or adenine deaminase.
[0002] Fusion proteins that link DNA binding proteins and deaminase enzymes enable the induction of DNA mutations, such as single nucleotide conversions in a targeted manner to replace nucleotides or correct bases in the genome without generating DNA double-strand breaks (DSBs), to correct point mutations that cause genetic disorders, or to introduce desired single nucleotide mutations in prokaryotes and eukaryotic cells such as humans.
[0003] Programmable genome editing tools, such as zinc-finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), clustered regularly interspaced short palindromic repeat (CRISPR) systems, and base editors composed of CRISPR-associated protein 9 (Cas9) variants and nucleotide deaminase proteins, have the potential to treat genetic diseases and improve crop traits through base sequence changes. However, these conventional genome editing tools are not suitable for correcting DNA bases in organelles such as mitochondria and chloroplasts, particularly because they cannot deliver the guide RNAs required to activate the most widely used CRISPR systems to these organelles. Mitochondria and chloroplasts encode several essential genes required for photosynthesis and cellular respiration. Methods or tools for correcting genes in these organelles would be useful for studying the function of these genes, treating mitochondrial genetic diseases, and improving crop productivity and traits.
[0004] bacterial toxin DddA tox DddA is the enzymatic component of a bacterial toxin derived from Burkholderia cenocepacia, which can deaminate cytosine in double-stranded DNA. tox Because it is toxic to cells, it is divided into two inactive splits to avoid toxicity in the host cell, and each split (or half) can be linked to a DNA binding protein designed to bind to DNA, so that it can be used as a cytosine base editor (DdCBE, DddA-derived cytosine base editor) that functions as a pair.
[0005] In principle, the enzymatic reaction of this deamination is activated when two inactive halves are present close to the target DNA by their respective DNA-binding proteins, and the base correction from cytosine to thymine (C-to-T) occurs between the DNA sites to which the two DNA-binding proteins each bind.
[0006] Meanwhile, DddA in DNA binding protein tox Adenine base correction can be achieved by linking cytosine deaminase with adenine deaminase, which can correct adenine to guanine (A-to-G) base pairs. Using TALE proteins as DNA-binding proteins and linking cytosine deaminase and adenine deaminase to a base editor (TALE, TALE-linked deaminase), unlike cytosine base editors that only correct C-to-T base pairs, TALE-linked deaminase can introduce a wide spectrum of base mutations (see WO 2022 / 060185).
[0007] Since conventional TALEDs have cytosine deaminase in addition to adenine deaminase, they result in the correction of not only adenine bases but also cytosine bases present in the target DNA. This has been experimentally confirmed to be particularly prominent in plant chloroplasts. Since there is a need to selectively correct only adenine bases while maintaining the wild-type sequence of cytosine bases, the inventors of the present invention sought to develop a method that can selectively correct only adenine bases to guanine bases without causing correction of unwanted cytosine bases. This method is characterized by the use of UDG.
[0008] One aspect of the present invention provides a DNA base correction composition comprising a DNA binding protein, a cytosine deaminase, an adenine deaminase, and UDG, or comprising polynucleotides encoding the proteins, or a DNA base correction method using the same. The DNA base correction composition and the correction method have the activity of selectively correcting only an adenine base without substantially causing correction of cytosine bases, thereby correcting the adenine base to guanine. To this end, UDG exists independently, without being fused to the DNA binding protein, cytosine deaminase, and / or adenine deaminase.
[0009] The present invention relates to a DNA base correction composition comprising UDG and a DNA base correction method using the same. The present invention is useful for correcting bases of nuclear DNA or organelle DNA, and is particularly useful for correcting bases of organelle DNA such as chloroplast or mitochondrion.
[0010] Figures 1 to 3 show the base correction results when using the previously known TALED structure without UDG. Figure 1 is a schematic diagram of a base editor targeting the plant (Arabidopsis) chloroplast gene psaA using the TALED base structure. CTS represents the chloroplast transit signal, AD represents adenine deaminase, NTD represents the N-terminal domain of TALE protein, and CTD represents the C-terminal domain of TALE protein. 1397N and 1397C represent the used DddA. toxRefers to the segments of TALE. "Right TALE repeats" and "Left TALE repeats" represent TALE arrays, respectively, and the underlined base sequences represent the binding sites of TALE proteins (consisting of an N-terminal domain, a TALE array, and a C-terminal domain). Deaminases act and base correction occurs in the region between the DNA base sequences where the two TALE proteins bind (spacer region). The positions of bases in the spacer region are indicated by subscript numbers, and the first base in the spacer region is counted as number 1. This base numbering remains the same in subsequent drawings unless otherwise indicated. Figure 2 shows the A-to-G base correction efficiency at A8 (the 8th adenine) and the C-to-T base correction efficiency at C2 (the 2nd cytosine) obtained by the base editor used, and plots the values measured from each individual used in the experiment. It should be noted that a significant amount of C-to-T editing occurred. Figure 3 shows the overall base editing results obtained from individual individuals that formed the basis for the plotting results in Figure 2. "Col-0" is a wild-type individual not transformed with a base editor. "L" indicates Left TALE, and "R" indicates Right TALE.
[0011] Figures 4 to 6 show the base correction results when UDG was used together with the same TALED configuration used in Figures 1 to 3. Figure 4 is a schematic diagram of the base editor used, showing the use of the TALED base editor used in Figures 1 to 3 together with a separately expressed UDG. Figure 5 shows the A-to-G base correction efficiency at A8 (the 8th adenine) and the C-to-T base correction efficiency at C2 (the 2nd cytosine) obtained by the base editor used, and plots the values measured from each individual used in the experiment. It should be noted that, unlike the results in Figure 2, when UDG was used together, the C-to-T correction efficiency was almost non-existent. Figure 6 shows the overall base correction results obtained from individual individuals that formed the basis of the plotting results in Figure 5.
[0012] Figure 7 is a schematic diagram of a vector construct in which the base editor system used in the experiments of Figures 4 to 6, i.e., the TALED base editor and the UDG expressed separately therefrom, are cloned together.
[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Generally, the terms used herein are those well known and commonly used in the art.
[0014] The embodiments described in this specification and the configurations depicted in the drawings are only one embodiment of how the present invention is realized and do not fully represent the technical idea of the present invention. Therefore, it should be understood that there may be various equivalents, modifications, and applicable examples that can replace them at the time of this application. In addition, the various aspects and embodiments described in this specification can be applied to other aspects and embodiments, and all combinations of the various elements described in the present invention fall within the scope of the present invention, and the scope of the present invention cannot be considered limited by the specific description described below.
[0015] In this specification, the use of the singular includes the plural unless specifically stated otherwise. As used herein, it should be noted that the singular form includes plural referents unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless otherwise stated.
[0016] The term “comprising” as used herein, unless otherwise specified, is understood to be an open-ended expression that essentially includes the described components, ingredients, steps, etc., but does not exclude the presence of other components, ingredients, steps, etc. Accordingly, the term “comprising” is interpreted to include the more limited meaning of “consisting of” or “consisting essentially of.”
[0017] The expression “for” as used in this specification and claims not only means that the composition or material is designed to be used for a particular use, but also includes the meaning that it has functions or characteristics that are suitable or useful for that use even if it is not actually used for that use.
[0018] The terms "correction," "editing," and "editing" as used herein are used interchangeably and refer to a method of altering a nucleic acid sequence by selectively modifying a specific genomic target. Such specific genomic target includes, but is not limited to, a gene, a promoter, an open reading frame, or any nucleic acid sequence.
[0019] As used herein, the terms "base editor," "base editing system," "base editor system," and "base correction system" are used interchangeably and refer to a material or composition suitable for modifying a nucleic acid sequence by inducing selective mutations in a genomic target. As used herein, the terms "base editor," "base editing system," "base editor system," or "base correction system" may be in the form of a polypeptide (which may be a fusion protein) or a polynucleotide, or a combination thereof, depending on the context, and may be a composition comprising one or more polypeptides (which may be fusion proteins) or polynucleotides, or a combination thereof.
[0020] As used herein, the term "conservative amino acid substitution" refers to the replacement of some amino acids with amino acids of different properties while maintaining structural or functional similarity within a protein. Specifically, it refers to substitutions between amino acids with similar physicochemical properties (e.g., charge, size, hydrophobicity, polarity, etc.), and includes substitutions that substantially maintain the structural stability or biological function of the protein.
[0021] For example, substitutions within the following amino acid groups may be conservative substitutions:
[0022] Hydrophobic amino acid group: Ala, Val, Leu, Ile, Met
[0023] Polar uncharged amino acid group: Ser, Thr, Gln, Asn
[0024] Acidic amino acid group: Asp, Glu
[0025] Basic amino acid group: Lys, Arg, His
[0026] Aromatic amino acid group: Phe, Tyr, Trp
[0027] Such determination of substitution can be performed based on the standard amino acid classification that takes into account the charge, polarity, hydrophobicity, structural similarity, etc. of the amino acid, and a person skilled in the art can objectively determine whether the substitution is conservative by utilizing sequence alignment tools such as BLAST and Clustal Omega and conservation matrices (BLOSUM, PAM, etc.). When a specific amino acid sequence is described in this specification, it is interpreted that a variant in which one or more amino acids in the sequence are substituted with another amino acid corresponding to the conservative substitution is also included in the technical scope of the present invention.
[0028] The term “sequence” in this specification may be interpreted as a nucleic acid (or polynucleotide) molecule or a protein (or polypeptide) molecule having a given sequence, depending on the context.
[0029] The term "other amino acid" as used herein means an amino acid selected from among alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, valine, aspartic acid, cysteine, glutamine, glycine, serine, threonine, tyrosine, aspartic acid, glutamic acid, arginine, histidine, lysine, and all known variants of the above amino acids, excluding the amino acid that the wild-type protein originally has at the mutation position.
[0030] The term “sequence identity” or “sequence homology” as used herein means the number of residues present at the same position when two amino acid sequences or nucleic acid sequences are aligned, expressed as a percentage of the total length of the sequences. When it is said herein that a specific sequence “has at least X% sequence identity,” X can be, for example, 85%, 90%, 95%, 98%, or 99%. Sequence identity is typically calculated using BLAST (Basic Local Alignment Search Tool), ClustalW, EMBOSS, or other known sequence alignment algorithms, and is based on default parameters. For example, when aligning amino acid sequences using BLASTP, the identity value calculated using the BLOSUM62 matrix and gap penalty as default values can be used as a basis. In addition, in this specification, “sequence identity” or “sequence homology” may include a value calculated according to an optimized global alignment or local alignment that takes into account insertions, deletions, substitutions, etc. during sequence alignment, and is also used as a standard for explaining the scope of functional equivalents that can maintain the technical effect of the invention.
[0031] The term "homolog" as used herein refers to a protein or nucleic acid that has homology or similarity to a specific protein or gene sequence and exhibits functionally similar biological activity. Homologs may perform the same or similar function, but may be of different species or may have some variation in the amino acid or base sequence. Homologs may be naturally occurring or artificially modified, and are considered to provide the same technical effect within the scope of the present invention.
[0032] The term "ortholog" as used herein refers to a gene or protein derived from two or more species that share a common ancestor, and thus corresponds to a corresponding gene or protein found in different species but with the same evolutionary origin. Generally, orthologs have a high degree of sequence homology between different species and are known to perform similar biological functions. As used herein, an ortholog of a specific protein may include a variant that performs essentially the same function as the original protein, even if a portion of the sequence contains amino acid substitutions, insertions, or deletions.
[0033] The term "functional variant" as used herein refers to a protein that substantially maintains the basic biological function of a polypeptide (e.g., a protein) having a specific amino acid sequence, or has an activity essentially similar to that of the protein, despite having one or more amino acids in the entire sequence conservatively or non-conservatively substituted, or modified, such as insertion, deletion, or substitution. For example, some differences in the sequence may be considered functional variants of the protein described herein, as long as the protein performs the effective function intended in the present invention, such as substrate recognition, catalytic activity, binding ability, or specificity for a target molecule. Such functional variants may be generated by spontaneous mutation, evolutionary modification, induced mutation, or genetic engineering methods.
[0034] As used herein, the terms "target" or "target site" refer to a pre-identified nucleic acid sequence of any composition and / or length. Such target sites include, but are not limited to, genes, promoters, or any nucleic acid sequence.
[0035] The term "fusion protein" as used herein refers to a protein in which two or more different protein (polypeptide) sequences or functional domains are combined into a single continuous polypeptide chain. Such a fusion protein may retain the original biological function of each component or be endowed with new functional properties, and may include a linker sequence between the components. When designating components of a fusion protein herein, unless otherwise specified, the left-to-right direction refers to the N-terminus to the C-terminus, respectively. Additionally, the linker sequence used may not be explicitly indicated.
[0036] The term “expression” as used herein refers to the process by which a polynucleotide (e.g., DNA or mRNA) encoding a DNA base editor or a component thereof is introduced into a cell, and the base editor or a component protein thereof is produced through the cell’s transcription and / or translation mechanisms. The expression may be transient, or stable when integrated into the genome of a target cell. The expression product may be a single protein, or may be a fusion protein in which multiple functional domains are fused. In addition, the expression may be performed in the cytoplasm or organelles (e.g., chloroplasts, mitochondria, nuclei, etc.), and for this purpose, the expression product may additionally include an organelle targeting sequence (MTS, CTS, etc.), a nuclear export signal (NLS), or an extranuclear export signal (NES). The expression level and location may vary depending on the sequence of the polynucleotide being introduced, the promoter, codon optimization, target cell type, introduction method, etc., and a person skilled in the art can appropriately adjust it according to the purpose.
[0037] The term "monomeric base editor" as used herein refers to a base editor comprising a DNA binding protein and a deaminase as a single fusion protein. The monomeric base editor does not exclude the presence of additional polypeptide or polynucleotide components other than the single fusion protein.
[0038] The term "dimeric base editor" as used herein refers to a base editor comprising a DNA binding protein and a deaminase as two fusion proteins. Among the two fusion proteins, the fusion protein that binds to a DNA sequence located 5' upstream of the spacer region may be referred to as a "first fusion protein" or a "left fusion protein," and the fusion protein that binds to a DNA sequence located 3' downstream of the spacer region may be referred to as a "second fusion protein" or a "right fusion protein." Similarly, the DNA binding protein included in the first fusion protein may be expressed by the modifier "first" or "left," and the DNA binding protein included in the second fusion protein may be expressed by the modifier "second" or "right." The dimeric base editor does not exclude the presence of additional polypeptide or polynucleotide components other than the two fusion proteins.
[0039] In some embodiments, the base editor may be a monomeric base editor, and in other embodiments, it may be a dimeric base editor comprising two fusion proteins.
[0040] In this specification, binding of a fusion protein to a given nucleotide sequence means that the DNA binding protein included in the fusion protein recognizes and binds to the nucleotide sequence.
[0041] The term “CRISPR-associated nuclease” as used herein, also called Cas, generally refers to a protein that is a component of the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) system, binds to a guide RNA, specifically binds to a target nucleotide sequence, and has the activity of cleaving or regulating the sequence. In this specification, “CRISPR-associated nuclease” and “Cas” are used interchangeably.
[0042] The term "organelle DNA" as used herein refers to the genetic material present in organelles other than the nucleus within a eukaryotic cell, including mitochondrial DNA or plastid DNA.
[0043] The term “plastid” as used herein refers to an intracellular organelle present in plant cells and some algal cells, including chloroplasts, chromoplasts, and leucoplasts.
[0044] 1. DNA base editor
[0045] One aspect of the present invention relates to a DNA base editor for correcting an adenine (A) base to a guanine (G) base, comprising a DNA binding protein, cytosine deaminase, adenine deaminase, and UDG. The cytosine deaminase exists in a full-length form or in the form of two fragments, and when present in the form of two fragments, the cytosine deaminase activity is exerted through dimerization.
[0046] The term "dimerization," as used herein, refers to the process by which two protein fragments are physically placed in close proximity to each other and form a functional protein complex through noncovalent interactions. For example, the cleaved form of cytosine deaminase exerts its overall enzymatic activity by dimerizing each fragment.
[0047] These DNA base editors have the ability to selectively correct specific DNA bases, and are particularly useful for converting adenine (A) to guanine (G) (A-to-G). In some embodiments, the DNA base editor according to the present invention is effective in selectively performing A-to-G base correction while reducing unwanted C-to-T base correction.
[0048] The DNA base editor of the present invention may comprise one or more fusion proteins, and in some embodiments, two fusion proteins. These fusion proteins may comprise one or more DNA binding proteins and a cytosine deaminase, and may also comprise an adenine deaminase. A linker sequence may be included between the components included in the fusion proteins.
[0049] A. DNA binding protein
[0050] A DNA base editor according to the present invention comprises one or more DNA binding proteins.
[0051] The "DNA binding protein" used in the DNA base editor according to the present invention means a protein that can recognize a specific base sequence and selectively bind to the sequence, and the binding specificity has a programmable characteristic according to the target sequence. In that sense, the DNA binding protein included in the DNA base editor according to the present invention is any programmable DNA binding protein suitable for use in DNA base editing. Those skilled in the art are well aware of the types of programmable DNA binding proteins suitable for use in DNA base editing. Examples of such DNA-binding proteins include zinc finger proteins and TALE (transcription activator-like effector, TALE) proteins, which can be designed to recognize specific base sequences through modular repeat sequences, dead Cas proteins (e.g., dCas9, dCas12a) that are targeted by guide RNA, and other artificially designed DNA recognition domains. These DNA-binding proteins can be fused to enzymatic proteins (e.g., nickases, cytosine deaminase, or adenine deaminase) to induce a desired biochemical reaction at a specific location in the genome.
[0052] The DNA binding protein used in the DNA base editor according to the present invention is not limited to a specific amino acid sequence or structure, and generally includes a programmable protein or functional variant thereof that has the ability to recognize a specific base sequence and selectively bind to that sequence. The DNA binding protein may be naturally occurring, a variant thereof, or an artificially designed protein. The DNA binding protein that can be used in the DNA base editor according to the present invention is understood to encompass all proteins capable of selectively binding to target DNA to achieve the purpose of the present invention, regardless of the specific sequence or origin.
[0053] The DNA base editor according to the present invention comprises one or more DNA binding proteins. The DNA binding proteins are proteins capable of selectively binding to a specific DNA sequence, enabling specific recognition of the target site for base editing.
[0054] In some embodiments, the DNA binding protein may be selected from the group consisting of a zinc finger protein, a TALE protein, a CRISPR-associated nuclease, or a combination thereof. Reference may be made to prior disclosures regarding zinc finger proteins, TALE proteins, and CRISPR-associated nucleases, including WO 2022 / 060185 and WO2022 / 017745, which are incorporated by reference herein in their entirety.
[0055] The above "zinc finger protein (ZFP)" generally refers to a protein or protein domain that forms a stabilized structure through the binding of zinc ions (Zn²) and has the function of binding to a specific DNA sequence. Such zinc finger proteins include one or more "zinc finger (ZF)" structures. Zinc finger proteins have sequence specificity that allows them to bind to a DNA sequence consisting of a specific 3-4 base pair, and by designing them by continuously combining multiple zinc finger domains, they can have high specificity and binding affinity for a long target DNA sequence.
[0056] Zinc finger proteins that can be used in some embodiments of the present invention include naturally occurring proteins or artificial recombinants, mutants, functional variants, or variants with improved specificity derived therefrom, and are not limited in origin or sequence composition, as long as they can specifically bind to a desired target DNA sequence. Those skilled in the art can design and produce a desired zinc finger protein using previously disclosed ZFP libraries, genetic engineering methods, and techniques for analyzing DNA-binding specificity.
[0057] Zinc finger proteins have a relatively small molecular weight and can bind to target sequences with only their pure protein structure without relying on an RNA guide sequence, so they have the advantage of being easy to apply even in delivery systems with vector size limitations or in environments where RNA is unstable.
[0058] The above "TALE protein" is generally based on a transcription activator-like factor derived from the plant pathogenic bacteria Xanthomonas genus, and refers to a DNA binding protein with sequence specificity that can bind to a specific DNA sequence. The TALE protein is composed of a series of repeat modules, each module consisting of about 34 amino acids, of which two amino acid residues at positions 12 and 13 (so-called RVD, repeat-variable diresidue) determine the binding specificity for a single DNA base. By designing a combination of these modules, a TALE protein with sequence specificity tailored to a desired target DNA sequence can be generated. As used herein, the TALE-repeat modules may be referred to as a "TALE array", a "TALE repeat sequence", etc., and the expression "TALE protein" means a configuration in which an N-terminal domain and a C-terminal domain (which may include a half domain) are included on both sides of the TALE array, respectively.
[0059] The term "N-terminal domain (NTD)" as used herein refers to a region located at the amino terminus of a TALE protein, which includes a sequence that contributes to the alignment of DNA binding sites or maintenance of protein stability. For example, in a TALE protein derived from Xanthomonas, the N-terminal domain may be composed of a sequence approximately between amino acids 1 and 150. However, such sequences are merely examples, and the "N-terminal domain" in the present specification also includes variants, homologs, or artificially designed sequences of the sequence as long as the sequence can perform the above function. A person skilled in the art can select or design a suitable N-terminal domain sequence based on the structure and function of a known TALE protein.
[0060] The term "C-terminal domain (CTD)" used herein refers to a region located at the carboxy terminus of a TALE protein, which includes a sequence that performs protein stability or other regulatory functions. For example, in Xanthomonas TALE, a sequence corresponding to amino acids 800 to 900 may be included. However, the "C-terminal domain" in the present specification is not limited to such sequence, and also includes other biological sequences or artificial sequences that can perform the same or similar functions. Such sequences can be easily selected or designed by a person skilled in the art based on publicly available TALE protein information.
[0061] TALE proteins that can be used in some embodiments of the present invention include naturally occurring TALE sequences or recombinants, mutants, functional variants, or forms with artificially controlled sequence specificity derived therefrom, and are not limited in sequence composition or origin, as long as they can bind to a desired DNA sequence. Those skilled in the art can utilize published TALE libraries and TALE design algorithms to create TALE proteins that specifically bind to various DNA target sequences. For example, Kim, Yongsub, et al. "A library of TAL effector nucleases spanning the human genome." Nature biotechnology 31.3 (2013): 251-258.
[0062] TALE proteins have the advantage of not requiring guide RNA, precisely recognizing target DNA sequences by directly assembling repeat modules at the protein level, and being free from constraints on PAM sequences. Therefore, TALE proteins are particularly advantageous in complex genomic environments or situations requiring flexible targeting.
[0063] The above "CRISPR-associated nuclease" is also called "Cas protein" and generally refers to a protein having nuclease activity capable of cleaving DNA or RNA derived from the CRISPR (clustered regularly interspaced short palindromic repeats)-Cas system, which is an acquired immune system of bacteria or archaea. These Cas proteins generally form a complex with a guide RNA, recognize a target nucleic acid sequence complementary to the base sequence of the guide RNA, and then induce cleavage (nicking or double-strand break) at the corresponding site. Representative examples include Cas9 (e.g., Streptococcus pyogenesCas9), Cas12a (Cpf1), Cas12b, Cas13, and Cas14.
[0064] CRISPR-associated nucleases that may be used in some embodiments of the present invention may include naturally occurring proteins, functional variants, conservative amino acid substitutions, truncated forms, or variants in which the enzymatic activity is altered or eliminated, and also include inactive forms (dead Cas or dCas), nickase forms (nCas), or forms that include fusions with various functional domains (e.g., deaminase, transcription factor, etc.).
[0065] When the base editor according to the present invention has two fusion proteins, the two fusion proteins each comprise a DNA binding protein, which may be the same or different from each other. That is, one of the two fusion proteins may be, for example, a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease, and the other may independently be, for example, a zinc finger protein, a TALE protein, or a CRISPR-associated nuclease.
[0066] B. Cytosine deaminase
[0067] The DNA base editor according to the present invention comprises cytosine deaminase.
[0068] The "cytosine deaminase" included in the DNA base editor according to the present invention generally refers to an enzyme that catalyzes a deamination reaction that converts the cytosine base in DNA to uracil. This enzyme converts cytosine to uracil by removing the amino group (-NH2) of cytosine, thereby inducing a C:G → T:A conversion in the corresponding base pair.
[0069] The cytosine deaminase that can be used in some embodiments of the present invention is not limited to a specific amino acid sequence, structural characteristics, biological species of origin, or name, and includes all wild-type proteins, artificial or evolutionary modifications, functional variants, orthologs, conservative amino acid substitutions, etc., as long as the enzyme has a biological activity that can act on a cytosine base on DNA to induce a deamination reaction. Those skilled in the art can easily access numerous literatures and public databases (e.g., GenBank, UniProt, REBASE, etc.) that are already known regarding the sequence, structure, and function of such enzymes, and can implement a cytosine deaminase for implementing selective DNA editing according to a specific purpose.
[0070] In some embodiments, the cytosine deaminase is preferably a double strand DNA-specific cytosine deaminase.
[0071] As used herein, the term "double-stranded DNA-specific cytosine deaminase" refers to an enzyme that deamines the cytosine base in double-stranded DNA, converting it to uracil. Unlike typical cytosine deaminases, this enzyme has the characteristic of directly recognizing and reacting with cytosine within the normal double-stranded DNA structure, rather than single-stranded DNA, as its substrate.
[0072] An example of such a double-stranded DNA-specific deaminase is DddA, a toxin protein from Burkholderia cenocepacia (DddA tox ) is known to have the activity of selectively recognizing cytosine in double-stranded DNA and converting it to uracil. Since the enzyme can generally cause cytotoxicity, it may be desirable to divide it into two split forms as needed and use it by fusing it with a DNA binding protein.
[0073] The double-stranded DNA-specific cytosine deaminase used in the present invention is not limited to the above examples, and is understood to include all proteins or functional variants thereof that have the activity of deaminating cytosine using double-stranded DNA as a substrate, regardless of its origin, amino acid sequence, structure, or name. Such double-stranded DNA-specific cytosine deaminase and variants thereof have already been described in various documents. For example, WO 022 / 060185, WO 2022 / 221337, WO 2022 / 155265, WO 2023 / 081855, WO 2023 / 097226, WO 2024 / 112441, WO 2024 / 107263, etc., which are incorporated by reference in their entirety by this application, and contents already known prior to the present application may be cited.
[0074] In some embodiments, the cytosine deaminase is DddA, a cytosine deaminase from Burkholderia cenocepacia. toxOr its variant. DddA tox The amino acid sequence is as follows.
[0075] wild-type DddA tox :
[0076] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNNSNSPKSPTKGGC (SEQ ID NO: 1)
[0077] DddA tox As variants of DddA, various variants characterized by the context of the cytosine (C) base to be base-edited, the editing efficiency, the improved specificity, etc. are also known, and these include DddA2, DddA3, DddA4, DddA5, DddA6, DddA7, DddA8, DddA9, DddA10, DddA11, etc. With regard to these variants, reference may be made to the contents already known prior to the present application, including, for example, WO 2022 / 221337, which is incorporated by reference in its entirety herein.
[0078] In some embodiments, the present invention provides DddA as a cytosine deaminase. tox Full-length or split forms (e.g., 1333N / 1333C or 1397N / 1397C) may be used, and the split forms may each be contained in separate fusion proteins and function cooperatively. As used herein, "cooperatively functioning" with respect to split forms of cytosine deaminase means that although they exist as separate proteins, they exert cytosine deaminase activity through dimerization.
[0079] DddA toxIt can be used alone, but because of its high toxicity and strong enzymatic activity, it is divided into N-terminal and C-terminal fragments (split DddA) for safer and more efficient DNA base editing. tox It is preferable to use it in the form of DddA. In this case, DddA tox It exists as two splits, and each is incorporated into or linked to an independent fusion protein to function. When the cytosine deaminase used in the present invention is used in the form of a first split and a second split, the first split and the second split do not have deamination activity, and the deaminization activity is exhibited only when the two splits are adjacent to each other. That is, in order for the cytosine deaminase to be used in the form of two splits, the full-length sequence of the cytosine deaminase must be formed when the sequences of the two splits are combined, and a person skilled in the art of base correction technology using cytosine deaminase is well aware of this point.
[0080] When DddAtox exists in the form of two fragments, one of the two fragments may include a sequence from the N-terminus to the 33rd, 44th, 54th, 68th, 82nd, 98th or 108th amino acid of the amino acid sequence of SEQ ID NO: 1, and the other of the two fragments may include a sequence from the 34th, 45th, 55th, 69th, 83rd, 99th or 109th amino acid of the amino acid sequence of SEQ ID NO: 1 to the C-terminus.
[0081] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of 1333N and 1333C having the following amino acid sequences.
[0082] 1333N:
[0083] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG (SEQ ID NO: 2)
[0084] 1333C:
[0085] PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNNSNSPKSPTKGGC (SEQ ID NO: 3)
[0086] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of 1397N and 1397C having the following amino acid sequences.
[0087] 1397N:
[0088] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG (SEQ ID NO: 4)
[0089] 1397C:
[0090] AIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO: 5)
[0091] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of DddA11 1333N and DddA11 1333C having the following amino acid sequences.
[0092] DddA11 1333N:
[0093] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGG
[0094] DddA11 1333C:
[0095] PTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVVPPEGAIPVKRGATGETKVFIGNNSNSPKSPTKGGC
[0096] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of DddA11 1397N and DddA11 1397C having the following amino acid sequences.
[0097] DddA11 1397N:
[0098] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVVPPEG
[0099] DddA11 1397C:
[0100] AIPVKRGATGETKVFIGNNSNSPKSPTKGGC
[0101] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of DddA11 1333N and DddA11 1333C having the following amino acid sequences.
[0102] In some embodiments, when cytosine deaminase is used in the form of two fragments, it may be in the form of FZY2 100N and FZY2 100C having the following amino acid sequences.
[0103] FZY2 100N:
[0104] MSLPEYDGTTTHGVLVLDDGTQIGFTSGNGDPRYTNYRNNGHVEQKSALYMRENNISNATVYHNNTNGTCGYCNTMIATFLPEGATLTVVPPENAVANNS
[0105] FZY2 100C:
[0106] RAIDYVKTYTGTSNDPKISPRYKGN
[0107] In some embodiments, DddA toxOr, when its variant exists in the form of splits of 1333N and 1333C, one or more amino acids selected from the group consisting of positions 3, 5, 10, 11, 13, 14, 15, 16, 17, 18, 19, 28, 30 and 31 of 1333N or one or more amino acids selected from the group consisting of positions 13, 16, 17, 20, 21, 28, 29, 30, 31, 32, 33, 56, 57, 58 and 60 of 1333C may be substituted with another amino acid. DddA tox Or, when its variant exists in the form of splits of 1397N and 1397C, one or more amino acids selected from the group consisting of positions 87, 88, 91, 92, 95, 100, 101, 102 and 103 of 1397N) or one or more amino acids selected from the group consisting of positions 13, 14, 15 and 16 of 1397C may be substituted with another amino acid. In some embodiments, the other amino acid is alanine.
[0108] In some embodiments, DddA tox Or, if its variant exists in the form of splits of 1333N and 1333C, the amino acid at position 56, 57 or 58 of 1333C may be substituted with another amino acid. DddA tox Alternatively, if its variant exists in the form of splits of 1397N and 1397C, the amino acid at positions 100, 101 or 102 of 1397N may be substituted with another amino acid. In some embodiments, the other amino acid is alanine.
[0109] Using these mutants, each linked to a DNA binding protein, DddA toxAlternatively, if two fragment pairs derived from its variant fail to bind to DNA, they may not function properly, resulting in highly efficient and precise C-to-T editing without causing undesirable off-target C-to-T editing. For this purpose, reference may be made to prior disclosures, including, for example, WO 2022 / 060185, which is incorporated herein by reference in its entirety.
[0110] In some embodiments, the cytosine deaminase used in the present invention may be used in a full-length form, and the full-length cytosine deaminase used in this case (e.g., DddA) tox ) are amino acid sequences that have been modified to have no or only low toxicity.
[0111] DddA tox The C-terminus of DNA has a specific concentration of positively charged amino acids. Since DNA is negatively charged, it binds to the positively charged amino acids of proteins. By replacing these positively charged amino acids, DddA is formed. tox By weakening the binding force of DddA to DNA, intracellular toxicity can be reduced or eliminated. In other words, if a positively charged amino acid is substituted to make it non-toxic, cloning using E. coli is possible, resulting in full-length DddA. tox can be secured. Based on this, the non-toxic full-length cytosine deaminase is DddA of sequence number 1. tox It can be provided by replacing one or more, two or more, three or more, four or more, or five or more amino acids in the amino acid sequence with other amino acids (e.g., alanine), and in this regard, reference may be made to contents already known prior to the present application, including, for example, WO 2022 / 060185, which is incorporated by reference in its entirety by this application.
[0112] In some embodiments, when a full-length cytosine deaminase is used as the cytosine deaminase, such cytosine deaminase may have one or more amino acid substitutions selected from the group consisting of a substitution of S at position 37 with G, a substitution of G at position 59 with S, a substitution of A at position 109 with V, and a substitution of S at position 129 with G in the amino acid sequence of SEQ ID NO: 1.
[0113] A cytosine deaminase mutant having all of the following substitutions: S to G at position 37, G to S at position 59, A to V at position 109, and S to G at position 129 in the amino acid sequence of SEQ ID NO: 1 is commonly referred to as "GSVG."
[0114] GSVG:
[0115] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGTPPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC (SEQ ID NO: 6)
[0116] In addition, the non-toxic full-length DddAtox may comprise an amino acid sequence selected from the group consisting of the following amino acid sequences.
[0117] A1341D KRKKA variant:
[0118] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYDNAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNNSNSPKSPTAGGC
[0119] AAAAA variant:
[0120] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTAGGC
[0121] AAAAK 변이체:
[0122] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETAVFTGNSNSPASPTKGGC
[0123] AAKAA 변이체:
[0124] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTAGGC
[0125] AAKAK 변이체:
[0126] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVAAGATGETKVFTGNSNSPASPTKGGC
[0127] KAAAA 변이체:
[0128] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKAGATGETAVFTGNSNSPASPTAGGC
[0129] E1347A 변이체:
[0130] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVAGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNNSNSPKSPTKGGC
[0131] SSVG variants:
[0132] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0133] GSAG mutants:
[0134] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGTPPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNGPKSPTKGGC
[0135] GSVS variants:
[0136] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLEGKVFSSGGPTPYPNYANAGHVESQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGVIPVKRGATGETKVFTGNNSNSPKSPTKGGC
[0137] C. adenine deaminase
[0138] The DNA base editor according to the present invention comprises an adenine deaminase.
[0139] The above "adenine deaminase" generally refers to an enzyme that catalyzes the deamination reaction that converts the adenine base in DNA to inosine. Inosine acts similarly to guanine during DNA replication or transcription, resulting in an A:T → G:C conversion in the corresponding base pair.
[0140] The adenine deaminase that can be used in some embodiments of the present invention is not limited to a specific amino acid sequence, structural characteristics, biological species or name, and includes all wild-type proteins, artificial or evolutionary modifications, functional variants, orthologs, conservative amino acid substitutions, etc., as long as the enzyme has a biological activity that can act on an adenine base on DNA to induce a deamination reaction. Those skilled in the art can implement an adenine deaminase suitable for a specific purpose by referring to various literature and public databases (e.g., GenBank, UniProt, etc.) regarding the sequence, structure, and function of such enzymes.
[0141] For example, TadA (tRNA-specific adenosine deaminase) from Escherichia coli is a well-known representative adenine deaminase that originally has the activity of deaminating adenosine in tRNA, and has been improved to acquire the activity that can act on DNA through specific mutations or protein engineering.
[0142] In some embodiments, the adenine deaminase has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to TadA or an ortholog thereof having the following amino acid sequence, a functional variant thereof, or a conservative amino acid substitution thereof.
[0143] TadA:
[0144] MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD
[0145] In some embodiments, the adenine deaminase has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to TadA8e having the amino acid sequence below, or a functional variant thereof, or a conservative amino acid substitution thereof. Such variants may include, for example, one or more amino acids selected from the group consisting of positions 28, 30, 46, 48, 49, 82, 84, 106, 108, 110, and 111 of the amino acid sequence of TadA8e (SEQ ID NO: 7) in which one or more amino acids are mutated to another amino acid or a conservative amino acid substitution thereof. For example, the amino acid variant may include one or more amino acid substitutions selected from the group consisting of V28Q, V28R, A48W, F84M, V106A, K110S, K110T, K110V, R111F, R111Q, R111S, R111T, and R111Y. V28Q means that the 28th valine (V) is mutated to glutamine (Q), and a person skilled in the art who is familiar with amino acid symbols can easily understand the meaning of the mutant notations. With regard to the composition of an adenine deaminase that can be used in the present invention, contents already known prior to the present application, including WO 2022 / 060185, WO 2023 / 086953, etc., which are incorporated by reference in their entirety by this application, may be cited.
[0146] TadA8e:
[0147] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0148] D. UDG
[0149] The base editor according to the present invention comprises UDG.
[0150] The UDG used in the present invention is not limited to a specific amino acid sequence, structure, origin or species, and is generally understood as a concept including all proteins having the activity of recognizing and removing uracil bases in DNA, truncated forms thereof (e.g., truncated forms in which the amino acid sequence corresponding to the N-terminal 1-106 amino acid sequence of Arabidopsis thaliana UDG is removed), or functional variants thereof. For example, human UDG (human UNG1, NCBI Reference Sequence: NP_003353; human UNG2, NCBI Reference Sequence: NP_550433), Escherichia coli UDG (E. coli UDG, NCBI Reference Sequence: NP_417075), Arabidopsis thaliana UDG (Arabidopsis thaliana UDG, NCBI Reference Sequence: NP_188493), archaeal UDG, or phage-derived UDG have different sequences and structures, but they all perform the common biological function of removing uracil, and are therefore included in the "UDG" referred to herein. UDG is also called UNG (uracil-N-glycosylase).
[0151] The structure and function of UDG are already well known (see Schormann, N., et al. "Uracil-DNA glycosylases―structural and functional perspectives on an essential family of DNA repair enzymes." Protein science 23.12 (2014): 1667-1685., Cordoba-Canero, D., et al. "Arabidopsis uracil DNA glycosylase (UNG) is required for base excision repair of uracil and increases plant sensitivity to 5-fluorouracil." Journal of Biological Chemistry 285.10 (2010): 7475-7483.).
[0152] In some embodiments, the UDG used in the DNA base editor according to the present invention is derived from UDG of Arabidopsis thaliana.
[0153] In some embodiments, the UDG used in the DNA base editor according to the present invention is derived from human UDG (i.e., UNG1 or UNG2).
[0154] The base editor according to the present invention comprises UDG, and is characterized in that the UDG exists in a protein form that is expressed separately from other components of the base editor (DNA binding protein, cytosine deaminase, and adenine deaminase).
[0155] E. Additional polypeptide elements
[0156] The DNA base editor according to the present invention may further comprise additional polypeptide or protein components in addition to (i) a DNA binding protein, (ii) a cytosine deaminase, (iii) an adenine deaminase, and (iv) UDG.
[0157] In addition to the DNA binding protein, cytosine deaminase, and UDG, the DNA base editor of the present invention may additionally include various additional protein or peptide sequences. These additional components are used to enhance base editing efficiency or to control intracellular delivery to the target sequence and localization within organelles.
[0158] In some embodiments, the DNA base editor of the present invention may comprise an NLS.
[0159] The above "NLS (nuclear localization signal)" refers to an amino acid sequence motif required for protein translocation from the cytoplasm to the nucleus. NLSs are generally composed of short sequences rich in basic amino acids such as lysine or arginine, and mediate protein entry into the nucleus through interaction with nuclear transport receptors (e.g., importins).
[0160] The NLSs that can be used in some embodiments of the present invention are not limited to a specific sequence, structure, or origin, and various forms of NLSs can be used as long as they maintain their functional characteristics. Information regarding the function and sequence of NLSs is widely known through prior literature and publicly available materials, and those skilled in the art can select or modify an NLS sequence suitable for the intended purpose based on this information.
[0161] In some embodiments, the DNA base editor of the present invention may comprise a NES.
[0162] The above "NES (nuclear export signal)" refers to a peptide sequence or a functional variant thereof capable of inducing protein transport into the cytoplasm. NESs generally bind to nuclear export receptors (exportins) to facilitate protein transport from the nucleus to the cytoplasm, and typically have a structural characteristic of 4-5 hydrophobic amino acids arranged at specific intervals. The NES used in some embodiments of the present invention is not limited to a specific sequence, structure, or origin, and may include various forms of nuclear export sequences as long as they maintain their functional characteristics.
[0163] The NES used in some embodiments of the present invention may be derived from various proteins existing in nature, and artificially designed sequences may also be utilized. For example, the NES derived from the HIV-1 Rev protein (e.g., LQLPPLERLTL), the NES derived from the PKI protein (e.g., LALKLAGLDI), or the NES derived from the NS2 protein of the mouse minivirus (MVM) (e.g., VDEMTKKFGTLTIHDTEK) are widely known as representative sequences, and various variants having similar nuclear export functions are also well known to those skilled in the art.
[0164] In some embodiments, the DNA base editor of the present invention may comprise MTS.
[0165] The above "MTS (mitochondrial targeting sequence)" refers to an amino acid sequence that induces the transport of a protein to the mitochondria after translation. MTS is generally located at the N-terminus and forms a unique α-helical structure with a repeating arrangement of basic and hydrophobic amino acids, which interacts with the mitochondrial inner membrane transport complex to achieve transport. The MTS that can be used in some embodiments of the present invention is not limited to a specific sequence, length, structure, or origin, and is understood to include any functional sequence that can effectively direct a protein to the mitochondria.
[0166] Information on MTS has been identified from various biological proteins, and relevant sequences are widely available through public databases such as UniProt, NCBI, and MitoCarta. Using this publicly available sequence information, those skilled in the art can select or combine MTS sequences appropriate for the desired protein to design it.
[0167] The MTS used in some embodiments of the present invention may be derived from various mitochondrial proteins existing in the natural world, and a sequence artificially designed to have a specific mitochondrial transport function may also be utilized. For example, MTS derived from human superoxide dismutase 2 (SOD2) protein (e.g., MALSRAVCGTSRQLAPVLGYLGSRQKHSLPD), MTS derived from cytochrome c oxidase subunit 8A (COX8A) (e.g., MASVLTPLLLRGLTGSARRLPVPRAKIHSL), or MTS derived from human mitochondrial ATP synthase F1β subunit (e.g., MLGFVGRVAAAPASGALRRLTPSASLPPAQLLLRAAPTAVHPVRDYAAQ) are widely known as representative sequences, and similarly, various variants having a protein transport function to mitochondria are well known to those skilled in the art.
[0168] In some embodiments, the DNA base editor of the present invention may comprise a CTS.
[0169] The above “chloroplast targeting sequence (CTS)” refers to an amino acid sequence that induces transport of a protein translated in the cytoplasm to the chloroplast. It is generally located at the N-terminus of the protein, and forms a structure in which positively charged amino acids and hydrophobic amino acids are repeatedly arranged to interact with a transport complex that passes through the outer and inner membranes of the chloroplast, thereby mediating transport to the chloroplast stroma. The CTS in the present specification is not limited to a specific sequence, length, structure, or biological origin, and is understood to include any functional sequence that can effectively direct a protein to the chloroplast.
[0170] CTS sequences have been identified from various plant-derived proteins (e.g., RuBisCO subunits, ferredoxin, plastocyanin, etc.) and are widely available in public databases such as UniProt, TAIR, and NCBI. Those skilled in the art can design a CTS suitable for a target protein by referencing or combining these published sequences. Therefore, the CTS that can be used in some embodiments of the present invention is not limited to a specific CTS sequence, but may include various homologs, functional analogs, or variants that can functionally induce chloroplast transport.
[0171] Additional polypeptide elements such as those described above may be present in the form of a fusion protein (including a DNA binding protein and a deaminase) or may be present separately from such a fusion protein.
[0172] When present in the form of a fusion protein, localization signal sequences such as NLS, NES, MTS, and CTS are preferably located at the N-terminus of the fusion protein, and can be appropriately designed depending on the cellular organelle in which the proofreading system is to function.
[0173] In some embodiments, the DNA base editor of the present invention may include additional sequences for biotechnology techniques, such as tags.
[0174] In some embodiments, the fusion protein and / or polypeptide comprised in the base editor according to the present invention may include a protein tag sequence, such as a His-tag, FLAG-tag, HA-tag, or Myc-tag, to facilitate purification or detection.
[0175] These tag sequences are additionally located at the N- or C-terminus of the fusion protein or polypeptide, providing experimental or manufacturing convenience. For example, the His-tag can be utilized for protein purification using metal columns, while the FLAG-tag or HA-tag are useful for antibody-based Western blotting or immunostaining.
[0176] The above tag can be designed and positioned within a range that does not affect protein function, and can also be configured with a removable cleavage sequence if necessary.
[0177] F. fusion protein
[0178] All or part of the components constituting the DNA base editor according to the present invention may be present in the form of a fusion protein.
[0179] F-1. Unit fusion protein A
[0180] In some embodiments, a DNA base editor according to the present invention may comprise (i) one or more DNA binding proteins (e.g., zinc finger proteins, TALE proteins, etc.) and (ii) a cytosine deaminase in the form of a single fusion protein.
[0181] The sequence of DNA binding proteins and cytosine deaminase included in the fusion protein can vary.
[0182] In some preferred embodiments, the sequence of the DNA binding protein and cytosine deaminase within the fusion protein is as follows. The sequence below is intended to indicate the relative positions of the indicated components, and other components that may be present, including linkers, are omitted.
[0183] (N-terminal) [DNA binding protein] - [cytosine deaminase] (C-terminal)
[0184] Various additional polypeptide components described in the “E. Additional Polypeptide Components” section above may be added to the arrangement described above. These may also be directly linked to other protein components or linked via one or more linkers.
[0185] The above-mentioned protein components may be directly linked to each other or linked via one or more linkers.
[0186] F-2. Unit fusion protein B
[0187] In some embodiments, a DNA base editor according to the present invention may comprise (i) one or more DNA binding proteins (e.g., zinc finger proteins, TALE proteins, etc.), (ii) cytosine deaminase, and (iii) adenine deaminase in the form of a single fusion protein.
[0188] The sequence of DNA binding protein, cytosine deaminase, and adenine deaminase included in the fusion protein can vary.
[0189] In some preferred embodiments, the sequence of the DNA binding protein, cytosine deaminase, and adenine deaminase within the fusion protein is as follows. The sequence below is intended to indicate the relative positions of the indicated components, and other components that may be present, including linkers, are omitted.
[0190] (N-terminal) [DNA binding protein] - [cytosine deaminase] - [adenine deaminase] (C-terminal)
[0191] (N-terminal) [DNA binding protein] - [adenine deaminase] - [cytosine deaminase] (C-terminal)
[0192] Various additional polypeptide components described in the “E. Additional Polypeptide Components” section above may be added to the arrangement described above. These may also be directly linked to other protein components or linked via one or more linkers.
[0193] The above-mentioned protein components may be directly linked to each other or linked via one or more linkers.
[0194] F-3. Linker
[0195] The term "linker" as used herein refers to an amino acid linker, which is an amino acid sequence that covalently connects two or more functional protein domains, peptides, or other biological molecular elements. Such linkers serve to provide sufficient flexibility, length, or spatial separation so that each connected component can maintain its own structural or functional activity, and may sometimes be designed to include a specific secondary structure (e.g., an α-helix) or recognition sequence. For example, a repeating sequence based on glycine (G) and serine (S) (e.g., GGGGS)n is known as a representative example that confers high flexibility and water solubility.
[0196] The linker that can be used in the DNA editing editor according to the present invention is not limited to a specific amino acid sequence, length, or structure. It can be a naturally occurring sequence or an artificially designed sequence, and any amino acid sequence having a variety of lengths, sequence combinations, or structural characteristics is encompassed, as long as the function of each component connected via the linker is substantially maintained. Those skilled in the art can select or design an appropriate linker based on the characteristics of the target domain and the intended application.
[0197] In some embodiments, the DNA editing editor according to the present invention may comprise one or more linkers selected from the following linkers:
[0198] 2a.a. Linker: GS
[0199] 4a.a. Linker: GSGS
[0200] 5a.a. Linker: TGEKQ
[0201] 10a.a. Linker: SGAQGSTLDF
[0202] 13a.a. Linker: AAEFGIRIPGEKP
[0203] 14a.a. Linker: AAEFGIHGVPAAMG
[0204] 16a.a. Linker: SGSETPGTSESATPES
[0205] 24a.a. Linker: SGTPHEVGVYTLSGTPHEVGVYTL
[0206] 32a.a. Linker: GSGGSSGGSSGSETPGTSESATPESSGGSSGGS
[0207] F-4. Monomeric Base Editor
[0208] In some embodiments, the DNA base editor according to the present invention comprises a DNA binding protein and a cytosine deaminase as a single fusion protein (monomeric base editor). That is, in some embodiments, the DNA base editor according to the present invention comprises a single fusion protein and UDG.
[0209] In some embodiments where the base editor according to the present invention is a monomeric base editor, DddA is used as the cytosine deaminase. tox When used, the cytosine deaminase is non-toxic full-length DddA tox It is desirable. For example, GSVG can be used.
[0210] In some embodiments in which the DNA base editor according to the present invention is a monomeric base editor, it is preferred that the fusion protein included in the base editor includes an adenine deaminase, i.e., a monomeric fusion protein B. That is, in some embodiments in which the DNA base editor according to the present invention is a monomeric base editor, the base editor includes a monomeric fusion protein B and UDG. In this case, UDG exists separately from the monomeric fusion protein B.
[0211] F-5. Dimeric Base Editor
[0212] In some embodiments, the DNA base editor according to the present invention comprises a DNA binding protein and a cytosine deaminase as two fusion proteins (dimeric base editor). That is, in some embodiments, the DNA base editor according to the present invention comprises two fusion proteins and UDG.
[0213] In certain embodiments where the DNA base editor according to the present invention is a dimeric base editor, the cytosine deaminase exists in two fragments, each of which can be divided into a first fusion protein and a second fusion protein. An example of a fragment is DddA. tox 1333N / 1333C or DddA tox Includes 1397N / 1397C.
[0214] In some embodiments where the DNA base editor according to the present invention is a dimeric base editor, the two fusion proteins included in the base editor may be as follows.
[0215] First fusion protein Second fusion protein Unit fusion protein A Unit fusion protein B Unit fusion protein B Unit fusion protein A Unit fusion protein B Unit fusion protein B
[0216] That is, in some embodiments where the DNA base editor according to the present invention is a monomeric base editor, the base editor comprises UDG in addition to the combination of the first fusion protein and the second fusion protein as described above. In this case, UDG exists separately from the monomeric fusion protein A or B.
[0217] In some embodiments where the DNA base editor according to the present invention is a dimeric base editor, DNA editing occurs in the region between the DNA sequences to which the two DNA binding proteins included in the two fusion proteins each bind, said region being referred to as a "spacer."
[0218] G. A-to-G base correction efficacy
[0219] The base editor of the present invention is useful for correcting adenine (A) bases to guanine (G) bases. The inventors of the present invention have confirmed that when performing A-to-G base correction using a DNA base editor containing cytosine deaminase, when UDG is used together, only adenine bases are selectively corrected while cytosine base correction practically does not occur.
[0220] The above A-to-G correction is A-to-G correction in nuclear or organelle DNA.
[0221] In some embodiments, the base editor of the present invention is suitable for A-to-G editing in nuclear DNA. In particular, it is suitable for selectively performing A-to-G editing while avoiding unwanted C-to-T editing in nuclear DNA.
[0222] In some embodiments, the base editor of the present invention is suitable for A-to-G editing in chloroplast DNA. In particular, it is suitable for selectively performing A-to-G editing while avoiding undesirable C-to-T editing in chloroplast DNA.
[0223] In some embodiments, the base editor of the present invention is suitable for A-to-G editing in mitochondrial DNA. In particular, it is suitable for selectively performing A-to-G editing while avoiding undesirable C-to-T editing in mitochondrial DNA.
[0224] In some embodiments, the base editor according to the present invention can reduce the frequency of C-to-T base correction by at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% compared to a base editor of the same configuration except that it does not include UDG. This is very useful for gene correction therapy, crop trait improvement, etc., where it is necessary to selectively correct only adenine bases while maintaining the wild-type sequence of cytosine bases.
[0225] 2. Polynucleotide
[0226] One aspect of the present invention relates to a polynucleotide encoding the DNA base editor of the present invention. With respect to the "DNA base editor of the present invention," the contents described in the "1. DNA base editor" section of this specification are incorporated herein by reference.
[0227] The polynucleotide according to the present invention is a polynucleotide encoding a protein component (fusion protein and / or polypeptide) constituting the DNA base editor of the present invention as described above, and the polynucleotide may be DNA or RNA. The DNA or RNA includes sequences contained within mRNA, cDNA, synthetic DNA, plasmid DNA, linear DNA, or viral vectors.
[0228] The polynucleotide according to the present invention may be two polynucleotides each encoding a fusion protein and UDG used in the monomeric base editor described above, or three polynucleotides each encoding two fusion proteins and UDG used in the dimeric base editor. In addition, it may further include nucleotide sequence(s) encoding additional polypeptide(s) as needed. These two or more polynucleotides can be combined and produced in the form of a single polynucleotide using known biotechnology techniques such as the Golden Gate technique (see FIG. 3).
[0229] A person skilled in the art can easily obtain the amino acid sequence of each constituent protein by referring to the contents disclosed in the present application specification and the published protein sequences registered in databases such as NCBI GenBank and UniProt, and produce a polynucleotide according to the present invention. Based on the obtained amino acid sequences, a nucleotide sequence can be designed considering the codon usage frequency suitable for the host organism, and further, it is obvious to a person skilled in the art or within the scope of routine experimental techniques to produce a polynucleotide containing the sequence using commercial services for custom production of synthetic genes (e.g., IDT, GenScript, etc.) and vector cloning, PCR amplification, and DNA assembly technologies (Gibson assembly, Golden Gate, etc.).
[0230] The polynucleotide(s) may include expression control sequences (e.g., promoter) in addition to the coding region so as to sufficiently express the function of the encoded protein(s).
[0231] In some embodiments of the present invention, the polynucleotide may be codon optimized to suit the type of expression system (e.g., bacteria, plants, mammalian cells, etc.). Codon optimization is a general technique for increasing gene expression efficiency, and can increase protein expression by adjusting the nucleotide sequence according to the tRNA utilization frequency of the target organism.
[0232] 3. Base correction composition
[0233] One aspect of the present invention relates to a base correction composition comprising the DNA base editor of the present invention or a polynucleotide encoding the same. With respect to the "DNA base editor of the present invention," the contents described in the "1. DNA base editor" section of this specification are incorporated herein by reference.
[0234] The base correction composition according to the present invention is for correcting the base adenine (A) to guanine (G) in the DNA of a nucleus or cell organelle.
[0235] In some embodiments, the base correction composition according to the present invention is suitable for correcting an adenine (A) base to a guanine (G) base in the DNA of a nucleus.
[0236] In some embodiments, the base correction composition according to the present invention is suitable for correcting an adenine (A) base to guanine (G) in the DNA of a chloroplast.
[0237] In some embodiments, the base correction composition according to the present invention is suitable for correcting an adenine (A) base to a guanine (G) base in mitochondrial DNA.
[0238] The base correction composition according to the present invention may comprise the DNA base editor of the present invention or a polynucleotide encoding the same as described above and a biocompatible carrier.
[0239] The above “biocompatible carrier” refers to a material that can be included in the composition according to the present invention, and which does not significantly inhibit physiological functions when in contact with a biological system—e.g., cells, tissues, organs, or entire organisms of humans, animals, or plants—and does not induce toxic or immune reactions. Such carriers can be selected in various ways depending on pharmaceutical, biological, or agricultural applications, and can play a role in improving the stability, permeability, and delivery efficiency of base-correction enzymes or related proteins, nucleic acids, auxiliary molecules, etc.
[0240] The biocompatible carrier that can be used in the base correction composition according to the present invention is not limited to a specific chemical structure, physical form or origin, and may include, for example, a buffer solution, a surfactant, a liposome, a nanoparticle, a hydrogel, a polymeric material (e.g., PEG, PVA, PLA, PLGA), a natural or synthetic polysaccharide (e.g., dextran, hyaluronic acid, chitosan), a plant- or microbial-derived polymer, a liposome, a micelle, a biopolymer, or a combination thereof. In addition, the biocompatible carrier may be appropriately selected or combined depending on a specific delivery route or application target (e.g., human tissue, animal tissue, plant tissue, etc.), and may also include a biologically acceptable solvent, preservative, stabilizer, buffer, surfactant, reducing agent, etc.
[0241] A person skilled in the art can easily select and combine a carrier suitable for a given application purpose, delivery route, or target organism based on information already widely known through various literature and public databases regarding the types, properties, and application methods of the biocompatible carriers.
[0242] The base correction composition according to the present invention may be provided in various physical forms. The physical form may be selected based on the intended application, route of administration, stability, storage conditions, or manufacturing process, and falls within the general pharmaceutical design criteria for enhancing the efficacy and ease of use of the composition.
[0243] In the present invention, UDG exists independently of the fusion protein comprising the DNA binding protein and the deaminase (cytosine deaminase and / or adenine deaminase). That UDG exists independently of the fusion protein means that UDG is expressed in a protein form separate from the DNA binding protein and the deaminase (cytosine deaminase and / or adenine deaminase).
[0244] In some embodiments, the base correction composition according to the present invention may be in the form of a liquid, suspension, gel, powder, lyophilisate, tablet, capsule, or injectable composition. It may also be formulated as a liposome, nanoparticle, lipid nanoparticle (LNP), or other delivery particle.
[0245] The base correction composition aspect of the present invention is not limited to the above physical form, and all formulation modifications that can be appropriately selected and manufactured by a person skilled in the art according to the purpose are included in the scope of the present invention.
[0246] The base correction composition according to the present invention can be applied in various ways to ensure effective delivery to the target cell, tissue, or organism. The application method may vary depending on the target species, cell type, delivery route, or formulation characteristics, and is selected based on the stability, efficacy, and biological compatibility of the composition.
[0247] The base correction composition of the present invention can be applied to plants or plant cells. Methods for applying to plants may include agroinfiltration, Agrobacterium-mediated delivery, gene gun delivery, electroporation, or direct intratissue injection. The composition to be applied may be prepared in the form of protein, DNA, mRNA, or ribonucleoprotein (RNP), and may be used with a delivery vehicle capable of penetrating plant cell walls (e.g., cell-penetrating peptide, non-targeting nanoparticle, etc.) depending on the purpose.
[0248] The base correction composition of the present invention can be flexibly applied to various biological subjects, and any in vivo, ex vivo, or in vitro delivery method that can be selected by a person skilled in the art to achieve the desired effect is included in the scope of application of the present invention.
[0249] The base correction composition of the present invention can be used to directly manipulate cells in an extracellular environment (in vitro), and can also be used for the purpose of directly inducing base correction in an intracellular environment (in vivo).
[0250] The base correction composition according to the present invention can be usefully applied to base correction technology that enables precise manipulation of genetic sequences by selectively converting specific DNA bases in vivo or in vitro. In particular, the composition can be used to precisely correct a desired target base sequence through A-to-G correction, which converts adenine (A) to guanine (G).
[0251] Because this correction reaction occurs without DNA double-strand breaks, it has the advantage of being less mutagenic and enhancing genome stability compared to existing gene editing technologies. Therefore, the composition of the present invention can be widely utilized in various fields, including plant variety improvement, industrial microorganism improvement, and biotechnology research.
[0252] 4. Transmitter
[0253] One aspect of the present invention relates to a delivery system comprising the DNA base editor of the present invention or a polynucleotide encoding the same. With respect to the "DNA base editor of the present invention," the contents described in the "1. DNA base editor" section of this specification are incorporated herein by reference.
[0254] The delivery vehicle according to the present invention refers to a means for effectively delivering the DNA base editor of the present invention or one or more polynucleotides (e.g., DNA, mRNA, etc.) encoding the same to a target site in a cell or a living body. Such a delivery vehicle may include various components to enhance the cell penetration efficiency, intracellular stability, organelle targeting ability, or in vivo distribution characteristics of the editor component, and preferably has acceptable properties such as biocompatibility, biodegradability, and non-immunogenicity. The delivery vehicle of the present invention aims to increase the efficiency and specificity of gene correction, while minimizing cytotoxicity and reducing the possibility of affecting non-target tissues.
[0255] The term "vector" or "delivery vehicle" as used herein refers to a biological or non-biological means capable of effectively delivering the DNA base editor of the present invention or one or more polynucleotides encoding it into cells. These delivery vehicles can be categorized into various types based on their structure, origin, mechanism of action, etc., and can be appropriately selected depending on the intended application target (e.g., plant cells, bacteria, etc.) and administration method.
[0256] In some embodiments, the vector may be a viral vector, including but not limited to adeno-associated virus (AAV), lentivirus, adenovirus, retrovirus, bacteriophage-based vector, and the like.
[0257] In other embodiments, the carrier may be a non-viral carrier, including, for example, lipid nanoparticles (LNPs), polymeric nanoparticles, cationic liposomes, lipofectins, peptide-based carriers, electroporation, or nanoneedle-based systems.
[0258] Additionally, in some embodiments for application to plants, the carrier may comprise a physical delivery means implemented by Agrobacterium tumefaciens strains, plant virus-based vectors, protoplast delivery systems, or gene gun technology.
[0259] A person skilled in the art can select and combine appropriate carriers based on known techniques, depending on the characteristics of a specific base editor system, the delivery route, and the type of cell or organism to which it is applied.
[0260] The delivery system according to the present invention is applicable to various cells and organisms and can be selectively adjusted according to the purpose. Specifically, the delivery target includes a eukaryotic cell or a prokaryotic cell, and in some embodiments, the delivery target may be a plant tissue, cell, or embryo.
[0261] 5. Base correction method
[0262] One aspect of the present invention relates to a method for correcting an adenine (A) base to a guanine (G) base, comprising introducing the DNA base editor of the present invention as described above into a cell containing target DNA for base correction, or expressing the DNA base editor within a cell containing target DNA for base correction. With respect to the “DNA base editor of the present invention,” the contents described in the “1. DNA base editor” section of the present specification are incorporated herein by reference.
[0263] In some embodiments, the method comprises introducing a DNA base editor, a base correction composition (e.g., comprising a polynucleotide(s) encoding a DNA base editor of the present invention), and a delivery vehicle (e.g., comprising a polynucleotide(s) encoding a DNA base editor of the present invention) into a cell containing target DNA for base correction, as described above.
[0264] The above method is a method of correcting the adenine (A) base in nuclear or organelle DNA to the guanine (G) base.
[0265] In some embodiments, the target DNA is nuclear DNA, and the method is a method of correcting an adenine (A) base in nuclear DNA to a guanine (G) base.
[0266] In some embodiments, the target DNA is chloroplast DNA, and the method is a method of correcting an adenine (A) base in chloroplast DNA to a guanine (G) base.
[0267] In some embodiments, the target DNA is mitochondrial DNA, and the method is a method of correcting an adenine (A) base in mitochondrial DNA to a guanine (G) base.
[0268] The above base correction method can be performed in vitro, ex vivo, or in vivo, and can be designed according to various application purposes such as research purposes and trait improvement purposes. In particular, the DNA base editor of the present invention can efficiently correct adenine bases in target DNA to guanine (A:T → G:C), and thus has wide applicability such as introduction of agriculturally useful traits and exploration of functional genes.
[0269] The components used in the base correction method of the present invention may include one or more of the DNA base editor described above, a polynucleotide encoding the same (e.g., mRNA or DNA), and a carrier (e.g., adeno-associated virus vector, lipid nanoparticle, etc.) containing the components. The components may be introduced into cells singly or in combination, and when present in the form of a fusion protein, stable expression and base correction activity can be provided through optimized binding between the component proteins. In addition, the components may additionally include an organelle targeting sequence (MTS, CTS, etc.), NES, or NLS, as needed.
[0270] The base correction method of the present invention is characterized by selectively correcting a specific base on DNA with another base, and the target of correction may be various, such as a mutation causing a genetic disease, an abnormal expression control region, or an artificial mutation for imparting a specific trait.
[0271] The method for expressing the DNA base editor or its components for implementing the base correction method of the present invention is not particularly limited and can be performed using technical means widely known to those skilled in the art. For example, a polynucleotide encoding a desired protein can be cloned into an appropriate expression vector, then introduced into a cell to induce transcription and / or translation, thereby causing expression. The expression can include both transient or stable expression within the cell, and can be implemented in various ways depending on the type of expression vector (e.g., plasmid, viral vector, etc.), the selection of promoter, the cell type, the introduction method, etc. Those skilled in the art can select and apply an appropriate expression system and conditions considering the desired cell type and base correction efficiency.
[0272] In the base correction method of the present invention, the DNA base editor or a composition comprising the same can be introduced into cells through various physical or chemical methods. For example, lipofection, electroporation, microinjection, viral vector delivery, nanoparticle delivery, Agrobacterium-mediated delivery, etc. can be used, and an appropriate delivery method can be selected depending on the type of cell being introduced, the target gene location, the target organism species, etc. In addition, the introduction conditions (e.g., pH, temperature, incubation time, introduction amount, etc.) can be easily optimized by a person skilled in the art, taking into account the desired base correction efficiency and cell viability.
[0273] The base correction method of the present invention is useful for correcting adenine (A) bases to guanine (G) bases. In some embodiments, the DNA base editor according to the present invention is effective in selectively performing A-to-G base corrections while reducing unwanted C-to-T base corrections.
[0274] In some embodiments, the base editing method of the present invention is suitable for A-to-G editing in nuclear DNA. In particular, it is suitable for selectively performing A-to-G editing while avoiding undesirable C-to-T editing in nuclear DNA.
[0275] In some embodiments, the base correction method of the present invention is suitable for A-to-G correction in chloroplast DNA. In particular, it is suitable for selectively performing A-to-G correction while avoiding undesired C-to-T correction in chloroplast DNA.
[0276] In some embodiments, the base correction method of the present invention is suitable for A-to-G correction in mitochondrial DNA. In particular, it is suitable for selectively performing A-to-G correction while avoiding undesirable C-to-T correction in mitochondrial DNA.
[0277] In some embodiments, the base correction method according to the present invention can reduce the frequency of C-to-T base correction by 50% or more, 60% or more, 70% or more, 80% or more, or 90% or more compared to a method using a base editor of the same configuration except that it does not include UDG. This is very useful for gene correction therapy, crop trait improvement, etc., where it is necessary to selectively correct only adenine bases while maintaining the wild-type sequence of cytosine bases.
[0278] 6. Corrective organisms
[0279] The DNA base correction composition of the present invention can be applied to the genomes of various eukaryotic cells, including plant cells. The target for correction includes not only nuclear DNA but also the genomes of cellular organelles such as mitochondrial DNA and chloroplast DNA.
[0280] The present invention is useful for regulating gene expression through substitution of specific bases in plant cells or protoplasts, or for producing plants having specific genotypes.
[0281] The DNA base editor or correction composition of the present invention can be introduced into plant cells or plant-derived protoplasts to selectively correct specific bases on a target base sequence. The base editor or correction composition can be directly introduced or can be delivered to plant cells, tissues, or protoplasts using techniques such as Agrobacterium-mediated delivery, PEG-mediated delivery, or gene gun, thereby inducing base correction of nuclear DNA, chloroplast DNA, or mitochondrial DNA. Target DNA may include nuclear DNA, mitochondrial DNA, and chloroplast DNA. In particular, when targeting chloroplast DNA, a chloroplast transit signal (CTS) and / or a nuclear export signal (NES) can be included in the fusion protein or transporter to induce movement into the chloroplast. Similarly, when targeting mitochondrial DNA, a mitochondrial targeting signal (MTS) can be added to facilitate target-specific organelle delivery.
[0282] Plant cells or protoplasts that have undergone base correction can be regenerated into whole plants under appropriate regeneration conditions. The regenerated plants contain the corrected base sequence, and the base correction is genetically stable and can be passed on to future generations. Progeny plants can be obtained through self-pollination or crossbreeding of the regenerated plants, and include F1, F2, and subsequent generations with the corrected genotype. These corrected plants and their progeny can be utilized in various fields, such as breeding, functional crop development, and molecular farming.
[0283] In some embodiments, the DNA base correction composition of the present invention can induce A-to-G base correction at a specific site within the plant cell genome. For example, an A present at a specific position in a wild-type sequence is converted to a G, which can serve as a useful tool for restoring abnormal gene sequences, regulating the expression of specific genes, or improving specific traits. The correction site can be located in the target DNA within the nucleus or various cellular organelles such as chloroplasts and mitochondria.
[0284] Plants obtained through the above DNA base editing technology can form seeds with corrected genotypes, and these seeds can stably transmit the corrected genotypes to future generations. The present invention encompasses such seeds, plants germinated from such seeds, and even tissues, cells, or biological byproducts obtained therefrom. Plants derived from such seeds maintain the corrected genotypes and can be continuously propagated into future generations with the base-edited traits in the same manner.
[0285] The base editing technology according to the present invention can be utilized as a powerful molecular biological tool for improving plant traits, modifying specific gene functions, or regulating target gene expression. For example, it can be applied to the development of functional crops or transgenic plants by targeting genes involved in the photosynthetic pathway, disease resistance genes, salt tolerance, or drought tolerance genes. In particular, chloroplast gene editing, which is inherited through cytoplasmic inheritance, offers advantageous characteristics in terms of trait stability and prevention of environmental transmission.
[0286] In some embodiments, plant cells, protoplasts, or plants containing a base sequence in which adenine (A) in the wild-type target DNA is substituted with guanine (G) can be generated. Such corrected cells or plants can be utilized in various plant biotechnology fields, such as functional analysis, metabolic regulation, crop improvement, and optimization of biosynthetic pathways.
[0287] In some embodiments, the base-correcting organism according to the present invention is a part of a plant selected from the group consisting of cereals, legumes, root vegetables, vegetables, fruits, and forage crops.
[0288] In some embodiments, the base-correcting organism according to the present invention is an algae or a part thereof.
[0289] In some embodiments, the base-correcting organism according to the present invention is a descendant or clone of a plant or algae, or a portion thereof.
[0290] The above “part of a progeny or clone” may include (i) a part of a cell or tissue derived from the progeny or clone, (ii) a part of a cell lineage or organ within the progeny or clone, or (iii) a part of an individual among the progeny or clone population having a corrected genotype. For example, when base correction is reflected only in some leaf tissue or flower tissue among plant tissues, or in some embodiments having a corrected genotype, the base correction organism according to the present invention is a seed obtained from a progeny or clone of the plant or algae.
[0291] In some embodiments, the base-correcting organism according to the present invention is a seed obtained from a progeny or clone of said plant or algae.
[0292] The present invention can be explained with specific embodiments exemplified below based on the above, but is not limited thereto.
[0293] The present invention can be explained with specific embodiments exemplified below based on the above, but is not limited thereto.
[0294] 1. A DNA base correction composition comprising a DNA binding protein, cytosine deaminase, adenine deaminase, and uracil DNA glycosylase (UDG), or polynucleotides encoding the proteins, wherein the cytosine deaminase exists in the form of a full-length or two fragments, and when it exists in the form of two fragments, it exhibits cytosine deaminase activity through dimerization thereof.
[0295] 2. In the above-mentioned first paragraph, the DNA binding protein, cytosine deaminase and adenine deaminase are present in the form of two fusion proteins,
[0296] The two fusion proteins above each independently comprise one DNA binding protein,
[0297] The two fusion proteins each contain two halves of cytosine deaminase,
[0298] At least one of the two fusion proteins comprises adenine deaminase,
[0299] A base correction composition wherein UDG exists independently and is not linked to the two fusion proteins.
[0300] 3. A base correction composition according to the above-mentioned first paragraph, wherein the DNA binding protein, cytosine deaminase, and adenine deaminase exist in the form of one fusion protein, and UDG exists independently without being linked to the fusion protein.
[0301] 4. A base correction composition according to any one of the above-mentioned claims 1 to 3, wherein the DNA binding protein is independently selected from the group consisting of a zinc finger protein, a TALE (transcription-activator-like effector) protein, and a CRISPR-associated nuclease.
[0302] 5. A base correction composition according to any one of the above-mentioned claims 1 to 4, wherein the cytosine deaminase is a double-strand DNA-specific cytosine deaminase.
[0303] 6. In any one of the above-mentioned clauses 1 to 5, the cytosine deaminase is DddA having the amino acid sequence of sequence number 1. tox Or a base correction composition which is a homolog thereof, or a variant thereof.
[0304] 7. A base correction composition according to any one of the above-mentioned claims 1 to 6, wherein the cytosine deaminase is included in the form of a first fragment and a second fragment, wherein the first fragment comprises a sequence from the N-terminus to the 33rd, 44th, 54th, 68th, 82nd, 98th or 108th amino acid in the amino acid sequence of SEQ ID NO: 1, and the second fragment comprises a sequence from the 34th, 45th, 55th, 69th, 99th or 109th amino acid in the amino acid sequence of SEQ ID NO: 1 to the C-terminus.
[0305] 8. A base correction composition according to the above-mentioned clause 7, wherein the first and second splitters each comprise amino acid sequences of SEQ ID NOs: 2 and 3, or amino acid sequences of SEQ ID NOs: 4 and 5.
[0306] 9. A base correction composition according to any one of the above-mentioned clauses 1 to 8, wherein the adenine deaminase is TadA or a variant thereof.
[0307] 10. A base correction composition according to any one of the above-mentioned claims 1 to 9, wherein the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 7.
[0308] 11. A base correction composition according to any one of the above-mentioned claims 1 to 10, wherein the UDG is derived from UDG of Arabidopsis thaliana.
[0309] 12. A base correction composition according to any one of the above-mentioned claims 1 to 11, wherein the UDG is derived from human UDG.
[0310] 13. In any one of the above-mentioned clauses 1 to 12, the DNA binding protein, cytosine deaminase and adenine deaminase are present in the form of two fusion proteins,
[0311] One of the two fusion proteins comprises a first fragment of a DNA binding protein and a cytosine deaminase,
[0312] The other of the two fusion proteins comprises a DNA binding protein, a second division of cytosine deaminase and an adenine deaminase,
[0313] A base correction composition wherein UDG exists independently and is not linked to the two fusion proteins.
[0314] 14. A base correction composition according to any one of the above-mentioned claims 1 to 13, wherein the DNA binding protein, cytosine deaminase, and adenine deaminase are present in the form of one fusion protein, and UDG is not linked to the two fusion proteins but exists independently.
[0315] 15. A base correction composition according to any one of the above-mentioned claims 1 to 14, wherein the DNA is DNA of a plant cell.
[0316] 16. A base correction composition according to any one of the above-mentioned clauses 1 to 14, wherein the DNA is DNA of an animal cell.
[0317] 17. A base correction composition according to any one of claims 1 to 16 mentioned above, wherein the DNA is nuclear DNA and optionally further comprises a nuclear localization signal (NLS).
[0318] 18. A base correction composition according to any one of claims 1 to 16 mentioned above, wherein the DNA is mitochondrial DNA and optionally further comprises a mitochondrial transfer signal (MTS) and / or a nuclear export signal (NES).
[0319] 19. A base correction composition according to any one of claims 1 to 14 mentioned above, wherein the DNA is chloroplast DNA and optionally further comprises a chloroplast transit signal (CTS) and / or NES.
[0320] 20. A base correction composition according to any one of the above-mentioned claims 1 to 19, which causes correction of an adenine (A) base to a guanine (G) base and substantially does not cause correction of a cytosine (C) base to a thymine (T) base.
[0321] 21. A plant cell or protoplast in which DNA base correction has been performed using a base correction composition according to any one of claims 1 to 15 and 17 to 20 mentioned above.
[0322] 22. A plant or part thereof grown or cultured from a plant cell or protoplast according to the above-mentioned Article 21.
[0323] 23. A plant or part thereof which is a descendant or clone of a plant according to the above-mentioned paragraph 30.
[0324] 24. Seeds obtained from the plant according to the above-mentioned item 22 or 23.
[0325] 25. A plant or its offspring, or a part thereof, grown from a seed according to the above-mentioned clause 24.
[0326] 26. A plant cell or protoplast according to the above-mentioned item 21, a plant or a part thereof according to item 22, 23 or 25, or a seed according to item 24, wherein the adenine (A) base in the wild-type target DNA sequence is corrected to a guanine (G) base.
[0327] In another aspect, the present invention can be described with specific embodiments exemplified below based on the above contents, but is not limited thereto.
[0328] 1. A method for correcting an adenine (C) base to a guanine (G) base, comprising introducing a DNA base editor into a cell containing target DNA for base correction or expressing the DNA base editor within a cell containing target DNA for base correction.
[0329] The DNA base editor comprises a DNA binding protein, cytosine deaminase, adenine deaminase and uracil DNA glycosylase (UDG), wherein the cytosine deaminase exists in a full-length form or in the form of two fragments, and when it exists in the form of two fragments, it exhibits cytosine deaminase activity through their dimerization.
[0330] Base editing method.
[0331] 2. In the above-mentioned first paragraph, the DNA binding protein, cytosine deaminase and adenine deaminase are present in the form of two fusion proteins,
[0332] The two fusion proteins above each independently comprise one DNA binding protein,
[0333] The two fusion proteins each contain two halves of cytosine deaminase,
[0334] At least one of the two fusion proteins comprises adenine deaminase,
[0335] UDG exists independently and is not linked to the two fusion proteins.
[0336] Base editing method.
[0337] 3. A base correction method according to the above-mentioned first paragraph, wherein the DNA binding protein, cytosine deaminase, and adenine deaminase exist in the form of a single fusion protein, and UDG exists independently without being linked to the fusion protein.
[0338] 4. A base correction method according to any one of the above-mentioned claims 1 to 3, wherein the DNA binding protein is independently selected from the group consisting of a zinc finger protein, a TALE (transcription-activator-like effector) protein, and a CRISPR-associated nuclease.
[0339] 5. A base correction method according to any one of the above-mentioned clauses 1 to 4, wherein the cytosine deaminase is a double-strand DNA-specific cytosine deaminase.
[0340] 6. In any one of the above-mentioned clauses 1 to 5, the cytosine deaminase is DddA having the amino acid sequence of sequence number 1. tox or its homolog, or its variant, a base-editing method.
[0341] 7. A base correction method according to any one of the above-mentioned claims 1 to 6, wherein the cytosine deaminase is included in the form of a first fragment and a second fragment, wherein the first fragment includes a sequence from the N-terminus to the 33rd, 44th, 54th, 68th, 82nd, 98th or 108th amino acid in the amino acid sequence of SEQ ID NO: 1, and the second fragment includes a sequence from the 34th, 45th, 55th, 69th, 99th or 109th amino acid in the amino acid sequence of SEQ ID NO: 1 to the C-terminus.
[0342] 8. A base correction method according to the above-mentioned clause 7, wherein the first and second segments each include the amino acid sequences of SEQ ID NOs: 2 and 3, or the amino acid sequences of SEQ ID NOs: 4 and 5.
[0343] 9. A base correction method according to any one of the above-mentioned clauses 1 to 8, wherein the adenine deaminase is TadA or a variant thereof.
[0344] 10. A base correction method according to any one of the above-mentioned claims 1 to 9, wherein the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 7.
[0345] 11. A base correction method according to any one of the above-mentioned clauses 1 to 10, wherein the UDG is derived from UDG of Arabidopsis thaliana.
[0346] 12. A base correction method according to any one of the above-mentioned claims 1 to 11, wherein the UDG is derived from human UDG.
[0347] 13. In any one of the above-mentioned clauses 1 to 12, the DNA binding protein, cytosine deaminase and adenine deaminase are present in the form of two fusion proteins,
[0348] One of the two fusion proteins comprises a first fragment of a DNA binding protein and a cytosine deaminase,
[0349] The other of the two fusion proteins comprises a DNA binding protein, a second division of cytosine deaminase and an adenine deaminase,
[0350] A base correction method in which UDG exists independently and is not linked to the above two fusion proteins.
[0351] 14. A base correction method according to any one of the above-mentioned claims 1 to 13, wherein the DNA binding protein, cytosine deaminase, and adenine deaminase are present in the form of one fusion protein, and UDG is not linked to the two fusion proteins but exists independently.
[0352] 15. A base correction method according to any one of the above-mentioned clauses 1 to 14, wherein the DNA is DNA of a plant cell.
[0353] 16. A base correction method according to any one of the above-mentioned clauses 1 to 14, wherein the DNA is DNA of an animal cell.
[0354] 17. A base editing method according to any one of the above-mentioned claims 1 to 16, wherein the DNA is nuclear DNA and the base editor optionally additionally includes a nuclear localization signal (NLS).
[0355] 18. A base correction method according to any one of the above-mentioned claims 1 to 16, wherein the DNA is mitochondrial DNA, and the base editor optionally further comprises a mitochondrial transfer signal (MTS) and / or a nuclear export signal (NES).
[0356] 19. A base correction method according to any one of the above-mentioned claims 1 to 14, wherein the DNA is chloroplast DNA and the base editor optionally further comprises a chloroplast transit signal (CTS) and / or NES.
[0357] 20. A base correction method according to any one of the above-mentioned items 1 to 19, which causes correction of an adenine (A) base to a guanine (G) base and substantially does not cause correction of a cytosine (C) base to a thymine (T) base.
[0358] 21. A plant cell or protoplast in which DNA base correction has been performed by a base correction method according to any one of the above-mentioned items 1 to 15 and items 17 to 20.
[0359] 22. A plant or part thereof grown or cultured from a plant cell or protoplast according to the above-mentioned Article 21.
[0360] 23. A plant or part thereof which is a descendant or clone of a plant according to the above-mentioned paragraph 30.
[0361] 24. Seeds obtained from the plant according to the above-mentioned item 22 or 23.
[0362] 25. A plant or its offspring, or a part thereof, grown from a seed according to the above-mentioned clause 24.
[0363] 26. A plant cell or protoplast according to the above-mentioned item 21, a plant or a part thereof according to item 22, 23 or 25, or a seed according to item 24, wherein the adenine (A) base in the wild-type target DNA sequence is corrected to a guanine (G) base.
[0364] Hereinafter, the present invention will be described in detail with reference to the following examples. However, the following examples are provided only to illustrate the present invention and the present invention is not limited thereto.
[0365]
[0366] Example
[0367] Transgenic plants were created by cloning DNA encoding nucleotide editors targeting the psaA gene in the chloroplast of Arabidopsis thaliana and transforming the plant using Agrobacterium. The composition of the nucleotide editors used is as follows, and the DNA sequences that bind via TALE proteins and the spacer regions where nucleotide editing takes place are as shown in Figures 1A and 2A.
[0368] Fusion protein numberFusion protein composition1CTS + 3xFlag + Light TALE + 4aa lin + 1397C + 16aa lin + AD2CTS + 3xFlag + Right TALE + 4aa lin + 1397N
[0369] 2aa lin and 16aa lin are linkers, and AD is adenine deaminase.
[0370] Specifically, genetic constructs encoding base editors (fusion proteins) as shown in the table above were designed to be positioned between the RPS5A promoter and the 35S terminator, and cloned using Gibson assembly, Golden Gate, restriction enzymes, etc. into a vector suitable for transforming Agrobacterium tumefaciens, a strain capable of delivering genetic constructs to the Arabidopsis nuclear genome using T-DNA.
[0371] Specifically, plasmids corresponding to target DNAs were selected from a set of TALE subarray plasmids consisting of a total of 424 (6 × 64 tripartite plasmids + 2 × 16 bipartite plasmids + 2 × 4 monopartite plasmids). These subarray plasmids encode repeat units required for TALE proteins to recognize specific DNA sequences and were designed to contain a BsaI restriction enzyme recognition site for Golden Gate assembly. Each repeat unit has specificity for a specific nucleotide (e.g., NI for A, HD for C, NN for G, NG for T) (Kim, Y. et al., 2013). The selected TALE subarray plasmids are cleaved with the Bsa I restriction enzyme and then ligated to form complementary sticky ends. This process results in the synthesis of a TALE array between the N-terminal and C-terminal domains.
[0372] The assembled gene sequence was finally inserted into a destination vector to generate a base editor plasmid targeting a specific sequence. The destination vector was prepared in advance through DNA synthesis and Gibson assembly, and can include components such as the RPS5A promoter, CTS sequence, 3XFlag tag, N-terminal domain, TALE repeat insertion site, C-terminal domain, DddAtox fragment, and UDG (uracil DNA glycosylase). This modular assembly method enables rapid and efficient construction of customized base editors for various DNA target sequences.
[0373] Since DddAtox is used in the form of two inactive fragments to avoid intracellular toxicity, two plasmids containing each TALE sequence were synthesized so that each inactive fragment can bind to the target DNA, and these two plasmids were ligated into a single plasmid using the Golden Gate method. The Golden Gate assembly method cleaves the SapI recognition site and allows for the ligation of complementary sticky ends. This allows the gene fragments to be efficiently inserted into the destination vector in the correct orientation and order. The final plasmid contains the spectinomycin resistance gene as a selectable marker and the right arm (RB) and left arm (LB) sequences of the T-DNA, which are essential for Agrobacterium-mediated plant transformation. This final plasmid has a structure that can express a complete base editor under the control of the RPS5A promoter. It is designed to perform base editing by targeting a specific DNA sequence in plant cells.
[0374] Meanwhile, the base editor system of the present invention, in which UDG is expressed separately, cloned AtUDG (SEQ ID NO: 8) into a separate vector, and then combined and cloned into one vector together with the fusion protein 1 and fusion protein 2 using the Golden Gate method (Fig. 7).
[0375] The finally obtained recombinant plasmid was transformed into Agrobacterium strain GV3101, and then Arabidopsis thaliana Colombia (Col-0) plants were transformed using the floral dipping technique according to the published method (Zhang et al., Nat. Protoc. 1, 641-646, 2006).
[0376] DNA was extracted from untransformed wild-type Col-0 plants and plants grown from first-generation seeds of transformed Arabidopsis, and sequences were analyzed using targeted deep sequencing. Base correction efficiency (frequency) was calculated as the percentage of sequencing reads that reflected the desired base correction among all sequencing reads.
[0377] As can be seen from the comparison of Figs. 1B and 2B, the TALED base editor resulted in a significant degree of C-to-T base correction, but when UDG was used together, almost no C-to-T base correction occurred, and only A-to-G base correction occurred. This suggests that the use of UDG is effective when only adenine base correction is desired using the TALED base editor. In addition, as confirmed from Fig. 2, even when UDG was expressed separately from the TALED base editor, the inhibition effect of C-to-T base correction was significantly observed. This result, which occurred even when UDG was not used as a part fused with the base editor, is an unexpected result.
[0378]
[0379] order
[0380] The amino acid sequences of the polypeptides used in the above examples are as follows.
[0381] CTS:
[0382] MDSQLVLSLKLNPSFTPLSPLFPFTPCSSFSPSLRFSSCYSRRLYSPVTVYAAK
[0383] 3xFlag:
[0384] DYKDHDGDYKDHDIDYKDDDDK
[0385] 1397N (DddAtox 1397N):
[0386] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSGSGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEG
[0387] 1397C (DddAtox 1397C):
[0388] AIPVKRGATGETKVFTGNSNSPKSPTKGGC
[0389] 4aa Links (4aa flax):
[0390] GSGS
[0391] 16aa Links (16aa flax):
[0392] SGSETPGTSESATPES
[0393] AD (TadA8e):
[0394] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN
[0395] AtUDG:
[0396] MASSTPKTLMDFFQPAKRLKASPSSSSFPAVSVAGGSRDLGSVANSPPRVTVTTSVADDSSGLTPEQIARAEFNKFVAKSKRNAVCSERVTKAKSEGNCYVPLSELLVEESWLKALPGEFHKPYAKSLSDFLEREIITDSKSPLIYPPQHLIFNALNTTPFDRVKTVIIGQDPYHGPGQAMGLSFSVPEGEKLPSSLLNIFKELHKDVGCSIPRHGNLQKWAVQGVLLNAVLTVRSKQPNSHAKKGWEQFTDAVIQSISQQKEGVVFLLWGRYAQEKSKLIDATKHHILTAAHPSGLSANRGFFDCRHFSRANQLLEEMGIPIDWQL
[0397] Left TALE:
[0398] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG
[0399] Right TALE:
[0400] DLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFFTAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQ LDTGQLLKIAKRGGVTAAVHAWRNALTGAPLNLTPEQVVAIASNNGGKQALETVQRLPVLCQAHGLTPDQVVAIAASNGGGKQALETVQRLLPVLCQDH GLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQA HGLTPDQVVAIASNNGGKQALETVQRLLPVLCQDHGLTPAQVVAIASNGGGKQALETVQRLPVLCQDHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQ AHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLPVLCQAHGLTPDQVVAIASNNGGKQALETVQRLLPVLC QDHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVL CQDHGLTPDQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNNGGKQALETVQRLLPV LCQDHGLTPAQVVAIASNNGGKQALETVQRLLPVLCQAHGLTPEQVVAIASNGGGKQALESIVQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLG
Claims
1. A method for correcting an adenine (C) base to a guanine (G) base, comprising introducing a DNA base editor into a cell containing target DNA for base correction or expressing the DNA base editor within a cell containing target DNA for base correction. The DNA base editor comprises a DNA binding protein, cytosine deaminase, adenine deaminase and uracil DNA glycosylase (UDG), wherein the cytosine deaminase exists in a full-length form or in the form of two fragments, and when it exists in the form of two fragments, it exhibits cytosine deaminase activity through their dimerization. Base editing method.
2. In paragraph 1, the DNA binding protein, cytosine deaminase and adenine deaminase exist in the form of two fusion proteins, The two fusion proteins above each independently comprise one DNA binding protein, The two fusion proteins each contain two halves of cytosine deaminase, At least one of the two fusion proteins comprises adenine deaminase, UDG exists independently and is not linked to the two fusion proteins. Base editing method.
3. A base correction method in claim 1, wherein the DNA binding protein, cytosine deaminase, and adenine deaminase exist in the form of a single fusion protein, and UDG exists independently without being linked to the fusion protein.
4. A base correction method according to claim 1, wherein the DNA binding protein is independently selected from the group consisting of a zinc finger protein, a TALE (transcription-activator-like effector) protein, and a CRISPR-associated nuclease.
5. A base correction method according to claim 1, wherein the cytosine deaminase is a double strand DNA-specific cytosine deaminase.
6. A base correction method in the first paragraph, wherein the DNA is nuclear DNA and the base editor optionally additionally includes a nuclear localization signal (NLS).
7. A base correction method according to claim 1, wherein the DNA is mitochondrial DNA, and the base editor optionally additionally includes a mitochondrial transfer signal (MTS) and / or a nuclear export signal (NES).
8. A base correction method according to claim 1, wherein the DNA is chloroplast DNA and the base editor optionally additionally includes a chloroplast transit signal (CTS) and / or NES.
9. A base correction method in paragraph 1 that causes correction of an adenine (A) base to a guanine (G) base and does not substantially cause correction of a cytosine (C) base to a thymine (T) base.
10. A plant cell or protoplast in which DNA base correction has been performed by the base correction method according to paragraph 1.
11. A plant or part thereof grown, cultured or propagated from a plant cell or protoplast according to Article 10.
12. Seeds obtained from the plant according to Article 11.
Citation Information
Patent Citations
Targeted deaminase and base editing using same
WO2022060185A1
Adenine-to-guanine base editing method of plant cell organelle DNA
WO2023182858A1