Single base editor using ADAR enzyme or variant thereof

The novel ADAR-Cas fusion protein guided by tagRNA enables precise single-base correction, overcoming bystander editing issues in conventional tools, enhancing gene editing efficacy.

WO2026014914A1PCT designated stage Publication Date: 2026-01-15SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/009942
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2025-07-09
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Conventional base correction tools face the challenge of bystander editing, where surrounding bases are inadvertently altered along with the target base, which can worsen genetic diseases and limit their effectiveness in gene therapies.

Method used

A novel base correction method using a fusion protein comprising an ADAR enzyme or its variant and a Cas protein, combined with a tagRNA, allows precise single-base correction by attaching the ADAR deaminase domain to the N-terminal side of a Cas9 nickase with specific mutations, guided by tagRNA to selectively replace only the designated base.

Benefits of technology

The method achieves high-efficiency single-base correction with minimal bystander effect, enabling effective gene editing in human cells and potential applications in gene therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025009942_15012026_PF_FP_ABST
    Figure KR2025009942_15012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a single base editor using adenosine deaminase acting on RNA (ADAR enzyme) and a single base editing method using same. Specifically, the present invention relates to a single base editor using an ADAR enzyme and a single base editing method using same, wherein the single base editor solves the problem of not only a target single base but also all single bases around the target single base being replaced, which is the biggest limitation of conventional base editing tools, and can selectively replace just a single base.
Need to check novelty before this filing date? Find Prior Art

Description

Single base proofreader using ADAR enzyme or its variant

[0001] The present invention relates to a single base corrector using an ADAR enzyme (Adenosine deaminase acting on RNA) or a variant thereof, and a single base corrector method using the same. Specifically, the present invention relates to a single base corrector using an ADAR enzyme or a variant thereof, which can selectively substitute only a single base, and a single base corrector method using the same, which solves the problem of replacing not only the target single base but also all single bases in the surrounding area, which is the biggest limitation of conventional base corrector tools.

[0002]

[0003] CRISPR (type II clustered regularly interspaced short palindromic repeats) constitute an adaptive immune system that provides acquired resistance to invading foreign nucleic acids in bacteria and archaea. It consists of an array of short conserved repeats separated by similarly sized, unique, variable DNA sequences called spacers derived from phage or plasmid DNA. The CRISPR-Cas system works by acquiring short foreign DNA fragments (spacers) that are inserted into the CRISPR domain, thereby providing immunity when subsequently exposed to phage and plasmids with matching sequences (Brouns SJJ, Jore MM, Lundgren M, Westra ER, Slijkhuis RJ, Snijders AP, et al. Small CRISPR RNAs guide antiviral defense in prokaryotes. Science. 2008;321:960-4).

[0004] The CRISPR (type II clustered regularly interspaced short palindromic repeats)-Cas9 system is a useful tool for genome modification. The CRISPR-Cas9 system uses the enzyme Cas9 to introduce DNA double-strand breaks (DSBs) into target DNA sequences complementary to a 20-nucleotide "spacer sequence" of a guide RNA (gRNA) noncovalently bound to the enzyme.

[0005] Meanwhile, a base editor (programmable deaminase) containing a DNA binding module and an adenine or cytidine deaminase enables targeted nucleotide substitution or base correction without generating DNA double-strand breaks.

[0006] Base editing technology, which allows for the substitution of target bases without cleaving DNA, has already been significantly developed. Liu's research team modified the existing CRISPR technology to develop a base editing technology capable of substituting target bases without cleaving DNA (Nature volume 533, pages 420-424 (2016)).

[0007] Unlike nucleases such as CRISPR-Cas9 and zinc-finger nucleases (ZFNs), which induce small insertions or deletions (indels) at the target site, deaminase enzymes convert cytosine (C) or adenine (A) at the target site.

[0008] By attaching cytosine deaminase or adenine deaminase to the existing CRISPR Cas9 nickase protein, it was implemented so that it can correct cytosine (C) → thymine (T) or adenine (A) → guanine (G), respectively, and is currently receiving a lot of attention because it can correct target bases with high efficiency (30-50%).

[0009] Thus, base editing technology has rapidly developed and become more sophisticated. However, a serious problem that remains unresolved is the problem of substituting other bases surrounding the target base (bystander editing).

[0010] In particular, when attempting to develop a strategy to treat genetic diseases through base correction, the problem of peripheral base substitutions is extremely serious, and can even worsen the disease. However, because deaminase enzymes operate freely within the exposed target window, it is extremely difficult to specifically substitute specific bases.

[0011] Current base editing tools have a very large operating window of several tens of nt, and correcting a single base is considered virtually impossible.

[0012] Under this technical background, the inventors of the present application have solved the problem of replacing not only the target single base but also all the single bases in the surrounding range (bystander effect), which is the biggest limitation of conventional base correction tools, and have confirmed that the base to be replaced can be freely controlled through a single base corrector using an ADAR enzyme or a variant thereof that can selectively replace only a single base freely from the bystander effect, and that a single base can be corrected using tagRNA (Target A Guide RNA), thereby completing the present invention.

[0013]

[0014] Summary of the invention

[0015] An object of the present invention is to provide a base correction composition comprising an ADAR (Adenosine deaminase acting on RNA) enzyme or a variant thereof and a tagRNA (Target A Guide RNA).

[0016] The purpose of the present invention is to provide a base correction method using the ADAR (Adenosine deaminase acting on RNA) enzyme or a variant thereof and tagRNA (Target A Guide RNA).

[0017] To achieve the above object, the present invention relates to a composition for base correction comprising a fusion protein comprising an ADAR (Adenosine deaminase acting on RNA) enzyme or a variant thereof and a Cas protein or a polynucleotide encoding the same; and a tagRNA (Target A Guide RNA) comprising a single guide RNA scaffold, a spacer, and a target annealing region.

[0018] The present invention relates to a base correction method comprising a step of processing a fusion protein comprising an ADAR (adenosine deaminase acting on RNA) enzyme or a variant thereof and a Cas protein or a polynucleotide encoding the same; and a tagRNA (Target A Guide RNA) comprising a single guide RNA scaffold, a spacer, and a target annealing region.

[0019]

[0020] Figure 1 is a schematic diagram showing (A) a conventional adenine substitution tool and (B) a new type of adenine substitution tool according to the present invention.

[0021] Figure 2 (A) Schematic diagram of sgRNA used in a conventional adenine substitution tool; (B) Schematic diagram of tagRNA used in a novel adenine substitution tool according to the present invention; and (C) Schematic diagram of the process of correcting human genes by combining a novel adenine substitution tool and tagRNA according to the present invention.

[0022] Figure 3 illustrates the results of single-base-level human genome editing using a novel base-editing tool and tagRNA. Figure 3a compares gene editing using the novel and conventional base-editing tools in the human HEK3 gene region, and Figure 3b compares gene editing using the human RNF2 gene region. Single adenines (A) designated using tagRNA are highlighted in red.

[0023] Figure 4 illustrates a novel base-editing tool with improved efficiency by linking MLH1dn. (A) Basic single-base editing tool. (B) Improved single-base editing tool with improved efficiency by linking the MLH1dn protein to the P2A peptide. (C) Schematic of the mismatch-repair system in operation in human cells using the basic novel single-base editing tool. (D) Schematic of successful gene editing by interfering with the mismatch-repair system using the improved novel single-base editing tool.

[0024] Figure 5 shows the successful single-base correction in human cells using a novel base-editing tool with improved efficiency by linking MLH1dn. Compared to the base form, the improved form showed superior efficiency (0.25% and 0.64%).

[0025] Figure 6 (A) Schematic diagram of a novel base-editing tool using various types of Cas proteins. (B) Data demonstrating that the tool can successfully perform single-base corrections in human cells.

[0026] Figure 7 shows (A) a novel base-editing tool that attaches ADAR deaminase to opposite sides of the Cas protein, and the results of single-base editing performed with it in human cells. A split-recruit system that brings the Cas protein and ADAR deaminase closer together, either by applying an antigen-antibody reaction (B) or an aptamer (C), and the results of single-base editing performed with it in human cells.

[0027] Figure 8 shows (A) a novel base-editing tool utilizing various Cas proteins from various species. (B) The results of single-base editing performed in human cells using the base-editing tool.

[0028] Figure 9 shows (A) a novel base-editing tool utilizing various ADAR deaminase enzymes from various species. (B) The results of single-base editing performed in human cells using this base-editing tool.

[0029] Figure 10 shows the results of verifying the base substitution efficiency using a single base correction tool produced by linking human-derived protein La(SSB) that prevents tagRNA degradation to a single base correction tool having a specific mutation (E438Q).

[0030] Figure 11 shows the base substitution efficiency observed in the AAVS1 gene of human-derived cells (HEK293T) using the improved single base correction tool according to the present invention.

[0031] [Correction pursuant to Rule 91, July 23, 2025][Deleted]

[0032]

[0033] Detailed description of the invention and preferred embodiments

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. In general, the nomenclature used herein is well known and commonly used in the art.

[0035] Conventional base editing tools, which attach adenine deaminase to the CRISPR-Cas9 nickase protein to correct adenine (A) to guanine (G), have the problem of bystander editing, which also replaces other bases surrounding the target base. If not only the disease-causing mutant adenine but also the surrounding normal adenines are replaced, not only will the treatment efficacy be significantly reduced, but it may even worsen the disease. Therefore, the application of existing adenine base editing tools to most gene therapies is extremely limited.

[0036] To overcome these problems, the newly designed adenine base correction tool of the present invention has a form in which the deaminase domain of the ADAR (Adenosine deaminase acting on RNA) enzyme is attached to the N-terminal side of nCas9 having the H840A mutation. The base correction tool can designate the base to be replaced (programmable) using tagRNA, a special type of RNA that is a modified version of the existing sgRNA. The innovative single base correction tool proposed in the present invention is not only free from the bystander effect because its operating range (window) is a single base, but also can freely control the base to be replaced through tagRNA (Target A Guide RNA) (programmable).

[0037] Using the base correction tool of the present invention, we experimentally confirmed that base correction occurred in actual human cells. Furthermore, in a gene where multiple adenines (A) are located close together, conventional adenine base correction tools replace all adenines, whereas the novel single-base correction tool of the present invention selectively replaces only a single designated target adenine (A). Furthermore, by applying various improvements to maximize efficiency, we demonstrated that a high level of base correction efficiency was achieved. This level is believed to be immediately applicable to gene therapy targeting actual patients.

[0038] Based on this, the present invention relates to a composition for base correction comprising a fusion protein comprising an ADAR (Adenosine deaminase acting on RNA) enzyme or a variant thereof and a Cas protein or a polynucleotide encoding the same; and a tagRNA (Target A Guide RNA) comprising a single guide RNA scaffold, a spacer, and a target annealing region.

[0039] In addition, the present invention relates to a base correction method comprising a step of processing a fusion protein comprising an ADAR (Adenosine deaminase acting on RNA) enzyme or a variant thereof and a Cas protein or a polynucleotide encoding the same; and a tagRNA (Target A Guide RNA) comprising a single guide RNA scaffold, a spacer, and a target annealing region.

[0040] The present invention includes a fusion protein comprising an ADAR (Adenosine deaminase acting on RNA) enzyme or a variant thereof and a Cas protein.

[0041] The above ADAR enzymes are deaminizing enzymes that recognize specific structural motifs in double-stranded RNA (dsRNA), bind to dsRNA, convert adenosine to inosine, and cause recoding of amino acid codons, potentially altering the encoded protein and its function. The nucleobases surrounding the editing site, particularly the nucleobases immediately 5' to the editing site and the triplets immediately 3' to the editing site, play a crucial role in adenosine deamination. Recruiting ADARs to specific sites in selected transcripts and deaminating adenosine regardless of adjacent nucleotides offers significant potential for disease treatment.

[0042] The above ADAR enzyme may be characterized as being derived from, but not limited to, a group consisting of, human (Homo sapiens), lice (Pediculus humanus), fruit fly (Drosophila melanogaster), squid (Doryteuthis opalescens), and nematode (Caenorhabditis elegans).

[0043] Specifically, the ADAR enzyme showed an effective effect not only in humans (Homo sapiens), but also in animals (Pediculus humanus), fruit flies (Drosophila melanogaster), squid (Doryteuthis opalescens), and nematodes (Caenorhabditis elegans).

[0044] The variant includes a polypeptide in which one or more amino acids are conservatively substituted and / or modified, thereby differing from the amino acid sequence of the variant before the mutation, but retaining functions or properties. Such variants can generally be identified by modifying one or more amino acids in the amino acid sequence of the polypeptide and evaluating the properties of the modified polypeptide. That is, the ability of the variant may be increased, unchanged, or decreased compared to the polypeptide before the mutation. Furthermore, some variants may include variants in which one or more portions, such as the N-terminal leader sequence or the transmembrane domain, are deleted. Other variants may include variants in which portions are deleted from the N- and / or C-terminus of the mature protein.

[0045] The ADAR enzyme variant according to the present invention may be characterized by being encoded by the sequence of SEQ ID NO: 1.

[0046] [Sequence number 1]

[0047]

[0048] AAATCCCCACCATCACCAATATCCACGCCACCAACGAGCCTCTGATGTACAACAAATGCAAAGAGAAGGCTACAGGATTTCAGGCCGCTAAAATGGCCCTGCTGAAGGGCTTTTCTAAGGCCAAGCTGGGCGAATGGATGAAAAAGCCCCCCGAGGTGGACCAGTTCGAGTGGGACGAGGACAGCTTGCTGAAACATATCATTTATAACAGCAGCTATGGAAGCATCAGC

[0049] In some cases, the ADAR enzyme variant may include a form in which 10 amino acids of the N-terminal portion of the ADAR enzyme of Pediculus humanus are truncated, including the E438Q mutation.

[0050] [Sequence number 10]

[0051]

[0052] Additionally, in addition to the E438Q mutation, the following mutations, D268E, A269K, A269L, A269N, A269Q, A269Y, A269N, D371H, G374L, N376S, P412T, P412Q, P412R, P412H, P412V, V478T, I496T, I536S, P558K, I597C, T603K, T603Q, N604Q, E614Q, E614F, K632R, E637C, E637F, E637H, E637L, E637M, E637Q, E637T, E637V, E637Y, E637F, E637N, were applied to increase efficiency. Can be.

[0053] It can be combined alone with the E438Q mutation, and two selected from the group consisting of D268E, A269K, A269L, A269N, A269Q, A269Y, A269N, D371H, G374L, N376S, P412T, P412Q, P412R, P412H, P412V, V478T, I496T, I536S, P558K, I597C, T603K, T603Q, N604Q, E614Q, E614F, K632R, E637C, E637F, E637H, E637L, E637M, E637Q, E637T, E637V, E637Y, E637F, E637N It can also be combined with the above.

[0054] The above Cas protein may include not only a wild-type Cas protein but also a variant thereof. The variant may be a mutant form in which an amino acid residue in the Cas protein is changed to any other amino acid.

[0055] The above Cas proteins are Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas13a, Cas13b, Cas13c, Cas13d, Cas14, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, CsMT2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, It may be, but is not limited to, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3 or Csf4.

[0056] The above Cas proteins are selected from the group consisting of Corynebacter, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus (Streptococcus pyogenes), Lactobacillus, Mycoplasma, Bacteroides, Flaviflaviivola, Flavobacterium, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus (Staphylococcus aureus), It may be derived from a genus of microorganisms containing an ortholog of a Cas protein selected from the group consisting of Nitratifractor, Corynebacterium and Campylobacter, and may be simply isolated or recombinant therefrom.

[0057] Specifically, the Cas protein may be, for example, a Cas9 protein variant or a Cas12a protein variant.

[0058] The above Cas9 may be derived from Streptococcus pyogenes or Staphylococcus aureus.

[0059] The above Cas9 protein variant may be a mutant form of Cas9 in which the catalytic aspartate residue in the RuvC domain of Cas9 is changed to any other amino acid (D10A). Or, it may be a mutant form of Cas9 in which the catalytic histidine residue in the HNH domain is changed to any other amino acid (H840A).

[0060] The above Cas9 protein variant may be nCas9 (Cas9 nickase) or dCas9 (catalytically-deficient Cas9) containing the following amino acid substitutions:

[0061] The above Cas9 protein variant may be a mutant form of Cas9 modified to have NGN-PAM.

[0062] (1) D10, H840, or D10 + H840;

[0063] (2) D1135, R1335, T1337, or D1135 + R1335 + T1337; or

[0064] (3) A61R, L1111R, D1135L, S1136W, G1218K, E1219Q, N1317R, A1322R, R1333P, R1335Q, T1337R or D1135L, S1136W, G1218K, E1219Q, R1335Q, T1337R;

[0065] (4) Includes all amino acid substitutions of (1), (2), and (3).

[0066] Specifically, the Cas9 protein variant may be nCas9 (Cas9 nickase) comprising a substitution of H840.

[0067] Preferably, the amino acid may be substituted with another amino acid, such as, but not limited to, alanine.

[0068] Among the above Cas9 nickases, a Cas9 nickase with improved efficiency may be included. The Cas9 nickase may include, for example, a combination of amino acid substitutions of R221K, N394K, and H840A.

[0069] The above Cas12a may be derived from Lachnospiraceae bacterium ND2006.

[0070] According to the present invention, it can be used in combination with various species of Cas proteins, not just SpCas9. It can be applied not only to Cas proteins of other species, such as SaCas9 derived from Streptococcus aureus, but also to other types of Cas proteins, such as LbCas12a derived from Lachnospiraceae bacterium ND2006.

[0071] The above fusion protein refers to an artificially synthesized protein comprising an ADAR enzyme and a Cas protein. The fusion protein or the polypeptide constituting the fusion protein can be produced using a chemical peptide synthesis method known in the art, or the gene encoding the domain can be amplified using PCR (polymerase chain reaction). Furthermore, the fusion protein can be produced by cloning an expression vector synthesized using a known method and then expressing it in a cell.

[0072] Conventional adenine base correction tools are in the form of attaching an adenine deaminase called TadA to the N-terminus of nCas9 with the D10A mutation (the front of Cas9 is referred to as N, and the back as C) (Figure 1-A). The newly designed adenine base correction tool of the present invention is in the form of attaching the deaminase domain of the ADAR (Adenosine deaminase acting on RNA) enzyme to the N-terminus of nCas9 with the H840A mutation (Figure 1-B).

[0073] The fusion protein comprising an ADAR enzyme and a Cas protein according to the present invention acts on a target, which may be, for example, DNA. The inventors of the present application have confirmed that, unlike conventional methods that apply ADAR enzymes to RNA targets, the fusion protein can act on a DNA target to achieve the desired base correction.

[0074] The present invention comprises a tagRNA (Target A Guide RNA) comprising a single guide RNA scaffold, a spacer, and a target annealing region.

[0075] In the above tagRNA (Target A Guide RNA), a single guide RNA scaffold includes a sequence portion that interacts with the Cas protein of the fusion protein.

[0076] The above spacer sequence is determined by considering the sequence of the target gene or target nucleic acid and the PAM sequence recognized by the Cas protein of the fusion protein.

[0077] The above tagRNA has a spacer linked to the 5' end of a single guide RNA scaffold and a target annealing region linked to the 3' end of the single guide RNA scaffold.

[0078] A newly designed form of adenine base correction tool uses a special form of RNA tagRNA that is a modified version of the existing sgRNA, making it possible to specify the base to be replaced (programmable).

[0079] The method for specifying the base to be replaced is as follows. Using tagRNA, a DNA-RNA double helix is ​​formed with the adenine (A) base to be corrected positioned in the middle. Cytosine (C), not uracil (U), is positioned as the base complementary to the adenine (A) base, thereby inducing an adenine (A)-cytosine (C) mismatch. Then, ADAR deaminase, which recognizes a single adenine (A) mismatch, recognizes it and replaces the mismatched single adenine (A) base (Figure 1-B).

[0080] The sgRNA used in existing adenine substitution tools consists of two parts: a spacer and an sgRNA scaffold (Figure 2-A). The tagRNA (Target A Guide RNA) used in the novel base correction tool of the present invention has a form in which the 3' end of the existing sgRNA is extended to include a target annealing region.

[0081] The tagRNA may have a spacer having 10 to 30 nt of bases, a single guide RNA scaffold having 70 to 80 nt of bases, and 18 to 120 nt of bases.

[0082] The tagRNA must have a base complementary to the adenine (A) base to be corrected, for example, and must be cytosine (C). Specifically, the tagRNA must have a sequence that can complementarily bind to a DNA strand containing the adenine (A) base to be corrected.

[0083] Specifically, the roles of each component of the tagRNA are as follows. (Fig. 2-C) The spacer portion consists of 20 bases and can be arbitrarily controlled to guide Cas9 to a specific gene region (Fig. 2-Bi). The sgRNA scaffold portion consists of 76 unique bases and forms a tagRNA structure to enable stable complexation with Cas9 (Fig. 2-B.ii). The target annealing region complementarily binds to the target gene DNA, forming a DNA-RNA duplex and intentionally creating a mismatch at the adenine (A) to be corrected, thereby inducing base substitution by ADAR deaminase.

[0084] In particular, since the target annealing region is effective over a wide range of lengths, such as 18 to 120 nt, a proportionally wide range of target adenines (A) can be specified.

[0085] In some cases, MLH1dn (dominant negative MLH1 protein) or a polynucleotide encoding it may be additionally included. In some cases, La(SSB) or a polynucleotide encoding it may be additionally included.

[0086] The above MLH1dn may be characterized by being encoded by the sequence of sequence number 2.

[0087] [Sequence number 2]

[0088]

[0089]

[0090] The above MLH1dn can be linked to the fusion protein by an internal ribosomal entry site (IRES) or a 2A peptide that causes the ribosome to skip forming a peptide bond upon expression.

[0091] The above 2A peptide may be, for example, one or more selected from the group consisting of P2A, T2A, E2A, and F2A. Specifically, the 2A peptide may be P2A. The P2A may be characterized by being encoded by the sequence of SEQ ID NO: 3.

[0092] [Sequence number 3]

[0093]

[0094] The above La(SSB) is a human-derived protein that prevents tagRNA degradation, and may be characterized by being encoded by the sequence of SEQ ID NO: 8.

[0095] [Sequence number 8]

[0096]

[0097]

[0098] The above MLH1dn or La(SSB) can be bound to the N-terminus or C-terminus of the fusion protein. Specifically, the above MLH1dn or La(SSB) can be bound to the C-terminus of the Cas protein in the fusion protein.

[0099] Because the base-editing tool uses Cas9 with the H840A mutation, it cleaves one strand of the DNA where the correction occurred (Figure 1B). Even if the base-editing tool operates normally, since the newly substituted base exists only on one strand, a base mismatch naturally occurs after the correction. When a break occurs in the DNA strand where the correction occurred, human cells interpret the base mismatch as an abnormal signal and activate the DNA repair system to remove it and restore it to its original state (Figure 4C). This is called mismatch repair (MMR).

[0100] This mismatch-repair mechanism can be a major factor in reducing the efficiency of new single-base editing tools. To improve this, we designed an improved single-base editing tool by linking a protein called MLH1dn to the end of the existing form (Figure 4-A) with a P2A peptide (Figure 4-B). Because the P2A peptide is cleaved during protein expression, this tool can achieve the same effect as expressing two separate proteins.

[0101] A composition according to the present invention comprises a nuclear localization sequence (NLS) in a fusion protein. Specifically, the NLS may be located at the N-terminus and / or C-terminus of the fusion protein.

[0102] The terms "polypeptide," "peptide," and "protein" are used interchangeably to refer to a polymer of amino acid residues, wherein the polymer can be conjugated to a moiety that is not comprised of amino acids. This term can apply to both naturally occurring and non-naturally occurring amino acid polymers, as well as to amino acid polymers in which one or more amino acid residues are artificial chemical mimics of the corresponding naturally occurring amino acids. A "fusion protein" can include a chimeric protein encoding two or more individual protein sequences that are recombinantly expressed as a single moiety.

[0103] "Amino acid" can refer to naturally occurring amino acids, synthetic amino acids, amino acid analogs and amino acid mimetics that function in a similar manner to naturally occurring amino acids. Naturally occurring amino acids include those encoded by the genetic code and those that have been later modified (e.g., hydroxyproline, γ-carboxyglutamate and O-phosphoserine). Amino acid analogs include compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., a hydrogen, a carboxyl group, an amino group and a carbon bonded to the R group, such as homoserine, norleucine, methionine sulfoxide and methionine methyl sulfonium. Such analogs may have modified R groups (e.g., norleucine) or modified peptide backbones but retain the same basic chemical structure as the naturally occurring amino acid. Amino acid mimetics include compounds that have a structure that differs from the general chemical structure of an amino acid but function in a similar manner to a naturally occurring amino acid. The terms "non-naturally occurring amino acid" and "unnatural amino acid" include amino acid analogs, synthetic amino acids and amino acid mimetics that are not found in nature.

[0104] “Polynucleotide,” “nucleic acid,” “nucleic acid molecule,” “nucleic acid oligomer,” “oligonucleotide,” “nucleic acid sequence,” and “nucleic acid fragment” are used interchangeably and may include, but are not limited to, nucleotides, deoxyribonucleotides, or ribonucleotides, or analogs, derivatives, or modifications thereof, in the form of polymers covalently linked to each other, which may have various lengths.

[0105] Different polynucleotides can have different three-dimensional structures and perform a variety of known or unknown functions. Polynucleotides can include, but are not limited to, genes, gene fragments, exons, introns, intergenic DNA (including but not limited to heterologous DNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA sequences, isolated RNA sequences, nucleic acid probes, and primers.

[0106] The above polynucleotide can generally include four nucleotide bases: adenine (A), cytosine (C), guanine (G), and thymine (T) (if the polynucleotide is RNA, thymine (T) is represented as uracil (U)).

[0107] A "sequence" is an alphabetical representation of a molecule, which can be entered into a database on a computer with a central processing unit and used in bioinformatics applications such as functional genomics and homology searching. A polynucleotide may optionally include one or more non-standard nucleotides, nucleotide analogs, and / or modified nucleotides.

[0108] A "nucleic acid" may comprise deoxyribonucleotides or ribonucleotides and polymers thereof or their complements in single, double, or multi-stranded form. A "polynucleotide" may comprise a linear sequence of nucleotides.

[0109] "Nucleotide" generally refers to a single unit of a polynucleotide. A nucleotide may be a ribonucleotide, a deoxyribonucleotide, or a modified form thereof. Examples of polynucleotides include single- and double-stranded DNA, single- and double-stranded RNA (including siRNA), and hybrid molecules containing a mixture of single- and double-stranded DNA and RNA.

[0110] The nucleic acid may be linear or branched. For example, the nucleic acid may be a linear chain of nucleotides, or the nucleic acid may be branched such that the nucleic acid comprises one or more nucleotide arms or branches.

[0111] Nucleic acids, including nucleic acids having a phosphothioate backbone, may contain one or more reactive moieties. The reactive moieties may include any group capable of reacting with another molecule, such as a nucleic acid or a polypeptide, through covalent, noncovalent, or other interactions. For example, the nucleic acid may contain an amino acid reactive moiety that reacts with an amino acid of a protein or polypeptide through covalent, noncovalent, or other interactions.

[0112] "Conservatively modified mutations" can be applied to nucleic acids. In the context of a specific nucleic acid sequence, a "conservatively modified mutation" refers to a nucleic acid that encodes the same or essentially the same amino acid sequence. Due to the degeneracy of the genetic code, multiple nucleic acid sequences encode a given protein. For example, the codons GCA, GCC, GCG, and GCU all encode the amino acid alanine. Therefore, at any position where a codon specifies alanine, the codon can be changed to any of the corresponding codons described without altering the encoded polypeptide. Such nucleic acid variations are called "silent mutations" and are a type of conservatively modified mutation. Any nucleic acid sequence encoding a polypeptide can also include all possible silent mutations of the nucleic acid. A skilled practitioner will recognize that each codon in a nucleic acid (except AUG, the sole codon for methionine, and TGG, the sole codon for tryptophan) can be modified to produce a functionally identical molecule.

[0113] When applied to nucleic acids or proteins, "isolation" indicates that the nucleic acid or protein is essentially free of other cellular components associated with it in its natural state. Purity and homogeneity can typically be determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high-performance liquid chromatography.

[0114] "Complementarity" or "complementary" refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence, either by the traditional Watson-Crick or other non-traditional type. For example, the AGT sequence is complementary to the TCA sequence. Complementarity refers to the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairs) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, and 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementarity, respectively). "Perfectly complementary" means that every adjacent residue in one nucleic acid sequence hydrogen bonds with the same number of adjacent residues in a second nucleic acid sequence. As used herein, "substantially complementary" means a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, and 95%. It may mean two nucleic acids that hybridize at 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% for a nucleotide region or under stringent conditions.

[0115] A "gene" is a segment of DNA involved in protein production, and may include the preceding and following coding regions (leader and trailer) and intermediate sequences (introns) between individual coding segments (exons). The leader, trailer, and introns contain regulatory elements necessary for the transcription and translation of a gene. Furthermore, a "protein gene product" includes the protein expressed from a specific gene.

[0116] The expression level of a DNA molecule within a cell can be determined based on the amount of corresponding mRNA present within the cell or the amount of protein encoded by that DNA produced by the cell. The expression level of a non-coding nucleic acid molecule (e.g., sgRNA) can be detected using standard PCR or Northern blot methods well known to those skilled in the art.

[0117] “CRISPR,” “CRISPR system,” and “CRISPR / Cas system” refer to molecular systems capable of performing genome editing, including modified versions thereof (e.g., systems incorporating dCas9). These include molecular systems capable of altering gene expression (e.g., transcription and / or translation).

[0118] "Altering" or "modifying" gene expression refers to any action or process capable of modulating the transcription and / or translation of a target nucleic acid (e.g., a gene). Thus, in one example, altering gene expression includes all transcriptional modulations, such as transcriptional activation and transcriptional repression. Altering gene expression includes translational activation and translational repression. In an embodiment, altering gene expression includes editing a nucleic acid sequence in genomic DNA. Thus, in embodiments, editing a nucleic acid sequence includes genomic editing. In embodiments, editing a nucleic acid sequence includes editing a sequence of non-genomic DNA or RNA (e.g., mRNA).

[0119] "Editing, proofreading, editing" refers to altering the DNA sequence of a genome. Genomic alteration can be accomplished by deleting portions of the genomic DNA sequence, inserting additional DNA sequences into the genome, or replacing portions of the genome with other DNA sequences. The present invention can specifically induce the conversion of a single nucleotide, adenine.

[0120] "Guide RNA" or "gRNA" refers to a ribonucleotide sequence capable of binding to a nuclear protein to form a nuclear protein complex. "Guide DNA" or "gDNA" refers to a deoxyribonucleic acid sequence capable of binding to a nuclear protein to form a deoxyribonucleoprotein complex by forming a deoxyribonucleoprotein sequence. The guide RNA may comprise one or more RNA molecules. The guide DNA may comprise one or more DNA molecules. The gRNA may comprise a nucleotide sequence complementary to a target site. The complementary nucleotide sequence can mediate binding of the ribonucleoprotein complex or deoxyribonucleoprotein complex to the target site, thereby providing sequence specificity for the ribonucleoprotein complex or deoxyribonucleoprotein complex. Therefore, the guide RNA or guide DNA comprises a sequence complementary to a target nucleic acid.

[0121] The guide RNA is complementary to a CRISPR nucleic acid sequence. The complement of the guide RNA or guide DNA comprises a sequence having about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the target nucleic acid (e.g., a regulator binding sequence).

[0122] The target sequence may be, for example, an endogenous nucleic acid sequence present in the cell. The target nucleic acid sequence may be, for example, an exogenous nucleic acid sequence.

[0123] 'PAM (protospacer adjacent motif)' is a sequence that is essential for Cas proteins to bind to target DNA, and refers to the sequence located following the target nucleic acid. Bacteria store part of the sequence of the invading virus in part of the bacterial genome, and this sequence is called a protospacer. Since the protospacer sequence is a part, the adjacent sequence becomes the original sequence of the bacteria. The role of the PAM site is to prevent the sequence from being cut if many sequences within the bacteria match the protospacer, but PAM is not attached next to it.

[0124] The optimal PAM sequence for Streptococcus pyrogenes Cas9 (SpCas9), the most commonly used CRISPR nuclease, is NGG. NAG and NGA can also be considered as PAM sequences.

[0125] The composition according to the present invention can be delivered to a cell containing a target nucleic acid via a delivery means.

[0126] The above delivery vehicle may be, for example, a viral delivery vehicle. The viral delivery vehicle may include a viral vector. The viral vector may include, but is not limited to, an Adeno-Associated Viral Vector (AAV), an Adenoviral Vector (AdV), a Lentiviral Vector (LV), or a Retroviral Vector (RV).

[0127] The delivery vehicle may be, for example, a non-viral delivery vehicle. Non-viral delivery vehicles may include, for example, episomal vectors containing viral replicons. Non-viral delivery vehicles may include delivery in the form of mRNA, ribonucleoproteins (RNPs), or carriers. Delivery in the form of mRNA may include synthetic modified mRNA or self-replicating RNA (srRNA). Carriers may include nanoparticles, cell-penetrating peptides (CPPs), or polymers.

[0128] Delivery means usable in the present invention may include, for example, vectors. A "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Examples thereof include nucleic acid molecules that are single-stranded, double-stranded, or partially double-stranded; nucleic acid molecules comprising one or more free ends, non-free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and various other polynucleotides known in the art.

[0129] For example, the vector is a "plasmid", which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, for example, by standard molecular cloning techniques.

[0130] Viral delivery vehicles usable in the present invention may include, for example, viral vectors. The viral vectors may include, for example, Adeno-Associated Viral Vector (AAV), Adenoviral Vector (AdV), Lentiviral Vector (LV), or Retroviral Vector (RV) (Acta Biochim Pol. 2005, 52(2): 285-291. and Biomolecules. 2020 Jun, 10(6): 839.).

[0131] For viral vectors, the virus-derived DNA or RNA sequences are present in the vector for packaging into the virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors comprise polynucleotides carried by a virus for transfection into a host cell.

[0132] In some cases, vectors can autonomously replicate in the host cell into which they are introduced (e.g., bacterial vectors with bacterial replicating Ori and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) integrate into the genome of the host cell upon introduction, thereby replicating along with the host genome.

[0133] Certain vectors can direct the expression of genes to which they are operably linked. Such vectors are referred to herein as "expression vectors." Common expression vectors useful in recombinant DNA technology are often in the form of plasmids.

[0134] A recombinant expression vector may contain a nucleic acid in a form suitable for expression of the nucleic acid in a host cell, meaning that the recombinant expression vector includes one or more regulatory elements that can be selected based on the host cell so that the recombinant expression vector is used for expression, i.e., is operably linked to the nucleic acid sequence to be expressed.

[0135] Within a recombinant expression vector, “operably linked” means that the nucleotide sequence of interest is linked to a regulatory element in a manner that permits expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell if the vector is introduced into the host cell).

[0136] "Regulatory elements" may include promoters, enhancers, internal ribosome entry sites (IRES), and other expression-regulating elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Regulatory elements include elements that direct inducible or constant expression of a nucleotide sequence in many types of host cells and elements that direct expression of a nucleotide sequence only in specific host cells (e.g., tissue-specific regulatory sequences). A tissue-specific promoter may direct expression primarily in a desired tissue of interest, such as muscle, neurons, bone, skin, blood, a specific organ (e.g., liver, pancreas), or a specific cell type (e.g., lymphocytes). Regulatory elements may also direct expression in a temporally-dependent manner, such as a cell-cycle-dependent or developmental stage-dependent manner, which may or may not be tissue- or cell-type-specific.

[0137] In some cases, the vector comprises one or more pol III promoters, one or more pol II promoters, one or more pol I promoters, or a combination thereof. Examples of pol III promoters include, but are not limited to, the U6 and H1 promoters. Examples of pol II promoters include, but are not limited to, the retrovirus Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer) (e.g., Boshart et al (1985) Cell 41:521-530), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter.

[0138] "Regulatory elements" may include enhancers, for example, WPRE; CMV enhancer; R-U5' segment in the LTR of HTLV-I; SV40 enhancer; and intronic sequences between exons 2 and 3 of rabbit β-globin. It will be appreciated by those skilled in the art that the design of the expression vector will depend on factors such as the choice of host cell to be transformed, the level of expression desired, and the like. The vector can be introduced into a host cell to produce a transcript, protein, or peptide, including a fusion protein or peptide encoded by a nucleic acid as described herein (e.g., a clustered regularly interspaced short palindromic repeats (CRISPR) transcript, protein, enzyme, mutants thereof, fusion proteins thereof, etc.). Beneficial vectors include lentiviruses and adeno-associated viruses, and the type of such vectors can also be selected to target specific types of cells.

[0139] "Vector" refers to a polynucleotide comprising a polynucleotide of the present invention, for example, one encoding a gRNA and / or a Cas9 protein. Vectors include, but are not limited to, plasmids, viral vectors, cosmids, artificial chromosomes, and phagemids. A vector is capable of replicating in a host cell and is characterized by one or more endonuclease restriction sites, which allow the vector to be cleaved and a desired nucleic acid sequence inserted therein.

[0140] The vector may contain one or more marker sequences suitable for use in the identification and / or selection of cells transformed or genetically modified, or not, with the vector. Markers include, for example, genes encoding proteins that increase or decrease either resistance or susceptibility to antibiotics (e.g., kanamycin, ampicillin) or other compounds, genes encoding enzymes (e.g., β-galactosidase, alkaline phosphatase, or luciferase) the activity of which is detectable by standard assays known in the art, and genes that visibly affect the phenotype of transformed or transfected cells, hosts, colonies, or plaques. Any vector, for example, a vector belonging to the pUC series, the pGEM series, the pET series, the pBAD series, the pTET series, or the pGEX series, is suitable for transformation of host cells (e.g., E. coli), mammalian cells such as CHO cells, insect cells, etc.) as encompassed by the present invention. In some embodiments, the vector is suitable for transformation of a host cell for production of a recombinant protein. Methods for selecting and manipulating vectors and host cells for expression of gRNA and / or proteins (e.g., those provided herein), transformation of cells, and expression / purification of recombinant proteins are well known in the art.

[0141] Vectors can be introduced and propagated in prokaryotes. In some embodiments, prokaryotes can be used to amplify copies of the vector to be introduced into eukaryotic cells or as intermediate vectors in the production of vectors to be introduced into eukaryotic cells (e.g., amplifying plasmids as part of a viral vector packaging system). Prokaryotes can be used to amplify copies of the vector and express one or more nucleic acids, for example, to provide a source of one or more proteins for delivery to a host cell or host organism. Expression of proteins in prokaryotes can be performed in E. coli containing vectors that contain constitutive or inducible promoters.

[0142] The above vectors can be delivered in vivo or into cells via electroporation, lipofection, viral vectors, nanoparticles, and PTD (Protein translocation domain) fusion protein methods, respectively.

[0143] When delivered in the form of mRNA, gene editing can be initiated more quickly than when delivered in the form of a DNA vector, as transcription into mRNA is unnecessary. Transient protein expression is also more likely.

[0144] In order to induce apoptosis via a vector, the following process must be followed: A cell is transfected with a plasmid containing a polynucleotide encoding a composition according to the present invention, the plasmid is translocated to the nucleus within the cell, where transcription begins. Then, mRNA for the nuclease is translocated outside the nucleus and translated into protein. The cleavage factor RNA is transcribed from the plasmid and translocated outside the nucleus. The nuclease and the cleavage factor RNA combine to form a ribonucleoprotein (RNP), which is then transported back to the nucleus.

[0145] Preformed RNPs can be introduced into cells to initiate editing of target genomic DNA. This RNP-based delivery method offers the advantage of reducing off-target effects compared to vector delivery and directly delivering nuclease and cleavage factor RNA to the cell nucleus.

[0146] To form ribonucleoproteins (RNPs), nucleases and cleavage factors can be synthesized and purified, respectively, to produce RNPs. In some cases, CL7-tagged nuclease and cleavage factor RNAs can be directly expressed via plasmids, and then the self-assembled RNPs can be purified.

[0147] The composition according to the present invention may be delivered via nanoparticles. With respect to nanoparticles, the composition according to the present invention may be delivered via polymer nanoparticles, metal nanoparticles, metal / inorganic nanoparticles, or lipid nanoparticles.

[0148] The components included in the composition according to the present invention can be used singly or in combination with the same or different components to deliver each of them.

[0149] The composition according to the present invention may be a pharmaceutical composition. In some cases, it may additionally include pharmaceutically acceptable excipients.

[0150] The formulation of the above pharmaceutical composition may be prepared by any known method or may be subsequently developed in the art of pharmacology. Typically, such a preparation method comprises the steps of combining the active ingredient(s) with excipients and / or one or more auxiliary ingredients, and then, if necessary and / or desirable, shaping and / or packaging the product into desired single- or multi-dose units.

[0151] The above formulations may further comprise pharmaceutically acceptable excipients, including any and all solvents, dispersion media, diluents, or other liquid vehicles, dispersion or suspension aids, surface active agents, isotonic agents, thickening or emulsifying agents, preservatives, solid binders, lubricants, and the like, as suitable for the particular dosage form desired as used herein.

[0152]

[0153] Hereinafter, the present invention will be described in more detail through examples. These examples are intended solely to illustrate the present invention, and it will be apparent to those skilled in the art that the scope of the present invention is not limited by these examples.

[0154]

[0155] Example 1: Design of base editing tools and tagRNA

[0156]

[0157] 1. Development of a new type of base editing tool with single nucleotide resolution using ADAR (Adenosine deaminase acting on RNA) enzyme or its variant.

[0158]

[0159] A new type of base correction tool was designed using ADAR (Adenosine deaminase acting on RNA) enzyme, and a schematic diagram comparing it with existing adenine base correction tools is described in (Fig. 1).

[0160] The existing adenine base correction tool is in the form of an adenine deaminase called TadA attached to the N-terminal side of nCas9 with the D10A mutation (the front of Cas9 is referred to as N, and the back as C) (Fig. 1-A). The adenine base correction tool according to the present invention is in the form of an adenosine deaminase acting on RNA (ADAR) enzyme deaminase domain attached to the N-terminal side of nCas9 with the H840A mutation (Fig. 1-B).

[0161] It may also include an ADAR (Adenosine deaminase acting on RNA) deaminase mutant from Pediculus humanus. The mutant has a mutation in which E (glutamic acid) is substituted with Q (glutamine) at position 438, and can be encoded by the sequence of SEQ ID NO: 1.

[0162] The structure contains the sequence of sequence number 4 as follows.

[0163] [Sequence number 4]

[0164]

[0165] The ADAR-nCas9 construct is composed as follows. The bpNLS sequences are located at both ends of the construct, which allow the construct to be localized inside the nucleus of the cell. The deaminase domain of the ADAR (adenosine deaminase acting on RNA) enzyme of Pediculus humanus is attached to the N-terminal portion of the construct. The deaminase domain has a mutation in which E (glutamic acid) is substituted with Q (glutamine) at position 438, and this mutation plays a role in increasing the proofreading efficiency of this tool. The C-terminal portion of the construct contains the Streptococcus pyrogenes Cas9 (SpCas9) with the H840A mutation. The linker connecting the two proteins described above is located in the center of the construct, and the linker is composed as follows: SGGSSGGS-SGSETPGTSESATPES(Xaa)-SGGSSGGS (SEQ ID NO: 5, where Xaa is XTEN).

[0166] A newly designed adenine base correction tool uses a special type of RNA that is a modified version of the existing sgRNA, tagRNA, to specify the base to be replaced (programmable). Using the tagRNA, a DNA-RNA double helix is ​​formed with the adenine (A) base to be corrected positioned in the middle. By positioning cytosine (C), not uracil (U), as the base complementary to the adenine (A), an adenine (A)-cytosine (C) mismatch is induced. Then, ADAR deaminase, which recognizes a single adenine (A) mismatch, recognizes it and replaces the mismatched single adenine (A) base (Figure 1-B).

[0167]

[0168] 2. Design of a programmable tagRNA that can specify the desired substitution base (Target A Guide RNA)

[0169]

[0170] The novel base correction tool according to the present invention is programmable to designate a desired substitution base when used with a specific type of RNA called tagRNA (Target A Guide RNA) (Fig. 2-C).

[0171] The sgRNA used in the existing adenine substitution tool is composed of two parts: a spacer and an sgRNA scaffold (Figure 2-A), but the tagRNA (Target A Guide RNA) used in the new base correction tool according to the present invention is in the form of a target annealing part added by extending the 3' side of the existing sgRNA (Figure 2-B).

[0172] The spacer portion of the tagRNA consists of 20 bases and can be arbitrarily adjusted by the user to guide Cas9 to a specific genomic region (Figure 2-Bi). The sgRNA scaffold portion consists of 76 unique bases and forms the tagRNA structure, enabling stable complexation with Cas9 (Figure 2-B.ii). The target annealing region complementarily binds to the target gene DNA, forming a DNA-RNA duplex and intentionally creating a mismatch at the adenine (A) to be corrected, thereby inducing a base substitution by ADAR deaminase.

[0173] The target annealing region is effective over a wide range of lengths, from 18 nt to over 100 nt, and a proportionally wide range of target adenines (A) can be specified.

[0174]

[0175] Example 2: Confirmation of human gene editing at the single-base level.

[0176]

[0177] Using the single-base editing tool of the present invention, the base substitution efficiency was observed in the HEK3 gene of human-derived cells (HEK293T). 0.5 x 105 HEK293T cells were seeded (DMEM, 10% FBS, 1% antibiotics) in a 48-well plate (SPL, #30048). After 24 hours, a total of 500 ng of plasmid (375 ng of the single-base editing tool or the existing base editing tool, 125 ng of tagRNA) was transfected into HEK293T using jetOPTIMUS® (Polyplus). After 36 hours, the medium was changed, and after 72 hours, the cells were harvested and gDNA was obtained using QuickExtract DNA Extraction Solution.

[0178] Afterwards, PCR was performed using the following pair of primers (AAVS1 target A: GGAGTTTTCCACACGGACAC, TTTCACTGATCCTGGTGCTG, AAVS1 target B: TACGAGACCAAGGTGCATCG, ACAGCCCAGGAGTCAGTAGT) and the PCR products were analyzed via NGS (Illumina® MiniSeq™).

[0179] The base substitution efficiency observed in the AAVS1 gene of human-derived cells (HEK293T) using the single base correction tool according to the present invention is shown in Fig. 3. It was confirmed that a human gene can be successfully edited at the single nucleotide level using a composition comprising the base correction tool according to the present invention and tagRNA (Fig. 3).

[0180] According to Figure 3, while existing base correction tools exhibit the problem of replacing adenine bases in the surrounding range including the target adenine (bystander effect), the new single base correction tool was confirmed to replace only the single target adenine designated by tagRNA with an efficiency of approximately 7.38%. (Figure 3-A)

[0181] Furthermore, by modifying and using tagRNA that matches the adenine base to be replaced, we confirmed that only a specific adenine base can be selectively substituted within the same position (Figure 3-B).

[0182]

[0183] Example 3: Production of a novel, improved single-base editing tool by linking MLH1dn.

[0184]

[0185] The novel single-base editing tool described above uses Cas9 with the H840A mutation, which causes a cleavage of one strand of DNA after the correction (Figure 1B). Even if the base editing tool operates normally, a base mismatch naturally occurs after the correction because the newly substituted base exists only on one strand. When a cleavage occurs in the corrected strand, human cells interpret the base mismatch as an abnormal signal and activate the DNA repair system to remove it and restore it to its original state (Figure 4C). This is called mismatch repair (MMR).

[0186] This mismatch-repair mechanism can be a major factor in reducing the efficiency of new single-base editing tools. To improve this, we designed an improved single-base editing tool by linking a protein called MLH1dn to the end of the existing form (Figure 4-A) with a P2A peptide (Figure 4-B). Because the P2A peptide is cleaved during protein expression, this tool can achieve the same effect as expressing two separate proteins.

[0187] The structure contains the sequence of sequence number 6 as follows.

[0188] [Sequence number 6]

[0189]

[0190] An improved form of ADAR-nCas9 construct using MLH1dn is composed as follows. It has MLH1dn attached to the C-terminal portion of the previously described ADAR-nCas9 construct. MLH1dn was previously invented by Liu's research team, and it was discovered that expression of MLH1dn can suppress the mismatch repair mechanism in cells. (Cell 184, 5635-5652.e29 (2021)) A P2A sequence is located between the attached MLH1dn protein and the ADAR-nCas9 construct, and the P2A sequence includes the following: GSGATNFSLLKQAGDVEENPGP (SEQ ID NO: 7)

[0191]

[0192] Example 4: Enhanced efficiency of single-base-level human gene editing using an improved single-base editing tool and tagRNA.

[0193]

[0194] Using the single-base editing tool of the present invention, the base substitution efficiency was observed in the HEK3 gene of human-derived cells (HEK293T). 0.5 x 105 HEK293T cells were seeded (DMEM, 10% FBS, 1% antibiotics) in a 48-well plate (SPL, #30048). After 24 hours, a total of 500 ng of plasmid (375 ng of the single-base editing tool or the existing base editing tool, 125 ng of tagRNA) was transfected into HEK293T using jetOPTIMUS® (Polyplus). After 36 hours, the medium was changed, and after 72 hours, the cells were harvested and gDNA was obtained using QuickExtract DNA Extraction Solution.

[0195] Afterwards, PCR was performed using the following pair of primers (AAVS1 target A: GGAGTTTTCCACACGGACAC, TTTCACTGATCCTGGTGCTG) and the PCR product was analyzed via NGS (Illumina® MiniSeq™).

[0196] The base substitution efficiency observed in the AAVS1 gene of human-derived cells (HEK293T) using the single base correction tool according to the present invention is shown in Fig. 5. Compared to the basic form of Example 1, single base correction was achieved at a higher efficiency in the improved form (8.83%, 10.64%).

[0197] In this way, using an improved form of single base editing tool and tagRNA, human genes were successfully edited at the single nucleotide level (Fig. 5).

[0198]

[0199] Example 5: A novel base editing tool platform applicable to various types of Cas proteins and various geometries.

[0200]

[0201] Using the single-base editing tool of the present invention, the base substitution efficiency was observed in the HEK3 gene of human-derived cells (HEK293T). 0.5 x 105 HEK293T cells were seeded (DMEM, 10% FBS, 1% antibiotics) in a 48-well plate (SPL, #30048). After 24 hours, a total of 500 ng of plasmid (375 ng of the single-base editing tool or the existing base editing tool, 125 ng of tagRNA) was transfected into HEK293T using jetOPTIMUS® (Polyplus). After 36 hours, the medium was changed, and after 72 hours, the cells were harvested and gDNA was obtained using QuickExtract DNA Extraction Solution.

[0202] Afterwards, PCR was performed using the following pair of primers (AAVS1 target A: GGAGTTTTCCACACGGACAC, TTTCACTGATCCTGGTGCTG) and the PCR product was analyzed via NGS (Illumina® MiniSeq™).

[0203] The novel base editing tool according to the present invention can be applied not only to the nickase type that cleaves only the DNA strand where proofreading has occurred, such as H840A described above, but also to various types of Cas, such as D10A (which cleaves one strand where proofreading has not occurred) and H840A / D10A (which does not cleave either DNA strand). (Fig. 6). The base substitution efficiencies observed in the AAVS1 gene region of derived cells (HEK293T) using various types of Cas are shown in Fig. 6-B. (The target adenine is labeled in red.) Single base substitution efficiencies of 7.38%, 1.98%, and 2.34% were confirmed for H840A, D10A, and H840A / D10A, respectively. (Fig. 6-B)

[0204] Also, as described above, the efficacy is still effective when ADAR deaminase is attached to a different orientation of the Cas protein, such as the C-terminus, rather than the N-terminus. In addition, a split-recruit system, which does not directly bind the proteins but expresses them separately and then uses aptamers or antigen-antibody reactions to bring the Cas protein and ADAR deaminase close together, is also effective (Fig. 7). The base substitution efficiencies observed in the AAVS1 gene region of derived cells (HEK293T) using various types of systems and strategies are shown on the right side of Fig. 7-A, B, and C. (The target adenine is labeled in red.) When ADAR deaminase was attached to the C-terminus of Cas, single base substitution efficiencies of 2.51%, 2.04%, and 2.38% were confirmed when the split-recruit system using antigen-antibody reaction or aptamer was used, respectively. (Right side of Fig. 7-A, B, C)

[0205]

[0206] Example 6. A novel base-editing tool platform applicable to various Cas proteins from various species.

[0207]

[0208] Using the single-base editing tool of the present invention, the base substitution efficiency was observed in the HEK3 gene of human-derived cells (HEK293T). 0.5 x 105 HEK293T cells were seeded (DMEM, 10% FBS, 1% antibiotics) in a 48-well plate (SPL, #30048). After 24 hours, a total of 500 ng of plasmid (375 ng of the single-base editing tool or the existing base editing tool, 125 ng of tagRNA) was transfected into HEK293T using jetOPTIMUS® (Polyplus). After 36 hours, the medium was changed, and after 72 hours, the cells were harvested and gDNA was obtained using QuickExtract DNA Extraction Solution.

[0209] Afterwards, PCR was performed using the following pair of primers (AAVS1: GGGCCGGTTAATGTGGCTCT, AGAGATGGCTCCAGGAAATG) and the PCR product was analyzed via NGS (Illumina® MiniSeq™).

[0210] The novel base editing tool according to the present invention can be used in combination with various Cas proteins, not just SpCas9. It can be applied not only to Cas proteins of other species, such as SaCas9 from Streptococcus aureus, but also to other types of Cas proteins, such as LbCas12a from Lachnospiraceae bacterium ND2006 (Fig. 8). The base substitution efficiencies observed in the AAVS1 gene region of derived cells (HEK293T) using various Cas species are shown in Fig. 8-B (target adenines are marked in red). Single-base substitution efficiencies of 7.67%, 3.72%, and 3.07% were confirmed for SpCas9, SaCas9, and LbCas12a, respectively (Fig. 8-B).

[0211] The novel base correction tool according to the present invention can utilize not only human-derived ADAR enzymes but also ADAR enzymes from other species. Effective effects were also observed using ADAR deaminases derived from humans (Homo sapiens), as well as from fish (Pediculus humanus), fruit flies (Drosophila melanogaster), squid (Doryteuthis opalescens), and nematode worms (Caenorhabditis elegans) (Fig. 9). The base substitution efficiencies observed in the AAVS1 gene region of derived cells (HEK293T) using ADAR enzymes from various species are shown in Fig. 9-B (target adenines are marked in red). Single nucleotide substitution efficiencies of 2.32%, 10.64%, 1.03%, 0.41%, and 0.82% were confirmed in humans (Homo sapiens), lice (Pediculus humanus), fruit flies (Drosophila melanogaster), squid (Doryteuthis opalescens), and nematode worms (Caenorhabditis elegans), respectively (Figure 9-B).

[0212]

[0213] Example 7: Production of a novel, improved single-base editing tool by linking the specific mutation E438Q of the ADAR protein of Pediculus humanus to the La(SSB) protein.

[0214] The novel single-base editing tool described above can be used with ADAR enzymes from other species, as well as the human ADAR protein. (Figure 9) By introducing a specific mutation (E438Q) that changes glutamic acid at position 438 to glutamine in the Pediculus humanus ADAR enzyme, a single-base editing tool with significantly improved efficiency was created. (Figure 10-A)

[0215] In addition, a single base editing tool with even improved efficiency was constructed by linking the human protein La(SSB), which prevents tagRNA degradation, to a single base editing tool with the aforementioned specific mutation (E438Q). (Figure 10-A)

[0216] The structure contains the sequence of sequence number 9 as follows.

[0217] [Sequence number 9]

[0218]

[0219] An improved form of ADAR-nCas9 construct that links a specific mutation (E438Q) in which the 438th glutamic acid of the Pediculus humanus ADAR enzyme is changed to glutamine, and the human protein La(SSB) that prevents tagRNA degradation is composed as follows. It has a form in which the La(SSB) protein is attached to the C-terminal part of the ADAR-nCas9 construct described above. The La(SSB) protein was invented by the Britt Adamson research team, and it was discovered that attaching La(SSB) can protect gRNA from degradation. (Nature volume 628, pages639-647 (2024)) A linker sequence is located between the attached La(SSB) protein and the ADAR-nCas9 construct, and the linker sequence is as follows: SGGSSGGSSGSETPGTSESATPESSGGSSGGS.

[0220] Using the single-base editing tool of the present invention, the base substitution efficiency in the AAVS1 gene of human-derived cells (HEK293T) was observed. 0.5 x 105 HEK293T cells were seeded (DMEM, 10% FBS, 1% antibiotics) in a 48-well plate (SPL, #30048). After 24 hours, a total of 500 ng of plasmid (375 ng of the single-base editing tool or the existing base editing tool, 125 ng of tagRNA) was transfected into HEK293T using jetOPTIMUS® (Polyplus). After 36 hours, the medium was changed, and after 72 hours, the cells were harvested and gDNA was obtained using QuickExtract DNA Extraction Solution.

[0221] Afterwards, PCR was performed using the following pair of primers (AAVS1 target A: GGAGTTTTCCACACGGACAC, TTTCACTGATCCTGGTGCTG) and the PCR product was analyzed via NGS (Illumina® MiniSeq™).

[0222] The base substitution efficiency observed in the AAVS1 gene of human-derived cells (HEK293T) using the improved single-base correction tool according to the present invention is shown in Fig. 10-B. Compared to the wild-type Pediculus humanus ADAR enzyme without mutation, the single-base correction tool with a specific mutation (E438Q) and the improved form linked to the La(SSB) protein achieved higher efficiency single-base correction (1.42%, 6.07%, 21.06%).

[0223]

[0224] Example 8: Development of a novel, improved single-base editing tool using the ADAR protein of Pediculus humanus.

[0225] The following work was performed on a novel single-base editing tool using the ADAR protein of Pediculus humanus to improve its shape and increase its efficiency. (Figure 1-A)

[0226] i) Ten amino acids from the N-terminal portion of the ADAR protein of this (Pediculus humanus) including the E438Q mutation were cut out.

[0227] ii) An improved linker was applied to the site connecting the ADAR protein and Cas9 protein. The linker sequence is as follows: (SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGS)

[0228] iii) Nickase Cas9 with improved efficiency (R221K, N394K, H840A) was used.

[0229] iv) Additional NLS (C-myc NLS) was applied.

[0230] v) La protein was linked.

[0231] vi) MLH1-small binder, which blocks the MMR pathway, was applied.

[0232]

[0233] Using the single-base editing tool of the present invention, the base substitution efficiency in the AAVS1 gene of human-derived cells (HEK293T) was observed. 0.5 x 105 HEK293T cells were seeded (DMEM, 10% FBS, 1% antibiotics) in a 48-well plate (SPL, #30048). After 24 hours, a total of 500 ng of plasmid (375 ng of single-base editing tool, 125 ng of tagRNA) was transfected into HEK293T using jetOPTIMUS® (Polyplus). After 36 hours, the medium was changed, and after 72 hours, the cells were harvested and gDNA was obtained using QuickExtract DNA Extraction Solution.

[0234] Afterwards, PCR was performed using the following pair of primers (AAVS1 target A: GGAGTTTTCCACACGGACAC, TTTCACTGATCCTGGTGCTG) and the PCR product was analyzed via NGS (Illumina® MiniSeq™).

[0235]

[0236] The base substitution efficiency observed in the AAVS1 gene of human-derived cells (HEK293T) using the improved single-base correction tool according to the present invention is shown in Figure 11B. Compared to previously proposed base correction tools, the improved base correction tool exhibited significantly higher correction efficiencies (5%, 20%).

[0237] Example 9: Mutations that enhance the efficiency of the ADAR protein of Pediculus humanus

[0238] The following mutations in the ADAR protein of Pediculus humanus, including the E438Q mutation, can increase efficiency. (Figure 12A)

[0239] - D268E, A269K, A269L, A269N, A269Q, A269Y, A269N, D371H, G374L, N376S, P412T, P412Q, P412R, P412H, P412V, V478T, I496T, I536S, P558K, I597C, T603K, T603Q, N604Q, E614Q, E614F, K632R, E637C, E637F, E637H, E637L, E637M, E637Q, E637T, E637V, E637Y, E637F, E637N

[0240] This mutation can be combined alone with the existing E438Q mutation (Figure 12B), or even in combination with multiple mutations, it shows a significant increase in efficiency (Figure 12C).

[0241]

[0242] Using the single-base editing tool of the present invention, the base substitution efficiency in the AAVS1 gene of human-derived cells (HEK293T) was observed. 0.5 x 105 HEK293T cells were seeded (DMEM, 10% FBS, 1% antibiotics) in a 48-well plate (SPL, #30048). After 24 hours, a total of 500 ng of plasmid (375 ng of single-base editing tool, 125 ng of tagRNA) was transfected into HEK293T using jetOPTIMUS® (Polyplus). After 36 hours, the medium was changed, and after 72 hours, the cells were harvested and gDNA was obtained using QuickExtract DNA Extraction Solution.

[0243] Afterwards, PCR was performed using the following pair of primers (AAVS1 target A: GGAGTTTTCCACACGGACAC, TTTCACTGATCCTGGTGCTG) and the PCR product was analyzed via NGS (Illumina® MiniSeq™).

[0244]

[0245] The base substitution efficiency observed in the AAVS1 gene of human-derived cells (HEK293T) using the single-base correction tool according to the present invention is shown in Figures 12B and 12C. Compared to an ADAR protein containing only the E438Q mutation, an efficiency increase of up to 4.5 times was observed when an ADAR protein with the additional mutation was used.

[0246]

[0247] Existing adenine base correction tools have a critical problem: they replace not only the target adenine but also all surrounding adenines (the bystander effect). If they replace not only the disease-causing mutant adenine but also the surrounding normal adenines, not only will the treatment's effectiveness be significantly reduced, but they could even worsen the disease. Therefore, the applicability of existing adenine base correction tools to most gene therapies is extremely limited.

[0248] The novel single-base correction tool according to the present invention overcomes the limitations of conventional adenine base correction tools. Because it can only substitute a single base, it can be applied to almost any genetic disease involving a single adenine mutation. Furthermore, because the target adenine to be corrected can be freely controlled at the same location via tagRNA, it is significantly superior to existing base correction tools with fixed correction ranges.

[0249] Another drawback of conventional adenine base correction tools is the off-target effect. This is a potentially fatal side effect, where the adenine base correction tool substitutes a base in a completely different, normal region, rather than the intended target. In severe cases, the base correction tool used to treat a disease can even cause the development of a new disease. The novel single-base correction tool of the present invention utilizes ADAR deaminase or a variant thereof, which specifically recognizes and substitutes single-base mismatches, thereby dramatically reducing the off-target effect. Since the tagRNA does not recognize its target at non-target DNA sites, the probability of unintended corrections is expected to be extremely low.

[0250]

[0251] While specific aspects of the present invention have been described in detail above, it will be apparent to those skilled in the art that these specific descriptions merely represent preferred embodiments and are not intended to limit the scope of the present invention. Therefore, the substantial scope of the present invention is defined by the appended claims and their equivalents.

[0252]

[0253] Electronic file attached.

Claims

1. A fusion protein comprising ADAR (Adenosine deaminase acting on RNA) enzyme or a variant thereof and a Cas protein or a polynucleotide encoding the same; and A composition for base correction comprising a tagRNA (Target A Guide RNA) comprising a guide RNA scaffold, a spacer, and a target annealing region.

2. A composition according to claim 1, characterized in that the ADAR enzyme is derived from a group consisting of human (Homo sapiens), lice (Pediculus humanus), fruit fly (Drosophila melanogaster), squid (Doryteuthis opalescens), and nematode (Caenorhabditis elegans).

3. In the first paragraph, the ADAR enzyme variant has the sequence of SEQ ID NO:

10. i) E438Q substitution; ii) E438Q substitution and truncation of 10 amino acids from the N-terminus; or iii) A composition characterized in that it is a variant further comprising an E438Q substitution and at least one substitution selected from the group consisting of: D268E, A269K, A269L, A269N, A269Q, A269Y, A269N, D371H, G374L, N376S, P412T, P412Q, P412R, P412H, P412V, V478T, I496T, I536S, P558K, I597C, T603K, T603Q, N604Q, E614Q, E614F, K632R, E637C, E637F, E637H, E637L, E637M, E637Q, E637T, E637V, E637Y, E637F and E637N 4. A composition according to claim 1, characterized in that the Cas protein comprises a Cas9 or Cas12a protein variant.

5. A composition according to claim 4, characterized in that the Cas9 protein is derived from Streptococcus pyogenes or Staphylococcus aureus.

6. In the fourth paragraph, the Cas9 variant comprises nCas9 (Cas9 nickase) or dCas9 (catalytically-deficient Cas9) comprising the following amino acid substitutions; or Composition characterized by SpGCas9 (NGN-PAM variants Cas9) or SpRYCas9 (NGN-PAM variants Cas9): (1) D10, H840, or D10 + H840; (2) D1135, R1335, T1337, or D1135 + R1335 + T1337; (3) A61R, L1111R, D1135L, S1136W, G1218K, E1219Q, N1317R, A1322R, R1333P, R1335Q, T1337R or D1135L, S1136W, G1218K, E1219Q, R1335Q, T1337R; or (4) R221K, N394K, H840A.

7. A composition according to claim 5, characterized in that the Cas12a protein is derived from Lachnospiraceae bacterium ND2006.

8. In the first paragraph, the tagRNA is a composition characterized in that the spacer is bound to the 5' end of the single guide RNA scaffold and the target annealing region is bound to the 3' end of the single guide RNA scaffold.

9. A composition according to claim 1, wherein the tagRNA has a spacer having bases of 10 to 30 nt, a single guide RNA scaffold having bases of 70 to 80 nt, and a base of 18 to 120 nt.

10. A composition according to claim 1, characterized in that it further comprises MLH1dn (dominant negative MLH1 protein) or a polynucleotide encoding it; or La (SSB) (Small RNA Binding Exonuclease Protection Factor La) or a polynucleotide encoding it.

11. A composition according to claim 10, wherein the MLH1dn is linked to a 2A peptide selected from the group consisting of P2A, T2A, E2A, and F2A in the fusion protein.

12. A fusion protein comprising ADAR (Adenosine deaminase acting on RNA) enzyme or a variant thereof and a Cas protein or a polynucleotide encoding the same; and A base editing method comprising the step of processing a tagRNA (Target A Guide RNA) comprising a guide RNA scaffold, a spacer, and a target annealing region.

Citation Information

Patent Citations

  • Method for forming Robust Common Spatial Patterns using Lp-norm for EEG based Motor Imagery BCI

    KR1020220051613A

  • Perfume induction device for beekeeping

    KR1020240000272A

  • Programmable DNA base editing by nme2CAS9-deaminase fusion proteins

    US20220290113A1

  • KR20240049138A