Multifunctional rna base editors with flexible editing windows

By designing the AIM tool, which utilizes circularizable guide RNA to bind with peptide constructs, the problem of inflexible RNA base editor editing windows was solved, enabling efficient and precise RNA base editing and expanding its application scope.

CN120932729BActive Publication Date: 2026-03-20PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing RNA base editors lack flexible and variable editing windows, making it difficult to meet the needs of different application scenarios for the number and location of edits, especially when multiple nucleotides need to be edited.

Method used

An AIM tool was developed that precisely regulates the RNA base editing window by designing a circularizable guide RNA (LF-gRNA) to bind to a peptide construct, enabling controllable number and location of base editing.

Benefits of technology

It enables efficient and precise RNA base editing within a user-defined editing window, expanding the application scope and flexibility of RNA base editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932729B_ABST
    Figure CN120932729B_ABST
Patent Text Reader

Abstract

The present application provides a RNA base editor with controllable editing window (e.g., adenine base editor, cytosine base editor, adenine / cytosine dual base editor) and its application. In addition, the present application also provides a series of novel TadA mutants for RNA base editing, and base editing tools based on the TadA mutants.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of base editing, in particular, the present application relates to a RNA base editor (e.g., adenine base editor, cytosine base editor, adenine / cytosine dual base editor) with controllable editing window and its application. In addition, the present application also provides a series of novel TadA mutants for RNA base editing, and a base editing tool based on the TadA mutants. BACKGROUND

[0002] 1. RNA base editing

[0003] Base editing on DNA or RNA is a powerful strategy that can be used for both basic research and clinical applications. Programmable RNA base editing can reversibly reprogram genetic information without changing the genomic DNA sequence, providing a safer and more flexible option compared to DNA base editing. In addition, RNA editing shows unique advantages in applications that require short-term and dynamic editing effects, such as signal transduction pathway regulation, pain relief, and viral infection treatment, etc.

[0004] Given such potential, a variety of different types of RNA base editors have been developed, which can respectively achieve A-to-I 1-5 , C-to-U 5,6 and U-to-Ψ 7-9 base conversion in mammalian cells. So far, editors based on exogenous or endogenous expression of RNA adenosine deaminase (ADAR) dominate the field of RNA base editing; on this basis, editing tools with higher editing activity and better specificity have been developed one after another, and their effectiveness has been verified from primary cell lines to animal models, which provides great hope for clinical application transformation. However, we still need new types of RNA base editors to expand the application range of RNA base editing.

[0005] 2. Concept of base editing window

[0006] A specific base editor can function within a given region of RNA or DNA and efficiently introduce point mutations, which is referred to as the "editing window" or "active window" 10-12 of the RNA or DNA base editor. In DNA base editing, the Cas9 complex unwinds double-stranded DNA (dsDNA) and forms a single-stranded DNA (ssDNA) region, which can be edited by APOBECs 13 and evolved tRNA adenosine deaminase (TadA) 14ssDNA deaminases are used and edits are performed. In this case, the editing window is usually 4-5 nucleotides wide, but can also be designed to be only 1-2 nucleotides wide. Different applications can require editing different numbers of nucleotides: while a narrow editing window is more suitable for correcting pathogenic SNPs, a wide editing window is required in cases where multiple nucleotides need to be edited in order to rescue gene function. For example, in the treatment of sickle cell disease and beta-thalassemia, disrupting the binding site (CCCCTTCCCC) of the transcriptional repressor LFR in the HBG promoter is an effective strategy to reactivate fetal hemoglobin (HbF) expression 15 . It has been shown that using a DNA cytosine base editor (CBE) to edit all 8 cytosines results in a significant reduction of LRF binding and higher HbF expression levels compared to editing only the 3' position of the tetra-cytosine 16 , which is consistent with the structural evidence that all 8 cytosines are important for LRF binding 17 . In a second example, we need to disrupt two consecutive oxidative activation sites (Met281 and Met282) on the CaMKII delta protein in order to prevent the oxidation and over-activation of CaMKII delta and provide additional cardiac protection 18-20 . To achieve this goal, we need to use a DNA adenine base editor (ABE) to edit both A841 and A844 sites of the CaMKII delta gene 21 . Therefore, the editing window is a very important parameter for DNA base editors.

[0007] In contrast, there is little discussion about the concept of editing window in the field of RNA base editing, because ADAR-based editing tools dominate in this field, which recognize double-stranded RNA rather than single-stranded RNA substrates. ADAR-based RNA editors prefer to act on a single nucleotide given mismatch at the target site, so ADAR is suitable for editing a single base; while consecutive base mismatches significantly reduce the editing efficiency of ADAR. Therefore, there is currently a lack of RNA base editors with flexible and variable editing windows.

[0008] 2. Single-stranded deaminase TadA and TadA-based DNA base editors

[0009] TadA (tRNA adenosine deaminase A) is a deaminase that can catalyze the deamination of adenine (A) at a specific position of tRNA to inosine (I) 22 . In Escherichia coli, TadA can specifically catalyze the deamination of tRNA Arg2The adenine at position 34 is deaminated to hypoxanthine, and this position corresponds to the third position of the codon. Since hypoxanthine can pair with A, C, or U, the anticodon of the same tRNA molecule can recognize three codons after the deamination reaction. Structural biology studies have shown that the activation of TadA deamination activity is based on the recognition of the stem-loop structure of the anticodon loop of the tRNA molecule and specific nucleotides by TadA. Only tRNA with a specific structure can be catalyzed by TadA to deaminate, which ensures the accuracy of TadA-catalyzed deamination.

[0010] The David R. Liu group first developed an adenine base editor (ABE) that can catalyze the deamination of adenine at a specific site on DNA to hypoxanthine using TadA 14 . Since there is no natural deaminase in nature that can deaminate single-stranded DNA, researchers linked wild-type Escherichia coli-derived TadA (EcTadA) to Cas9 protein to form a fusion protein, and through bacterial selection, EcTadA evolved the ability to deaminate single-stranded DNA after introducing several amino acid mutations. The optimized version ABE7.10 is composed of Cas9 protein, evolved TadA*, and wild-type wt-EcTadA, which can achieve high-efficiency and accurate adenine base editing on DNA.

[0011] Based on ABE7.10, the David R. Liu group obtained the ABE8e tool with significantly improved editing efficiency compared to ABE7.10 through phage-assisted non-continuous and continuous evolution (PANCE and PACE) technology 23 . At the same time, the study also noticed that with the improvement of editing efficiency, ABE8e would cause relatively more adenine off-target editing on DNA and RNA transcripts. Subsequent structural biology studies revealed the structure of ABE8e capturing single-stranded DNA substrates, and confirmed that ABE8e and miniABE-max both have catalytic activity for deamination of adenine on single-stranded RNA 24 . Subsequent work has also introduced amino acid mutations into TadAmax and TadA8e to reduce adenine off-target editing of miniABE-max and ABE8e on DNA and RNA transcripts 25-28Beam Therapeutics’ researchers further optimized the bacterial selection strategy employed in the evolution of ABE7.10, and obtained a series of ABE8s by continuing to direct molecular evolution of ABE7.10, among which the best-performing tool is ABE8.20 29 Compared with ABE7.10, ABE8.20 has a wider editing window and higher editing efficiency (~1.5-4.2-fold improvement), but also causes more off-target editing on the genome and transcriptome.

[0012] 3. TadA-based A-to-I base editing tool on RNA

[0013] In July 2024, Weixin Tang’s team published a research online, which realized programmable A-to-I base editing on RNA using TadA deaminase DECOR 30 This is the first RNA A-to-I base editor that does not rely on ADAR deaminase. DECOR is obtained by linking a highly active TadA8e (this TadA variant has been reported in DNA base editors 23 ) with dRfxCas13d to form a fusion protein. DECOR can achieve about 20%-50% A-to-G conversion efficiency on endogenous transcripts, but also causes a large number of bystander editing on targeted transcripts and 840-880 whole-transcriptome off-target editing. SUMMARY

[0014] Based on in-depth research, the inventors of the present application have developed a new RNA base editing toolkit, named AIM, which can be used to achieve controllable number and location of RNA base editing within a user-defined editing window. Specifically, AIM induces a loop structure on the target RNA by using a specially designed guide RNA (the guide RNA is named loop-forming guide RNA or “LF-gRNA”), so that by adjusting the size of the loop, the editing window of AIM can be accurately regulated Figure 1 A and 1B).

[0015] Based on this, in order to further expand the application range of the AIM tool, the inventors of the present application rationally designed single-stranded deaminase TadA and developed a series of novel mutant-based AIM tools that can achieve precise and efficient A-to-I, C-to-U or simultaneous A&C editing.

[0016] I. Base editing system based on loop-forming guide RNA

[0017] Therefore, in a first aspect, the present application provides a composition comprising:

[0018] (a) a first component selected from the group consisting of: (i) a guide RNA molecule comprising a backbone sequence and a guide sequence capable of hybridizing to a target RNA sequence comprising a base to be edited; (ii) a nucleotide sequence encoding the guide RNA molecule; (iii) any combination of (i) and (ii);

[0019] and,

[0020] (b) a second component selected from the group consisting of: (i) a polypeptide construct comprising a base deaminase and a guide RNA molecule binding domain linked to the base deaminase, the guide RNA molecule binding domain being capable of binding to the backbone sequence of the guide RNA molecule; (ii) a nucleotide sequence encoding the polypeptide construct; (iii) any combination of (i) and (ii);

[0021] wherein the guide sequence of the guide RNA molecule is designed to cause the target RNA sequence to form a single-stranded RNA loop region comprising the base to be edited upon hybridization to the target RNA sequence.

[0022] Based on the disclosure herein, one of skill in the art would readily appreciate that, in embodiments where the first component is a nucleotide sequence encoding the guide RNA molecule and the second component is a nucleotide sequence encoding the polypeptide construct, the nucleotide sequence of the first component and the nucleotide sequence of the second component can be present in different nucleic acid molecules or vectors, respectively, or in the same nucleic acid molecule or vector.

[0023] In certain embodiments, the target RNA sequence comprises: a target loopable region comprising the base to be edited, a first region located 5' to the target loopable region, and a second region located 3' to the target loopable region;

[0024] the guide sequence of the guide RNA molecule is designed to comprise a first region, a second region, and optionally a third region located between the first region and the second region, the second region being located upstream of the first region; and the first region is capable of complementarity to the first region of the target RNA sequence, the second region is capable of complementarity to the second region of the target RNA sequence, and the third region is absent or the third region is not capable of complementarity to the target loopable region of the target RNA sequence;

[0025] thereby, upon hybridization of the guide RNA molecule to the target RNA sequence, the target loopable region of the target RNA sequence located between the first region and the second region is present as a single strand to form a single-stranded RNA loop region comprising the base to be edited.

[0026] Further, since the polypeptide construct of the present application comprises a guide RNA molecule binding domain capable of recognizing and binding to a guide RNA molecule linked to a base deaminase, the guide RNA molecule can guide the polypeptide construct comprising a base deaminase to a target RNA, thereby causing deamination of the base to be edited in the single-stranded RNA loop region of the target RNA.

[0027] In certain embodiments, the third region of the guide RNA molecule is absent.

[0028] In certain embodiments, the third region of the guide RNA molecule has a number of nucleotide residues less than that of the target RNA sequence in the target loop-forming region.

[0029] In certain embodiments, the first region and the second region of the guide RNA molecule have the same or different number of nucleotide residues.

[0030] Based on the disclosure of the present application, one skilled in the art would readily understand that the nucleotide composition of the single-stranded RNA loop region formed by the target RNA sequence hybridized with the guide RNA molecule can be regulated by the design of the guide RNA molecule. Specifically, according to the disclosure of the present application, the target loop-forming region between the first region and the second region of the target RNA sequence corresponds to the single-stranded loop region formed after the target RNA sequence hybridizes with the guide RNA, and thus, the first region and the second region can be selected on both sides of the target loop-forming region of the target RNA sequence (the target loop-forming region with a target nucleotide composition; for example, in the target loop-forming region, the base to be edited is located at a target position and / or the target loop-forming region has a selected number of nucleotide residues) according to specific editing requirements, and a guide RNA molecule comprising a first region and a second region capable of complementarily binding to the first region and the second region, respectively, is designed based on this, so that the target RNA sequence hybridizes with the guide RNA molecule to form a single-stranded RNA loop region with a target nucleotide composition. Thus, the compositions provided by the present application can achieve regulated base editing with a self-defined editing window (i.e., the number of residues in the editing window can be regulated, and the position of the base to be edited in the editing window can be regulated).

[0031] In certain embodiments, the guide RNA molecule binding domain is selected from the group consisting of: a Cas effector protein of a CRISPR system lacking nuclease activity, a MS2 bacteriophage coat protein (MCP), a Bacteriophage Lambda N protein (a N-peptide).

[0032] Based on the disclosure herein, one of skill in the art would readily appreciate that the design of the guide RNA molecule backbone sequence in the compositions of the present application corresponds to the selection of the guide RNA molecule binding domain of the polypeptide construct, i.e., the guide RNA molecule backbone sequence is designed to be recognized and bound by the guide RNA molecule binding domain comprised by the polypeptide construct. For example, when the guide RNA molecule binding domain is a Cas effector protein of a CRISPR system lacking nuclease activity, the guide RNA molecule backbone sequence is selected from the group consisting of the backbone sequence of a crRNA (or a crRNA / tracrRNA complex or construct) that is recognized and bound by the Cas effector protein.

[0033] In certain embodiments, the guide RNA molecule binding domain is selected from the group consisting of: a CRISPR Cas13 effector protein lacking nuclease activity. In certain embodiments, the guide RNA molecule backbone sequence is selected from the group consisting of the backbone sequence of a crRNA that is recognized and bound by the CRISPR Cas13 effector protein.

[0034] In certain embodiments, the guide RNA molecule binding domain is selected from the group consisting of a MS2 bacteriophage coat protein (MCP). In certain embodiments, the guide RNA molecule backbone sequence is selected from the group consisting of a MS2 aptamer that is recognized and bound by the MS2 bacteriophage coat protein (MCP).

[0035] As generally understood by one of skill in the art, a MS2 aptamer sequence is an RNA sequence derived from bacteriophage MS2 that is capable of forming a stem-loop structure via intramolecular base interactions that is specifically recognized and bound by the MS2 coat protein (MCP). Various natural or artificially modified mutant MS2 aptamer sequences and the corresponding sequences of their MCP proteins are known in the art. Any MS2 aptamer and the corresponding MCP protein thereof can be used to practice the embodiments of the present application.

[0036] In certain embodiments, the guide RNA molecule binding domain is selected from the group consisting of a Bacteriophage Lambda N protein. In certain embodiments, the guide RNA molecule backbone sequence is selected from the group consisting of a boxB sequence that is recognized and bound by the Bacteriophage Lambda N protein.

[0037] As generally understood by one skilled in the art, a boxB sequence (or boxB element) refers to a class of RNA motifs that can form a stem-loop structure within the nut site (N utilization site) in the genome of a bacteriophage, which stem-loop structure can be recognized and bound by the bacteriophage Lambda N protein.

[0038] In certain embodiments, the guide RNA molecule binding domain is a PspCas13b that lacks nuclease activity. In certain embodiments, the guide RNA molecule binding domain comprises an amino acid sequence as set forth in SEQ ID NO: 10. In certain embodiments, the backbone sequence of the guide RNA molecule comprises an amino acid sequence as set forth in SEQ ID NO: 13.

[0039] The present application is not limited to the number of bases to be edited in a target RNA sequence. Without being limited by theory, for a given target RNA sequence and the composition provided herein designed therefor, the bases to be edited in the target RNA sequence within the editing window are theoretically all subjected to the catalysis of the base deaminase and thus converted to another base.

[0040] In certain embodiments, the single-stranded RNA loop region formed after the hybridization of the target RNA sequence and the guide RNA molecule comprises one or more bases to be edited. In certain embodiments, the single-stranded RNA loop region formed after the hybridization of the target RNA sequence and the guide RNA molecule comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases to be edited.

[0041] In certain embodiments, the bases to be edited are selected from the group consisting of adenine, cytosine, and any combination thereof.

[0042] In certain embodiments, the first region of the guide RNA molecule consists of 3-50 (e.g., 5-50, 5-40, 5-30, 5-25, 5-20, 5-18, 5-15, 10-50, 10-40, 10-30, 10-25, 10-20, 10-18, 12-50, 12-40, 12-30, 12-25, 12-20, 12-18) nucleotide residues.

[0043] In certain embodiments, the first region of the guide RNA molecule consists of 3-50 (e.g., 5-50, 5-40, 5-30, 5-25, 5-20, 5-18, 5-15, 10-50, 10-40, 10-30, 10-25, 10-20, 10-18, 12-50, 12-40, 12-30, 12-25, 12-20, 12-18) nucleotide residues.

[0044] In certain embodiments, the single-stranded RNA loop region formed upon hybridization of the target RNA sequence to the guide RNA molecule consists of at least 3, at least 4, at least 5, or at least 6 nucleotide residues. In certain embodiments, the single-stranded RNA loop region formed upon hybridization of the target RNA sequence to the guide RNA molecule consists of 3-30 (e.g., 6-30, 6-25, 6-20, 6-19, 6-18, 6-17) nucleotide residues. In certain embodiments, the single-stranded RNA loop region formed upon hybridization of the target RNA sequence to the guide RNA molecule consists of 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17 nucleotide residues.

[0045] In certain embodiments, the base to be edited in the target RNA sequence is separated from the first region by at least 1, at least 2, at least 3, at least 4, or at least 5 nucleotide residues, and / or, is separated from the second region by at least 1, at least 2, at least 3, at least 4, or at least 5 nucleotide residues.

[0046] In certain embodiments, the base deaminase is an adenosine deaminase or a cytidine deaminase, or the base deaminase comprises an adenosine deaminase and a cytidine deaminase or the base deaminase has both adenosine deaminase activity and cytidine deaminase activity.

[0047] In certain embodiments, the base to be edited is adenine and the base deaminase is an adenosine deaminase.

[0048] In certain embodiments, the base to be edited is cytosine and the base deaminase is a cytidine deaminase.

[0049] In certain embodiments, the base to be edited is adenine, cytosine, or a combination thereof, and the base deaminase comprises an adenosine deaminase and a cytidine deaminase or the base deaminase has both adenosine deaminase activity and cytidine deaminase activity.

[0050] In certain embodiments, in the polypeptide construct, the guide RNA molecule binding domain is linked to the N-terminus and / or the C-terminus of the base deaminase.

[0051] In certain embodiments, the polypeptide construct further comprises a nuclear localization signal (NLS) and / or a nuclear export signal (NES). In certain embodiments, the polypeptide construct further comprises one or more NLS and / or one or more NES. In certain embodiments, the NLS comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 11, 22-23. In certain embodiments, the NES comprises an amino acid sequence as set forth in SEQ ID NO: 12.

[0052] In certain embodiments, the polypeptide construct is a fusion protein, e.g., a fusion protein comprising the base deaminase and the guide RNA molecule binding domain.

[0053] In certain embodiments, the polypeptide construct is a fusion protein, and each adjacent domain of the fusion protein is connected to each other independently with or without a linker (e.g., a peptide linker, e.g., a peptide linker comprising one or more glycine and / or one or more serine). In certain embodiments, the linker comprises the amino acid sequence GS or the amino acid sequence set forth in SEQ ID NO: 21.

[0054] I-I. Base editing system based on a guideable loop-forming gRNA - adenine base editor

[0055] In certain embodiments, the base to be edited is adenine, and the polypeptide construct comprises a base deaminase having adenosine deaminase activity. In certain embodiments, the base deaminase does not have cytidine deaminase activity, or has significantly lower cytidine deaminase activity (e.g., has at most 30%, at most 20%, at most 10%, at most 5%, at most 2%, or at most 1% of the cytidine deaminase activity of a base deaminase having the sequence set forth in SEQ ID NO: 6).

[0056] In certain embodiments, the base deaminase is selected from the group consisting of: TadA comprising a sequence as set forth in SEQ ID NO: 1 or 2, a TadA mutant, and a functional fragment of the TadA or the TadA mutant.

[0057] wherein the TadA mutant comprises a mutation selected from the group consisting of:

[0058] (1) V28R, (2) V30A, (3) I76Y, (4) V82T, (5) F84H, F84V, F84N, F84S, F84I, F84T, F84C or F84A, (6) R98E, (7) N108H, (8) R111C or R111L, (9) D147R, (10) Q154R, (11) any combination of (1) to (10).

[0059] As used herein, the mutation "V28R" means that the residue in the TadA mutant which is in the corresponding position of the residue "V" at position 28 of SEQ ID NO: 1 is replaced with the residue "R"; wherein the residue in the TadA mutant which is in the corresponding position of the residue "V" at position 28 of SEQ ID NO: 1 means the amino acid site / residue in the sequence under comparison which is in the equivalent position of the amino acid residue "V" at position 28 of SEQ ID NO: 1 when the amino acid sequence of the TadA mutant is optimally aligned with SEQ ID NO: 1, i.e. when the amino acid sequence of the TadA mutant is aligned with SEQ ID NO: 1 to obtain the highest percentage identity.

[0060] Unless specifically indicated otherwise or clearly contradicted by context, the meaning of the remaining similar expressions herein is defined in a manner analogous to the above.

[0061] In certain embodiments, the base deaminase is selected from the group consisting of the TadA mutants and functional fragments thereof.

[0062] In certain embodiments, the functional fragment possesses the adenosine deaminase activity of the base deaminase from which it is derived.

[0063] In certain embodiments, the TadA mutant comprises a mutation selected from the group consisting of:

[0064] (i) (a) F84H, F84V or F84I, (b) V82T, (c) D147R, (d) I76Y, and / or, (e) Q154R;

[0065] (ii) (a) F84H, F84V or F84I, and (b) V82T;

[0066] (iii) (a) F84H, F84V or F84I, (b) V82T, and (c) D147R;

[0067] (iv) (a) F84H, F84V or F84I, (b) V82T, (c) D147R, (d) I76Y, and (e) Q154R;

[0068] (v) F84I, V82T, D147R, I76Y, and / or, Q154R;

[0069] (vi) F84I and V82T;

[0070] (vii) F84I, V82T and D147R;

[0071] (viii) F84I, V82T, D147R, I76Y and Q154R.

[0072] In certain embodiments, the TadA mutant has mutations F84I and V82T compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0073] In certain embodiments, the TadA mutant has mutations F84I, V82T and D147R compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0074] In certain embodiments, the TadA mutant has mutations F84I, V82T, D147R, I76Y and Q154R compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0075] In certain embodiments, the guide RNA molecule binding domain comprises an amino acid sequence as set forth in SEQ ID NO: 10; and / or, the guide RNA molecule has a backbone sequence comprising a nucleotide sequence as set forth in SEQ ID NO: 13.

[0076] In certain embodiments, the base deaminase comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 1-2, 4-5.

[0077] In certain embodiments, the polypeptide construct is a fusion protein comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 14-16.

[0078] I-II. Base editing system based on guideable loop-forming gRNAs - cytosine base editors

[0079] In certain embodiments, the base to be edited is cytosine, and the base deaminase comprised by the polypeptide construct has cytidine deaminase activity. In certain embodiments, the base deaminase does not have adenosine deaminase activity, or has significantly lower adenosine deaminase activity (e.g., has at most 30%, at most 20%, at most 10%, at most 5%, at most 2%, or at most 1% of the adenosine deaminase activity of a base deaminase as set forth in SEQ ID NO: 4).

[0080] In certain embodiments, the base deaminase is selected from the group consisting of TadA mutants and functional fragments thereof; wherein the TadA mutant comprises mutations: (i) N46P or N46V, and (ii) A48G, as compared to TadA comprising the sequence set forth in SEQ ID NO: 1, and further comprises a mutation selected from the group consisting of:

[0081] V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P, H96N, and any combination thereof.

[0082] In certain embodiments, the functional fragment possesses cytosine deaminase activity of the base deaminase from which it is derived.

[0083] In certain embodiments, the TadA mutant comprises a mutation selected from the group consisting of:

[0084] (i) (a) N46P and A48G, and (b) V82T, I76Y, D147R, and / or Q154R;

[0085] (ii) (a) N46P and A48G, (b) V82T, I76Y, D147R, and / or Q154R, and (c) R26G, V28A, Y73P, and / or H96N;

[0086] (iii) N46P, A48G, and V82T;

[0087] (iv) N46P, A48G, and I76Y;

[0088] (v) N46P, A48G, and D147R;

[0089] (vi) N46P, A48G, and Q154R;

[0090] (vii) N46P, A48G, V82T, I76Y, D147R, and Q154R;

[0091] (viii) N46P, A48G, V82T, I76Y, D147R, Q154R, and R26G;

[0092] (ix) N46P, A48G, V82T, I76Y, D147R, Q154R, and V28A;

[0093] (x) N46P, A48G, V82T, I76Y, D147R, Q154R, and Y73P;

[0094] (xi) N46P, A48G, V82T, I76Y, D147R, Q154R and H96N;

[0095] (xii) N46P, A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P and H96N;

[0096] (xiii) (a) N46V and A48G, and (b) V82T, I76Y, D147R and / or Q154R;

[0097] (xiv) (a) N46V and A48G, (b) V82T, I76Y, D147R and / or Q154R, and (c) R26G, V28A, Y73P and / or H96N;

[0098] (xv) N46V, A48G and V82T;

[0099] (xvi) N46V, A48G and I76Y;

[0100] (xvii) N46V, A48G and D147R;

[0101] (xviii) N46V, A48G and Q154R;

[0102] (xix) N46V, A48G, V82T, I76Y, D147R and Q154R;

[0103] (xx) N46V, A48G, V82T, I76Y, D147R, Q154R and R26G;

[0104] (xxi) N46V, A48G, V82T, I76Y, D147R, Q154R and V28A;

[0105] (xxii) N46V, A48G, V82T, I76Y, D147R, Q154R and Y73P;

[0106] (xxiii) N46V, A48G, V82T, I76Y, D147R, Q154R and H96N;

[0107] (xxiv) N46V, A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P and H96N.

[0108] In certain embodiments, the TadA mutant has the mutations N46P, A48G and V82T compared to a TadA comprising the sequence set forth in SEQ ID NO: 1.

[0109] In certain embodiments, the TadA mutant has mutations N46P, A48G, and I76Y compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0110] In certain embodiments, the TadA mutant has mutations N46P, A48G, and D147R compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0111] In certain embodiments, the TadA mutant has mutations N46P, A48G, and Q154R compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0112] In certain embodiments, the TadA mutant has mutations N46P, A48G, V82T, I76Y, D147R, and Q154R compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0113] In certain embodiments, the TadA mutant has mutations N46P, A48G, V82T, I76Y, D147R, Q154R, and R26G compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0114] In certain embodiments, the TadA mutant has mutations N46P, A48G, V82T, I76Y, D147R, Q154R, and V28A compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0115] In certain embodiments, the TadA mutant has mutations N46P, A48G, V82T, I76Y, D147R, Q154R, and Y73P compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0116] In certain embodiments, the TadA mutant has mutations N46P, A48G, V82T, I76Y, D147R, Q154R, and H96N compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0117] In certain embodiments, the TadA mutant has mutations N46P, A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P, and H96N compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0118] In some embodiments, the guide RNA molecule binding domain comprises an amino acid sequence as set forth in SEQ ID NO: 10; and / or, the backbone sequence of the guide RNA molecule comprises a nucleotide sequence as set forth in SEQ ID NO: 13.

[0119] In some embodiments, the base deaminase comprises an amino acid sequence as set forth in SEQ ID NO: 6 or 7.

[0120] In some embodiments, the polypeptide construct is a fusion protein comprising an amino acid sequence as set forth in SEQ ID NO: 17 or 18.

[0121] I-III. Base editing system based on guideable loop-forming gRNA - adenine / cytosine dual base editor

[0122] In some embodiments, the base to be edited is cytosine, adenine, or a combination thereof, the polypeptide construct comprises an adenosine deaminase and a cytidine deaminase, or the polypeptide construct comprises a base deaminase that has both adenosine deaminase activity and cytidine deaminase activity.

[0123] In some embodiments, the base deaminase is selected from a TadA mutant and a functional fragment thereof; wherein the TadA mutant comprises a mutation A48G, A48M, or A48H, and further comprises a mutation selected from:

[0124] (1) N46C;

[0125] (2) V82T, I76Y, D147R, and / or Q154R;

[0126] (3) R26G, V28A, Y73P, and / or H96N;

[0127] (4) (a) N46C, and (b) V82T, I76Y, D147R, and / or Q154R;

[0128] (5) (a) N46C, and (b) R26G, V28A, Y73P, and / or H96N;

[0129] (6) (a) V82T, I76Y, D147R, and / or Q154R, and (b) R26G, V28A, Y73P, and / or H96N, and (c) the TadA mutant does not comprise a mutation at the N46 position.

[0130] In some embodiments, the functional fragment possesses the adenosine deaminase activity and the cytidine deaminase activity of the base deaminase from which it is derived.

[0131] In certain embodiments, the TadA mutant comprises a mutation selected from the group consisting of:

[0132] (i) A48G and N46C;

[0133] (ii) A48G, N46C, V82T, I76Y, D147R and Q154R;

[0134] (iii) A48G, N46C, R26G, V28A, Y73P and H96N;

[0135] (iv) A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P and H96N, and the TadA mutant does not comprise a mutation at the N46 position;

[0136] (v) A48M and N46C;

[0137] (vi) A48M, N46C, V82T, I76Y, D147R and Q154R;

[0138] (vii) A48M, N46C, R26G, V28A, Y73P and H96N;

[0139] (viii) A48M, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P and H96N, and the TadA mutant does not comprise a mutation at the N46 position;

[0140] (ix) A48H and N46C;

[0141] (x) A48H, N46C, V82T, I76Y, D147R and Q154R;

[0142] (xi) A48H, N46C, R26G, V28A, Y73P and H96N;

[0143] (xii) A48H, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P and H96N, and the TadA mutant does not comprise a mutation at the N46 position.

[0144] In certain embodiments, the TadA mutant has the mutations A48G and N46C compared to a TadA comprising the sequence set forth in SEQ ID NO: 1.

[0145] In certain embodiments, the TadA mutant has mutations A48G, N46C, V82T, I76Y, D147R, and Q154R compared to TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0146] In certain embodiments, the TadA mutant has mutations A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P, and H96N compared to TadA comprising a sequence as set forth in SEQ ID NO: 1. In certain embodiments, the TadA mutant has mutations A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P, and H96N compared to TadA comprising a sequence as set forth in SEQ ID NO: 1, and the TadA mutant does not comprise a mutation at position N46.

[0147] In certain embodiments, the guide RNA molecule binding domain comprises an amino acid sequence as set forth in SEQ ID NO: 10; and / or, the backbone sequence of the guide RNA molecule comprises a nucleotide sequence as set forth in SEQ ID NO: 13.

[0148] In certain embodiments, the base deaminase comprises an amino acid sequence as set forth in SEQ ID NO: 8 or 9.

[0149] In certain embodiments, the polypeptide construct is a fusion protein comprising an amino acid sequence as set forth in SEQ ID NO: 19 or 20.

[0150] In a second aspect, the present application provides a complex comprising:

[0151] (a) a guide RNA molecule as defined in the first aspect;

[0152] and,

[0153] (b) a polypeptide construct as defined in the first aspect.

[0154] In a third aspect, the present application provides a vector system comprising one or more vectors comprising:

[0155] (1) a first nucleotide sequence encoding a guide RNA molecule as defined in the first aspect; optionally, the first nucleotide sequence is operably linked to a first regulatory element; and,

[0156] (2) a second nucleotide sequence encoding a polypeptide construct as defined in the first aspect; optionally, the second nucleotide sequence is operably linked to a second regulatory element;

[0157] wherein the first nucleotide sequence and the second nucleotide sequence are present on the same or different vectors.

[0158] In certain embodiments, the first regulatory element is a promoter, e.g., an inducible promoter.

[0159] In certain embodiments, the second regulatory element is a promoter, e.g., an inducible promoter.

[0160] In a fourth aspect, the present application provides an engineered RNA molecule comprising a sequence of a guide RNA molecule as defined in the first aspect.

[0161] In another aspect, the present application also provides a nucleic acid molecule or a vector or a host cell encoding the engineered RNA molecule. In certain embodiments, the cell is a non-plant cell. In certain embodiments, the cell is a microbial cell.

[0162] In a fifth aspect, the present application provides a kit comprising a composition as described in the first aspect, a complex as described in the second aspect, a vector system as described in the third aspect, an engineered RNA molecule as described in the fourth aspect, or a nucleic acid molecule or a vector encoding the engineered RNA molecule or a host cell.

[0163] In certain embodiments, the kit further comprises instructions for using the composition, complex, vector system, engineered RNA molecule, nucleic acid molecule or vector encoding the engineered RNA molecule or host cell.

[0164] In a sixth aspect, the present application provides a delivery composition comprising a delivery vehicle, and one or more selected from the group consisting of a composition as described in the first aspect, a complex as described in the second aspect, a vector system as described in the third aspect, an engineered RNA molecule of the fourth aspect, or a nucleic acid molecule or a vector encoding the engineered RNA molecule.

[0165] In certain embodiments, the delivery vehicle is a particle.

[0166] In certain embodiments, the delivery vehicle is selected from the group consisting of a lipid particle, a sugar particle, a metal particle, a protein particle, a liposome, an exosome, a microvesicle, a gene gun, or a viral vector (e.g., a replication-defective retrovirus, a lentivirus, an adenovirus, or an adeno-associated virus).

[0167] In another aspect, the application provides a cell comprising: the composition of the first aspect, the complex of the second aspect, the vector system of the third aspect, the engineered RNA molecule of the fourth aspect, the nucleic acid molecule or vector encoding the engineered RNA molecule, or the delivery composition of the sixth aspect. In certain embodiments, the cell is a non-plant cell. In certain embodiments, the cell is a microbial cell.

[0168] In a seventh aspect, the application provides a method of editing a target RNA extracellularly, comprising contacting the target RNA with one or more of the composition of the first aspect, the complex of the second aspect, the engineered RNA molecule of the fourth aspect, the delivery composition of the sixth aspect, under conditions suitable for editing of the target RNA, thereby inducing deamination of a base to be edited in the target RNA.

[0169] In an eighth aspect, the application provides a method of editing a target RNA intracellularly, comprising delivering the composition of the first aspect, the complex of the second aspect, the vector system of the third aspect, the engineered RNA molecule of the fourth aspect, the nucleic acid molecule or vector encoding the engineered RNA molecule, or the delivery composition of the sixth aspect into a cell containing the target RNA, thereby inducing deamination of a base to be edited at the target site.

[0170] In a ninth aspect, the application provides a cell obtained from the method of the eighth aspect, or a progeny thereof, wherein the cell comprises a modification that is not present in its wild type. In certain embodiments, the cell is a non-plant cell. In certain embodiments, the cell is a microbial cell.

[0171] In a tenth aspect, the application provides the composition of the first aspect, the complex of the second aspect, the vector system of the third aspect, the engineered RNA molecule of the fourth aspect, the nucleic acid molecule or vector encoding the engineered RNA molecule, the kit of the fifth aspect, or the delivery composition of the sixth aspect, for use in RNA editing, or for use in the manufacture of a medicament for RNA editing.

[0172] In certain embodiments, the RNA editing is base editing performed on an RNA.

[0173] II. Base editing system based on novel TadA mutants

[0174] II-I. Base editing system based on novel TadA mutants - TadA mutants for adenine base editing

[0175] In an eleventh aspect, the present application provides a TadA mutant or a functional fragment thereof, wherein the TadA mutant comprises a mutation selected from the group consisting of:

[0176] (1) V28R, (2) V30A, (3) I76Y, (4) V82T, (5) F84H, F84V, F84N, F84S, F84I, F84T, F84C or F84A, (6) R98E, (7) N108H, (8) R111C or R111L, (9) D147R, (10) Q154R, (11) any combination of (1) to (10).

[0177] In certain embodiments, the functional fragment possesses the adenosine deaminase activity of the TadA mutant from which it is derived.

[0178] In certain embodiments, the TadA mutant comprises a mutation selected from the group consisting of:

[0179] (i) (a) F84H, F84V or F84I, (b) V82T, (c) D147R, (d) I76Y, and / or, (e) Q154R;

[0180] (ii) (a) F84H, F84V or F84I, and (b) V82T;

[0181] (iii) (a) F84H, F84V or F84I, (b) V82T, and (c) D147R;

[0182] (iv) (a) F84H, F84V or F84I, (b) V82T, (c) D147R, (d) I76Y, and (e) Q154R;

[0183] (v) F84I, V82T, D147R, I76Y, and / or, Q154R;

[0184] (vi) F84I and V82T;

[0185] (vii) F84I, V82T and D147R;

[0186] (viii) F84I, V82T, D147R, I76Y and Q154R.

[0187] In certain embodiments, the TadA mutant has the mutations F84I and V82T compared to a TadA comprising the sequence set forth in SEQ ID NO: 1.

[0188] In certain embodiments, the TadA mutant has mutations F84I, V82T, and D147R compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0189] In certain embodiments, the TadA mutant has mutations F84I, V82T, D147R, I76Y, and Q154R compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0190] In certain embodiments, the TadA mutant comprises an amino acid sequence as set forth in SEQ ID NO: 4 or 5.

[0191] In certain embodiments, the TadA mutant has adenosine deaminase activity. In certain embodiments, the TadA mutant has no cytidine deaminase activity, or, has significantly lower cytidine deaminase activity (e.g., has at most 30%, at most 20%, at most 10%, at most 5%, at most 2%, or at most 1% of the cytidine deaminase activity of a deaminase as set forth in SEQ ID NO: 6).

[0192] In certain embodiments, the TadA mutant is used for adenine base editing.

[0193] II-II. Base editing systems based on novel TadA mutants - TadA mutants for cytosine base editing

[0194] In a twelfth aspect, the present application provides a TadA mutant or a functional fragment thereof, wherein the TadA mutant comprises, compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1, mutations: (i) N46P or N46V, and (ii) A48G, and further comprises a mutation selected from:

[0195] V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P, H96N, and any combination thereof.

[0196] In certain embodiments, the functional fragment has cytidine deaminase activity of the TadA mutant from which it is derived.

[0197] In certain embodiments, the TadA mutant comprises, compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1, a mutation selected from:

[0198] (i) (a) N46P and A48G, and (b) V82T, I76Y, D147R, and / or Q154R;

[0199] (ii) (a) N46P and A48G, (b) V82T, I76Y, D147R and / or Q154R, and (c) R26G, V28A, Y73P and / or H96N;

[0200] (iii) N46P, A48G and V82T;

[0201] (iv) N46P, A48G and I76Y;

[0202] (v) N46P, A48G and D147R;

[0203] (vi) N46P, A48G and Q154R;

[0204] (vii) N46P, A48G, V82T, I76Y, D147R and Q154R;

[0205] (viii) N46P, A48G, V82T, I76Y, D147R, Q154R and R26G;

[0206] (ix) N46P, A48G, V82T, I76Y, D147R, Q154R and V28A;

[0207] (x) N46P, A48G, V82T, I76Y, D147R, Q154R and Y73P;

[0208] (xi) N46P, A48G, V82T, I76Y, D147R, Q154R and H96N;

[0209] (xii) N46P, A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P and H96N;

[0210] (xiii) (a) N46V and A48G, and (b) V82T, I76Y, D147R and / or Q154R;

[0211] (xiv) (a) N46V and A48G, (b) V82T, I76Y, D147R and / or Q154R, and (c) R26G, V28A, Y73P and / or H96N;

[0212] (xv) N46V, A48G and V82T;

[0213] (xvi) N46V, A48G and I76Y;

[0214] (xvii) N46V, A48G and D147R;

[0215] (xviii) N46V, A48G and Q154R;

[0216] (xix) N46V, A48G, V82T, I76Y, D147R and Q154R;

[0217] (xx) N46V, A48G, V82T, I76Y, D147R, Q154R and R26G;

[0218] (xxi) N46V, A48G, V82T, I76Y, D147R, Q154R and V28A;

[0219] (xxii) N46V, A48G, V82T, I76Y, D147R, Q154R and Y73P;

[0220] (xxiii) N46V, A48G, V82T, I76Y, D147R, Q154R and H96N;

[0221] (xxiv) N46V, A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P and H96N.

[0222] In certain embodiments, the TadA mutant has the mutations N46P, A48G and V82T, as compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0223] In certain embodiments, the TadA mutant has the mutations N46P, A48G and I76Y, as compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0224] In certain embodiments, the TadA mutant has the mutations N46P, A48G and D147R, as compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0225] In certain embodiments, the TadA mutant has the mutations N46P, A48G and Q154R, as compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0226] In certain embodiments, the TadA mutant has the mutations N46P, A48G, V82T, I76Y, D147R and Q154R, as compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0227] In certain embodiments, the TadA mutant has mutations N46P, A48G, V82T, I76Y, D147R, Q154R, and R26G compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0228] In certain embodiments, the TadA mutant has mutations N46P, A48G, V82T, I76Y, D147R, Q154R, and V28A compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0229] In certain embodiments, the TadA mutant has mutations N46P, A48G, V82T, I76Y, D147R, Q154R, and Y73P compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0230] In certain embodiments, the TadA mutant has mutations N46P, A48G, V82T, I76Y, D147R, Q154R, and H96N compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0231] In certain embodiments, the TadA mutant has mutations N46P, A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P, and H96N compared to a TadA comprising a sequence as set forth in SEQ ID NO: 1.

[0232] In certain embodiments, the TadA mutant comprises an amino acid sequence as set forth in SEQ ID NO: 6 or 7.

[0233] In certain embodiments, the TadA mutant has cytidine deaminase activity. In certain embodiments, the TadA mutant does not have adenosine deaminase activity, or, has significantly lower adenosine deaminase activity (e.g., has at most 30%, at most 20%, at most 10%, at most 5%, at most 2%, or at most 1% of the adenosine deaminase activity of a deaminase as set forth in SEQ ID NO: 4).

[0234] In certain embodiments, the TadA mutant is used for cytosine base editing.

[0235] II-III. Base editing systems based on novel TadA mutants - TadA mutants for adenine / cytosine dual base editing

[0236] In a thirteenth aspect, the present application provides a TadA mutant or a functional fragment thereof, wherein the TadA mutant comprises a mutation A48G, A48M or A48H, and further comprises a mutation selected from the group consisting of:

[0237] (1) N46C;

[0238] (2) V82T, I76Y, D147R and / or Q154R;

[0239] (3) R26G, V28A, Y73P and / or H96N;

[0240] (4) (a) N46C, and (b) V82T, I76Y, D147R and / or Q154R;

[0241] (5) (a) N46C, and (b) R26G, V28A, Y73P and / or H96N;

[0242] (6) (a) V82T, I76Y, D147R and / or Q154R, and (b) R26G, V28A, Y73P and / or H96N, and (c) the TadA mutant does not comprise a mutation at the N46 position.

[0243] In certain embodiments, the functional fragment possesses the adenosine deaminase activity and the cytidine deaminase activity of the TadA mutant from which it is derived.

[0244] In certain embodiments, the TadA mutant comprises a mutation selected from the group consisting of:

[0245] (i) A48G and N46C;

[0246] (ii) A48G, N46C, V82T, I76Y, D147R and Q154R;

[0247] (iii) A48G, N46C, R26G, V28A, Y73P and H96N;

[0248] (iv) A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P and H96N, and the TadA mutant does not comprise a mutation at the N46 position;

[0249] (v) A48M and N46C;

[0250] (vi) A48M, N46C, V82T, I76Y, D147R, and Q154R;

[0251] (vii) A48M, N46C, R26G, V28A, Y73P, and H96N;

[0252] (viii) A48M, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P, and H96N, and the TadA mutant does not comprise a mutation at position N46;

[0253] (ix) A48H and N46C;

[0254] (x) A48H, N46C, V82T, I76Y, D147R, and Q154R;

[0255] (xi) A48H, N46C, R26G, V28A, Y73P, and H96N;

[0256] (xi) A48H, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P, and H96N, and the TadA mutant does not comprise a mutation at position N46.

[0257] In certain embodiments, the TadA mutant has the mutations A48G and N46C, as compared to a TadA comprising the sequence set forth in SEQ ID NO: 1.

[0258] In certain embodiments, the TadA mutant has the mutations A48G, N46C, V82T, I76Y, D147R, and Q154R, as compared to a TadA comprising the sequence set forth in SEQ ID NO: 1.

[0259] In certain embodiments, the TadA mutant has the mutations A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P, and H96N, as compared to a TadA comprising the sequence set forth in SEQ ID NO: 1. In certain embodiments, the TadA mutant has the mutations A48G, V82T, I76Y, D147R, Q154R, R26G, V28A, Y73P, and H96N, as compared to a TadA comprising the sequence set forth in SEQ ID NO: 1, and the TadA mutant does not comprise a mutation at position N46.

[0260] In certain embodiments, the TadA mutant comprises the amino acid sequence set forth in SEQ ID NO: 8 or 9.

[0261] In certain embodiments, the TadA mutant is used for adenine / cytosine double base editing.

[0262] In a fourteenth aspect, the present application provides a polypeptide construct comprising a base deaminase and a guide RNA molecule binding domain linked to the base deaminase, the guide RNA molecule binding domain being capable of binding to a backbone sequence of the guide RNA molecule;

[0263] In certain embodiments, the base deaminase is selected from the TadA mutant or functional fragment thereof of any one of the eleventh aspect to the thirteenth aspect.

[0264] In certain embodiments, the guide RNA molecule binding domain is selected from: a Cas effector protein of a CRISPR system lacking nuclease activity, a MS2 bacteriophage coat protein (MCP), a Bacteriophage Lambda N protein (λ N-peptide).

[0265] In certain embodiments, the guide RNA molecule binding domain is selected from: a CRISPR Cas13 effector protein lacking nuclease activity. In certain embodiments, the backbone sequence of the guide RNA molecule is selected from a backbone sequence of a crRNA that the CRISPR Cas13 effector protein is capable of recognizing and binding to.

[0266] In certain embodiments, the guide RNA molecule binding domain is selected from a MS2 bacteriophage coat protein (MCP). In certain embodiments, the backbone sequence of the guide RNA molecule is selected from a MS2 aptamer that the MS2 bacteriophage coat protein (MCP) is capable of recognizing and binding to.

[0267] In certain embodiments, the guide RNA molecule binding domain is selected from a Bacteriophage Lambda N protein. In certain embodiments, the backbone sequence of the guide RNA molecule is selected from a boxB sequence that the Bacteriophage Lambda N protein is capable of recognizing and binding to.

[0268] In certain embodiments, the guide RNA molecule binding domain is a PspCas13b lacking nuclease activity. In certain embodiments, the guide RNA molecule binding domain comprises an amino acid sequence as set forth in SEQ ID NO: 10.

[0269] In certain embodiments, in the polypeptide construct, the guide RNA molecule binding domain is linked to the N-terminus and / or the C-terminus of the base deaminase.

[0270] In certain embodiments, the polypeptide construct further comprises a nuclear localization sequence (NLS) and / or a nuclear export sequence (NES). In certain embodiments, the polypeptide construct further comprises one or more NLS and / or one or more NES. In certain embodiments, the NLS comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 11, 22-23. In certain embodiments, the NES comprises an amino acid sequence as set forth in SEQ ID NO: 12.

[0271] In certain embodiments, the polypeptide construct is a fusion protein, e.g., a fusion protein comprising the base deaminase and the guide RNA molecule binding domain.

[0272] In certain embodiments, the polypeptide construct is a fusion protein, and each adjacent domain of the fusion protein is connected to each other independently with or without a linker (e.g., a peptide linker, e.g., a peptide linker comprising one or more glycine and / or one or more serine). In certain embodiments, the linker comprises the amino acid sequence GS or the amino acid sequence set forth in SEQ ID NO: 21.

[0273] In certain embodiments, the guide RNA molecule binding domain comprises an amino acid sequence as set forth in SEQ ID NO: 10; and / or, the backbone sequence of the guide RNA molecule comprises a nucleotide sequence as set forth in SEQ ID NO: 13.

[0274] In certain embodiments, the base to be edited is adenine, and the base deaminase is selected from the TadA mutants or functional fragments thereof of the eleventh aspect. In certain embodiments, the base deaminase is selected from the TadA mutants or functional fragments thereof comprising an amino acid sequence as set forth in SEQ ID NO: 4 or 5. In certain embodiments, the polypeptide construct is a fusion protein comprising an amino acid sequence as set forth in SEQ ID NO: 15 or 16.

[0275] In certain embodiments, the base to be edited is cytosine, and the base deaminase is selected from the TadA mutants or functional fragments thereof of the twelfth aspect. In certain embodiments, the base deaminase is selected from the TadA mutants or functional fragments thereof comprising an amino acid sequence as set forth in SEQ ID NO: 6 or 7. In certain embodiments, the polypeptide construct is a fusion protein comprising an amino acid sequence as set forth in SEQ ID NO: 17 or 18.

[0276] In certain embodiments, the base to be edited is an adenine, a cytosine, or a combination thereof, and the base deaminase is selected from a TadA mutant of the thirteenth aspect or a functional fragment thereof. In certain embodiments, the base deaminase is selected from a TadA mutant comprising an amino acid sequence as set forth in SEQ ID NO: 8 or 9, or a functional fragment thereof. In certain embodiments, the polypeptide construct is a fusion protein comprising an amino acid sequence as set forth in SEQ ID NO: 19 or 20.

[0277] In a fifteenth aspect, the present application provides a nucleic acid molecule encoding a TadA mutant of any one of the eleventh to thirteenth aspects or a functional fragment thereof, or a polypeptide construct of the fourteenth aspect.

[0278] In a sixteenth aspect, the present application provides a vector comprising the nucleic acid molecule of the fifteenth aspect. In certain embodiments, the vector is a cloning vector or an expression vector.

[0279] In a seventeenth aspect, the present application provides a host cell comprising the nucleic acid molecule of the fifteenth aspect or the vector of the sixteenth aspect. In certain embodiments, the cell is not a plant cell. In certain embodiments, the cell is a microbial cell.

[0280] In an eighteenth aspect, the present application provides a method of making a TadA mutant of any one of the eleventh to thirteenth aspects or a functional fragment thereof, or a polypeptide construct of the fourteenth aspect, comprising culturing the host cell of the seventeenth aspect under conditions permitting protein expression, and recovering the TadA mutant or functional fragment thereof or the polypeptide construct from the cultured host cell culture.

[0281] In a nineteenth aspect, the present application provides a composition comprising:

[0282] (a) a first component selected from the group consisting of: (i) a guide RNA molecule comprising a backbone sequence and a guide sequence, the guide sequence being capable of hybridizing to a target RNA sequence, the target RNA sequence comprising a base to be edited; (ii) a nucleotide sequence encoding the guide RNA molecule; (iii) any combination of (i) and (ii);

[0283] and,

[0284] (b) a second component selected from the group consisting of: (i) a polypeptide construct comprising a base deaminase and a guide RNA molecule binding domain linked to the base deaminase, the guide RNA molecule binding domain being capable of binding to the backbone sequence of the guide RNA molecule; (ii) a nucleotide sequence encoding the polypeptide construct; (iii) any combination of (i) and (ii);

[0285] wherein the base deaminase is selected from the TadA mutant or functional fragment thereof of any one of the eleventh to thirteenth aspects.

[0286] In embodiments where the first component is a nucleotide sequence encoding the guide RNA molecule and the second component is a nucleotide sequence encoding the polypeptide construct, the nucleotide sequence of the first component and the nucleotide sequence of the second component can be present in different nucleic acid molecules or vectors, respectively, or in the same nucleic acid molecule or vector.

[0287] In certain embodiments, the guide sequence of the guide RNA molecule is designed to cause the target RNA sequence to form a single-stranded RNA loop region comprising the base to be edited upon hybridization with the target RNA sequence.

[0288] In certain embodiments, the target RNA sequence comprises: a target loopable region comprising the base to be edited, a first region located 5' to the target loopable region, and a second region located 3' to the target loopable region;

[0289] the guide sequence of the guide RNA molecule is designed to comprise a first region, a second region, and optionally a third region located between the first region and the second region, the second region being located upstream of the first region; and the first region is capable of complementarity to the first region of the target RNA sequence, the second region is capable of complementarity to the second region of the target RNA sequence, and the third region is absent or the third region is not capable of complementarity to the target loopable region of the target RNA sequence;

[0290] thereby, upon hybridization of the guide RNA molecule with the target RNA sequence, the target loopable region of the target RNA sequence located between the first region and the second region exists as a single strand to form a single-stranded RNA loop region comprising the base to be edited.

[0291] In certain embodiments, the third region of the guide RNA molecule is absent.

[0292] In certain embodiments, the third region of the guide RNA molecule has a number of nucleotide residues less than that of the target loopable region of the target RNA sequence.

[0293] In certain embodiments, the first region and the second region of the guide RNA molecule have the same or different number of nucleotide residues.

[0294] In certain embodiments, the guide RNA molecule binding domain is selected from the group consisting of: a Cas effector protein of a CRISPR system lacking nuclease activity, MS2 bacteriophage coat protein (MCP), Bacteriophage Lambda N protein (λ N-peptide).

[0295] Based on the disclosure herein, one of skill in the art would readily appreciate that the design of the guide RNA molecule backbone sequence in the compositions of the present application corresponds to the selection of the guide RNA molecule binding domain of the polypeptide construct, i.e., the guide RNA molecule backbone sequence is designed to be recognized and bound by the guide RNA molecule binding domain comprised by the polypeptide construct.

[0296] In certain embodiments, the guide RNA molecule backbone sequence is selected from the group consisting of a crRNA (or crRNA / tracrRNA complex or construct) that is recognized and bound by the Cas effector protein.

[0297] In certain embodiments, the guide RNA molecule binding domain is selected from the group consisting of: a CRISPR Cas13 effector protein lacking nuclease activity. In certain embodiments, the guide RNA molecule backbone sequence is selected from the group consisting of a crRNA that is recognized and bound by the CRISPR Cas13 effector protein.

[0298] In certain embodiments, the guide RNA molecule binding domain is selected from the group consisting of MS2 bacteriophage coat protein (MCP). In certain embodiments, the guide RNA molecule backbone sequence is selected from the group consisting of a MS2 aptamer that is recognized and bound by the MS2 bacteriophage coat protein (MCP).

[0299] In certain embodiments, the guide RNA molecule binding domain is selected from the group consisting of Bacteriophage Lambda N protein. In certain embodiments, the guide RNA molecule backbone sequence is selected from the group consisting of a boxB sequence that is recognized and bound by the Bacteriophage Lambda N protein.

[0300] In certain embodiments, the guide RNA molecule binding domain is PspCas13b lacking nuclease activity. In certain embodiments, the guide RNA molecule binding domain comprises an amino acid sequence as set forth in SEQ ID NO: 10. In certain embodiments, the guide RNA molecule backbone sequence comprises an amino acid sequence as set forth in SEQ ID NO: 13.

[0301] In certain embodiments, the single-stranded RNA loop region formed upon hybridization of the target RNA sequence to the guide RNA molecule comprises one or more bases to be edited. In certain embodiments, the single-stranded RNA loop region formed upon hybridization of the target RNA sequence to the guide RNA molecule comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases to be edited.

[0302] In certain embodiments, the bases to be edited are selected from the group consisting of adenine, cytosine, and any combination thereof.

[0303] In certain embodiments, the first region of the guide RNA molecule consists of 3-50 (e.g., 5-50, 5-40, 5-30, 5-25, 5-20, 5-18, 5-15, 10-50, 10-40, 10-30, 10-25, 10-20, 10-18, 12-50, 12-40, 12-30, 12-25, 12-20, 12-18) nucleotide residues.

[0304] In certain embodiments, the second region of the guide RNA molecule consists of 3-50 (e.g., 5-50, 5-40, 5-30, 5-25, 5-20, 5-18, 5-15, 10-50, 10-40, 10-30, 10-25, 10-20, 10-18, 12-50, 12-40, 12-30, 12-25, 12-20, 12-18) nucleotide residues.

[0305] In certain embodiments, the single-stranded RNA loop region formed upon hybridization of the target RNA sequence to the guide RNA molecule consists of at least 3, at least 4, at least 5, or at least 6 nucleotide residues. In certain embodiments, the single-stranded RNA loop region formed upon hybridization of the target RNA sequence to the guide RNA molecule consists of 3-30 (e.g., 6-30, 6-25, 6-20, 6-19, 6-18, 6-17) nucleotide residues. In certain embodiments, the single-stranded RNA loop region formed upon hybridization of the target RNA sequence to the guide RNA molecule consists of 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17 nucleotide residues.

[0306] In certain embodiments, the bases to be edited in the target RNA sequence are spaced at least 1, at least 2, at least 3, at least 4, or at least 5 nucleotide residues from the first region and / or at least 1, at least 2, at least 3, at least 4, or at least 5 nucleotide residues from the second region.

[0307] In certain embodiments, in the polypeptide construct, the guide RNA molecule binding domain is linked to the N-terminus and / or C-terminus of the base deaminase.

[0308] In certain embodiments, the polypeptide construct further comprises a nuclear localization sequence (NLS) and / or a nuclear export sequence (NES). In certain embodiments, the NLS comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 11, 22-23. In certain embodiments, the NES comprises an amino acid sequence as set forth in SEQ ID NO: 12.

[0309] In certain embodiments, the polypeptide construct is a fusion protein, e.g., a fusion protein comprising the base deaminase and the guide RNA molecule binding domain.

[0310] In certain embodiments, the polypeptide construct is a fusion protein, and each adjacent domain of the fusion protein is linked to each other independently with or without a linker (e.g., a peptide linker, e.g., a peptide linker comprising one or more glycine and / or one or more serine). In certain embodiments, the linker comprises the amino acid sequence GS or the amino acid sequence as set forth in SEQ ID NO: 21.

[0311] In certain embodiments, the guide RNA molecule binding domain comprises an amino acid sequence as set forth in SEQ ID NO: 10; and / or, the backbone sequence of the guide RNA molecule comprises a nucleotide sequence as set forth in SEQ ID NO: 13.

[0312] In certain embodiments, the base to be edited is adenine, and the base deaminase is selected from the TadA mutants or functional fragments thereof of the eleventh aspect. In certain embodiments, the base deaminase is selected from the TadA mutants or functional fragments thereof comprising an amino acid sequence as set forth in SEQ ID NO: 4 or 5. In certain embodiments, the polypeptide construct is a fusion protein comprising an amino acid sequence as set forth in SEQ ID NO: 15 or 16.

[0313] In certain embodiments, the base to be edited is cytosine, and the base deaminase is selected from the TadA mutants or functional fragments thereof of the twelfth aspect. In certain embodiments, the base deaminase is selected from the TadA mutants or functional fragments thereof comprising an amino acid sequence as set forth in SEQ ID NO: 6 or 7. In certain embodiments, the polypeptide construct is a fusion protein comprising an amino acid sequence as set forth in SEQ ID NO: 17 or 18.

[0314] In certain embodiments, the base to be edited is adenine, cytosine, or a combination thereof, and the base deaminase is selected from the TadA mutants of the thirteenth aspect or a functional fragment thereof. In certain embodiments, the base deaminase is selected from the TadA mutants comprising the amino acid sequence set forth in SEQ ID NO: 8 or 9, or a functional fragment thereof. In certain embodiments, the polypeptide construct is a fusion protein comprising the amino acid sequence set forth in SEQ ID NO: 19 or 20.

[0315] In a twentieth aspect, the present application provides a complex comprising:

[0316] (a) a guide RNA molecule as defined in the nineteenth aspect;

[0317] and,

[0318] (b) a polypeptide construct as defined in the nineteenth aspect.

[0319] In a twenty-first aspect, the present application provides a vector system comprising one or more vectors comprising:

[0320] (1) a first nucleotide sequence encoding a guide RNA molecule as defined in the nineteenth aspect; optionally, the first nucleotide sequence is operably linked to a first regulatory element; and,

[0321] (2) a second nucleotide sequence encoding a polypeptide construct as defined in the nineteenth aspect; optionally, the second nucleotide sequence is operably linked to a second regulatory element;

[0322] wherein the first nucleotide sequence and the second nucleotide sequence are present on the same or different vectors.

[0323] In certain embodiments, the first regulatory element is a promoter, e.g., an inducible promoter.

[0324] In certain embodiments, the second regulatory element is a promoter, e.g., an inducible promoter.

[0325] In a twenty-second aspect, the present application provides a kit comprising the TadA mutant or a functional fragment thereof of any one of the eleventh to thirteenth aspects, the polypeptide construct of the fourteenth aspect, the nucleic acid molecule of the fifteenth aspect, the vector of the sixteenth aspect, the host cell of the seventeenth aspect, the composition of the nineteenth aspect, the complex of the twentieth aspect, or the vector system of the twenty-first aspect.

[0326] In a twenty-third aspect, the present application provides a delivery composition comprising a delivery vehicle, and one or more selected from the group consisting of: a TadA mutant or functional fragment thereof according to any one of the eleventh to thirteenth aspects, a polypeptide construct according to the fourteenth aspect, a nucleic acid molecule according to the fifteenth aspect, a vector according to the sixteenth aspect, a composition according to the nineteenth aspect, a complex according to the twentieth aspect, and a vector system according to the twenty-first aspect.

[0327] In certain embodiments, the delivery vehicle is a particle.

[0328] In certain embodiments, the delivery vehicle is selected from the group consisting of a lipid particle, a sugar particle, a metal particle, a protein particle, a liposome, an exosome, a microvesicle, a gene gun, or a viral vector (e.g., a replication-defective retrovirus, a lentivirus, an adenovirus, or an adeno-associated virus).

[0329] In another aspect, the present application provides a cell comprising: a TadA mutant or functional fragment thereof according to any one of the eleventh to thirteenth aspects, a polypeptide construct according to the fourteenth aspect, a nucleic acid molecule according to the fifteenth aspect, a vector according to the sixteenth aspect, a composition according to the nineteenth aspect, a complex according to the twentieth aspect, a vector system according to the twenty-first aspect, or a delivery composition according to the twenty-third aspect. In certain embodiments, the cell is a non-plant cell. In certain embodiments, the cell is a microbial cell.

[0330] In a twenty-fourth aspect, the present application provides a method of editing a target RNA extracellularly, comprising contacting the target RNA with one or more of: a TadA mutant or functional fragment thereof according to any one of the eleventh to thirteenth aspects, a polypeptide construct according to the fourteenth aspect, a nucleic acid molecule according to the fifteenth aspect, a vector according to the sixteenth aspect, a composition according to the nineteenth aspect, a complex according to the twentieth aspect, or a delivery composition according to the twenty-third aspect, under conditions suitable for editing the target RNA, thereby inducing deamination of a base to be edited in the target RNA.

[0331] In a twenty-fifth aspect, the present application provides a method of editing a target RNA intracellularly, comprising delivering into a cell containing the target RNA: a TadA mutant or functional fragment thereof according to any one of the eleventh to thirteenth aspects, a polypeptide construct according to the fourteenth aspect, a nucleic acid molecule according to the fifteenth aspect, a vector according to the sixteenth aspect, a composition according to the nineteenth aspect, a complex according to the twentieth aspect, a vector system according to the twenty-first aspect, or a delivery composition according to the twenty-third aspect, thereby inducing deamination of a base to be edited at the target site.

[0332] In a twenty-sixth aspect, the present application provides a cell obtained from the method of the twenty-fifth aspect, or a progeny thereof, wherein the cell comprises a modification that is not present in its wild type. In certain embodiments, the cell is a non-plant cell. In certain embodiments, the cell is a microbial cell.

[0333] In a twenty-seventh aspect, the present application provides the TadA mutant or functional fragment thereof of any one of aspects eleven to thirteen, the polypeptide construct of aspect fourteen, the nucleic acid molecule of aspect fifteen, the vector of aspect sixteen, the host cell of aspect seventeen, the composition of aspect nineteen, the complex of aspect twenty, the vector system of aspect twenty-one, the kit of aspect twenty-two, or the delivery composition of aspect twenty-three, for use in RNA editing, or in the manufacture of a medicament for RNA editing.

[0334] In certain embodiments, the RNA editing is base editing performed on an RNA.

[0335] In another aspect, the present application provides the following embodiments:

[0336] Embodiment 1. A TadA mutant or functional fragment thereof, wherein the TadA mutant comprises a mutation selected from the group consisting of:

[0337] (1) V28R, (2) V30A, (3) I76Y, (4) V82T, (5) F84H, F84V, F84N, F84S, F84I, F84T, F84C, or F84A, (6) R98E, (7) N108H, (8) R111C or R111L, (9) D147R, (10) Q154R, (11) any combination of (1) to (10).

[0338] Embodiment 2. The TadA mutant or functional fragment thereof of embodiment 1, wherein the TadA mutant comprises a mutation selected from the group consisting of:

[0339] (i) (a) F84H, F84V, or F84I, (b) V82T, (c) D147R, (d) I76Y, and / or, (e) Q154R;

[0340] (ii) (a) F84H, F84V, or F84I, and (b) V82T;

[0341] (iii) (a) F84H, F84V, or F84I, (b) V82T, and (c) D147R;

[0342] (iv) (a) F84H, F84V or F84I, (b) V82T, (c) D147R, (d) I76Y, and (e) Q154R;

[0343] (v) F84I, V82T, D147R, I76Y, and / or, Q154R;

[0344] (vi) F84I and V82T;

[0345] (vii) F84I, V82T and D147R;

[0346] (viii) F84I, V82T, D147R, I76Y and Q154R.

[0347] Embodiment 3. The TadA mutant or functional fragment thereof of embodiment 1, wherein the TadA mutant comprises a mutation selected from the group consisting of:

[0348] (i) F84I and V82T;

[0349] (ii) F84I, V82T and D147R;

[0350] (iii) F84I, V82T, D147R, I76Y and Q154R.

[0351] Embodiment 4. The TadA mutant or functional fragment thereof of embodiment 1, wherein the TadA mutant comprises an amino acid sequence as set forth in SEQ ID NO: 4 or 5.

[0352] Embodiment 5. A polypeptide construct comprising a base deaminase and a guide RNA molecule binding domain linked to the base deaminase.

[0353] wherein the base deaminase is selected from the group consisting of the TadA mutant or functional fragment thereof of any one of embodiments 1-4.

[0354] Embodiment 6. The polypeptide construct of embodiment 5, wherein the guide RNA molecule binding domain is selected from the group consisting of: a Cas effector protein of a CRISPR system lacking nuclease activity, a MS2 bacteriophage coat protein (MCP), a Bacteriophage Lambda N peptide (λN-peptide).

[0355] Embodiment 7. The polypeptide construct of embodiment 5, having one or more features selected from:

[0356] (1) in the polypeptide construct, the guide RNA molecule binding domain is linked to the N-terminus and / or C-terminus of the base deaminase;

[0357] (2) the polypeptide construct further comprises a nuclear localization sequence (NLS) and / or a nuclear export sequence (NES);

[0358] (3) the guide RNA molecule binding domain comprises an amino acid sequence as set forth in SEQ ID NO: 10; and / or, the backbone sequence of the guide RNA molecule comprises a nucleotide sequence as set forth in SEQ ID NO: 13;

[0359] (4) the polypeptide construct is a fusion protein;

[0360] (5) the polypeptide construct is a fusion protein comprising an amino acid sequence as set forth in SEQ ID NO: 15 or 16.

[0361] Embodiment 8. A nucleic acid molecule encoding the TadA mutant or functional fragment thereof of any one of embodiments 1-4, or the polypeptide construct of any one of embodiments 5-7.

[0362] Embodiment 9. A vector comprising the nucleic acid molecule of embodiment 8.

[0363] Embodiment 10. A host cell comprising the nucleic acid molecule of embodiment 8 or the vector of embodiment 9.

[0364] Embodiment 11. A composition comprising:

[0365] (a) a first component selected from: (i) a guide RNA molecule comprising a backbone sequence and a guide sequence, the guide sequence being capable of hybridizing to a target RNA sequence, the target RNA sequence comprising a base to be edited; (ii) a nucleotide sequence encoding the guide RNA molecule; (iii) any combination of (i) and (ii);

[0366] and,

[0367] (b) a second component selected from: (i) a polypeptide construct comprising a base deaminase and a guide RNA molecule binding domain linked to the base deaminase, the guide RNA molecule binding domain being capable of binding to a backbone sequence of the guide RNA molecule; (ii) a nucleotide sequence encoding the polypeptide construct; (iii) any combination of (i) and (ii);

[0368] wherein the base deaminase is selected from the TadA mutants or functional fragments thereof according to any one of embodiments 1-4.

[0369] Embodiment 12. The composition of embodiment 11, wherein the guide sequence of the guide RNA molecule is designed to cause the target RNA sequence to form a single-stranded RNA loop region comprising the base to be edited upon hybridization with the target RNA sequence.

[0370] Embodiment 13. The composition of embodiment 12, wherein the target RNA sequence comprises: a target loopable region comprising the base to be edited, a first region located 5’ to the target loopable region, and a second region located 3’ to the target loopable region.

[0371] the guide sequence of the guide RNA molecule is designed to comprise a first region, a second region, and optionally a third region located between the first and second regions, the second region being located upstream of the first region; and the first region is capable of complementarity to the first region of the target RNA sequence, the second region is capable of complementarity to the second region of the target RNA sequence, and the third region is absent or the third region is not capable of complementarity to the target loopable region of the target RNA sequence.

[0372] Embodiment 14. The composition of embodiment 13, having one or more features selected from:

[0373] (1) the third region of the guide RNA molecule is absent;

[0374] (2) the third region of the guide RNA molecule has a number of nucleotide residues that is less than the number of nucleotide residues of the target loopable region of the target RNA sequence;

[0375] (3) the first region and the second region of the guide RNA molecule have the same or different number of nucleotide residues;

[0376] (4) the single-stranded RNA loop region formed upon hybridization of the target RNA sequence with the guide RNA molecule comprises one or more bases to be edited;

[0377] (5) the single-stranded RNA loop region formed upon hybridization of the target RNA sequence with the guide RNA molecule consists of at least 3, at least 4, at least 5, or at least 6 nucleotide residues;

[0378] (6) the polypeptide construct is as defined in any one of embodiments 5-7.

[0379] Embodiment 15. A complex comprising:

[0380] (a) a guide RNA molecule as defined in any one of embodiments 11-14;

[0381] and,

[0382] (b) a polypeptide construct as defined in any one of embodiments 5-7.

[0383] Embodiment 16. A vector system comprising one or more vectors comprising:

[0384] (1) a first nucleotide sequence encoding a guide RNA molecule as defined in any one of embodiments 11-14; and,

[0385] (2) a second nucleotide sequence encoding a polypeptide construct as defined in any one of embodiments 5-7;

[0386] wherein the first nucleotide sequence and the second nucleotide sequence are present on the same or different vectors.

[0387] Embodiment 17. A kit comprising a TadA mutant or functional fragment thereof of any one of embodiments 1-4, a polypeptide construct of any one of embodiments 5-7, a nucleic acid molecule of embodiment 8, a vector of embodiment 9, a composition of any one of embodiments 11-14, a complex of embodiment 15, or a vector system of embodiment 16.

[0388] Embodiment 18. A delivery composition comprising a delivery vehicle, and one or more selected from the group consisting of a TadA mutant or functional fragment thereof of any one of embodiments 1-4, a polypeptide construct of any one of embodiments 5-7, a nucleic acid molecule of embodiment 8, a vector of embodiment 9, a composition of any one of embodiments 11-14, a complex of embodiment 15, a vector system of embodiment 16.

[0389] Embodiment 19. A cell comprising a TadA mutant or functional fragment thereof of any one of embodiments 1-4, a polypeptide construct of any one of embodiments 5-7, a nucleic acid molecule of embodiment 8, a vector of embodiment 9, a composition of any one of embodiments 11-14, a complex of embodiment 15, a vector system of embodiment 16, or a delivery composition of embodiment 18.

[0390] Embodiment 20. A method of editing a target RNA extracellularly, comprising contacting the target RNA with one or more of the TadA mutant or functional fragment thereof of any one of embodiments 1-4, the polypeptide construct of any one of embodiments 5-7, the nucleic acid molecule of embodiment 8, the vector of embodiment 9, the composition of any one of embodiments 11-14, the complex of embodiment 15, the vector system of embodiment 16, and / or the delivery composition of embodiment 18, under conditions suitable for editing the target RNA, thereby inducing deamination of a base to be edited in the target RNA.

[0391] Embodiment 21. A method of editing a target RNA intracellularly, comprising delivering the TadA mutant or functional fragment thereof of any one of embodiments 1-4, the polypeptide construct of any one of embodiments 5-7, the nucleic acid molecule of embodiment 8, the vector of embodiment 9, the composition of any one of embodiments 11-14, the complex of embodiment 15, the vector system of embodiment 16, and / or the delivery composition of embodiment 18 into a cell containing the target RNA, thereby inducing deamination of a base to be edited at the target site.

[0392] Embodiment 22. A cell obtained from the method of embodiment 21, or a progeny thereof, wherein the cell comprises a modification that is not present in its wild type.

[0393] Embodiment 23. The TadA mutant or functional fragment thereof of any one of embodiments 1-4, the polypeptide construct of any one of embodiments 5-7, the nucleic acid molecule of embodiment 8, the vector of embodiment 9, the composition of any one of embodiments 11-14, the complex of embodiment 15, the vector system of embodiment 16, the kit of embodiment 17, or the delivery composition of embodiment 18, for use in RNA editing, or for use in the manufacture of a medicament for RNA editing.

[0394] Definitions of terms

[0395] In the present application, unless otherwise indicated, the scientific and technical terms used herein have the meanings that would be generally understood by one of ordinary skill in the art. Also, the viral, biochemical, immunological laboratory procedures described herein are in accordance with conventional techniques of the corresponding fields. In addition, for better understanding of the present application, the definitions and explanations of relevant terms are provided as follows.

[0396] When the terms “for example,” “for instance,” “such as,” “including,” “containing,” “consisting of,” or variations thereof are used herein, these terms are not to be interpreted in an exclusive or exhaustive sense, i.e., there can be additional examples of enzymes, cells, vectors, compositions, kits, and / or methods that are not expressly mentioned.

[0397] The terms "a" and "an," and the like, as used in describing the present application, (especially in the context of the following claims) should not be construed to cover only a singular form of what is described, but rather the terms "a" and "an," and the like, are to be construed to cover both the singular and plural. The use of the term "about" in describing the present application is intended to cover variations that are outside the precise value, but are within the recognized reasonable range for the particular dimension or value.

[0398] As used herein, the term "base editing" has the meaning commonly understood by those of skill in the art, and is used to refer to a genome editing technology that involves conversion of a particular nucleic acid base to another nucleic acid base at a site of a target nucleic acid (e.g., including in an RNA).

[0399] As used herein, the term "deaminase" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is an adenosine (or adenine) deaminase that catalyzes hydrolytic deamination of adenine or adenosine. In some embodiments, the adenosine deaminase catalyzes hydrolytic deamination of adenine or adenosine to inosine in a ribonucleic acid (RNA). In certain embodiments, the deaminase is a cytidine (or cytosine) deaminase that catalyzes hydrolytic deamination of cytosine or cytidine. In some embodiments, the cytidine deaminase catalyzes hydrolytic deamination of cytosine or cytidine to uridine in a ribonucleic acid (RNA). In certain embodiments, the deaminase possesses both the adenosine deaminase activity and the cytidine deaminase activity, which is capable of catalyzing both hydrolytic deamination of adenine or adenosine and hydrolytic deamination of cytosine or cytidine. In certain embodiments, the deaminase is a single-stranded RNA deaminase, or is modified, evolved, or otherwise altered to be able to utilize single-stranded RNA as a substrate for deamination.

[0400] As used herein, the term "TadA", tRNA adenosine deaminase A, refers to a deaminase that is capable of catalyzing deamination of adenine (A) to inosine (I) using RNA as a substrate. Specific amino acid sequences of various naturally-occurring or artificially-engineered TadA (e.g., the amino acid sequence set forth in any one of SEQ ID NOs: 1-3) are available from public databases (e.g., GenBank database).

[0401] Based on the disclosure herein, one of skill in the art will readily appreciate that the present application provides a series of mutations for modifying TadA deaminase activity that are applicable to existing TadA (including not only natural TadA directly derived from an organism, but also non-natural TadA variants artificially engineered / screened) to have altered base deaminase activity (e.g., enhanced or reduced adenosine deaminase activity, and / or, enhanced or reduced cytidine deaminase activity). Thus, as used herein, the term "TadA mutant" includes not only TadA mutants obtained by performing the mutations of the present application on natural TadA directly derived from an organism, but also TadA mutants obtained by performing the mutations of the present application on non-natural TadA artificially engineered / screened (i.e., engineered / screened non-natural TadA indirectly derived from an organism).

[0402] In the present context, the term "polypeptide construct" is used in its broadest sense. Generally, the polypeptide construct is intended to mean a construct comprising one or more polypeptide or protein components, which can each independently have different origins or different biological activities or functions, and are linked by covalent and / or non-covalent means (e.g., covalent linkage by covalent bonds comprising peptide bonds, isopeptide bonds, and / or disulfide bonds, and / or non-covalent linkage by hydrogen bonds). The polypeptide construct of the present application is not limited in the number of molecular chains (e.g., peptide chains) it contains, e.g., the polypeptide construct of the present application can contain only one molecular chain (e.g., peptide chain), or two or more molecular chains (e.g., peptide chains) covalently and / or non-covalently linked (e.g., covalent linkage by covalent bonds comprising peptide bonds, isopeptide bonds, and / or disulfide bonds, and / or non-covalent linkage by hydrogen bonds) between them. Likewise, one of skill in the art will readily appreciate that in embodiments comprising multiple polypeptide or protein components, the polypeptide or protein components comprised by the polypeptide construct of the present application can be all or partially located in the same molecular chain (e.g., peptide chain), or each in a different molecular chain (e.g., peptide chain). In certain preferred embodiments, the polypeptide construct of the present application is a fusion protein (e.g., a fusion protein comprising a base deaminase and a guide RNA molecule binding domain).

[0403] As used herein, the term "CRISPR system" has the meaning generally understood by those of skill in the art, and generally includes a transcript or other element associated with the expression of a CRISPR-associated ("Cas") gene, or a transcript or other element capable of directing the activity of said Cas gene. Such transcripts or other elements can include a sequence encoding a Cas effector protein and a guide RNA (crRNA) or other sequence or transcript from a CRISPR locus.

[0404] As used herein, the term "guide RNA" has the meaning generally understood by those of skill in the art. Generally, a guide RNA can comprise or essentially consist of a repeat sequence (also referred to as a backbone sequence, which generally is capable of forming a stem loop structure to be specifically recognized and bound by a particular effector protein) and a guide sequence. In certain instances, a guide sequence is any polynucleotide sequence that has sufficient complementarity to a target sequence to hybridize to said target sequence and direct specific binding of an effector protein and guide RNA complex to said target sequence. In certain embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence, when optimally aligned, is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Determining optimal alignment is within the capabilities of those of skill in the art. For example, there are published and commercially available alignment algorithms and programs, such as, but not limited to, ClustalW, Smith-Waterman in matlab, Bowtie, Geneious, Biopython, and SeqMan.

[0405] In certain instances, the guide sequence is at least 5, at least 10, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 35, at least 40, at least 45, or at least 50 nucleotides in length. In certain instances, the guide sequence is no more than 50, 45, 40, 35, 30, 25, 24, 23, 22, 21, 20, 15, 10, or fewer nucleotides in length. In certain embodiments, the guide sequence is 10-35 in length.

[0406] In certain instances, the repeat sequence is at least 10, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, or at least 70 nucleotides in length. In certain instances, the repeat sequence is no more than 70, 65, 64, 63, 62, 61, 60, 59, 58, 57, 56, 55, 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 15, 10, or fewer nucleotides in length. In certain embodiments, the repeat sequence is 55-70 nucleotides in length. In certain embodiments, the repeat sequence is 15-50 nucleotides in length.

[0407] In the present application, the expression "target sequence" can be any endogenous or exogenous polynucleotide to a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. In certain instances, the target sequence or target polynucleotide can include disease-associated polynucleotides as well as signaling biochemical pathway-associated polynucleotides.

[0408] As used herein, the terms "non-naturally occurring" or "engineered" are used interchangeably and indicate human intervention. When these terms are used to describe a nucleic acid molecule or polypeptide or composition or complex, it indicates that the nucleic acid molecule or polypeptide or composition or complex is at least substantially isolated from at least another component with which it is associated in nature or as it is found in nature, or, it indicates that the nucleic acid molecule or polypeptide comprises at least one component or moiety that was formed with artificial involvement (e.g., artificial design or modification).

[0409] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences), and detailed descriptions of the same can be found in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). In certain instances, regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells as well as those that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in a desired tissue of interest, such as muscle, neuronal, bone, skin, blood, a particular organ (e.g., liver, pancreas), or a particular cell type (e.g., lymphocytes). In certain instances, regulatory elements can also direct expression in a temporal manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which can or can not be tissue- or cell type-specific. In certain instances, the term "regulatory element" encompasses enhancer elements, such as the WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I (Mol. Cell. Biol., vol. 8(1), pp. 466-472, 1988); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit beta-globin (Proc. Natl. Acad. Sci. USA., vol. 78(3), pp. 1527-31, 1981).

[0410] As used herein, the term "promoter" has its art-understood meaning and refers to a non-coding nucleotide sequence located upstream of a gene that initiates expression of the downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked with a polynucleotide encoding or specifying a gene product, results in production of the gene product in a cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked with a polynucleotide encoding or specifying a gene product, results in production of the gene product in a cell essentially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked with a polynucleotide encoding or specifying a gene product, results in production of the gene product in a cell essentially only when the cell is of the tissue type to which the promoter corresponds.

[0411] As used herein, the term "operably linked" is intended to mean that a nucleotide sequence of interest is linked to the regulatory element(s) in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or, when the vector is introduced into a host cell, in the host cell).

[0412] As used herein, the term "hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via the hydrogen bonding between the bases of the nucleotide residues. The hydrogen bonding can occur by means of Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner. The complex can comprise two strands forming a duplex, three or more strands forming a multiwelled complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction can constitute a step in a more extensive process, such as the initiation of PCR, or cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing to a given sequence is referred to as the "complement" of the given sequence.

[0413] As used herein, the term "upstream" is used to describe the relative position of two nucleic acid sequences (or two nucleic acid molecules) and has the meaning generally understood by those skilled in the art. For example, the expression "a nucleic acid sequence is upstream of another nucleic acid sequence" means that, when arranged in the 5' to 3' direction, the former is located more toward the 5' end (i.e., is located more proximal to the 5' end) than the latter. As used herein, the term "downstream" has the opposite meaning of "upstream."

[0414] The terms "complementary" or "complementarity," as described herein, are used in reference to the sequence of nucleotides which, in a double-stranded polynucleotide, form hydrogen bonds with opposite sequences in the other polynucleotide strand. For example, the sequence A-G-T is complementary to the sequence 3'-T-C-A-5'. Complementarity can be "partial," in which only some of the nucleic acid bases match according to the base pairing rules of A with T or G with C (or G with A), or "complete" or "total," where all the matching nucleic acid bases are pairable. Most of the description herein refers to the complementarity of A-T and G-C base pairs; however, when addressing the complementarity of alternative base pairs such as A-U and G-U, the same principles apply.

[0415] As used herein, the term "mutant," in the context of polypeptides (including polypeptides), also refers to a polypeptide or peptide comprising an amino acid sequence that has been altered by the introduction of an amino acid residue substitution, deletion, or addition. In certain instances, the term "mutant" also refers to a polypeptide or peptide that has been modified (i.e., by covalently linking any type of molecule to the polypeptide or peptide). For example, but not by way of limitation, a polypeptide can be modified, e.g., by glycosylation, acetylation, pegylation, phosphorylation, amidation, derivatization by known protecting / blocking groups, proteolytic cleavage, attachment to a cellular ligand or other protein, etc. A derivatized polypeptide or peptide can be produced by chemical modification using techniques known to those of ordinary skill in the art, including, but not limited to, specific chemical cleavage, acetylation, formylation, metabolic synthesis in the presence of tunicamycin, etc. Furthermore, a mutant has similar, identical, or improved function as compared to the polypeptide or peptide from which it is derived.

[0416] As used herein, the term "vector" refers to a nucleic acid vehicle into which a polynucleotide can be inserted. When the vector is capable of directing the expression of a polynucleotide inserted into it, the vector is referred to as an expression vector. A vector can be introduced into a host cell by transformation, transduction or transfection, and directs the expression of elements of genetic material it carries in the host cell. Vectors are well known to those of ordinary skill in the art and include, but are not limited to, plasmids; phagemids; cosmids; artificial chromosomes, such as yeast artificial chromosomes (YACs), bacterial artificial chromosomes (BACs), or P1 -derived artificial chromosomes (PACs); bacteriophages, such as lambda phage or M13 phage; and animal viruses. Animal viruses that can be used as vectors include, but are not limited to, retroviruses (including lentiviruses), adenoviruses, adeno-associated viruses, herpesviruses (such as herpes simplex virus), poxviruses, baculoviruses, papillomaviruses, papova viruses (such as SV40). A vector can contain a variety of elements that control expression, including, but not limited to, promoter sequences, transcriptional start sequences, enhancer sequences, selection elements, and reporter genes. In addition, a vector can contain a replication origin.

[0417] In the present application, the terms "polypeptide" and "protein" have the same meaning and are used interchangeably. Also in the present application, amino acids are generally represented by the one-letter and three-letter abbreviations well known in the art. For example, alanine can be represented by A or Ala.

[0418] It is known to those skilled in the art that, during the process of translation of mRNA, the first amino acid of the polypeptide chain produced is often the amino acid encoded by the start codon (e.g., methionine (M)) due to the role of the start codon. Therefore, the polypeptides or proteins involved in the present application (e.g., the base deaminase of the present application, the TadA of the present application, the TadA mutant of the present application, the polypeptide construct of the present application) not only encompass the amino acid sequences that do not comprise the amino acid encoded by the start codon (e.g., methionine) at the N-terminus, but also encompass the amino acid sequences that comprise the amino acid encoded by the start codon (e.g., methionine) at the N-terminus.

[0419] Unless otherwise indicated herein or clearly contradicted by context, "A, B, and / or C" or similar expressions shall be understood as "A, B, C, or any combination thereof", e.g., as being understood as being selected from any one of A, B, C, A and B, A and C, B and C, A and B and C.

[0420] Advantages of the invention

[0421] The RNA base editing tool provided by the present application has one or more advantages selected from the following:

[0422] (1) Existing RNA editing tools (including DECOR) are all for editing at a single site, without controllable editing window. The present application develops, for the first time, an RNA editing tool with controllable editing window by using single-strand deaminase (e.g., TadA) and special LF-gRNA;

[0423] (2) The present application provides novel TadA mutants by rational design and mutation of TadA8e, significantly reduces bystander editing and off-target of the whole transcriptome, thereby realizing A-to-I base editing on RNA with less off-target and high precision (i.e., AIM-A and AIM-Amax systems);

[0424] (3) The present application provides novel TadA mutants by rational design and mutation of TadA8e, thereby realizing C-to-U base editing on RNA with high efficiency and no residual A-to-I activity (i.e., AIM-C and AIM-Cmax systems);

[0425] (4) There is currently a lack of bifunctional base editors for RNA in the field. The present application provides novel TadA mutants by rational design and mutation of TadA8e, thereby realizing simultaneous A-to-I and C-to-U base editing on a single RNA molecule (i.e., AIM-A&C and AIM-A&Cmax systems).

[0426] (5) The AIM-A, AIM-C and AIM-A&C systems provided by the application can achieve a controllable number of RNA base editing within a certain range.

[0427] Embodiments of the application will be described in detail below with reference to the accompanying drawings and examples, but those skilled in the art will understand that the following drawings and examples are only used to illustrate the application, and are not a limitation on the scope of the application. According to the following detailed description of the drawings and preferred embodiments, various objects and advantages of the application will become apparent to those skilled in the art. BRIEF DESCRIPTION OF DRAWINGS

[0428] Figure 1 AIM design schematic diagram. A. Main components of AIM: single-stranded deaminase TadA; dead Psp Cas13b and specially designed LF-gRNA that can generate a loop at the target site. B. Schematic diagram of AIM to achieve a controllable window. By designing different LF-gRNA to generate loops of different sizes, different width editing windows are achieved.

[0429] Figure 2 Using single-stranded deaminase TadA to achieve programmable A-to-I editing. A. TadA and its substrate tRNA structure schematic diagram. B. Schematic diagram of AIM tool structure. C. Reporter-UAG structure schematic diagram, and crRNA (-38) and crRNA (+40) position explanation targeting Reporter-UAG. D. Editing efficiency of different TadA variants on Reporter-UAG.

[0430] Figure 3 Using loop-forming guide RNA (LF-gRNA) to construct editing window. A. LF-gRNA-10 nt induces a 10-nucleotide single-stranded RNA loop on the targeted mRNA. B. EGFP expression signal of traditional crRNA and LF-gRNA-10 nt. C. A-to-I editing efficiency of LF-gRNA and traditional crRNA combined with TadA7.10 and TadA8e.

[0431] Figure 4 Parameter exploration of loop-forming guide RNA (LF-gRNA). A. Editing efficiency of various LF-gRNAs with different arm length combinations. B. Editing efficiency of LF-gRNAs with asymmetric arms. C. A series of LF-gRNAs were designed with the target base A located at positions 1 to 10 of the RNA loop, and the A-to-I editing efficiency was evaluated.

[0432] Figure 5A. "Reporter vector-4A" and the design of different LF-gRNAs. B. TadA7.10 can achieve different sizes of editing window in the design of different LF-gRNAs. The larger the loop ring caused by LF-gRNA at the target site, the wider the editing window. C. TadA8e can achieve different sizes of editing window in the design of different LF-gRNAs. The larger the loop ring caused by LF-gRNA at the target site, the wider the editing window.

[0433] Figure 6 Novel TadA mutants achieve precise and efficient A-to-I editing. A. Analysis of A-to-I editing events and proportions caused by TadA8e in the 100 bp region of the reporter transcript (Reporter vector-TAG). Red represents the editing efficiency of A in the loop caused by LF-gRNA, blue represents the editing efficiency of A outside the loop caused by LF-gRNA, and blue is off-target editing. B. Structure of ABE8e in single-stranded DNA binding state, aligned with the co-crystal structure of Staphylococcus aureus TadA-tRNA. Key amino acid sites are divided into 4 groups according to different functional regions. C. Analysis of A-to-I editing events and proportions caused by TadA8e (F84I) in the 100 bp region of the reporter transcript (Reporter vector-TAG). D. Editing efficiency of TadA8e (F84I) and TadA8e (F84I-V82T) on Reporter vector-TAG. E. Comparison of mutation sites of different TadA variants. F, G. Analysis of A-to-I editing events and proportions caused by AIM-A (F) and AIM-Amax (G) in the 100 bp region of the reporter transcript (Reporter vector-TAG).

[0434] Figure 7 AIM-A can achieve a controllable number and position of editing within a certain region. A. Representative sequence display of novel reporter genes containing isolated and continuous adenosine. LF-gRNAs corresponding to them are displayed, with loop sizes ranging from 4 nucleotides to 14 nucleotides. B. A-to-I editing efficiency mediated by AIM-A under the design of LF-gRNAs with different loop sizes. C. Representative sequence display of novel reporter genes containing isolated and continuous adenosine. LF-gRNA designs with different loop positions are displayed. D. A-to-I editing efficiency mediated by AIM-A under the design of LF-gRNAs with different loop positions.

[0435] Figure 8New TadA mutants enable C-to-U editing. A. Structure display of Reporter-CAC. B. Analysis of A-to-I and C-to-U editing events and ratios caused by TadA8e in the 100 bp region of the reporter transcript (Reporter-TAG). C. Structure of ABE8e in single-stranded DNA binding state, aligned with the co-crystal structure of S. aureus TadA-tRNA. Key amino acids that can affect C-to-U editing efficiency are labeled in dashed circles. D. A-to-I and C-to-U editing efficiency of N46 and A48 combined mutants. E. A-to-I and C-to-U editing efficiency of activation mutants based on N46P-A48G.

[0436] Figure 9 New TadA mutants enable precise and efficient C-to-U editing. A. Analysis of A-to-I and C-to-U editing events and ratios caused by AIM-C in the 100 bp region of the reporter transcript (Reporter-TAG). B. Sequence alignment of TadA variants for DNA CBE with TadA8e and AIM-C. C. A-to-I and C-to-U editing efficiency of activation mutants based on AIM-C. D. Analysis of A-to-I and C-to-U editing events and ratios caused by AIM-Cmax in the 100 bp region of the reporter transcript (Reporter-TAG).

[0437] Figure 10 New TadA mutants enable precise and efficient C-to-U editing. Left, a set of Reporter sequences, with which loops of 5-17 nt can be generated using the same LF-gRNA. Right, editing efficiency of AIM-C with different loop sizes.

[0438] Figure 11 New TadA mutants enable simultaneous A-to-I and C-to-U editing. A. Saturation mutagenesis of N46 site of TadA8e (A48G), statistics of A-to-I and C-to-U editing efficiency of mutants. B. Statistics of A-to-I and C-to-U editing efficiency of activation mutants based on TadA8e (A48 series mutations-N46C). HE represents 4 mutations: R26G / V28A / Y73P / H96N. C. Ratio of A-to-I and C-to-U editing caused by AIM-A&C in the same transcript. D. Further activation of AIM-A&C to obtain AIM-A&Cmax. E. A set of Reporter sequences, with which loops of 5-17 nt can be generated using the same LF-gRNA. F. Editing efficiency of AIM-&C with different loop sizes.

[0439] Sequence information

[0440] The description of the sequences involved in the present application is provided in the following table.

[0441] Table 1: Sequence information

[0442]

[0443] Note: The amino acid residues different from TadA8e in each TadA mutant are underlined. DETAILED DESCRIPTION

[0444] The present application will now be described with reference to the following examples, which are intended to be illustrative, but not limiting, of the present application. It will be known to the person skilled in the art that the examples describe the present application by way of example, and are not intended to limit the scope of the present application as claimed.

[0445] Experimental methods

[0446] 1. Construction of TadA

[0447] For TadA expression plasmids, the coding sequences of inactivated PspCas13b (amino acid sequence as SEQ ID NO: 10) and TadA (amino acid sequences of TadA8e and each TadA mutant based thereon are shown in Table 1) were synthesized according to the information provided by Addgene, and were ligated into pcDNA3.1 backbone (transcription driven by CMV promoter, terminated by BGH poly-A signal) by Gibson cloning method (exemplary amino acid sequences of fusion proteins are as SEQ ID NOs: 14-20) with synthetic linker sequence containing NLS / NES in the middle of the fragments. Sequence PCR amplification used high-fidelity enzyme 2x Phanta Max MasterMix (Vazyme). The DNA fragments obtained by PCR amplification were purified and recovered using TIANprep Midi Purification Kit (Tiangen). The concentrations of plasmids, purified DNA fragments and extracted RNA were measured using a Nanodrop spectrophotometer.

[0448] For crRNA expression plasmids, a PspCasl3b crRNA backbone plasmid was constructed according to the information provided by Addgene (No. 103854). The spacer sequence of crRNA was synthesized by fragment synthesis and inserted into the PspCasl3b crRNA backbone (SEQ ID NO: 13) by Goldengate cloning method (transcription driven by human U6 promoter, terminated by poly-U signal). Goldengate assembly was performed with Type IIS restriction enzyme Esp3I (Thermo Fisher Scientific) and T4 DNA ligase (New England Biolabs).

[0449] All assembled plasmid vectors were transformed into Trans1-T1 phage resistance chemically competent bacteria and cultured on Luria-Bertani (LB) agar medium containing penicillin. The plasmid sequences were verified using Sanger sequencing and the verified bacteria containing plasmids were cultured in LB medium containing penicillin. Endotoxin-free plasmid DNA was extracted using TIANprep Midi Plasmid Kit kit for subsequent cell transfection.

[0450] 2. Cell culture and liposome transfection

[0451] HEK293T cells were cultured at 37 °C, 5% CO2, in DMEM medium containing 10% FBS and 1% penicillin-streptomycin. When the cells were subcultured, the cells were washed with PBS and treated with 0.25% trypsin and incubated at 37 °C for 2 minutes. Then, the trypsin was neutralized with FBS-containing medium. After centrifugation at 600 rpm for 3 minutes, the cells were counted and plated. The mycoplasma contamination of the cell line was negative.

[0452] HEK293T cells were plated in 24-well plates at a density of 1.75 x 10 5 cells / well and transfected when they reached approximately 70% confluency after 16-18 hours of culture. Transfection was performed using Lipofectamine LTX with Plus Reagent (Thermo Fisher Scientific) according to the manufacturer's instructions.

[0453] To evaluate the RNA editing efficiency on the exogenous reporter gene, HEK293T cells were transfected with 500 ng of AIM expression vector (which contains the coding sequence of the fusion protein of dCasl3b and TadA), 500 ng of LF-gRNA plasmid and 125 ng of exogenous reporter gene plasmid. After 72 hours, RNA was extracted from the cells and the editing efficiency was evaluated by Sanger or high-throughput sequencing.

[0454] 3. RNA extraction and reverse transcription

[0455] After 72 hours of cell transfection, the cell culture medium in the wells of the culture dish was aspirated, 500 μΐ of Trizol total RNA extraction reagent (Thermo Fisher Scientific) was added to each well, and the mixture was incubated at room temperature for 10 minutes, followed by the addition of 100 μΐ of chloroform. After the mixture was shaken, it was layered at room temperature for three times, and then centrifuged at 12,000 rpm at 4°C for 15 minutes. The supernatant was transferred to a new tube, and then an equal volume of isopropanol was added. After mixing well, the mixture was placed at -20°C for at least 1 hour, and then centrifuged at 14,500 rpm at 4°C for 30 minutes. After centrifugation, a flocculent gelatinous precipitate, i.e., an RNA precipitate, was observed on the wall and bottom of the tube. The RNA precipitate was washed twice with 1 ml of 75% ethanol, and then the supernatant was removed. After the precipitate was dried in air for 5-10 minutes, the RNA precipitate was resuspended with RNase-free water, and the concentration was determined. The cDNA was generated from the extracted RNA using random primers and oligo-dT primers using the HiScript III 1st Strand cDNA Synthesis Kit (Vazyme) according to the manufacturer's instructions, and downstream experiments were performed.

[0456] 4. Amplicon library construction for next-generation sequencing

[0457] The amplicon library was obtained through two rounds of PCR amplification reactions. The first round of PCR reaction (total reaction system 10 μΐ) was amplified using Q5 Hot Start High-Fidelity 2X Master Mix (NEB, #M0543), and 0.5 μΜ of gene-specific primers carrying Illumina adapter sequences were added. After 10 cycles of PCR amplification reaction, the PCR product was purified using 1-fold volume of AMPure XP magnetic beads (BECKMAN, #A63882). Next, the second round of PCR reaction (30 μΐ) was performed using unique Illumina barcode primers (each 0.5 μΜ) containing adapter sequences. After 15 cycles of the second round of PCR amplification reaction, the product was purified using 0.8-fold volume of XP magnetic beads. The fragment size and distribution of the library were evaluated by the Agilent 4150 TapeStation system (Agilent, G2992AA). The qualified library was subjected to 150 bp paired-end sequencing using the Illumina NovaSeq platform.

[0458] Experimental results

[0459] 1. Programmable A-to-I editing using the single-stranded deaminase TadA.

[0460] Inspired by CRISPR-based DNA base editors, which utilize single-stranded DNA deaminases to introduce point mutations in specific single-stranded DNA regions, the inventors of this application sought a single-stranded RNA deaminase to achieve RNA editing. Currently, the APOBEC family and TadA are known to possess single-stranded RNA deamination activity. However, while APOBEC1 and APOBEC3A have been used for RNA editing, they exhibit significant editing activity against single-stranded DNA and even detectable activity against double-stranded DNA, resulting in off-target editing at the genome level. Bacterial TadA catalyzes the editing of tRNA... Arg2 Deamination reaction on the stem-ring structure of the anticodon ( Figure 2 A); while the evolved Escherichia coli TadA (EcTadA) can act on single-stranded DNA and has low activity on single-stranded RNA, but no activity on double-stranded DNA.

[0461] The inventors of this application attempt to utilize TadA as an effector protein for a programmable RNA editor. To perform programmable RNA base editing using TadA, the inventors of this application combine a single TadA with a Cas13b (dCas13b) that targets RNA but lacks nuclease activity. Figure 2 B) were fused together. To assess the efficiency of RNA A-to-I editing, a reporter vector containing UAG was constructed ( Figure 2 C), which contains a UAG premature stop codon located within the protein linker sequence connecting two fluorescent proteins (mCherry and EGFP). The adenine-specific deamination reaction within the UAG stop codon reactivates EGFP expression, thus allowing for the quantification of A-to-I editing efficiency. Following standard Cas13b crRNA design principles, two crRNAs were generated, located upstream (-38) or downstream (+40) of the adenine in the stop codon, respectively. Figure 2 C). The programmable RNA A-to-I editing capabilities of wild-type EcTadA (SEQ ID NO: 3), TadA7.10 (SEQ ID NO: 2), and TadA8e (SEQ ID NO: 1) were then evaluated. The results showed that TadA7.10 and TadA8e exhibited A-to-I conversion rates of 20%–48% at the target sites, while WT-TadA and completely inactivated TadA (E59A) showed only background-level A-to-I editing. Figure 2 D). Therefore, it is concluded that TadA7.10 and TadA8e, rather than WT-TadA, can be used for programmable RNA A-to-I editing.

[0462] 2. Constructing an editing window with loop-forming guide RNA (LF-gRNA)

[0463] Based on the discovery that TadA7.10 and TadA8e can be used for programmable RNA A-to-I editing, further, the present inventors tried to construct an editing window for TadA. The present inventors designed a crRNA to form a single-stranded RNA loop surrounded by double-stranded RNA in the targeted transcript. The present inventors named such non-standard crRNA as loop-forming guide RNA (LF-gRNA, Figure 1 A and Figure 2 B). And further confirmed that the induced single-stranded RNA loop can be used as a high-efficiency substrate for TadA, while the adjacent nucleotides in the double-stranded RNA region are expected to be shielded and protected from deamination.

[0464] Firstly, the editing efficiency mediated by LF-gRNA was evaluated. LF-gRNA-10 nt was designed to induce a 10-nucleotide single-stranded RNA loop on the targeted mRNA, and the targeted A is expected to be located at the sixth position. This single-stranded RNA loop is surrounded by symmetrically complementary double-stranded RNA arms (left 15 nucleotides + right 15 nucleotides, Figure 3 A). Compared with traditional crRNA, LF-gRNA-10 nt showed an average 2.3-fold increase in EGFP expression signal ( Figure 3 B). Next-generation sequencing (NGS) analysis confirmed that the combination of LF-gRNA with TadA7.10 and TadA8e achieved A-to-I editing efficiencies of up to 62% and 73%, respectively ( Figure 3 C). Further testing of various LF-gRNAs with different arm length combinations found that symmetric RNA complementary arms from 24 bp to 36 bp all achieved high-efficiency editing (>30%, Figure 4 A). In addition, LF-gRNAs with asymmetric arms also showed good editing efficiency (20%-65%, Figure 4 B).

[0465] Next, the position of base A in the 10-nucleotide single-stranded RNA loop was studied to see how it would affect editing efficiency. A series of LF-gRNAs were designed to have the targeted base A located at positions 1 to 10 of the RNA loop ( Figure 4C). The data showed that TadA7.10 exhibited high efficient A-to-I conversion (>10%) on the bases A located at positions 4-8 of the 10-nucleotide loop, with the major window (A-to-I editing rate >30%) located at positions 5-8. TadA-8e exhibited a wider window, covering positions 3-9, with the major window located at positions 4-9 (A-to-I editing rate >30%). Figure 4 C).

[0466] 3. Control the width of editing window by changing loop size

[0467] Further attempts were made to control the width of editing window by manipulating the size of RNA loop induced by LF-gRNA. To facilitate the observation of editing window, a “reporter vector-4A” containing a targeting site with 4 consecutive adenines (A) was constructed (Figure 3A). Figure 5 A). Various LF-gRNAs for reporter vector-4A were designed, inducing loop sizes ranging from 4 nucleotides (LF-gRNA-4 nt) to 14 nucleotides (LF-gRNA-14 nt). The editing efficiency of A located within the ssRNA loop and adjacent dsRNA arms was detected using NGS.

[0468] The results showed that all LF-gRNAs that induced RNA loops no smaller than 6 nucleotides in size exhibited significant A-to-I editing, with TadA7.10 (5.6%-23.3% of A with the highest efficiency, Figure 5 B) and TadA8e (11.9%-33.8% of A with the highest efficiency, Figure 5 C). LF-gRNA-4 nt exhibited low A-to-I editing efficiency (TadA7.10: 0.6%-2.9%; TadA8e: 1.2%-6.1%), revealing a potential lower limit of RNA loop size. As the loop size increased, the editing window expanded accordingly. Specifically, LF-gRNA-6 nt formed a 1-2 nucleotide narrow window, with the highest efficient editing site located at position 5 of the loop. LF-gRNA-8 nt expanded the window size to 3-4 nucleotides, with the highest efficient editing site located at position 5. Moreover, LF-gRNA-14 nt could edit multiple A within the loop, forming a 7-nucleotide and a 10-nucleotide wide editing window in TadA7.10 and TadA8e, respectively (Figure 3B and Figure 5 B and Figure 5 C). Therefore, the present application successfully controlled the width of editing window by manipulating the size of ssRNA loop induced by LF-gRNA.

[0469] 4. Identify novel TadA mutants to achieve precise and efficient A-to-I editing

[0470] To investigate the editing specificity of TadA7.10 and TadA8e under the LF-gRNA design, we analyzed A-to-I editing events located within 100 bp of the reporter transcript. We found a large number of A-to-I off-target sites, which were located outside the double-stranded RNA arm region, and even off-targets were detected in the double-stranded RNA arm region Figure 6 A). Therefore, there is an urgent need to identify new TadA mutants to reduce off-target reactions.

[0471] The present inventors tried to design TadA structural guide mutants for RNA base editors. Since there is no structural information to elucidate how the evolved TadA recognizes and catalyzes the cyclic RNA substrate, the present inventors used the structure of ABE8e in a single-stranded DNA binding state, aligned with the structure of S. aureus TadA-tRNA, to guide the redesign of TadA8e, aiming to reduce off-target editing of RNA. To evaluate the precision of the mutants, the efficiency of targeting the target was divided by the average editing efficiency on the transcript within a 100 bp range around the targeted target, where the precision of TadA8e was normalized to 1. The present inventors designed four strategies to affect the substrate binding and catalytic activity of TadA Figure 6 B).

[0472] (1) First, residues interacting with the RNA substrate (R74, R98, R129, R64, V88, H96, V120) were focused on. These residues were substituted with representative amino acids with different properties in order to weaken the non-specific RNA binding affinity. Among them, R98A retained a relatively high target editing activity (>50%) and the precision was improved compared to the original TadA8e (Table 2).

[0473] (2) Next, residues in the TadA dimer interface (R111, L145, Y149, Q154, V155) were focused on. These residues were substituted with representative amino acids with different properties in order to weaken the ability of TadA to form dimers. Studies have shown that the formation of TadA dimers can promote the deamination activity of TadA. Among them, the mutation of the R111 site can greatly reduce the bystander editing activity, although the targeted editing activity is weakened (Table 1).

[0474] (3) Subsequently, N108 and V106 were saturated mutated, which appeared in the early stage of TadA7.10 (for A-to-I editing on single-stranded DNA) directed evolution. The present inventors found that most N108 mutations significantly reduced A-to-I editing activity, while most V106 variants changed little (Table 1). This indicates that N108D mutation not only initiates single-stranded DNA editing activity during directed evolution, but also has a significant impact on ssRNA activity.

[0475] (4) Subsequently, residues in the adenine binding pocket (V28, V30, N46, F84) were tested, which might affect the catalytic activity of evolved TadA. As expected, several mutations at the edge of the active pocket (V28D; V30W, Q, R and all N46 variants) significantly reduced A-to-I editing activity. Interestingly, four mutations of F84 showed a clear reduction in bystander editing, but retained partial targeting efficiency, thus obtaining a higher specificity value (Table 2). Inspired by these observations, F84 was saturated mutated, and it was found that F84I, F84V and F84H could well balance efficiency and specificity (Table 2). Further analysis of A-to-I editing events in all amplicon sequencing (amplicon sequencing of the region about 100 bp upstream and downstream of the targeted site) again demonstrated the high specificity of these F84 variants (Table 2). Figure 6 C). In summary, by screening mutations at about 20 residue sites of TadA, the present application determines that F84 is a key residue that determines the specificity of TadA RNA editing.

[0476] Table 2 Targeting efficiency and precision evaluation of TadA mutants.

[0477]

[0478] To further enhance its activity, a mutation V82T was superimposed on the basis of TadA8e (F84I), showing ~1.5-fold higher activity (Table 3). Figure 6 D). In addition, we analyzed TadA-8.20, a super-active TadA mutant sequence (Table 3). Figure 6E), and introducing I76Y, D147R, and Q154R into TadA8e (F84I-V82T). The results showed that the triple mutant F84I-V82T-D147R and the quintuple mutant F84I-V82T-I76Y-D147R-Q154R reached 62% and 76% of targeted editing efficiency, respectively. Further analysis of A-to-I editing events in amplicon sequencing (regions around 100 bp upstream and downstream of the targeted site were amplified for amplicon sequencing) again demonstrated the high specificity of these F84 variants. Figure 6 F and G). Therefore, we named the new TadA8e mutants carrying the triple mutation (F84I-V82T-D147R) and the quintuple mutation (F84I-V82T-I76Y-D147R-Q154R) connected with dCas13b as AIM-A and AIM-Amax, respectively.

[0479] To further evaluate whether AIM-A has controllable editing ability within a certain range, a new reporter gene containing isolated and consecutive adenosines was constructed Figure 7 A). Various LF-gRNAs were designed against this reporter gene to induce loops ranging from 4 nucleotides (LF-gRNA-4 nt) to 14 nucleotides (LF-gRNA-14 nt). The results showed that LF-gRNAs inducing loops no smaller than 6 nucleotides could all mediate significant A-to-I editing by AIM-A (the highest adenosine editing rate was 11.4%-28.0%, Figure 7 B). LF-gRNA-4 nt resulted in a lower A-to-I editing rate, revealing the lower limit of loop size. The number of editable bases increased as the loop expanded: LF-gRNA-6 nt mediated the editing of two adenosines, with the most efficient site located at the 4th position of the loop; while LF-gRNA-14 nt edited 6 adenosines within the loop, covering a window as wide as 10 nucleotides Figure 7 B). The background editing levels of all LF-gRNAs within the double-stranded RNA arm regions were low, indicating that the bases in these regions were effectively protected.

[0480] To demonstrate selective editing of adenosines in the reporter gene, various LF-gRNAs with a fixed 7 nt loop were designed, with the loop gradually moving one nucleotide towards the 5' end Figure 7 C). The results showed that based on different designs of LF-gRNAs, it was possible to selectively edit one isolated adenosine, or preferentially edit two or more consecutive adenosines Figure 7 D). It was thus concluded that AIM could edit a specified number and position of bases.

[0481] 5. Identification of novel TadA mutants for precise and efficient C-to-U editing

[0482] Further, the inventors of the present application performed protein engineering of TadA proteins in an attempt to achieve precise and efficient C-to-U editing. First, Reporter-CAC was constructed to quantify C-to-U editing levels (Fig. 1A). Figure 8 A). However, the results showed that TadA8e had very weak C-to-U activity on RNA (~2%, Figure 8 B).

[0483] Subsequently, the inventors of the present application attempted to find new TadA variants for efficient C-to-U RNA editing. First, it was hypothesized that a smaller cytosine loop (compared to a larger adenine) should be more deeply seated in the active pocket to achieve efficient deamination reactions; by analyzing the structure of the bound substrates (EcTadA-ssDNA and SaTadA-tRNA), three loop regions were identified that could potentially facilitate cytosine deep-seating (Fig. 1C). Figure 8 C). Mutations were made to E27, V28, P29, V30 (in loop I), N46, A48 (in loop II), and K110, R111 (in loop III), and the target C-to-U and the adjacent A-to-I editing efficiencies were detected. Through screening of over 130 variants, A48F and A48G were found to have higher C-to-U activity than TadA-8e (4.6% and 2.6%, respectively, Table 3). However, these variants still had significant A-to-I activity (~40%).

[0484] To obtain a pure C-to-U editor, the A-to-I activity must be eliminated. During the screening process, the inventors of the present application noted that N46 was critical for the A-to-I RNA activity of TadA8e, as 18 out of 20 N46 mutations lost activity (<1%, Table 3). Notably, there were four N46 variants that had no A-to-I activity but still had weak and detectable C-to-U activity (>0.2%); among them, N46P performed the best, with almost no A-to-I activity (0.04%). Therefore, combining A48F or A48G with N46P to achieve efficient and pure C-to-U RNA editing is very attractive. N46P-A48F and N46P-A48G showed 0.68% and 7.4% C-to-U activity, respectively, and no detectable A-to-I editing (Fig. 2B). Figure 8D). To enhance C-to-U activity, we checked whether the hyperactive mutations of AIM-A are compatible with the N46P-A48G double mutations. We found that adding V82T, I76Y, D147R or Q154R mutations alone can increase the efficiency by about 2-fold; adding the four mutations together in N46P-A48G achieved about 38.1% C-to-U editing, equivalent to about 5-fold improvement in C-to-U efficiency. Figure 8 E). This new variant (N46P-A48G-V82T-I76Y-D147R-Q154R) is named AIM-C by the present application. Further analysis of A-to-I and C-to-U editing events in all amplicon sequencing (amplicon sequencing of the region about 100 bp upstream and downstream of the target site) proves that AIM-C has no A-to-I activity and has high-precision C-to-U editing. Figure 9 A).

[0485] To further enhance the activity of AIM-C, we performed sequence alignment of TadA variants for DNA CBE and found multiple regions with different amino acid residues compared to TadA8e, especially in the R26-V28 region and the TadA dimerization interface (H96 and Y73). Figure 9 B). Through single and combined mutations, we determined that the four mutations (R26G, V28A, Y73P and H96N) alone showed a 1.2 to 1.4-fold improvement. Figure 9 C). After combining the four mutations, the efficiency was improved by 2.1-fold compared to the original AIM-C, achieving an editing efficiency of 66% in Reporter-CAC. Figure 9 C). This new variant (adding R26G / V28A / Y73P / H96N mutations based on AIM-C) is named AIM-Cmax by the present application. Further analysis of A-to-I and C-to-U editing events in all amplicon sequencing (amplicon sequencing of the region about 100 bp upstream and downstream of the target site) proves that AIM-Cmax has no A-to-I activity and has high-precision C-to-U editing. Figure 9 D).

[0486] Table 3. A-to-I and C-to-U efficiency statistics of new TadA mutants

[0487]

[0488] Furthermore, AIM-C was further tested for its ability to achieve controllable C-to-U editing within a certain range. The results showed that a loop size ranging from 7 nt to 17 nt was able to achieve C-to-U editing (with the highest editing rate of 7.5%~23.5%); while a 5 nt loop showed lower activity (0.5%). Consistent with AIM-A, the number of editable bases increased with the expansion of the loop: a 7 nt loop mediated editing of a single cytosine at position 5, while expanding the loop size to 17 nt enabled editing between positions 4 to 13, forming a 10 nt wide editing window Figure 10 ). Overall, these results demonstrated that AIM-C was able to achieve controllable C-to-U editing within a user-defined region.

[0489] 6. Identification of novel TadA mutants for simultaneous A-to-I and C-to-U editing on a single transcript

[0490] Simultaneous editing of A and C bases provides new opportunities for RNA information manipulation. For example, it allows for an additional 30 codon changes compared to single-function editing, resulting in 10 amino acid substitutions. Therefore, it was a further goal of the present application to develop a dual-function AIM system capable of achieving both A-to-I and C-to-U changes.

[0491] During the screening of AIM-C, the present inventors found that TadA8e (A48G) exhibited strong A-to-I activity and weak but clear C-to-U activity (Table 3), which provided a good starting point for developing a dual-function base editor. Since N46 is particularly important for A-to-I activity, N46 site of TadA8e (A48G) was subjected to saturation mutagenesis in an attempt to balance the A and C activities. N46C-A48G was identified, which exhibited moderate but most balanced A&C activities (A-to-I: 6.8%; C-to-U: 3.4%, Figure 11 A). To improve the editing efficiency, high-activity mutations were integrated into N46C-A48G, resulting in a six-mutant (N46C-A48G-V82T-I76Y-D147R-Q154R) with an editing rate of about 60% A-to-I and about 40% C-to-U Figure 11 B). These six-mutants fused with dCas13b were named as AIM-A&C system. The frequency of simultaneous A&C editing was further investigated. The results showed that AIM-A&C exhibited high levels of simultaneous A&C editing within a single RNA molecule, accounting for 40.2% of all transcripts Figure 11 C). Therefore, AIM-A&C was highly efficient in simultaneous A&C editing.

[0492] Given that C-to-U editing efficiency in AIM-A&C system is generally lower than A-to-I efficiency, the present inventors tried to further boost C-to-U activity. Four mutations (R26G, V28A, Y73P and H96N) in AIM-Cmax system were integrated into AIM-A&C, resulting in 1.4-fold increase in C-to-U efficiency (from 40% to 56%, D), but significant decrease in A-to-I activity (from 63% to 9%). As mentioned above, N46 is critical for A-to-I activity (Table 3), thus, the N46 mutation in AIM-A&C(G)-R26G / V28A / Y73P / H96N mutant was reverted to its original Asn residue. Interestingly, at this time we observed that A-to-I and C-to-U editing efficiency reached a balance (about 68.7% and 71.5%, respectively), which was 1.1-fold and 1.7-fold increase compared to original AIM-A&C, respectively (D). The present inventors named this optimized system as AIM-A&Cmax. The present inventors further confirmed that AIM-A&Cmax can also achieve controllable editing. With a 7-nt loop, AIM-A&Cmax can precisely edit A5 and C6. Further expanding the loop size to 17 nt, AIM-A&Cmax edited 4 A and 4 C across 11 nt (D). Figure 11 Figure 11 Figure 11 Figure 11 Figure 11

[0493] While the specific embodiments of the application have been described in detail, those skilled in the art will appreciate that various modifications and alterations to the details can be made within the scope of the application as disclosed in the teachings of the present application. The entire disclosure of the application is set out in the accompanying claims and any equivalents thereof.

[0494] References

[0495] ADDIN EN.REFLIST 1. Cox, D.B.T., Gootenberg, J.S., Abudayyeh, O.O., Franklin, B., Kellner, M.J., Joung, J., and Zhang, F. (2017). RNA editing with CRISPR-Cas13. Science (New York, N.Y.) 358, 1019-1027. 10.1126 / science.aaq0180.

[0496] ​​​​2. Qu, L., Yi, Z., Zhu, S., Wang, C., Cao, Z., Zhou, Z., Yuan, P., Yu, Y., Tian, F., Liu, Z., et al. (2019). Programmable RNA editing by recruiting endogenous ADAR using engineered RNAs. Nature biotechnology 37, 1059-1069. 10.1038 / s41587-019-0178-z.

[0497] 3. Yi, Z., Qu, L., Tang, H., Liu, Z., Liu, Y., Tian, F., Wang, C., Zhang, X., Feng, Z., Yu, Y., et al. (2022). Engineered circular ADAR-recruiting RNAs increase the efficiency and fidelity of RNA editing in vitro and in vivo. Nature biotechnology. 10.1038 / s41587-021-01180-3.

[0498] 4. Vogel, P., Moschref, M., Li, Q., Merkle, T., Selvasaravanan, K. D., Li, J. B., and Stafforst, T. (2018). Efficient and precise editing of endogenous transcripts with SNAP-tagged ADARs. Nature methods 15, 535-538. 10.1038 / s41592-018-0017-z.

[0499] 5. Katrekar, D., Chen, G., Meluzzi, D., Ganesh, A., Worlikar, A., Shih, Y.R., Varghese, S., and Mali, P. (2019). In vivo RNA editing of point mutations via RNA-guided adenosine deaminases. Nature methods 16, 239-242. 10.1038 / s41592-019-0323-0.

[0500] 6. Huang, X., Lv, J., Li, Y., Mao, S., Li, Z., Jing, Z., Sun, Y., Zhang, X., Shen, S., Wang, X., et al. (2020). Programmable C-to-U RNA editing using the human APOBEC3A deaminase. Embo j 39, e104741. 10.15252 / embj.2020104741.

[0501] 7. Song, J., Dong, L., Sun, H., Luo, N., Huang, Q., Li, K., Shen, X., Jiang, Z., Lv, Z., Peng, L., et al. (2023). CRISPR-free, programmable RNA pseudouridylation to suppress premature termination codons. Molecular cell 83, 139-155.e139. 10.1016 / j.molcel.2022.11.011.

[0502] 8. Luo, N., Huang, Q., Dong, L., Liu, W., Song, J., Sun, H., Wu, H., Gao, Y., and Yi, C. (2024). Near-cognate tRNAs increase the efficiency and precision of pseudouridine-mediated readthrough of premature termination codons. Nature biotechnology. 10.1038 / s41587-024-02165-8.

[0503] 9. Adachi, H., Pan, Y., He, X., Chen, J.L., Klein, B., Platenburg, G., Morais, P., Boutz, P., and Yu, Y.T. (2023). Targeted pseudouridylation: An approach for suppressing nonsense mutations in disease genes. Molecular cell 83, 637-651.e639. 10.1016 / j.molcel.2023.01.009.

[0504] 10. Rallapalli, K.L., and Komor, A.C. (2023). The Design and Application of DNA-Editing Enzymes as Base Editors. Annual Review of Biochemistry 92, 43-79. 10.1146 / annurev-biochem-052521-013938.

[0505] 11. Rees, H.A., and Liu, D.R. (2018). Base editing: precision chemistry on the genome and transcriptome of living cells. Nature reviews. Genetics 19, 770-788. 10.1038 / s41576-018-0059-1.

[0506] 12. Porto, E.M., Komor, A.C., Slaymaker, I.M., and Yeo, G.W. (2020). Base editing: advances and therapeutic opportunities. Nature Reviews Drug Discovery 19, 839-859. 10.1038 / s41573-020-0084-6.

[0507] 13. Komor, A.C., Kim, Y.B., Packer, M.S., Zuris, J.A., and Liu, D.R. (2016). Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-+. 10.1038 / nature17946.

[0508] 14. Gaudelli, N.M., Komor, A.C., Rees, H.A., Packer, M.S., Badran, A.H., Bryson, D.I., and Liu, D.R. (2017). Programmable base editing of A•T to G•C in genomic DNA without DNA cleavage. Nature 551, 464-471. 10.1038 / nature24644.

[0509] 15. Weber, L., Frati, G., Felix, T., Hardouin, G., Casini, A., Wollenschlaeger, C., Meneghini, V., Masson, C., De Cian, A., Chalumeau, A., et al. (2020). Editing a γ-globin repressor binding site restores fetal hemoglobin synthesis and corrects the sickle cell disease phenotype. Sci Adv 6. ARTN eaay9392 10.1126 / sciadv.aay9392.

[0510] 16. Antoniou, P., Hardouin, G., Martinucci, P., Frati, G., Felix, T., Chalumeau, A., Fontana, L., Martin, J., Masson, C., Brusson, M., et al. (2022). Base-editing-mediated dissection of a γ-globin -regulatory element for the therapeutic reactivation of fetal hemoglobin expression. Nature communications 13. ARTN 6618 10.1038 / s41467-022-34493-1.

[0511] 17. Yang, Y., Ren, R., Ly, L.C., Horton, J.R., Li, F.D., Quinlan, K.G.R., Crossley, M., Shi, Y.Y., and Cheng, X.D. (2021). Structural basis for human ZBTB7A action at the fetal globin promoter. Cell Reports 36. ARTN 109759 10.1016 / j.celrep.2021.109759.

[0512] 18. Erickson, J.R., Joiner, M.L.A., Guan, X., Kutschke, W., Yang, J.Y., Oddis, C.V., Bartlett, R.K., Lowe, J.S., O'Donnell, S.E., Aykin-Burns, N., et al. (2008). A dynamic pathway for calcium-independent activation of CaMKII by methionine oxidation. Cell 133, 462-474. 10.1016 / j.cell.2008.02.048.

[0513] 19. Luo, M., Guan, X.Q., Luczak, E.D., Lang, D., Kutschke, W., Gao, Z., Yang, J.Y., Glynn, P., Sossalla, S., Swaminathan, P.D., et al. (2013). Diabetes increases mortality after myocardial infarction by oxidizing CaMKII (vol 123, pg 1262, 2013). Journal of Clinical Investigation 123, 2333-2333. 10.1172 / Jci70180.

[0514] 20. Zhang, T., Zhang, Y., Cui, M.Y., Jin, L., Wang, Y.M., Lv, F.X., Liu, Y.L., Zheng, W., Shang, H.B., Zhang, J., et al. (2016). CaMKII is a RIP3 substrate mediating ischemia- and oxidative stress-induced myocardial necroptosis. Nature medicine 22, 175-182. 10.1038 / nm.4017.

[0515] 21. Lebek, S., Chemello, F., Caravia, X.M., Tan, W., Li, H., Chen, K.N., Xu, L., Liu, N., Bassel-Duby, R., and Olson, E.N. (2023). Ablation of CaMKIId oxidation by CRISPR-Cas9 base editing as a therapy for cardiac disease. Science (New York, N.Y.) 379, 179-185. 10.1126 / science.ade1105.

[0516] 22. Wolf, J., Gerber, A.P., and Keller, W. (2002). tadA, an essential tRNA-specific adenosine deaminase from Escherichia coli. Embo Journal 21, 3841-3851. DOI 10.1093 / emboj / cdf362.

[0517] 23. Richter, M.F., Zhao, K.T., Eton, E., Lapinaite, A., Newby, G.A., Thuronyi, B.W., Wilson, C., Koblan, L.W., Zeng, J., Bauer, D.E., Doudna, J.A., and Liu, D.V.R. (2020). Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity (vol 15, pg 891, 2020). Nature biotechnology 38, 901-901. 10.1038 / s41587-020-0562-8.

[0518] 24. Lapinaite, A., Knott, G.J., Palumbo, C.M., Lin-Shiao, E., Richter, M.F., Zhao, K.T., Beal, P.A., Liu, D.R., and Doudna, J.A. (2020). DNA capture by a CRISPR-Cas9-guided adenine base editor. Science (New York, N.Y.) 369, 566-+. 10.1126 / science.abb1390.

[0519] 25. Grunewald, J., Zhou, R.H., Iyer, S., Lareau, C.A., Garcia, S.P., Aryee, M.J., and Joung, J.K. (2019). CRISPR DNA base editors with reduced RNA off-target and self-editing activities. Nature biotechnology 37, 1041-+. 10.1038 / s41587-019-0236-6.

[0520] 26. Rees, H.A., Wilson, C., Doman, J.L., and Liu, D.R. (2019). Analysis and minimization of cellular RNA editing by DNA adenine base editors. Sci Adv 5. ARTN eaax5717 10.1126 / sciadv.aax5717.

[0521] 27. Zhou, C.Y., Sun, Y.D., Yan, R., Liu, Y.J., Zuo, E.W., Gu, C., Han, L.X., Wei, Y., Hu, X.D., Zeng, R., et al. (2019). Off-target RNA mutation induced by DNA base editing and its elimination by mutagenesis. Nature 571, 275-+. 10.1038 / s41586-019-1314-0.

[0522] 28. Li, J.A., Yu, W.X., Huang, S.S., Wu, S.S., Li, L.P., Zhou, J.K., Cao, Y., Huang, X.X., and Qiao, Y.B. (2021). Structure-guided engineering of adenine base editor with minimized RNA off-targeting activity. Nature communications 12. ARTN 2287 10.1038 / s41467-021-22519-z.

[0523] 29. Gaudelli, N.M., Lam, D.K., Rees, H.A., Sola-Esteves, N.M., Barrera, L.A., Born, D.A., Edwards, A., Gehrke, J.M., Lee, S.J., Liquori, A.J., et al. (2020). Directed evolution of adenine base editors with increased activity and therapeutic application. Nature biotechnology 38, 892-U899. 10.1038 / s41587-020-0491-6.

[0524] 30. Yan, H., and Tang, W. (2024). Programmed RNA editing with an evolved bacterial adenosine deaminase. Nature chemical biology. 10.1038 / s41589-024-01661-x.

Claims

1. A TadA mutant, characterized in that, The TadA mutant differs from the TadA sequence shown in SEQ ID NO: 1 in that: (i) F84I, V82T and D147R; (ii) F84I, V82T, D147R and I76Y; (iii) F84I, V82T, D147R and Q154R; or (iv) F84I, V82T, D147R, I76Y and Q154R.

2. The TadA mutant according to claim 1, characterized in that, The amino acid sequence of the TadA mutant is shown in SEQ ID NO: 4 or 5.

3. A polypeptide construct, characterized in that, The polypeptide construct includes a base deaminase and a guide RNA molecule binding domain linked to the base deaminase. The base deaminase is selected from the TadA mutant described in claim 1 or 2.

4. The polypeptide construct according to claim 3, characterized in that, The guide RNA molecule binding domain is selected from: the Cas effector protein of the CRISPR system lacking nuclease activity, the MS2 phage coat protein, and the phage Lambda N protein.

5. The polypeptide construct according to claim 3, characterized in that, The polypeptide construct possesses one or more of the following characteristics: (1) In the polypeptide construct, the guide RNA molecule binding domain is linked to the N-terminus and / or C-terminus of the base deaminase; (2) The polypeptide construct further comprises a nuclear localization sequence (NLS) and / or a nuclear export sequence (NES). (3) The binding domain of the guide RNA molecule contains an amino acid sequence as shown in SEQ ID NO: 10; and / or, the backbone sequence of the guide RNA molecule contains a nucleotide sequence as shown in SEQ ID NO: 13; (4) The polypeptide construct is a fusion protein; (5) The polypeptide construct is a fusion protein, and its amino acid sequence is shown in SEQ ID NO: 15 or 16.

6. A nucleic acid molecule, characterized in that, The nucleic acid molecule encodes the TadA mutant as described in claim 1 or 2, or the polypeptide construct as described in any one of claims 3-5.

7. A carrier, characterized in that, The carrier comprises the nucleic acid molecule of claim 6.

8. A host cell, characterized in that, The host cell comprises the nucleic acid molecule of claim 6 or the vector of claim 7.

9. A composition, characterized in that, The composition comprises: (a) A first component, selected from: (i) a guide RNA molecule comprising a backbone sequence and a guide sequence, the guide sequence being capable of hybridizing with a target RNA sequence comprising bases to be edited; (ii) a nucleotide sequence encoding the guide RNA molecule; (iii) any combination of (i) and (ii); as well as, (b) A second component selected from: (i) a polypeptide construct comprising a base deaminase and a guide RNA molecule binding domain linked to the base deaminase, the guide RNA molecule binding domain being capable of binding to the backbone sequence of the guide RNA molecule; (ii) a nucleotide sequence encoding the polypeptide construct; (iii) any combination of (i) and (ii); The base deaminase is selected from the TadA mutant described in claim 1 or 2.

10. The composition of claim 9, characterized in that, The guide sequence of the guide RNA molecule is designed to cause the target RNA sequence to form a single-stranded RNA loop region containing the bases to be edited after hybridization with the target RNA sequence. The target RNA sequence comprises: a target circular region containing the bases to be edited, a first region located at the 5' end of the target circular region, and a second region located at the 3' end of the target circular region; The guide sequence of the guide RNA molecule is designed to include a first region, a second region, and an optional third region located between the first and second regions, the second region being upstream of the first region; and the first region being complementary to a first region of the target RNA sequence, the second region being complementary to a second region of the target RNA sequence, and the third region being absent or not complementary to a target circularizable region of the target RNA sequence.

11. The composition of claim 10, characterized in that, The composition has one or more of the following characteristics: (1) The third region of the guide RNA molecule does not exist; (2) The number of nucleotide residues in the third region of the guide RNA molecule is less than the number of nucleotide residues in the target circular region of the target RNA sequence; (3) The first region and the second region of the guide RNA molecule have the same or different numbers of nucleotide residues; (4) The single-stranded RNA loop region formed after the target RNA sequence hybridizes with the guide RNA molecule contains one or more bases to be edited; (5) The single-stranded RNA loop region formed after the target RNA sequence hybridizes with the guide RNA molecule consists of at least 3, at least 4, at least 5, or at least 6 nucleotide residues; (6) The polypeptide construct is defined as in any one of claims 3-5.

12. A complex, characterized in that, The complex comprises: (a) A guide RNA molecule, as defined in any one of claims 9-11; as well as, (b) A polypeptide construct as defined in any one of claims 3-5.

13. A carrier system comprising one or more carriers, said one or more carriers comprising: (1) A first nucleotide sequence encoding a guide RNA molecule as defined in any one of claims 9-11; and, (2) A second nucleotide sequence encoding a polypeptide construct as defined in any one of claims 3-5; in, The first nucleotide sequence and the second nucleotide sequence may be present on the same or different vectors.

14. A reagent kit, characterized in that, The kit comprises the TadA mutant of claim 1 or 2, the polypeptide construct of any one of claims 3-5, the nucleic acid molecule of claim 6, the vector of claim 7, the composition of any one of claims 9-11, the complex of claim 12, or the vector system of claim 13.

15. A delivery composition, characterized in that, The composition comprises a delivery vector and one or more of the following: the TadA mutant of claim 1 or 2, the polypeptide construct of any one of claims 3-5, the nucleic acid molecule of claim 6, the vector of claim 7, the composition of any one of claims 9-11, the complex of claim 12, and the vector system of claim 13.

16. A cell characterized in that, The cell comprises: the TadA mutant of claim 1 or 2, the polypeptide construct of any one of claims 3-5, the nucleic acid molecule of claim 6, the vector of claim 7, the composition of any one of claims 9-11, the complex of claim 12, the vector system of claim 13, or the delivery composition of claim 15.

17. A method for extracellular editing of target RNA for non-therapeutic purposes, characterized in that, The method includes, under conditions suitable for editing the target RNA, contacting the target RNA with one or more of the following: the TadA mutant of claim 1 or 2, the polypeptide construct of any one of claims 3-5, the nucleic acid molecule of claim 6, the vector of claim 7, the composition of any one of claims 9-11, the complex of claim 12, the vector system of claim 13, and the delivery composition of claim 15, thereby inducing deamination of the bases to be edited in the target RNA.

18. A method for editing target RNA intracellularly for non-therapeutic purposes, characterized in that, The method comprises delivering the TadA mutant of claim 1 or 2, the polypeptide construct of any one of claims 3-5, the nucleic acid molecule of claim 6, the vector of claim 7, the composition of any one of claims 9-11, the complex of claim 12, the vector system of claim 13, and / or the delivery composition of claim 15 into a cell containing the target RNA, thereby inducing deamination of the base to be edited at the target site.

19. Use of the TadA mutant of claim 1 or 2, the polypeptide construct of any one of claims 3-5, the nucleic acid molecule of claim 6, the vector of claim 7, the composition of any one of claims 9-11, the complex of claim 12, the vector system of claim 13, the kit of claim 14, or the delivery composition of claim 15 for RNA editing for non-therapeutic purposes, or for use in the preparation of formulations for RNA editing.