Highly efficient and highly specific rna editing system
By designing a fusion protein containing RNA-specific adenosine deaminase and cytidine deaminase inhibitory domains, and utilizing protease cleavage sites to activate targeted editing, the problems of high off-target editing rate and low editing efficiency in existing RNA editing systems have been solved, achieving efficient and precise RNA editing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI TECH UNIV
- Filing Date
- 2024-09-20
- Publication Date
- 2026-06-16
AI Technical Summary
Existing RNA editing systems suffer from high off-target editing rates, low editing efficiency, and ineffectiveness in cells with low ADAR background expression levels, and may also trigger inflammatory responses.
A fusion protein was designed containing inhibitory domains of RNA-specific adenosine deaminase and cytidine deaminase. By designing protease cleavage sites, it is designed to activate editing when the target RNA is bound and remain inactive when it is not bound, thereby reducing off-target editing.
It achieves efficient and precise RNA editing, reduces or eliminates off-target editing, is applicable to different cells and species, improves editing efficiency, and avoids inflammatory responses.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Background Technology
[0001] Targeted RNA editing, broadly defined, can be defined as any site-specific modification of RNA molecules transcribed from a DNA template. Gene expression alterations induced by RNA editing have been observed in a wide range of organisms, from single-celled protozoa to humans, affecting mRNA, tRNA, and rRNA in different cellular compartments. RNA editing is rapidly developing due to its ability to produce transient and reversible modifications, offering a more flexible option for repairing disease-related mutations compared to permanent DNA editing. RNA editing can also circumvent the potential genotoxicity posed by genome editing tools, including the CRISPR / Cas9 system.
[0002] The vast majority of RNA editing in the human body is catalyzed by RNA adenosine deaminases (ADARs). These enzymes deaminate adenosine in double-stranded RNA regions to generate inosine nucleosides (A-to-I editing), thereby affecting RNA structure and function. Currently, five ADAR isoforms encoded by three different genes have been identified in humans, including ADAR1 (containing ADAR1p110 and ADAR1p150 isoforms), ADAR2 (containing ADAR2a and ADAR2b isoforms), and ADAR3. These ADARs edit double-stranded RNA molecules in a structure-dependent rather than sequence-dependent manner, thereby regulating regulatory non-coding RNAs, activating antiviral immunity, or suppressing abnormal innate immune responses against their own transcripts.
[0003] To achieve targeted RNA editing, ADAR proteins or their catalytic domains can be fused with RNA-binding proteins (such as λN peptides or MCP peptides), SNAP tags, or CRISPR / Cas proteins (such as catalytically inactivating Cas13). At the same time, ADAR recruitment guide RNAs (agRNAs) can be designed to recruit chimeric ADAR proteins to specific target sites (see, for example, Cox et al., Science 358, 1019-1027 (2017), Vogel et al., Nat Methods 15, 535-538 (2018), Abudayyeh et al., Science 365, 382-386 (2019), Katrekar et al., Nat Methods 16, 239-242 (2019), and Kannan et al., Nat Biotechnol 40, 194-197 (2022)). In addition, other RNA editing systems have been developed that recruit endogenous ADAR proteins using designed guide RNAs.
[0004] Guided by agRNA, ADAR-based RNA editors can trigger site-specific A-to-I editing. During cellular protein translation, hypoxanthine is often recognized as guanine due to structural similarity. Therefore, ADAR-mediated A-to-I editing can be viewed as A-to-G mutations at the RNA level. This technology has broad applications in multiple fields, including repairing disease-related mutations, regulating gene expression, and altering protein-protein interactions.
[0005] However, existing RNA editing systems have significant drawbacks. For example, RNA editing systems based on exogenous ADAR expression suffer from high off-target editing rates. On the other hand, RNA editing systems based on recruiting endogenous ADAR have low editing efficiency and may not function effectively in cells with low ADAR basal expression levels. Furthermore, insufficient editing of endogenous substrates by ADAR may trigger inflammatory responses. Therefore, developing novel RNA editing systems with high efficiency, low off-target mutation rates, and broad applicability is crucial. Summary of the Invention
[0006] This disclosure describes fusion proteins and related molecules capable of specific RNA editing against target RNA molecules and exhibiting minimal or no off-target editing. The fusion protein may contain an RNA-specific adenosine deaminase or its deaminase domain (e.g., the ADAR2 deaminase domain, ADAR2dd), which is cleavably linked to the repressive domain of a cytidine deaminase (e.g., apolipoprotein B mRNA editing enzyme catalyzed polypeptide-like protein, APOBEC). This disclosure unexpectedly reveals that the repressive domain of the cytidine deaminase can inhibit the activity of the RNA adenosine deaminase. Upon binding to the target RNA molecule, the repressive domain can be cleaved by the corresponding protease, thereby enabling the RNA-specific adenosine deaminase to efficiently edit the RNA molecule. In the absence of the target RNA molecule, the RNA-specific adenosine deaminase in the fusion protein remains inactive, thereby inhibiting or preventing off-target editing.
[0007] The protease can be fused with an RNA recognition peptide, such as lambda(λN) peptide, PP7 capsid protein (PCP), or MS2 capsid protein (MCP). Simultaneously, a corresponding RNA recognition site (called an RNA tag) can be added to the guide RNA to facilitate the recruitment of the protease to the target RNA molecule. In some embodiments, the guide RNA already contains an RNA tag (e.g., MS2) that can be used to recruit an RNA-specific adenosine deaminase fused with the MS2 capsid protein (MCP).
[0008] The mutant RNA-specific adenosine deaminase was also investigated, and as shown in the examples, specific mutations and combinations of mutations can help to further reduce or even completely eliminate off-target RNA editing.
[0009] Compared to existing RNA editing systems, the RNA editing system developed in this invention possesses highly efficient editing capabilities, while exhibiting no detectable off-target editing compared to the unedited control group. This novel RNA editing system can also achieve efficient and precise editing in different cells and species.
[0010] Therefore, one embodiment of this disclosure provides a fusion protein comprising: a first fragment containing an RNA-specific adenosine deaminase or a deaminase domain thereof; a second fragment containing an inhibitory domain of cytidine deaminase; and a protease cleavage site located between the first fragment and the second fragment.
[0011] In some embodiments, the RNA-specific adenosine deaminase is an RNA adenosine deaminase (ADAR), optionally selected from the group consisting of ADAR1, ADAR2, and ADAR3. In some embodiments, the ADAR is human ADAR2, comprising the amino acid sequence shown in SEQ ID NO: 1, or an amino acid sequence having at least 85% sequence identity with SEQ ID NO: 1. In some embodiments, relative to SEQ ID NO: 1, the ADAR2 comprises one or more mutations selected from the group consisting of: R348V, V351G, T375S, N473S, K475I, K475Q, S486A, E488Q, K594I, E620G, and Q696F. In some embodiments, relative to SEQ ID NO: 1, the ADAR2 comprises one or more mutations selected from the following: V351G, T375S, N473S, K475I, K475Q, S486A, E488Q, K594I, E620G, and Q696F. In some embodiments, relative to SEQ ID NO: 1, the ADAR2 comprises mutations selected from combinations of the following: K475Q and E488Q, K475I and S486A / E488Q, and K475I and S486A.
[0012] In some embodiments, the deaminase domain of ADAR2 comprises amino acid residues 316-700 of SEQ ID NO: 1. In some embodiments, the cytidine deaminase is APOBEC (apolipoprotein B mRNA editing enzyme, catalyzing a polypeptide-like protein). In some embodiments, the inhibitory domain comprises an amino acid sequence selected from SEQ ID NO: 2-91, or an amino acid sequence having at least 85% sequence identity with any amino acid sequence selected from SEQ ID NO: 2-92. In some embodiments, the inhibitory domain is selected from the group consisting of mA3-CDA2, hA3B-CDA1, hA3D-CDA1, hA3F-CDA1, and hA3G-CDA1. In some embodiments, the nucleobase deaminase inhibitor comprises an amino acid sequence as shown in SEQ ID NO: 2 or amino acid residues 128-223 of SEQ ID NO: 2.
[0013] In some embodiments, the protease cleavage site is not recognized by endogenous proteases in human cells. In some embodiments, the protease cleavage site is recognized by a protease selected from the group consisting of TuMV, PPV, PVY, ZIKV, and WNV proteases. In some embodiments, the protease cleavage site is a TEV protease cleavage site. In some embodiments, the second fragment comprises at least two inhibitory domains; optionally, the inhibitory domains are the same or different inhibitory domains.
[0014] In one embodiment, the present disclosure also provides a fusion protein comprising: (a) an N-terminal domain (TEVn) of a TEV protease; (b) a C-terminal domain (TEVc) of a TEV protease; (c) a self-cleavage site located between TEVn and TEVc; and (d) an RNA recognition peptide.
[0015] In some embodiments, TEVn is located at the C-terminus of TEVc. In some embodiments, the RNA recognition peptide is a lambda (λN) peptide, PP7 capsid protein (PCP), or MS2 capsid protein (MCP). In some embodiments, the RNA recognition peptide is a lambda (λN) peptide.
[0016] This disclosure also provides a mutant human ADAR2 deaminase domain (ADAR2dd) or a mutant ADAR2 containing said ADAR2dd, wherein said ADAR2dd contains mutations selected from the following: R348A; N473S; K475I; Q696F; E488Q / V351G; E488Q / T375S; E488Q / N473S; E488Q / K475I; E488Q / S486A; E488Q / K594I; E488Q / E620G; E488Q / Q696F; E488Q / K475Q; E488Q / K475I / S486A; K475I / S486A, and combinations thereof; wherein the amino acid positions are as per SEQ ID NO: 1. In some implementations, the ADAR2dd contains E488Q / K475Q, E488Q / K475I / S486A, or K475I / S486A mutations. Also provided is a mutant human ADAR2 deaminase domain (ADAR2dd) or a mutant ADAR2 containing said ADAR2dd, wherein said ADAR2dd contains mutations selected from the following: N473S; K475I; Q696F; E488Q / V351G; E488Q / T375S; E488Q / N473S; E488Q / K475I; E488Q / S486A; E488Q / K594I; E488Q / E620G; E488Q / Q696F; E488Q / K475Q; E488Q / K475I / S486A; K475I / S486A, and combinations thereof; wherein the amino acid positions are as per SEQ ID NO: 1. In some implementations, the ADAR2dd contains E488Q / K475Q, E488Q / K475I / S486A, or K475I / S486A mutations.
[0017] Also provided is a guide RNA or DNA encoding the guide RNA for recognizing a target RNA sequence, wherein the guide RNA comprises an antisense fragment capable of hybridizing with the target RNA sequence, the antisense fragment being flanked by two MS2 aptamers, and an RNA tag that can be recognized by an RNA recognition peptide that is not an MS2 coat protein (MCP).
[0018] In some embodiments, the RNA tag is a BoxB aptamer or a PP7 aptamer. In some embodiments, the guide RNA contains two copies of the RNA tag. In some embodiments, the two RNA tags are located outside the two MS2 aptamers. In some embodiments, each of the two RNA tags is separated from its corresponding MS2 aptamer by one or three nucleotides.
[0019] In another embodiment, a polynucleotide is also provided that encodes the fusion protein described herein or mutant ADAR2dd or mutant ADAR2.
[0020] A method for editing a target RNA molecule in a cell is also provided, comprising introducing into the cell: a first polynucleotide encoding a fusion protein, the fusion protein comprising: a first fragment containing an RNA-specific adenosine deaminase or a deaminase domain thereof, a second fragment containing an inhibitory domain of cytidine deaminase, a protease cleavage site located between the first fragment and the second fragment, and an RNA recognition peptide; a second polynucleotide encoding a protease capable of cleaving the protease cleavage site, and a guide RNA or DNA encoding the guide RNA, wherein the guide RNA comprises an antisense fragment capable of hybridizing with the target RNA molecule and an RNA tag recognizable by the RNA recognition peptide.
[0021] In some embodiments, the RNA tag comprises an MS2 aptamer, and the RNA recognition peptide comprises an MS2 capsid protein (MCP). In some embodiments, the protease comprises: (a) an N-terminal domain (TEVn) of the TEV protease; (b) a C-terminal domain (TEVc) of the TEV protease; and (c) a self-cleavage site located between TEVn and TEVc. In some embodiments, the protease is fused to a second RNA recognition peptide, and the guide RNA further comprises a second RNA tag that can be recognized by the second RNA recognition peptide. In some embodiments, the second RNA recognition peptide comprises a lambda (λN) peptide or a PP7 capsid protein (PCP), and the second RNA tag comprises a BoxB aptamer or a PP7 aptamer. In some embodiments, the guide RNA comprises the antisense fragment flanked by two copies of the RNA tag, and the RNA tag flanked by two copies of the second RNA tag.
[0022] In some embodiments, the cells are mammalian cells, preferably human cells. In some embodiments, the method is performed in vitro, ex vivo, or in vivo.
[0023] A kit or package is also provided, comprising: a first fragment containing an RNA-specific adenosine deaminase or its deaminase domain; a second fragment containing an inhibitory domain of cytidine deaminase; a protease cleavage site located between the first fragment and the second fragment; and an RNA recognition peptide; and a second polynucleotide encoding a protease capable of cleaving the protease cleavage site.
[0024] In some embodiments, the kit or package further comprises guide RNA or DNA encoding the guide RNA, wherein the guide RNA comprises an antisense fragment capable of hybridizing with the target RNA molecule and an RNA tag recognizable by the RNA recognition peptide. Attached Figure Description
[0025] Figure 1 The results show that MCP-XTEN-ADAR2dd and MCP-XTEN-ADAR2dd (E488Q) generate a large number of off-target A-to-I mutations during targeted editing. (a) Schematic diagram of a conventional RNA editing system consisting of two main components: (1) an effector protein composed of MS2 coat protein (MCP), XTEN linker peptide, and ADAR2dd or ADAR2dd (E488Q); (2) a site-specific ADAR recruitment guide RNA (agRNA) composed of two MS2 aptamers and an antisense (AS) region. (b) Figure 1 (a) shows the A-to-I editing efficiency of the RNA editing system induced at four different target sites. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (c) Figure 1 (b) shows the Sanger sequencing results for quantitative analysis. (d) Figure 1 Figure a shows the efficiency of off-target RNA A-to-I editing induced at five different off-target sites by the RNA editing system. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (e) Figure 1 The d-values show the Sanger sequencing results of quantitative analysis.
[0026] Figure 2 This demonstrates the inhibitory effect of mA3CDA2 on the adenosine deamination activity of MCP-XTEN-ADAR2dd (E488Q). (a) Schematic diagram of MCP-XTEN-ADAR2dd (E488Q) fusion with different types of proteins. (b) Figure 2 (a) shows the A-to-I editing efficiency of the RNA editing system induced at two different target sites. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (c) Figure 2 Figure b shows the Sanger sequencing results of quantitative analysis.
[0027] Figure 3 The inhibitory effects of different CDA on the adenosine deamination activity of MCP-XTEN-ADAR2dd (E488Q) are shown. (a) Schematic diagram of the fusion of MCP-XTEN-ADAR2dd (E488Q) with different CDA that have a potential inhibitory effect on adenosine deamination activity. (b) Figure 3 (a) The RNA editing system shown in figure a induced A-to-I editing efficiency of targeted RNA at six different target sites. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (c) Figure 3 Figure b shows the Sanger sequencing results of quantitative analysis.
[0028] Figure 4 The inhibitory effect of 2×A3DNDI on the adenosine deamination activity of MCP-XTEN-ADAR2dd (E488Q) is shown. (a) Schematic diagram of the fusion of MCP-XTEN-ADAR2dd (E488Q) and 2×A3DNDI. (b) Figure 4 (a) The RNA editing system shown in figure a induced A-to-I editing efficiency of targeted RNA at six different target sites. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (c) Figure 4 Figure b shows the Sanger sequencing results of quantitative analysis.
[0029] Figure 5 The diagram shows that modifying ADAR2dd can reduce off-target editing. (a) Schematic diagram of introducing different amino acid mutations into ADAR2dd. (b) Figure 5 (a) shows the A-to-I editing efficiency of the effector protein at the CTNNB1-1 target site inducing targeted RNA. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (c) Figure 5 Figure b shows the Sanger sequencing results of quantitative analysis.
[0030] Figure 6 The expression of free TEVp cleavage of 2×A3DNDI restored the targeted editing activity of MCP-XTEN-ADAR2dd(E488Q)-TS-2×A3DNDI. (a) Schematic diagram of the RNA editing system, which consists of three main components: (1) effector protein (MCP-XTEN-ADAR2dd(E488Q)-TS-2×A3DNDI); (2) intact TEV protease (TEVp); and (3) site-specific agRNA. (b) RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (c) Figure 6 Figure b shows the Sanger sequencing results of quantitative analysis.
[0031] Figure 7The diagram shows that splitting the TEV protease into two segments reduces off-target editing. (a) Schematic diagram of the RNA editing system, which consists of three main components: (1) an effector protein (MCP-XTEN-ADAR2dd(E488Q)-TS-2×A3DNDI); (2) two fragments of the TEV protease linked by a T2A self-cleaving peptide (TEVc-T2A-TEVn or TEVn-T2A-TEVc); and (3) a site-specific agRNA. (b) Figure 7 (a) shows the A-to-I editing efficiency of the RNA editing system induced at two target sites. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (c) Figure 7 (b) shows the Sanger sequencing results for quantitative analysis. (d) Figure 7 (a) shows the off-target RNA A-to-I editing efficiency induced at the AP3D1 off-target site by the RNA editing system. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (e) Figure 7 (d) shows the Sanger sequencing results of quantitative analysis. (f) Figure 7 The off-target RNA A-to-I editing efficiency induced at the MDK off-target site by the RNA editing system shown in (a) was analyzed. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (g) Figure 7 The results of the quantitative analysis of Sanger sequencing are shown in f.
[0032] Figure 8 The diagram shows that modifying agRNA can reduce off-target editing. (a) Schematic diagram of the fusion protein of MCP-XTEN-ADAR2dd(E488Q)-TS-2×A3DNDI and λN-XTEN-TEVc-T2A-TEVn. (b) Modifying agRNA by ligating BoxB aptamers to the 5' end of the original agRNA or to both the 5' and 3' ends, and choosing whether to use an adenine sequence as the linking nucleotide. (c) Figure 8 a and Figure 8 (b) shows the A-to-I editing efficiency of the RNA editing system induced at the CTNNB1-1 target site. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (d) Figure 8 a and Figure 8 Figure b shows the efficiency of off-target RNA A-to-I editing induced at two off-target sites by the RNA editing system. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (e) Figure 8 c and Figure 8The quantified Sanger sequencing results in d.
[0033] Figure 9 The effect of linker peptides on targeted and off-target editing is shown. (a) Schematic diagram of replacing the XTEN linker peptide with the GS linker peptide in the λN-XTEN-TEVc-T2A-TEVn and MCP-XTEN-ADAR2dd(E488Q)-TS-2×A3DNDI fusion proteins. (b) Figure 9 (a) shows the A-to-I editing efficiency of the RNA editing system induced at two target sites. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (c) Figure 9 (b) shows the Sanger sequencing results for quantitative analysis. (d) Figure 9 (a) shows the off-target RNA A-to-I editing efficiency induced at the AP3D1 off-target site by the RNA editing system. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (e) Figure 9 (d) shows the Sanger sequencing results of quantitative analysis. (f) Figure 9 The off-target RNA A-to-I editing efficiency induced at the EIF4G1 off-target site by the RNA editing system shown in (a) was analyzed. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (g) Figure 9 The results of the quantitative analysis of Sanger sequencing are shown in f.
[0034] Figure 10 The diagram shows that introducing amino acid mutations in ADAR2dd can eliminate off-target mutations. (a) Schematic diagram of introducing amino acid mutations in ADAR2dd. (b) Figure 10 (a) shows the A-to-I editing efficiency of the RNA editing system induced at two target sites. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (c) Figure 10 (b) shows the Sanger sequencing results for quantitative analysis. (d) Figure 10 (a) shows the off-target RNA A-to-I editing efficiency induced at the AP3D1 off-target site by the RNA editing system. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (e) Figure 10 (d) shows the Sanger sequencing results of quantitative analysis. (f) Figure 10 The off-target RNA A-to-I editing efficiency induced at the KDELR1 off-target site by the RNA editing system shown in (a) was analyzed. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (g) Figure 10The results of the quantitative analysis of Sanger sequencing are shown in f.
[0035] Figure 11 The mechanism of action of the novel RNA base editing system is illustrated by a schematic diagram.
[0036] Figure 12 The diagram shows that modifying ADAR2dd (introducing the K475I / S486A mutation) can reduce off-target editing. (a) Schematic diagram of introducing the K475I and S486A amino acid mutations into ADAR2dd (amino acid residues Q316-T700). (b) Figure 12 (a) shows the A-to-I editing efficiency of the effector protein at the CTNNB1-1 target site inducing targeted RNA. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (c) Figure 12 Figure b shows the Sanger sequencing results for quantitative analysis. The modified ADAR2dd (K475I / S486A) exhibited lower off-target effects compared to wild-type ADAR2dd.
[0037] Figure 13 The diagram shows that using a single copy of A3DNDI can still effectively suppress off-target editing, and after its excision, it can achieve efficient RNA editing at target sites against the background of 5'-CAN, 5'-GAN, and 5'-AAN sequences. (a) Schematic diagram of introducing amino acid mutations in ADAR2dd. (b) Figure 13 (a) shows the A-to-I editing efficiency of the RNA editing system induced at three target sites (PAICS-1, STAT1-3, and CCNI-1) with 5'-CAN, 5'-GAN, and 5'-AAN sequence backgrounds. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (c) Figure 13 (b) shows the Sanger sequencing results for quantitative analysis. (d) Figure 13 (a) shows the off-target RNA A-to-I editing efficiency induced at the AP3D1 off-target site by the RNA editing system. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (e) Figure 13 (d) shows the Sanger sequencing results of quantitative analysis. (f) Figure 13 The off-target RNA A-to-I editing efficiency induced at the EIF4G1 off-target site by the RNA editing system shown in (a) was analyzed. RNA editing efficiency was detected by RT-PCR and quantified based on Sanger sequencing results. (g) Figure 13 The results of the quantitative analysis of Sanger sequencing are shown in f. Detailed Implementation
[0038] definition
[0039] The term "a" refers to one or more of the same thing; for example, "an antibody" should be understood to refer to one or more antibodies. Therefore, the terms "a," "one or more," and "at least one" are used interchangeably in this document.
[0040] As used herein, the term "polypeptide" encompasses both the singular and plural forms of "polypeptide," referring to a molecule composed of monomers (amino acids) linearly linked by amide bonds (also known as peptide bonds). The term "polypeptide" refers to any one or more chains of two or more amino acids, without specifying a particular length of the product. Therefore, peptide, dipeptide, tripeptide, oligopeptide, "protein," "amino acid chain," or any other term used to refer to any one or more chains of two or more amino acids are all included within the definition of "polypeptide," and the term "polypeptide" may be used in place of or interchangeably with any of the aforementioned terms. The term "polypeptide" also refers to products of post-expression modification of a polypeptide, including but not limited to glycosylation, acetylation, phosphorylation, amidation, derivatization by known protecting / blocking groups, proteolytic cleavage, or modification by non-naturally occurring amino acids. Polypeptides may be derived from natural biological sources or prepared through recombinant technologies, but need not be produced by translation of a specified nucleic acid sequence. They can be generated in any manner, including chemical synthesis.
[0041] "Homology," "identity," or "similarity" refers to the sequence similarity between two peptides or two nucleic acid molecules. Homology can be determined by comparing corresponding bases or amino acids in sequences to be compared. When a position in the compared sequences is occupied by the same base or amino acid, the molecules are homologous at that position. The degree of homology between sequences is a function of the number of shared matching sites or homologous sites. "Unrelated" or "non-homologous" sequences have less than 40% identity with one of the sequences described in this disclosure, preferably less than 25%.
[0042] "Sequence identity" of a polynucleotide or polynucleotide region (or polypeptide or polypeptide region) with another sequence at a certain percentage (e.g., 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99%) means that, after alignment, that percentage of bases (or amino acids) are identical in the two sequences. Such alignments and homology percentages or sequence identity can be determined using software programs known in the art, such as Ausubel. et alThe procedure described in Current Protocols in Molecular Biology (eds. (2007)) is preferred for alignment. Default parameters are preferred for comparison. One such alignment procedure is the BLAST procedure using default parameters.
[0043] The term "equivalent nucleic acid or polynucleotide" refers to a nucleic acid whose nucleotide sequence shares a degree of homology or sequence identity with the nucleotide sequence of a given nucleic acid or its complementary sequence. Homologs of double-stranded nucleic acids are intended to include nucleic acids whose nucleotide sequence shares a degree of homology with that nucleic acid or its complementary sequence. In one aspect, a homolog of a nucleic acid is capable of hybridizing with that nucleic acid or its complementary sequence. Similarly, "equivalent polypeptide" refers to a polypeptide whose amino acid sequence shares a degree of homology or sequence identity with a reference polypeptide. In some aspects, the sequence identity is at least about 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99%. In some aspects, the equivalent polypeptide or polynucleotide has one, two, three, four, or five additions, deletions, substitutions, or combinations thereof compared to a reference polypeptide or polynucleotide. In some aspects, the equivalent sequence retains the activity (e.g., epitope binding) or structure (e.g., salt bridge) of the reference sequence.
[0044] When the term "encodes" is used for polynucleotides, it refers to a polynucleotide that, in its natural state or after manipulation by methods well known to those skilled in the art, can be transcribed and / or translated to produce messenger RNA and / or fragments thereof of the polypeptide. The antisense strand is the complementary strand of this type of nucleic acid, from which the coding sequence can be derived.
[0045] RNA editing system
[0046] This article describes an RNA editing system and method capable of specifically performing RNA editing on target RNA molecules with minimal or no off-target editing. The design of this novel system stemmed at least in part from an unexpected discovery: the inhibitory domain of cytidine deaminase can inhibit the activity of RNA adenosine deaminase.
[0047] Many cytidine deaminases (such as apolipoprotein B mRNA editing enzyme catalyzing a polypeptide-like protein, APOBEC) contain two cytidine deaminase (CDA) domains, one of which is catalytically active and the other inactive. Studies have found that the inactive CDA domain can inhibit the homologous CDA domain or the active CDA domain of other cytidine deaminases. This type of inhibition had not previously been found to extend to RNA editing enzymes. For example, Wang... et al. Nature Cell BiologyAs reported in 2021 (23):552–563, the deaminase-CDA inhibitory domain fusion protein (tBE-V5-rA1) actually induces higher levels of RNA off-target mutations compared to the conventional base editor BE3 (page 560, paragraph 1).
[0048] In a surprising discovery of this disclosure, the inventors demonstrated that such CDA repressive domains can effectively inhibit the RNA editing activity of ADAR enzymes (Example 2). Fusing a CDA repressive domain (e.g., mA3-CDA2) with an ADAR2 deaminase domain (ADAR2dd) or a mutant ADAR2dd (E488Q) significantly reduces the editing efficiency of the deaminase. Interestingly, the inhibitory effect is even more pronounced when using a double-copy CDA repressive domain.
[0049] Based on this discovery, an RNA editing system was designed comprising a fusion protein containing an RNA-specific adenosine deaminase or its deaminase domain (e.g., ADAR2dd), which is cleavably linked to a repressive domain of a cytidine deaminase (e.g., APOBEC). Such a repressive domain is referred herein as a "nucleoside deaminase inhibitor domain" or "NDI domain".
[0050] Upon binding to the target RNA molecule, the NDI domain can be cleaved by the corresponding protease, thereby enabling RNA-specific adenosine deaminase to efficiently edit the RNA molecule. In the absence of the target RNA molecule, the RNA-specific adenosine deaminase in the fusion protein remains inactive, thereby inhibiting or preventing off-target editing.
[0051] In some embodiments, the protease is fused to an RNA recognition peptide, such as lambda(λN) peptide, PP7 capsid protein (PCP), or MS2 capsid protein (MCP). Simultaneously, a corresponding RNA recognition site (referred to as an RNA tag) may be added to the guide RNA to facilitate the recruitment of the protease to the target RNA molecule. In some embodiments, the guide RNA already contains an RNA tag (e.g., MS2) that can be used to recruit an RNA-specific adenosine deaminase fused to the MCP. Guide RNAs containing additional RNA tags (e.g., BoxB or PP7) in addition to the MS2 tag are referred to herein as “engineered agRNA” or “eagRNA”.
[0052] This article also investigated mutant RNA-specific adenosine deaminases, and as shown in the examples, specific mutations and combinations of mutations can help to further reduce or even completely eliminate off-target RNA editing.
[0053] fusion protein of RNA-specific adenosine deaminase
[0054] Based on the aforementioned unexpected research findings, this disclosure presents a fusion protein that can be used to construct an RNA editor with improved RNA editing specificity and efficiency. In one embodiment, this disclosure provides a fusion protein comprising a first fragment and a second fragment, the first fragment comprising an RNA-specific adenosine deaminase or its deaminase domain, and the second fragment comprising an inhibitory domain of a nucleobase deaminase (e.g., cytidine deaminase), the two fragments being fused in a cleavable manner.
[0055] As previously mentioned, these fusion proteins exhibit very low or no RNA editing activity due to the inhibitory effect of the fused nucleobase deaminase repressor domain (NDI). Once the repressor domain is dissociated and removed, adenosine deaminase or its deaminase domain is activated. Dissociation and release can be achieved via a protease that recognizes the protease cleavage site between the two domains.
[0056] RNA-specific adenosine deaminases, or RNA adenosine deaminases, are a class of enzymes that bind to double-stranded RNA (dsRNA) and convert adenosine to inosine via deamination. Inosine is structurally similar to guanine (G), leading to inosine-cytosine (I:C) pairing; therefore, this editing process produces A-to-G mutations.
[0057] Non-restricted examples of RNA-specific adenosine deaminases include ADAR1 (including isoforms ADAR1p110 and ADAR1p150), ADAR2 (including isoforms ADAR2a and ADAR2b), and ADAR3. An example GenBank gene identifier for human ADAR2 is 104 (entry name: ADARB1 adenosine deaminase RNA specificB1 [Homo sapiens (human)]). Table A lists more exemplary RNA-specific adenosine deaminases from different species.
[0058] Table A: Exemplary RNA-specific adenosine deaminases
[0059] RNA-specific adenosine deaminases typically contain a deaminase domain and one or more RNA-binding domains. For the purposes of this technology, the RNA-binding domain is not essential. Therefore, in some embodiments, the fusion protein contains only the deaminase domain of the RNA-specific adenosine deaminase. For example, the deaminase domain of human ADAR2 (SEQ ID NO: 1) contains amino acid residues 316-700.
[0060] In some embodiments, the RNA-specific adenosine deaminase or its deaminase domain is a biological equivalent of any RNA-specific adenosine deaminase or its deaminase domain disclosed herein (e.g., having at least about 80%, 85%, 90%, 95%, 97%, 98%, 99%, 99.5% sequence identity, or having one, two, or three amino acid additions / deletions / substitutions, while retaining adenosine deaminase activity).
[0061] As used in this article, "nucleobase deaminase" refers to a class of enzymes that catalyze the hydrolytic deamination of nucleobases such as cytidine, deoxycytidine, adenosine, and deoxyadenosine. Non-limiting examples of nucleobase deaminases include cytidine deaminase and adenosine deaminase.
[0062] Adenosine deaminase, also known as adenosine aminohydrolase (ADA), is an enzyme involved in purine metabolism (EC 3.5.4.4). This enzyme breaks down adenosine from food sources and participates in the turnover metabolism of nucleic acids within tissues.
[0063] Non-restrictive examples of adenosine deaminases include: tRNA-specific adenosine deaminase (TadA), tRNA-specific adenosine deaminase 1 (ADAT1), tRNA-specific adenosine deaminase 2 (ADAT2), tRNA-specific adenosine deaminase 3 (ADAT3), RNA-specific adenosine deaminase B1 (ADARB1), RNA-specific adenosine deaminase B2 (ADARB2), adenosine-phosphate deaminase 1 (AMPD1), adenosine-phosphate deaminase 2 (AMPD2), adenosine-phosphate deaminase 3 (AMPD3), adenosine deaminase (ADA), adenosine deaminase 2 (ADA2), adenosine deaminase-like protein (ADAL), adenosine deaminase domain-containing protein 1 (ADAD1), adenosine deaminase domain-containing protein 2 (ADAD2), RNA-specific adenosine deaminase (ADAR), and RNA-specific adenosine deaminase B1 (ADARB1).
[0064] "Cytidine deaminase" refers to enzymes that irreversibly catalyze the hydrolysis and deamination of cytidine and deoxycytidine to produce uridine and deoxyuridine, respectively. Cytidine deaminases are involved in maintaining intracellular pyrimidine pool homeostasis. One family of cytidine deaminases is the APOBEC family (apolipoprotein B mRNA editing enzyme catalytic polypeptide-like protein family). Members of this family are C-to-U editing enzymes. Some APOBEC family members contain two domains, one a catalytic domain and the other a pseudocatalytic domain. More specifically, the catalytic domain is a zinc ion-dependent cytidine deaminase domain, which is crucial for the cytidine deamination reaction. APOBEC-1-mediated RNA editing requires the formation of a homodimer, which interacts with RNA-binding proteins to form the edit body.
[0065] Non-restricted examples of APOBEC proteins include APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-inducible (cytidine) deaminase (AID).
[0066] Non-restrictive examples of cytidine deaminase inhibitory domains include mA3-CDA2 (SEQ ID NO: 2, or a fragment containing at least amino acid residues 128-223 of SEQ ID NO: 2), hA3B-CDA1 (SEQ ID NO: 3), hA3D-CDA1 (SEQ ID NO: 85), hA3F-CDA1 (SEQ ID NO: 54), and hA3G-CDA1 (SEQ ID NO: 92).
[0067] Table B provides a more detailed list of inhibitory domains, and this disclosure also covers their biological equivalents (e.g., having at least about 80%, 85%, 90%, 95%, 97%, 98%, 99%, 99.5% sequence identity, or having one, two, or three amino acid additions / deletions / substitutions, and retaining nucleobase deaminase inhibitory activity).
[0068] As demonstrated in the experimental examples, the inhibitory effect is more significant when two tandem repressive domains are used. Therefore, in some embodiments, the fusion protein comprises 2, 3, 4, 5, or 6 repressive domains, which may be the same or different repressive domains. In some embodiments, the repressive domains are grouped at the N-terminus or C-terminus of an RNA-specific adenosine deaminase or its deaminase domain. In a preferred embodiment, one or more repressive domains are arranged at the C-terminus of an RNA-specific adenosine deaminase or its deaminase domain.
[0069] Table B: Exemplary Nucleoside Deaminase Inhibition (NDI) Domains
[0070] The protease cleavage site located between the first fragment and the second fragment can be any known protease cleavage site (peptide) corresponding to any protease. In a preferred embodiment, the protease is not an endogenous protease in the target cell (e.g., human cells). Non-limiting examples of proteases include TEV protease, TuMV protease, PPV protease, PVY protease, ZIKV protease, and WNV protease. The protein sequences of exemplary proteases and their corresponding cleavage sites are shown in Table C.
[0071] Table C Related Sequences
[0072] In some embodiments, the fusion protein further comprises an RNA recognition peptide for recognizing a paired guide RNA. Non-limiting examples of RNA recognition peptides include: the MS2 coat protein (MCP, SEQ ID NO: 107) that recognizes the MS2 aptamer (SEQ ID NO: 106), the λN peptide (SEQ ID NO: 109) that recognizes the BoxB element (SEQ ID NO: 108), and the PP7 coat protein (PCP, SEQ ID NO: 111) that recognizes the PP7 aptamer (SEQ ID NO: 110). In a preferred embodiment, the fused RNA recognition peptide is MCP.
[0073] In some embodiments, the RNA recognition peptide is linked to an RNA-specific adenosine deaminase or its deaminase domain via a linker peptide. An exemplary linker peptide is the XTEN linker peptide.
[0074] Mutant ADAR2 and mutant ADAR2dd
[0075] Through screening, the inventors discovered several ADAR2 protein mutations, which, whether used alone or in combination, can enhance the editing specificity of the RNA editing system.
[0076] In some embodiments, the one or more mutations are located at amino acid residue sites selected from the following: R348, V351, T375, N473, K475, S486, E488, K594, E620, and Q696 (amino acid positions refer to SEQ ID NO: 1).
[0077] In some embodiments, at least one mutation is located at an amino acid residue site selected from R348, N473 and Q696 (amino acid position refers to SEQ ID NO: 1).
[0078] Suitable mutation types have been identified in this disclosure, such as V351G, T375S, N473S, K475I, K475Q, S486A, E488Q, K594I, E620G, and Q696F. In some embodiments, at least one mutation in ADAR2 or ADAR2dd is N475S, K475I, or Q696F.
[0079] In some implementations, the ADAR2 or ADAR2dd comprises a combination of mutations selected from the following: E488Q / V351G, E488Q / T375S, E488Q / N473S, E488Q / K475I, E488Q / S486A, E488Q / K594I, E488Q / E620G, E488Q / Q696F, E488Q / K475Q, E488Q / K475I / S486A, or K475I / S486A.
[0080] In some embodiments, the mutant ADAR2 or mutant ADAR2dd is used in a conventional RNA editing system. In some embodiments, the mutant ADAR2 or mutant ADAR2dd is used in the fusion protein described in this disclosure.
[0081] Fusion protein that splits protease
[0082] One advantage of this disclosed RNA editing system is that, after the deaminase is recruited via ADAR-recruited guide RNA (agRNA), system activation is achieved by cleaving and removing the fused repressive domain (NDI). Therefore, in one embodiment, this disclosure provides a fusion protein comprising a protease and an RNA recognition peptide capable of recognizing the guide RNA.
[0083] A conventional guide RNA contains an antisense RNA fragment (AS) complementary to the target RNA molecule, preferably with a mismatched base (e.g., a U-to-C mutation) at the expected editing site of the target RNA (e.g., A-to-I editing). Furthermore, the guide RNA contains two RNA tags (e.g., MS2) flanking the antisense RNA fragment (AS) for recognition by the MCP fused to the RNA-specific adenosine deaminase.
[0084] In some embodiments, the protease fusion protein comprises an MCP that recognizes the MS2 element in the guide RNA. In a preferred embodiment, the protease fusion protein comprises an RNA recognition peptide other than an MCP, examples of which include: a λN peptide (SEQ ID NO: 109) that recognizes BoxB (SEQ ID NO: 108), and a PP7 coat protein (PCP, SEQ ID NO: 111) that recognizes PP7 (SEQ ID NO: 110). In some embodiments, the RNA recognition peptide is a λN peptide (SEQ ID NO: 109).
[0085] To further reduce off-target editing, the protease can be split into two fragments, neither of which can cleave the protease cleavage site on its own. The TEV protease is such an example, which can be split into an N-terminal domain (TEVn, SEQ ID NO: 93) and a C-terminal domain (TEVc, SEQ ID NO: 94). Therefore, according to one embodiment of this disclosure, the protease in the fusion protein comprises TEVn and TEVc, which are linked by a self-cleaving peptide. The use of the self-cleaving peptide ensures post-translational dissociation, thereby allowing at least one of them to separate from the RNA recognition peptide.
[0086] In some embodiments, the protease cleavage site is a self-cleaving peptide, such as a 2A peptide. A "2A peptide" is a viral oligopeptide of 18 to 22 amino acids in length that mediates polypeptide cleavage during eukaryotic translation. The "2A" designation originates from a specific region of the viral genome, and 2A peptides from different viral sources are usually named after their original virus. The first discovered 2A peptide was the foot-and-mouth disease virus 2A peptide (F2A), followed by the equine rhinitis A virus 2A peptide (E2A), porcine teschovirus-12A (P2A), and Thorsea asigna virus 2A peptide (T2A). SEQ ID NO: 112-114 shows several non-limiting examples of 2A peptides.
[0087] In some embodiments, the TEVn is located at the N-terminus of the TEVc. In some embodiments, the TEVn is located at the C-terminus of the TEVc. In some embodiments, the RNA recognition peptide is fused to the TEVn, optionally via a peptide linker (e.g., XTEN). In some embodiments, the RNA recognition peptide is fused to the TEVc, optionally via a peptide linker (e.g., XTEN).
[0088] In one specific embodiment, the RNA recognition peptide is located at the N-terminus of the TEVc, and the TEVc is then linked to the N-terminus of the TEVn via a self-cleavage site. In some embodiments, the TEVn and the TEVc may each be additionally fused with a nuclear export signal (NES).
[0089] In some embodiments, peptide linkers may optionally be provided between the fragments of the fusion protein. In some embodiments, the peptide linker comprises 1-100 amino acid residues (or 3-20, 4-15, or any number). In some embodiments, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the amino acid residues in the peptide linker are selected from alanine, glycine, cysteine, and serine.
[0090] For any fusion protein described in this disclosure, this disclosure also covers its bioequivalence. In some embodiments, the bioequivalence has at least about 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with a reference fusion protein. Preferably, the bioequivalence retains the target activity of the reference fusion protein. In some embodiments, the bioequivalence is obtained by modifying the reference sequence with one, two, three, four, five, or more amino acid additions, deletions, substitutions, or combinations thereof. In some embodiments, the substitutions are conserved amino acid substitutions.
[0091] "Conservative amino acid substitution" refers to replacing a certain amino acid residue with another amino acid residue with a similar side chain structure. Families of amino acid residues with similar side chains have been classified in the art, including: basic side chain amino acids (such as lysine, arginine, histidine), acidic side chain amino acids (such as aspartic acid, glutamic acid), uncharged polar side chain amino acids (such as glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), nonpolar side chain amino acids (such as alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), β-branched side chain amino acids (such as threonine, valine, isoleucine), and aromatic side chain amino acids (such as tyrosine, phenylalanine, tryptophan, histidine). Therefore, non-essential amino acid residues in immunoglobulin polypeptides are preferably replaced with another amino acid residue from the same side chain family. In another embodiment, a string of amino acids may be replaced with a string of structurally similar amino acids that differ in sequence and / or composition from the members of the side chain family.
[0092] Engineered ADAR recruitment guide RNA (eagRNA)
[0093] It also provides engineered agRNAs that can work in conjunction with the novel RNA editing system disclosed herein.
[0094] As described above, a conventional guide RNA contains an antisense RNA fragment (AS) complementary to the target RNA molecule, preferably with a mismatched base (e.g., U=>C) at the intended editing site (e.g., A-to-I) of the corresponding target RNA. Furthermore, the guide RNA contains two RNA tags (e.g., MS2) flanking the antisense RNA fragment (AS) for recognition by the MCP fused to the RNA-specific adenosine deaminase.
[0095] In the engineered agRNA (eagRNA), one or two copies of another RNA tag are added, which can be recognized by an RNA recognition peptide of a fusion protein containing a protease (or a cleaving protease). In one embodiment, the RNA tag is BoxB (SEQ ID NO: 108), which is recognized by the lambda (λN) peptide (SEQ ID NO: 109). In another embodiment, the RNA tag is PP7 (SEQ ID NO: 110), which is recognized by the PP7 coat protein (PCP, SEQ ID NO: 111).
[0096] The positions of additional RNA tags and MS2 aptamers can be designed as needed, some of which are such as Figure 8 As shown in b. In one embodiment, the eagRNA comprises two MS2 segments flanking the antisense fragment, and a BoxB or PP7. In some embodiments, the BoxB or PP7 is located at the 5' end of the 5' MS2. In some embodiments, the BoxB or PP7 is located at the 3' end of the 3' MS2.
[0097] In some embodiments, BoxB or PP7 is positioned between 5' MS2 and the antisense fragment. In some embodiments, BoxB or PP7 is positioned between 3' MS2 and the antisense fragment.
[0098] In some embodiments, the eagRNA comprises two copies of BoxB or PP7. In some embodiments, both copies of BoxB or PP7 are located outside the two MS2 segments (one on each side). In some embodiments, both copies of BoxB or PP7 are positioned between the respective MS2 segment and the antisense fragment. In a preferred embodiment, both copies of the λN peptide are located outside the two MS2 segments, arranged as 5'-λN peptide-MS2-antisense fragment-MS2-λN peptide-3'.
[0099] In some embodiments, each BoxB or PP7 is directly linked to the adjacent MS2. In some embodiments, one, two, three, four, five, or six nucleotides (e.g., A or AAA) are inserted between each BoxB or PP7 and the adjacent MS2. In some embodiments, the number of spacer nucleotides is odd (e.g., 1, 3, or 5, such as A, AAA, and AAAAA).
[0100] RNA editing
[0101] RNA editing methods utilizing one or more molecules disclosed herein are also provided, particularly methods for RNA editing in target cells.
[0102] For example, the method requires introducing into target cells (A) a first fragment containing an RNA-specific adenosine deaminase or its deaminase domain, a second fragment containing an inhibitory domain of cytidine deaminase, a protease cleavage site located between the first and second fragments, and an RNA recognition peptide; (B) a second polynucleotide encoding a protease capable of cleaving the protease cleavage site; and (C) a guide RNA or DNA encoding the guide RNA, wherein the guide RNA contains an antisense fragment capable of hybridizing with the target RNA molecule and an RNA tag recognizable by the RNA recognition peptide.
[0103] After being introduced into target cells, the guide RNA recognizes and binds to the target RNA molecule to be edited. Using its contained RNA tag, the guide RNA can recruit fusion proteins containing RNA-specific adenosine deaminase or its deaminase domain to the target RNA molecule. However, the RNA-specific adenosine deaminase or its deaminase domain in the fusion protein remains inactive until the protease is present simultaneously.
[0104] In one embodiment, the protease is fused to a second RNA recognition peptide, and the guide RNA further comprises a second RNA tag that can be recognized by the second RNA recognition peptide. For example, the first RNA tag may be an MS2 aptamer, and the corresponding RNA recognition peptide is an MS2 capsid protein (MCP); the second RNA tag may be a BoxB or PP7 aptamer, and the corresponding RNA recognition peptide is a lambda (λN) peptide or a PP7 capsid protein (PCP).
[0105] In this configuration, the guide RNA can also recruit the protease to the target RNA molecule where the fusion protein is present. The protease then cleaves the inhibitory domain of cytidine deaminase, thereby activating RNA-specific adenosine deaminase or its deaminase domain.
[0106] In some implementations, the protease is split into two segments separated by a self-cleavage site. Neither segment alone possesses protease activity. During intracellular expression, the self-cleavage site self-cleaves, separating the two segments. One segment fuses with a second RNA recognition peptide, recruiting it to the target RNA molecule; the other segment exists as a separate protein and can still bind to the first segment via free diffusion, thereby initiating the protease reaction. As illustrated in the examples, this configuration further inhibits off-target editing.
[0107] In some embodiments, the protease is a TEV protease, whose two cleavage fragments are an N-terminal domain (TEVn) and a C-terminal domain (TEVc). In some embodiments, the guide RNA is the eagRNA described above in this disclosure.
[0108] In some embodiments, the fusion protein comprises a single copy of a cytidine deaminase repressor domain (e.g., mA3-CDA2). In some embodiments, the fusion protein comprises two or more repressor domains. Exemplary structural configurations of fusion proteins containing repressor domains have been described above and are incorporated herein by reference.
[0109] In some embodiments, the RNA-specific adenosine deaminase or its deaminase domain contains one or more mutations as described in this disclosure, which can further enhance editing specificity. For example, the mutation site is selected from the following amino acid residues: R348, V351, T375, N473, K475, S486, E488, K594, E620, and Q696 (amino acid positions refer to SEQ ID NO: 1). In some embodiments, the mutation site is selected from the following amino acid residues: V351, T375, N473, K475, S486, E488, K594, E620, and Q696 (amino acid positions refer to SEQ ID NO: 1).
[0110] In some embodiments, the mutation includes one or more of V351G, T375S, N473S, K475I, K475Q, S486A, E488Q, K594I, E620G, and Q696F. In some embodiments, at least one mutation in ADAR2 or ADAR2dd is N473S, K475I, or Q696F.
[0111] In some implementations, the modified ADAR2 or ADAR2dd contains a combination of mutations selected from the following: E488Q / V351G, E488Q / T375S, E488Q / N473S, E488Q / K475I, E488Q / S486A, E488Q / K594I, E488Q / E620G, E488Q / Q696F, E488Q / K475Q, E488Q / K475I / S486A, and K475I / S486A.
[0112] The editing can be performed in vitro, particularly in cell culture systems. When the editing is performed in vitro or in vivo, the RNA editing system disclosed herein has clinical / therapeutic value. The introduction of relevant substances into cells or vivo can be achieved by administering them to living subjects, including but not limited to humans, animals, yeast, plants, bacteria, and viruses.
[0113] This disclosure also provides kits and packaging suitable for carrying out the methods described above. For example, the kit or packaging comprises: (A) a first polynucleotide encoding a fusion protein, said fusion protein comprising: a first fragment containing an RNA-specific adenosine deaminase or a deaminase domain thereof, a second fragment containing an inhibitory domain of cytidine deaminase, a protease cleavage site located between the first fragment and the second fragment, and an RNA recognition peptide; and (B) a second polynucleotide encoding a protease capable of cleaving said protease cleavage site.
[0114] In some embodiments, the kit or package further comprises: (C) guide RNA or DNA encoding the guide RNA, wherein the guide RNA comprises an antisense fragment capable of hybridizing with the target RNA molecule and an RNA tag recognizable by the RNA recognition peptide.
[0115] The components of the kit or package have been described in detail above, and this section is incorporated herein by reference.
[0116] Example
[0117] Example 1: Off-target editing of ADAR2dd and ADAR2dd (E488Q)
[0118] This embodiment examines the degree of off-target editing of the ADAR2 deaminase domain (ADAR2dd) and its mutant ADAR2dd (E488Q).
[0119] ADAR2dd and ADAR2dd (E488Q) were fused to MS2-binding protein (MCP) via XTEN linker peptides, respectively. Figure 1a). The above fusion protein was targeted using ADAR recruitment guide RNA (agRNA); the agRNA contained MS2 aptamers at both the 5' and 3' ends of the antisense (AS) targeting region. Figure 1 (a) In this study, six agRNAs targeting four target sites (CTNNB1-1, CTNNB1-2, PPIB-1, and STAT1-1) were co-transfected into HEK293FT cells with either MCP-XTEN-ADAR2dd or MCP-XTEN-ADAR2dd (E488Q). After 48 hours of culture, cellular RNA was extracted for reverse transcription, and the RNA editing effect was then detected by PCR amplification of target sites and potential off-target sites.
[0120] This study found that MCP-ADAR2dd and MCP-ADAR2dd (E488Q) achieved approximately 40-80% A-to-I editing efficiency at the target site in four tests. Figure 1 (b and c). However, this study also found that when co-transfected with agRNA (e.g., agCTNNB1-1), MCP-XTEN-ADAR2dd and MCP-XTEN-ADAR2dd (E488Q) induced significant off-target editing (approximately 30-80%) at all five detected off-target sites (KDELR1, MYBL2, MDK, PSMD1, and AP3D1). Figure 1 (d and e). In contrast, no significant off-target editing was detected in the corresponding untransfected (NT) group. Figure 1 (d and e). Interestingly, these five detected off-target sites showed no sequence similarity to each other or to the target sites. These results indicate that although MCP-XTEN-ADAR2dd and MCP-XTEN-ADAR2dd (E488Q) are capable of mediating efficient AI editing at the target RNA sites, they both induce transcriptome-wide off-target mutations.
[0121] Example 2: Inhibition of editing by the nucleoside deaminase inhibitor (NDI) domain
[0122] This embodiment attempts to reduce off-target mutations by inhibiting the deamination activity of these ADAR2 enzymes.
[0123] Some APOBEC cytidine deaminase family members contain a cytidine deaminase (CDA) domain, one of which is non-catalytically active and inhibits the activity of the flanking catalytically active domains. Previous studies have shown that although these CDA inhibitory domains effectively suppress the DNA-editing activity of deaminases, they lead to increased RNA editing. For example, Wang... et al. Nature Cell Biology2021 (23):552–563 reported that the deaminase-CDA repressor domain fusion (tBE-V5-rA1) actually induced a higher level of RNA off-target mutations compared to the conventional base editor BE3 (page 560, paragraph 1).
[0124] In a surprising discovery of this disclosure, the inventors demonstrate below that these CDA repressive domains are actually capable of effectively inhibiting the RNA editing activity of ADAR enzymes.
[0125] Here, this study fuses a CDA inhibitory domain—mouse APOBEC3 cytidine deamination domain 2 (mA3-CDA2)—into MCP-XTEN-ADAR2dd (E488Q). Figure 2 (a). For reference, this study also tested three other protein / domains, namely EGFP (an unrelated protein), ADAR3 (a deaminase-deficient member of the ADAR family that binds to a portion of RNA and inhibits the editing of ADAR1 and ADAR2), and Vif239 (an E3 ubiquitin ligase targeting a portion of cytidine deaminases), fused into MCP-XTEN-ADAR2dd (E488Q) ( Figure 2 (a). After co-transfecting these fusion proteins with two agRNAs, this study found that, compared with the EGFP control, the fusion of mA3-CDA2 and Vif239 significantly reduced the editing frequency of ADAR2dd (E488Q), while the fusion of ADAR3 only slightly reduced the editing frequency of ADAR2dd (E488Q). Figure 2 (b and c). These experiments show that CDA repressive domains (e.g., mA3-CDA2) also have an inhibitory effect on the activity of adenosine deaminases (e.g., ADAR).
[0126] To further explore whether other non-catalytically active CDA inhibitory domains can inhibit the activity of ADAR2dd (E488Q), this study fused the CDA inhibitory domains of human APOBEC3B (hA3B-CDA1, SEQ ID NO: 3), human APOBEC3D (hA3D-CDA1, SEQ ID NO: 85), human APOBEC3F (hA3F-CDA1, SEQ ID NO: 54), and human APOBEC3G (hA3G-CDA1, SEQ ID NO: 92) into MCP-XTEN-ADAR2dd (E488Q). Figure 3(a). Then, this study measured the editing efficiency of these constructed fusion proteins by co-expressing six agRNAs (agPPIB-1, agPPIB-2, agSTAT1-1, agSTAT1-2, agCTNNB1-1, and agCTNNB1-3) targeting six different RNA sites. The results showed that all tested CDA repressive domains inhibited the adenosine deaminase activity of ADAR2dd (E488Q), with A3DCDA2 exhibiting the highest inhibitory effect. Figure 3 (b and c).
[0127] These results indicate that, in addition to the previously reported effect of inhibiting cytidine deaminase activity, the inhibitory CDA domain can also inhibit the adenosine deaminase activity of ADAR. Therefore, these CDA inhibitory domains are referred to as nucleoside deaminase inhibitor (NDI) domains.
[0128] Example 3: The two NDI domains further inhibit deaminase activity.
[0129] This embodiment tested whether fusing an additional copy of the NDI domain could further enhance the inhibitory activity of the NDI domain.
[0130] This study fuses two copies of human APOBEC3D NDI (A3DNDI) into MCP-XTEN-ADAR2dd (E488Q). Figure 4 (a), and compared the editing efficiency of the obtained MCP-XTEN-ADAR2dd(E488Q)-2×A3DNDI with that of MCP-XTEN-ADAR2dd(E488Q)-A3DNDI. Figure 4 (b and c). For example Figure 4 As shown in b and c, the fusion of two copies of A3DNDI exhibits a stronger suppression effect than the fusion of one copy of A3DNDI.
[0131] However, further testing showed that even a single copy of NDI could still effectively suppress off-target editing. Figure 13 Figure a provides a schematic diagram of the RNA editing systems tested here. As shown, the three RNA editing systems tested each contained only one copy of A3DNDI, which was fused to an engineered ADAR2dd with different mutations via a TEV protease cleavage site (TS). These mutations in the ADAR2dd protein included K475Q / E488Q, K475I / S486A / E488Q, and K475I / S486A, respectively.
[0132] Targeted editing efficiency was tested at three target sites with 5'-CAN, 5'-GAN, and 5'-AAN sequence backgrounds (PAICS-1, STAT1-3, and CCNI-1, respectively). RNA editing efficiency was measured using RT-PCR and quantified based on Sanger sequencing results. Figure 13 (c). Sanger sequencing results quantified... Figure 13 In section b, as shown in the figure, all three editing systems demonstrated excellent editing efficiency.
[0133] Off-target A-to-I editing was tested at the AP3D1 site, measured by RT-PCR and quantified based on Sanger sequencing results. Figure 13 (e). For example Figure 13 The results summarized by d show that off-target rates are very low, especially in cases with K475I / S486A / E488Q and K475 / S486A mutations. Similar results were also observed at the EIF4G1 off-target site. Figure 13 f and g).
[0134] Example 4: Repressive Mutation
[0135] This embodiment tested whether mutations in the ADAR2 sequence could further reduce its deamination activity, which may help reduce off-target editing.
[0136] This study engineered ADAR2dd by introducing certain amino acid alterations, including E488Q, S486A, E488Q / R348A, E488Q / V351G, E488Q / T375S, E488Q / N473S, E488Q / K475I, E488Q / K475Q, E488Q / S486A, E488Q / K594I, E488Q / E620G, and E488Q / Q696F. Figure 5 (a and b). Then, this study examined the editing frequency of the engineered proteins with or without fusion of 2×A3DNDI (a and b). Figure 5 (b and c).
[0137] Table 1 ADAR2dd sequences
[0138] Except for E488Q / R348A, other amino acid changes showed relatively low editing frequencies when fused with 2×A3DNDI compared to ADAR2dd (E488Q), while maintaining relatively high editing efficiency when not fused with 2×A3DNDI. Figure 5 (b and c).
[0139] Example 5 Targeted activation of RNA base editor
[0140] This embodiment aims to develop a novel RNA adenine base editing system that activates the editing activity of an inactive RNA adenine base editor at a target site.
[0141] First, this study inserted a TEV protease cleavage site (TS) between ADAR2dd (E488Q) and 2×A3DNDI. Figure 6 (a), and found that the inhibitory effect of 2×A3DNDI still existed ( Figure 6 (b and c). Co-expression of the complete TEV protease, MCP-ADAR2dd(E488Q)-TS-2×A3DNDI, and the corresponding agRNA restored the editing efficiency of the target site. Figure 6 (b and c), because the TEV protease can cleave TS and remove 2×A3DNDI, thereby activating ADAR2dd (E488Q).
[0142] Unfortunately, the complete TEV protease ( Figure 7 Co-expression of ac also restored off-target editing ( Figure 7 The dg indicates that the intact TEV protease can approach TS and cleave NDI regardless of the RNA site, which may be due to the free diffusion of the intact TEV protease.
[0143] To reduce off-target effects of the RNA base editing system in this study, the TEV protease was split into two segments: a C-terminal fragment (TEVc) and an N-terminal fragment (TEVn). Nuclear export signals (NES) were added to the C-terminus of TEVc and the N-terminus of TEVn, respectively. The two TEV protease fragments were linked by a T2A self-cleaving peptide, ensuring that the two TEV fragments were produced as two independent proteins. Figure 7 (a). After co-expression of agRNA and MCP-XTEN-ADAR2dd(E488Q)-2×A3DNDI with TEVc-T2A-TEVn or TEVn-T2A-TEVn, this study found a significant reduction in off-target editing, especially when TEVc-T2A-TEVn was co-expressed. Figure 7 dg), although targeted editing also declined ( Figure 7 (b and c).
[0144] Then, this study attempted to engineer agRNA to enhance targeted editing efficiency. In this study, the λN peptide (which can bind to the BoxB aptamer) was fused to the N-terminus of TEVc via an XTEN linker. Figure 8 (a), and BoxB aptamers were added to the 5' end of the original agRNA or simultaneously to its 5' and 3' ends. Figure 8 (b). After co-transfection with MCP-XTEN-ADAR2dd(E488Q)-TS-2×A3DNDI, λN-XTEN-TEVc-T2A-TEVn and engineered agRNA, this study found that all of these agRNAs induced highly efficient editing at the target site. Figure 8 (c and e). At two confirmed off-target sites, this study found that BoxB-MS2-AS-MS2-BoxB-agRNA induced relatively low levels of off-target mutations (c and e). Figure 8 (d and e). Therefore, in subsequent studies, this study selected BoxB-MS2-AS-MS2-BoxB-agRNA (referred to as engineered agRNA, eagRNA).
[0145] This study also tested the effect of linker peptides on the performance of the RNA base editing system used in this study. In this study, the XTEN linker peptides in MCP-XTEN-ADAR2dd(E488Q)-TS-2×A3DNDI and / or λN-XTEN-TEVc-T2A-TEVn were replaced with GS dipeptides ( Figure 9 λN-GS-TEVc-T2A-TEVn and MCP-XTEN-ADAR2dd(E488Q)-TS-2×A3DNDI were co-transfected into HEK293FT cells. This study found that both λN-GS-TEVc-T2A-TEVn and MCP-XTEN-ADAR2dd(E488Q)-TS-2×A3DNDI induced highly efficient targeted editing (a). Figure 9 (b and c), while simultaneously inducing low-level off-target editing at two off-target sites ( Figure 9 (dg).
[0146] To completely suppress off-target editing of the RNA base editing system in this study, further amino acid alterations were introduced into ADAR2dd, namely K475Q / E488Q, K475I / S486A / E488Q, and K475I / S486A (…). Figure 10 (a), and tested their effects on targeted and off-target editing (a). Figure 10 (bg). This study found that MCP-XTEN-ADAR2dd(K475Q / E488Q)-TS-2×A3DNDI induced the highest on-target editing efficiency ( Figure 10 (b and c), while still triggering low-level off-target editing ( Figure 10 The targeted editing efficiency induced by MCP-XTEN-ADAR2dd(K475I / S486A / E488Q)-TS-2×A3DNDI is slightly lower than that induced by MCP-XTEN-ADAR2dd(K475I / E488Q)-TS-2×A3DNDI. Figure 10 (b and c), but off-target editing was significantly suppressed ( Figure 10 Notably, compared to the non-transfected (NT) control, MCP-XTEN-ADAR2dd(K475I / S486A)-TS-2×A3DNDI did not induce observable off-target editing (dg). Figure 10 (dg), although its targeted editing efficiency is relatively lower than that induced by MCP-XTEN-ADAR2dd(K475Q / E488Q)-TS-2×A3DNDI and MCP-XTEN-ADAR2dd(K475I / S486A / E488Q)-TS-2×A3DNDI. Figure 10 (b and c).
[0147] Figure 11 This study demonstrates the proposed working mechanism of this novel RNA base editor. The editing activity of MCP-ADAR2dd is inhibited by the fused 2×A3DNDI, thereby preventing off-target mutations caused by free MCP-ADAR2dd binding to off-target sites of RNA that can interact with MCP or ADAR2dd. The newly developed eagRNA contains BoxB and MS2 to recruit effector proteins. After the MCP-ADAR2dd-TS-2×A3DNDI fusion protein is recruited to the eagRNA, the 2×A3DNDI is cleaved by the recombinant TEV protease. The MCP-ADAR2dd-eagRNA complex can then bind to the target RNA and induce efficient A-to-I targeted editing. Specific mutations (e.g., K475I / S486A / E488Q or K475I / S486A) can be incorporated to further reduce or even eliminate off-target editing.
[0148] Further testing was performed on ADAR2dd with the K475I / S486A mutation (explained in...). Figure 12 (a). The results are shown in Figure 12 c and summarized in Figure 12 b. For example Figure 12 As shown in bc, the engineered ADAR2dd (K475I / S486A) exhibits reduced off-target effects compared to the wild-type ADAR2dd.
[0149]
[0150] The scope of this disclosure is not limited to the specific embodiments described, which are intended as a single illustration of various aspects of this disclosure, and any functionally equivalent compositions or methods are within the scope of this disclosure. It will be apparent to those skilled in the art that various modifications and variations can be made to the methods and compositions of this disclosure without departing from the spirit or scope of this disclosure. Therefore, this disclosure is intended to cover modifications and variations thereof, provided they fall within the scope of the appended claims and their equivalents.
[0151] All publications and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication or patent application is specifically and individually indicated to be incorporated by reference.
Claims
1. A fusion protein comprising: The first segment containing RNA-specific adenosine deaminase or its deaminase domain; The second segment contains the inhibitory domain of cytidine deaminase; and The protease cleavage site located between the first fragment and the second fragment.
2. The fusion protein as described in claim 1, characterized in that, The RNA-specific adenosine deaminase is RNA adenosine deaminase (ADAR); optionally, the ADAR is selected from the group consisting of ADAR1, ADAR2 and ADAR3.
3. The fusion protein of claim 2, wherein the ADAR is human ADAR2, comprising the amino acid sequence shown in SEQ ID NO: 1, or an amino acid sequence having at least 85% sequence identity with SEQ ID NO:
1.
4. The fusion protein of claim 3, wherein, Relative to SEQ ID NO: 1, the ADAR2 contains one or more mutations selected from the following: R348V, V351G, T375S, N473S, K475I, K475Q, S486A, E488Q, K594I, E620G and Q696F.
5. The fusion protein of claim 4, wherein, Relative to SEQ ID NO: 1, the ADAR2 contains mutations selected from the following combinations: K475Q and E488Q, K475I and S486A / E488Q, and K475I and S486A.
6. The fusion protein according to any one of claims 3-5, wherein the deaminase domain of said ADAR2 comprises amino acid residues 316-700 of SEQ ID NO:
1.
7. The fusion protein as claimed in any of the preceding claims, wherein the cytidine deaminase is APOBEC (apolipoprotein B mRNA editing enzyme that catalyzes polypeptide-like proteins).
8. The fusion protein of claim 7, wherein the repressor domain comprises an amino acid sequence selected from SEQ ID NO: 2-91, or an amino acid sequence having at least 85% sequence identity with any amino acid sequence selected from SEQ ID NO: 2-92.
9. The fusion protein of claim 7, wherein the repressive domain is selected from the group consisting of mA3-CDA2, hA3B-CDA1, hA3D-CDA1, hA3F-CDA1 and hA3G-CDA1.
10. The fusion protein of claim 7, wherein the nucleobase deaminase inhibitor comprises the amino acid sequence shown in SEQ ID NO: 2 or amino acid residues 128-223 of SEQ ID NO:
2.
11. The fusion protein of any of the preceding claims, wherein the protease cleavage site is not recognized by endogenous proteases in human cells.
12. The fusion protein of any of the preceding claims, wherein the protease cleavage site is recognized by a protease selected from the group consisting of TuMV protease, PPV protease, PVY protease, ZIKV protease and WNV protease.
13. The fusion protein of claim 12, wherein the protease cleavage site is a TEV protease cleavage site.
14. The fusion protein of any of the preceding claims, wherein the second fragment comprises at least two repressive domains; optionally, the repressive domains are the same repressive domains or different repressive domains.
15. A fusion protein comprising: (a) an N-terminal domain (TEVn) of a TEV protease; (b) a C-terminal domain (TEVc) of a TEV protease; (c) a self-cleavage site located between TEVn and TEVc; and (d) an RNA recognition peptide.
16. The fusion protein of claim 15, wherein TEVn is located at the C-terminus of TEVc.
17. The fusion protein of claim 16, wherein the RNA recognition peptide is a lambda (λN) peptide, PP7 capsid protein (PCP), or MS2 capsid protein (MCP).
18. The fusion protein of claim 16, wherein the RNA recognition peptide is a lambda (λN) peptide.
19. A mutant human ADAR2 deaminase domain (ADAR2dd) or a mutant ADAR2 containing said ADAR2dd, wherein, The ADAR2dd contains mutations selected from the following: R348A; N473S; K475I; Q696F; E488Q / V351G; E488Q / T375S; E488Q / N473S; E488Q / K475I; E488Q / S486A; E488Q / K594I; E488Q / E620G; E488Q / Q696F; E488Q / K475Q; E488Q / K475I / S486A; K475I / S486A, and their combinations, The amino acid positions are referenced in SEQ ID NO:
1.
20. The mutant ADAR2dd or ADAR2 as described in claim 19, wherein the ADAR2dd comprises an E488Q / K475Q, E488Q / K475I / S486A, or K475I / S486A mutation.
21. A guide RNA or DNA encoding the guide RNA for recognizing an RNA sequence, wherein the guide RNA comprises an antisense fragment capable of hybridizing with a target RNA sequence, the antisense fragment being flanked by two MS2 aptamers, and an RNA tag that can be recognized by an RNA recognition peptide that is not an MS2 coat protein (MCP).
22. The guide RNA of claim 21 or the DNA encoding the guide RNA, wherein the RNA tag is a BoxB aptamer or a PP7 aptamer.
23. The guide RNA or DNA encoding the guide RNA as claimed in claim 21 or 22, wherein the guide RNA comprises two copies of the RNA tag.
24. The guide RNA of claim 23 or the DNA encoding the guide RNA, wherein the two RNA tags are located outside the two MS2 aptamers.
25. The guide RNA of claim 24 or the DNA encoding the guide RNA, wherein each of the two RNA tags is separated from the corresponding MS2 aptamer by one or three nucleotides.
26. A polynucleotide encoding a fusion protein as described in any one of claims 1-18, or a mutant ADAR2dd or ADAR2 as described in claim 19 or 20.
27. A method for editing a target RNA molecule in a cell, comprising introducing into said cell: A first polynucleotide encoding a fusion protein comprising: a first segment containing an RNA-specific adenosine deaminase or its deaminase domain; a second segment containing an inhibitory domain of cytidine deaminase; a protease cleavage site located between the first and second segments; and an RNA recognition peptide. The second polynucleotide encodes a protease capable of cleaving the protease cleavage site, and a guide RNA or DNA encoding the guide RNA, wherein the guide RNA comprises an antisense fragment capable of hybridizing with the target RNA molecule and an RNA tag that can be recognized by the RNA recognition peptide.
28. The method of claim 27, wherein the RNA tag comprises an MS2 aptamer and the RNA recognition peptide comprises an MS2 coat protein (MCP).
29. The method of claim 27 or 28, wherein the protease comprises: (a) an N-terminal domain (TEVn) of the TEV protease; (b) a C-terminal domain (TEVc) of the TEV protease; and (c) a self-cleaving site located between TEVn and TEVc.
30. The method of claim 29, wherein the protease is fused to the second RNA recognition peptide, and the guide RNA further comprises a second RNA tag that can be recognized by the second RNA recognition peptide.
31. The method of claim 30, wherein the second RNA recognition peptide comprises a lambda (λN) peptide or a PP7 capsid protein (PCP), and the second RNA tag comprises a BoxB aptamer or a PP7 aptamer.
32. The method of claim 31, wherein the guide RNA comprises the antisense fragment, the antisense fragment is flanked by two copies of the RNA tag, and the RNA tag is flanked by two copies of the second RNA tag.
33. The method according to any one of claims 27-32, wherein the cell is a mammalian cell, preferably a human cell.
34. The method according to any one of claims 27-33, wherein it is performed in vitro, ex vivo, or in vivo.
35. A reagent kit or package comprising: A first polynucleotide encoding a fusion protein comprising: a first segment containing an RNA-specific adenosine deaminase or its deaminase domain; a second segment containing an inhibitory domain of cytidine deaminase; a protease cleavage site located between the first and second segments; and an RNA recognition peptide. The second polynucleotide encodes a protease capable of cleaving the protease cleavage site.
36. The kit or package of claim 35 further comprises guide RNA or DNA encoding the guide RNA, wherein the guide RNA comprises an antisense fragment capable of hybridizing with a target RNA molecule and an RNA tag recognizable by the RNA recognition peptide.