An engineered cytosine deaminase and methods of making and uses thereof
By designing and engineering cytosine deaminase ProACD, the CU-REWIRE system was constructed, solving the problems of low efficiency and high off-target rate of existing RNA base editing tools. This system enables efficient and precise RNA base editing, which is suitable for gene therapy and the treatment of genetic diseases.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SONGJIANG HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIVERSITY SCHOOL OF MEDICINE
- Filing Date
- 2025-01-20
- Publication Date
- 2026-05-29
AI Technical Summary
Existing RNA base editing tools suffer from problems such as complex editing systems, low efficiency, interference with non-target RNA, and toxicity to immune responses. In particular, APOBEC3A protein is inefficient in editing specific sequences, which limits the application of C-to-U base editors.
An engineered cytosine deaminase, ProACD, was designed. By splitting the cytosine deaminase protein into N-terminal and C-terminal domains and fusing them with the deamination domain, the CU-REWIRE system was constructed. The ePUF10 protein was used to target RNA and bind to the ProACD protein, achieving efficient C-to-U base editing independent of gRNA.
It achieves efficient and precise RNA base editing, reduces off-target rate, stabilizes expression, is suitable for clinical applications, and is safe and efficient, making it suitable for gene therapy and basic research.
Smart Images

Figure BDA0005249490910000281 
Figure HDA0005249490920000011 
Figure HDA0005249490920000012
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene editing, and in particular to an engineered cytosine deaminase, its preparation method, and its uses. Background Technology
[0002] I. Scientific Research and Gene Therapy Value of Gene Editing Tools
[0003] Genes are the most important hereditary material in living organisms. How to correctly understand, predict, and ultimately control gene expression so that cells can function normally is a major opportunity and challenge for the further development of biology now and in the future.
[0004] Base editing, a ubiquitous gene modification phenomenon in organisms, alters DNA or RNA sequence information to rewrite codons, create new RNA splicing sites, and process a single DNA or RNA molecule into multiple mature mRNAs that can be translated into different functional proteins. This diversification of genetic information provides an essential basis for the organism's evolutionary complexity. Furthermore, a significant portion of known human diseases are caused by gene mutations, with approximately 58% of human hereditary gene mutation diseases resulting from single-base mutations.
[0005] Therefore, designing and developing efficient and precise base editing enzymes to modify the sequence of single bases in DNA or RNA to manipulate the expression of target genes is of great research value. At the same time, using base editing enzymes to precisely repair pathogenic single-base mutations in target genes is very meaningful for the research and treatment of such gene mutation diseases.
[0006] Current gene therapy strategies for diseases caused by single-base mutations primarily involve directly rewriting the sequence information of the target gene at the DNA or RNA level through base editing to repair or replace the mutated gene and thus treat the disease. DNA base editing, as a traditional gene manipulation strategy, can directly edit bases on DNA to rewrite the target DNA sequence and correct gene expression. However, because changes to genomic DNA persist throughout a cell's lifespan, the long-term safety concerns surrounding genomic DNA editing remain a significant issue.
[0007] II. Advantages and Immense Application Value of RNA Base Editing
[0008] Compared to DNA base editing therapies, RNA base editing directly modifies the RNA sequence by editing target bases, without altering the DNA sequence. It is a non-permanent, reversible gene expression regulation strategy that does not produce heritable edited products and is relatively easy to manipulate. As a potential therapeutic approach, it is safer than DNA editing because it does not produce heritable edited products. Therefore, manipulating genes at the RNA level offers better controllability and safety, making this type of gene therapy more conducive to the translation of basic research into clinical applications.
[0009] RNA base editing, a ubiquitous gene modification phenomenon in organisms, alters RNA sequences to rewrite codons and create new RNA splicing sites. It directly regulates alternative splicing, translation, and degradation of RNA, manipulating the expression of target genes and ultimately diversifying genetic information. This provides organisms with a better foundation for complex evolution, making it an indispensable part of biological processes. RNA base editing in organisms can be broadly classified into two categories based on their molecular targeting mechanisms: guide RNA (gRNA)-mediated RNA base editing and RNA-binding protein (RBP)-mediated base editing. The products of these targeted editing processes can exert their respective physiological functions in different environments.
[0010] In eukaryotes, existing base-editing proteins mainly fall into two categories: ADARs and APOBECs. Adenosine deaminases (ADARs) primarily act on double-stranded RNA, while cytidine deaminases (APOBECs) primarily act on single-stranded RNA. ADAR proteins contain an RNA-binding protein domain at their N-terminus, responsible for recognizing and locating the target RNA. Their C-terminus contains an adenosine deamination catalytic domain, responsible for catalyzing the hydrolysis and deamination of target adenine, converting adenosine to inosine. Inosine is then recognized as guanine during translation, achieving adenosine-to-guanine (A-to-G) base editing. APOBEC protein catalyzes the hydrolysis and deamination of RNA cytosine, converting cytidine to uridine, thus achieving base editing from cytidine to uridine (C-to-U).
[0011] Abnormalities in the RNA base editing process can affect the normal growth and development of organisms, leading to various diseases in higher organisms. Statistics show that approximately 58% of human hereditary gene mutation diseases are caused by single-base mutations. Among these, 21,593 guanosine-to-adenosine (G-to-A) and 6,945 thymine-to-cytosine (T-to-C) single-base pathogenic mutations have been identified (as of March 2022). Theoretically, these types of diseases can be repaired and corrected through DNA or RNA base editing.
[0012] Given that RNA base editing directly or indirectly participates in most life activities, and that its editing process is closely related to disease, using ADAR or APOBEC as RNA base editing enzyme donors and recruiting them to the target RNA editing site using RBP or gRNA can achieve precise base editing at the target site. RNA base editing gene therapy is a strategy that directly repairs or replaces mutated genes at the RNA level, thereby correcting abnormal RNA processing and metabolism, and ultimately treating diseases at their root. This strategy does not alter the genome sequence and does not produce heritable byproducts, making it a potential gene therapy method that can be applied in vivo.
[0013] III. Research Progress and Design Flaws of Existing RNA Base Editing Tools
[0014] In recent years, several adenosine-to-guanosine (A-to-I) RNA base editing tools have been developed. Their main editing strategy is to use gRNA to target the target RNA and recruit the adenosine deaminase ADAR2 or Cas13-ADAR2 effector protein to form a target RNA / gRNA / ADAR ternary complex at the site to be edited, thus achieving A-to-I base editing. This has a good foundation for widespread application, but it still has some shortcomings: (1) The editing system is complex. In this system, gRNA, ADAR2 effector protein and target RNA cofold and assemble into a ternary complex in the cell, which is the basis for base editing. (2) The editing efficiency is low. The assembly efficiency and ratio of the gRNA / ADAR effector protein / target RNA ternary complex in this system are the rate-limiting steps for efficient base editing. The whole process requires the colocalization and folding assembly of gRNA and ADAR effector protein, and gRNA and target RNA, respectively. (3) Disrupting the expression of non-target RNAs: Overexpressed gRNA components pair with non-target RNAs, altering their structure and interfering with their normal expression, thereby disrupting the normal expression of transcriptome RNAs. (4) As the core protein of the CRISRP system, the Cas protein, being derived from bacterial proteins, may induce immune responses and toxicity in clinical applications.
[0015] In recent years, the development of tools for cytidine-to-uridine (C-to-U) base editing has remained in a relatively early stage, with the main editing strategy utilizing natural APOBEC cytidine deaminases. APOBEC proteins are effector domains that function within RNA-binding proteins (RBPs). Natural APOBEC proteins, as a large class of cytidine deaminases acting on single-stranded RNA in the human body, can achieve C-to-U base editing. Most APOBEC proteins are only 200 amino acids long, much smaller than ADAR, giving them a significant advantage in in vivo delivery.
[0016] There are two main categories of C-to-U base editing tools: The first category employs a strategy similar to A-to-I base editing, utilizing gRNA to target the target RNA and recruiting the cytidine deaminase APOBEC3A-Cas13 effector protein. This forms a target RNA / gRNA / APOBEC3A-Cas13 ternary complex at the RNA editing site, allowing for C-to-U base editing. However, this system suffers from a significant drawback: APOBEC3A is a single-stranded RNA cytidine deaminase. The gRNA targets and binds to the target RNA, forming a gRNA / target RNA double strand at the target site. This targeting eliminates the single-stranded RNA environment required for APOBEC to function, leading to decreased editing efficiency and precision, making it difficult to apply in real-world scenarios. The second type: To overcome the difficulty of CRISPR systems in utilizing APOBEC proteins for C-to-U base editing, a gRNA-independent base editing system (CU-REWIRE) was developed. This system uses RNA-binding proteins (PUFs) to recognize and bind to target RNA, and targets the APOBEC3A protein to the target RNA for C-to-U base editing. This process does not require the participation of gRNA, thus solving the problem encountered by the first type of tools and achieving efficient and precise base editing of target RNA. IV. The Necessity and Immense Application Value of Artificially Designed and Engineered Cytosine Base Editing Enzymes
[0017] However, existing natural APOBEC proteins exhibit relatively low activity in RNA base editing systems. Currently, only human APOBEC3A can efficiently perform site-specific cytidine-to-uridine (C-to-U) base editing at the RNA level. However, existing APOBEC3A proteins exhibit strong sequence bias when editing target bases; they can only edit the "C" base in the "UC" sequence. APOBEC3A cannot edit the "C" base in the "AC", "CC", or "GC" sequences, resulting in extremely low editing efficiency in these sequences. This editing property of APOBEC3A limits the scope of C-to-U base editors and becomes a rate-limiting step for the further application of C-to-U base editing tools.
[0018] In summary, the currently highly active APOBEC3A protein can only edit the "C" base in the "UC" sequence; other natural APOBEC proteins have relatively low editing activity, and their editing performance cannot meet the requirements for use, making them difficult to apply in C-to-U site-directed base editing tools.
[0019] Therefore, it is essential to design non-natural cytosine deaminases with highly efficient editing activity that can edit the base "C" in the "UC", "AC", "CC", and "GC" sequences respectively. This achievement also has extremely high application value.
[0020] Utilizing artificially designed non-natural cytosine deaminases to design next-generation DNA or RNA base editing tools can improve editing efficiency and precision, thereby enabling more accurate and efficient C-to-U base editing. This is a necessary alternative to existing tools and can be further applied to basic research and gene therapy. This achievement also has extremely high application value. Summary of the Invention
[0021] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide an engineered cytosine deaminase, its preparation method and uses, to solve the problems in the prior art.
[0022] To achieve the above and other related objectives, the present invention provides an engineered cytosine deaminase, comprising a backbone and a deamination domain disposed within the backbone. The backbone of the cytosine deaminase protein is composed of an N-terminal domain and a C-terminal domain of the cytosine deaminase protein, wherein the N-terminal and C-terminal domains are derived from one cytosine deaminase protein, and the deamination domain is derived from another cytosine deaminase protein.
[0023] The present invention also provides a method for preparing the engineered cytosine deaminase, the method comprising linking a nucleotide encoding the backbone of one cytosine deaminase protein with a nucleotide encoding the deamination domain of another cytosine deaminase protein to obtain a nucleotide encoding the engineered cytosine deaminase fragment, and then cloning the nucleotide encoding the engineered cytosine deaminase fragment into an expression vector and expressing it to obtain the engineered cytosine deaminase.
[0024] The present invention also provides an ePUF10 protein, wherein the ePUF10 protein is obtained by integrating leucine and proline into the fourth recognition unit of the PUF10 domain.
[0025] The present invention also provides the use of the engineered cytosine deaminase or the ePUF10 protein in the preparation of a base editing system.
[0026] The present invention also provides a DNA base editing system, wherein the DNA base editing system includes the cytosine deaminase.
[0027] The present invention also provides an RNA base editing system, wherein the RNA base editing system includes the engineered cytosine deaminase.
[0028] The present invention also provides an isolated polynucleotide encoding any one or a fragment thereof: the engineered cytosine deaminase, the ePUF10 protein, or the base editing system.
[0029] The present invention also provides a nucleic acid construct comprising the isolated polynucleotides described above.
[0030] The present invention also provides a viral vector system, the viral vector system comprising the nucleic acid construct.
[0031] The present invention also provides an adeno-associated virus (AAV), which is formed by viral packaging of the viral vector system.
[0032] The present invention also provides a cell containing the aforementioned nucleic acid construct or a genome in which exogenous polynucleotides are integrated.
[0033] The present invention also provides a base editing method, wherein a target gene is brought into contact with the adeno-associated virus or the base editing system to achieve single base editing on the target gene.
[0034] As described above, the engineered cytosine deaminase, its preparation method, and its uses of the present invention have the following beneficial effects: it is a gRNA-independent, simple, efficient, and low-off-target enzyme that can be stably expressed in cells, meeting the high-efficiency delivery requirements for later clinical applications. It can serve as a potential, relatively safe in vivo editing tool, and is expected to play an important supplementary role in certain areas of gene therapy and basic research, such as the treatment of Duchenne muscular dystrophy, premature aging, and other genetic diseases. Attached Figure Description
[0035] Figure 1 Crystal structure-based optimization of the ePUF10 protein enhanced the protein expression intensity of CU-REWIRE. Figure a shows the crystal structures of PUF10 and ePUF10 predicted using the AlphaFold2 tool (left), and the schematic diagram of ePUF10 protein binding to substrate RNA (right). Figure b shows a schematic diagram of the CU-REWIRE editing system designed using different versions of PUF10. The original CU-REWIRE 3.0 (top) incorporates the PUF10 protein, and the upgraded CU-REWIRE 4.0 (bottom) incorporates the ePUF10 protein, containing 10 RNA recognition units targeting 10-nt RNA sequences. Figure c shows the protein expression level analysis of different versions of the CU-REWIRE editing system.
[0036] Figure 2Crystal-structure-based optimization of the ePUF10 protein enhanced the editing efficiency of CU-REWIRE. Figure a shows the effect of CU-REWIRE 3.0 and CU-REWIRE 4.0 on CL in EGFP mRNA. 459 Efficient editing. The PUF10 binding site in green fluorescent protein is marked with a blue underline, and adjacent cytosine is marked with red (top image). The heatmap shows the target site C for each CU-REWIRE and control group. 459 And the editing efficiency of all nearby cytosines. C 459 The editing efficiency is shown on the right, determined by RNA-seq. Figure b shows the off-target effects of different versions of the CU-REWIRE editing enzyme. Transcriptome-level C-to-U base editing events were determined by RNA-seq experiments. Orange diamonds represent the target site EGFP-C. 459 The right side shows the average efficiency of total off-target editing events. "n" represents the total number of cytosine off-target editing events detected. Figure c shows the effect of CU-REWIRE4 on the target site C in different EGFP transcripts. 459 Editing activity analysis. Figure d shows the editing window analysis of the CU-REWIRE4 system. Efficiency analysis of cytosine at different positions downstream of the editing binding site in CU-REWIRE4.0, where n-bp represents the number of bases between the editing site and the 3' end binding site of PUF10.
[0037] Figure 3 Optimizing the amino acid composition of APOBEC3A to reduce off-target effects of the CU-REWIRE editing system. Figure a shows a schematic diagram of the functional domains of different structural domains of human APOBEC3A. Figure b shows a schematic diagram of different versions of CU-REWIRE4.x recognizing and binding target EGFP mRNA. "X" represents versions of APOBEC3A carrying different point mutations. Figure c shows the base editing efficiency analysis of EGFP mRNA by CU-REWIRE4.X. The ePUF10 binding site in EGFP is marked with a blue underline, and adjacent cytosine is marked with red. The heatmap shows the target site C for each CU-REWIRE and ePUF10 control group. 459 Nearby cytosine editing efficiency, detailed display on the right shows C 459 Editing efficiency. Editing rate was determined by three biological replicates using RNA-seq. Figure d shows the transcriptome-level RNA base editing efficiency and off-target effects in samples treated with different CU-REWIRE4 protocols. Orange diamonds represent the target site EGFP-C. 459The average editing efficiency of off-target editing events is labeled next to the dot plot. "n" represents the total number of cytosine base edits detected. Figure e shows the correlation analysis between the on-target site editing efficiency (y-axis) of different CU-REWIRE4.X and the whole transcriptome C-to-U off-target editing events (x-axis).
[0038] Figure 4 Schematic diagrams illustrating the breakdown of different domains of APOBEC proteins. Figure a shows the crystal structure and domain breakdown of human APOBEC3A (PDB:4XXO). Through sequence analysis, APOBEC3A protein was broken down into an N-terminal domain (NTD), a C-to-U deaminase domain, and a C-terminal domain (CTD), each with independent functions. Figures be show the crystal structures and domain breakdowns of natural APOBEC proteins predicted by AlphaFold2, including mouse APOBEC3 (b), human APOBEC1 (c), mouse APOBEC1 (d), and rat APOBEC1 (e).
[0039] Figure 5 The application of artificial intelligence to develop a programmable engineered cytosine base editing enzyme, ProACD. Figure a shows a schematic diagram of the splitting of different domains of the natural AID / APOBEC protein (the splitting strategy is derived from...). Figure 4 Figure b shows the assembly diagram of the programmable engineered cytosine base editing enzyme ProACD developed with the assistance of artificial intelligence. Figure c shows the phylogenetic tree analysis of AID / APOBEC family proteins, with major branches marked in different colors. Figure d shows the artificial synthesis and production of engineered ProACD based on AI prediction. Figures ef and ef show the crystal structures of natural AID / APOBEC proteins and engineered ProACD proteins predicted by AlphaFold2. Top: Crystal structure of natural AID / APOBECs predicted by artificial intelligence. Bottom: Crystal structure of ProACD predicted by artificial intelligence.
[0040] Figure 6A new generation of programmable RNA base editing system, CU-REWIRE5, was developed using ProACD. Figure a shows a schematic diagram of the programmable CU-REWIRE5 system designed using the ProACD protein. The top figure shows a schematic diagram of the upgraded CU-REWIRE5 structure for C-to-U base editing. The middle figure shows a schematic diagram of the programmable assembly of the modules in the CU-REWIRE5 system. The deamination domain of ProACD is provided by natural APOBEC deaminases from different sources for specific C-to-U base editing. The RNA recognition domain ePUF10 includes 10 RNA recognition units, each designed for specific RNA base recognition. The bottom figure shows a schematic diagram of the precise binding of CU-REWIRE5 to the target EGFP mRNA; the ePUF10 binding sequence is AACGCUCUAUA, and the editing target site is C. 459 Figure b shows the donor source table for the deamination domain of ProACD in different versions of CU-REWIRE5. Figure c shows the protein expression level analysis of different versions of CU-REWIRE5 in HEK293T cells. Figure d shows the expression level of different versions of CU-REWIRE5 at the EGFP mRNA target site C. 459 Editing efficiency analysis.
[0041] Figure 7 Targeting range and specificity analysis of the CU-REWIRE5 base editing system. Figure a shows the targeting range and specificity of CU-REWIRE5s in the target EGFP mRNA. 459 Schematic diagram of target site recognition. Original U 458 The C sequence was manually changed to A. 458 C, G 458 C and C 458 Figure C is used to test the base editing sequence preference of CU-REWIRE5. Figure b shows the statistical analysis of EGFP mRNA target sites C. 459 In different base sequence environments (A) 458 C,C 458 C,U 458 C,G 458 Editing efficiency in C). The binding sequence and editing target of ePUF10 are the same. Figure 7 Figures a and c show the overall off-target effects of CU-REWIRE5s at the transcriptome level, assessed using RNA-seq experiments. (Samples were obtained from...) Figure 7 b). The orange diamond represents EGFPC. 459 The right side shows the average editing efficiency of off-target editing events.
[0042] Figure 8Analysis of editing sequence preference and editing accuracy of CU-REWIRE 5.1 and 5.8 for editing GC motifs. Figure a shows the editing sequence preference and editing accuracy of different versions of CU-REWIRE5s at the EGFP mRNA target site C. 459 The editing efficiency and accuracy analysis are shown in Figure b. Figure c shows the editing preference analysis of the CU-REWIRE 5.1 editing enzyme at on-target sites. Figure b shows the editing preference analysis of the CU-REWIRE 5.1 editing enzyme in all off-target editing events in the transcriptome (data from...). Figure 7 c). d shows the editing preference analysis of the CU-REWIRE5.8 editing enzyme at on-target sites. e shows the editing preference analysis of the CU-REWIRE5.8 editing enzyme in all off-target editing events in the transcriptome (data from...). Figure 7 c).
[0043] Figure 9 Editing window analysis and gene therapy applications of CU-REWIRE5.1 (CU5.1). Figure a shows the editing of EGFPC by CU-REWIRE5.1. 459 Analysis of the editing window size at the target site; the figure shows C at different distances from the CU-REWIRE5.1 binding site to PUF. 459 Editing efficiency of target sites. Figure b shows the pathogenic target sites of APOE4 mRNA edited using CU5.1. 388 Editing efficiency was determined by Sanger sequencing. ePUF10 binding sites are marked in blue. Figure c shows the pathogenic target sites C for editing APOE4 mRNA using CU5.1. 526 Editing efficiency was determined by Sanger sequencing. ePUF10 binding sites are marked in blue. Figure d shows the pathogenic target site C of Rhodopsin mRNA edited using CU5.1. 76 Efficiency analysis. Figure e shows the pathogenic target site C of SOD1 mRNA edited using CU5.1. 143 Efficiency analysis. Figure f summarizes the editing window sizes of the target sites for editing EGFP, APOE4Rhodopsin, and SOD1 mRNA by CU-REWIRE5.1. The x-axis represents the distance from the editing site to the PUF binding site, and the y-axis represents the C-to-U editing efficiency, which was determined by Sanger sequencing.
[0044] Figure 10 Analysis of editing sequence preference and editing accuracy of CU-REWIRE5.15 for editing CC motifs. Figure a shows the editing sequence preference and editing accuracy of different versions of CU-REWIRE5s at the EGFP mRNA target site C. 459 The editing efficiency and accuracy were analyzed. Figure b shows the CU-REWIRE 5.15 editing enzyme in C. 459Editing preference analysis at the site. Figure c shows the editing preference analysis of the CU-REWIRE5.15 editing enzyme in all off-target editing events in the transcriptome (data from...). Figure 7 c).
[0045] Figure 11 Editing window analysis and gene therapy applications of CU-REWIRE5.15 (CU5.15). Figure a shows the editing of EGFPC by CU-REWIRE5.15. 459 Analysis of the editing window size at the target site. Figure b shows the pathogenic target site C of Mef2cmRNA edited using CU5.15. 104 Editing efficiency was determined by Sanger sequencing. ePUF10 binding sites are marked in blue. Figure c summarizes the editing window sizes of target sites for editing EGFP and Mef2c mRNA using CU-REWIRE 5.15.
[0046] Figure 12 Editing sequence preference and editing accuracy analysis of CU-REWIRE5.16 for editing AC motifs. Figure a shows the editing sequence preference and editing accuracy of different versions of CU-REWIRE5s at the EGFP mRNA target site C. 460 The editing efficiency and accuracy were analyzed. Figure b shows the CU-REWIRE 5.16 editing enzyme at target site C. 460 Editing preference analysis. Figure c shows the editing preference analysis of the CU-REWIRE 5.16 editing enzyme in all off-target editing events in the transcriptome (data from...). Figure 7 c). d shows EGFPC edited by CU-REWIRE 5.16. 460 Analysis of editing window size at target sites. Figure e summarizes the editing window size at target sites for editing different mutant EGFP mRNAs using CU-REWIRE 5.16.
[0047] Figure 13 Analysis of the editing enzyme characteristics of CU-REWIRE5s, which can edit cytosine C in any AC, CC, GC, or UC motif. Figure a shows the different versions of CU-REWIRE5s at the EGFP mRNA target site C. 459 The editing efficiency and accuracy were analyzed. Figure b shows the CU-REWIRE 5.17 editing enzyme at target site C. 459 Editing preference analysis. Figure c shows the editing preference analysis of the CU-REWIRE 5.17 editing enzyme in all off-target editing events in the transcriptome (data from...). Figure 7 c). d shows the CU-REWIRE5.3 editing enzyme at target site C. 459 Editing preference analysis. Figure e shows the editing preference analysis of the CU-REWIRE 5.3 editing enzyme in all off-target editing events in the transcriptome (data from...). Figure 7 c).
[0048] Figure 14 Editing window analysis and gene therapy applications of CU-REWIRE 5.17 (CU5.17). Figure a shows the editing of EGFPC by CU-REWIRE 5.17. 459 Analysis of the editing window size at the target site. Figure b shows the pathogenic target site C of DDX3X mRNA edited using CU5.17. 1541 Editing efficiency was determined by Sanger sequencing. ePUF10 binding sites are marked in blue. Figure c summarizes the editing window sizes of target sites for different transcripts of EGFP and DDX3X edited by CU-REWIRE 5.17.
[0049] Figure 15 A new generation base editing system, CU-REWIRE5s, was constructed using an engineered ProACD enzyme derived from mammals. Figure 1 shows the crystal structure of the engineered ProACD protein predicted by AlphaFold2. The design and structure prediction methods for the ProACD protein are the same as those used in other systems. Figure 5 .
[0050] Figure 16 To construct a new generation of base editing system CU-REWIRE5s using mammalian-derived synthetic ProACD enzymes. Figure a shows a schematic diagram of the programmable CU-REWIRE5 system designed using mammalian-derived ProACD protein. Figure b shows a schematic diagram of the precise binding of CU-REWIRE5 to the target EGFP mRNA; the ePUF10 binding sequence is AACGUCUAUA, and the editing target site is C. 459 The figure shows the target sites C of different versions of CU-REWIRE5 on different EGFP transcripts. 459 Editing efficiency analysis. Figure g shows the base editing preference analysis of different versions of CU-REWIRE5 on different EGFP transcripts.
[0051] Figure 17 Analysis of protein expression and editing efficiency of the AAV-CU-REWIRE5 system in the mouse brain. Figure a shows a schematic diagram of CU-REWIRE5.15 (CU5.15) packaged in PHP.eb serotype AAV that can cross the blood-brain barrier and injected into mice via intravenous injection. Figure b shows the immunofluorescence protein expression analysis of CU5.15 (green) and DAPI (blue) in the mouse cortex and hippocampus 4 weeks after CU5.15 viral injection. Scale bar: 200 μm. Figure c shows the effect of CU5.15 on the target site C in mouse Mef2c mRNA. 104 Editing efficiency analysis. The binding sites of ePUF10 are shown in blue, C 104Marked in red (top). Editing efficiency was determined by three replicates using three sequencing sequences. Figure d shows Mef2c WT or L35P edited in ePUF10 or CU5.15. + / - In mice, U+ levels were measured in the prefrontal cortex and hippocampus. 104 The percentage. Figure e shows the editing efficiency analysis of CU5.15 on the C-to-U site of Mef2c C104 in each group. Editing efficiency formula: Cedited = (U CU5.15 -U ePUF10 ) / C total (1-U ePUF10 Figure f shows the MEF2C WT or L35P levels in the ePUF10 or CU5.15 administration groups. + / - Analysis of MEF2C protein expression levels in the prefrontal cortex and hippocampus of mice. The antibody used was the MEF2C antibody. GAPDH was used as an internal control. Figure g shows the quantitative analysis of MEF2C protein expression levels using ImageJ (data from...). Figure 17 f).
[0052] Figure 18 Protein expression analysis of MEF2C protein in the brain of Mef2c L35P mice. Figure a shows the expression of MEF2C WT and L35P in different drug administration groups. + / - Analysis of in situ protein expression levels of MEF2C (red) and DAPI (blue) in the splenial cortex (RSC) and hippocampus of mice. Results were determined by immunohistochemical staining. Figure b shows the quantitative fluorescence density of MEF2C protein in the RSC (ePUF10 vs. CU5.15, P = 0.0040) (left). The quantitative fluorescence density of MEF2C protein in the hippocampus (ePUF10 vs. CU5.15, P < 0.0001) (right). Figure c shows the expression levels of Mef2c L35P in the ePUF10 or CU5.15 treatment groups. + / - Survival curve analysis of mice. 23 WT (ePUF10) mice and 17 L35P+ / - (ePUF10) mice were included. + / - (CU5.15) 23 mice).
[0053] Figure 19Phenotypic analysis of autism-related social behaviors in Mef2c L35P mice. Figure a shows the motion heatmap analysis of social activities in Mef2c WT or L35P+ / - mice in the three-box test in different drug administration groups. Figure b shows the time analysis of social interaction between Mef2c mice and unfamiliar mice in the three-box test in different drug administration groups. Figure c shows the cumulative social interaction time analysis between Mef2c mice and unfamiliar or familiar mice in different drug administration groups. Figure d shows a schematic diagram of the unfamiliar mouse intrusion test. T0: The experimental mice were isolated and housed for three days. T1-T4: Data analysis of social interaction tests between experimental mice and their partners. T5: Data analysis of social interaction between experimental mice and their new partner mice. The recording time was the sniffing time between mice within 2 minutes. Figure e shows the sniffing time analysis of Mef2c WT+ePUF10 in the social intruder test. Figure f shows the Mef2c L35P... + / - Analysis of sniffing time of +ePUF10 in social intruder tests. Figure g shows Mef2c L35P. + / - Analysis of sniffing time in the social intruder test in the +CU5.15 experimental group. Detailed Implementation
[0054] This invention provides a method for preparing and using an artificially designed programmable cytosine deaminase. The method includes splitting a cytosine deaminase protein into three segments, and fusing the backbone composed of the N-terminal functional domain (NTD) and C-terminal functional domain (CTD) of the split cytosine deaminase protein with the deaminase domain of another cytosine deaminase protein to prepare an engineered cytosine deaminase, named ProACD. Pro grammable A rtificial C ytosine D (Cytosine deaminase). This cytosine deaminase can be used in RNA and DNA base editing systems.
[0055] This invention constructs a novel gRNA base editing tool, CU-REWIRE (C-to-U RNA editing with individual RNA-binding enzyme), by fusing the cytosine deaminase ProACD and the programmable RNA-binding protein ePUF10 (enhanced PUF10). CU-REWIRE targets RNA editing based on ePUF10. Specifically, it fuses the RNA recognition domain (ePUF10) and the effector domain (ProACD) to construct the new base editing protein CU-REWIRE. This protein specifically targets target RNA through the ePUF10 domain and utilizes the cytosine deaminase ProACD domain to achieve C-to-U (Cytosine to Uridine) base editing at the target RNA site. By performing gene editing at the RNA level, it can eliminate pathogenic RNA to correct pathogenic point mutations; or edit expressed RNA molecules to specifically manipulate gene function.
[0056] Specifically, this invention first provides an engineered cytosine deaminase, which includes a backbone and a deamination domain disposed within the backbone. The backbone of the cytosine deaminase protein is composed of an N-terminal domain and a C-terminal domain of the cytosine deaminase protein. The N-terminal and C-terminal domains are derived from one cytosine deaminase protein, and the deamination domain is derived from another cytosine deaminase protein. That is, the N-terminal and C-terminal domains are derived from the same cytosine deaminase protein, while the deamination domain is derived from a cytosine deaminase protein different from the N-terminal and C-terminal domains.
[0057] The terms "one cytosine deaminase protein" and "another cytosine deaminase protein" refer to cytosine deaminase proteins derived from different individuals. Different individuals can refer to individuals from different classes (e.g., mammals), or individuals from the same species or genus.
[0058] In some embodiments of the present invention, the backbone of the cytosine deaminase protein is derived from the backbone of the human cytosine deaminase protein. In one embodiment, the backbone of the cytosine deaminase protein can be the backbone of the natural cytosine deaminase protein, or it can be a backbone obtained by mutation based on the natural cytosine deaminase protein.
[0059] In some embodiments of the present invention, the backbone of the cytosine deaminase protein is the C-terminal domain (CTD) and N-terminal domain (NTD) of the cytosine deaminase protein.
[0060] The N-terminal and C-terminal domains are prepared by cleaving a cytosine deaminase protein at the H70 and M153 sites, and the resulting N-terminal and C-terminal fragments are the N-terminal and C-terminal domains, respectively.
[0061] That is, the N-terminal domain is a segment from the N-terminus of the cytosine deaminase protein to the H70 site or any site within 5 amino acids above and below the H70 site (e.g., positions 65, 66, 67, 68, 69, 71, 72, 73, 74, and 75), and the C-terminal domain is a segment from the M153 site of the cytosine deaminase protein or any site within 5 amino acids above and below the M153 site (e.g., positions 148, 149, 150, 151, 152, 154, 155, 156, 157, and 158) to the C-terminus.
[0062] In one embodiment, the N-terminal domain is a fragment from the first amino acid at the N-terminus of the cytosine deaminase protein to the H70 site or any site within 5 amino acids above and below the H70 site, and the C-terminal domain is a fragment from the M153 site of the cytosine deaminase protein or any site within 5 amino acids above and below the M153 site to the first amino acid at the C-terminus.
[0063] The deamination domain is the segment between the H70 site of the cytosine deaminase protein or any site within 5 amino acids above and below it, and the M153 site or any site within 5 amino acids above and below the M153 site.
[0064] In some embodiments of the present invention, the backbone of APOBEC3A includes: an N-terminal domain obtained by mutating one or more of the following sites based on the N-terminal domain of wild-type APOBEC3A as shown in SEQ ID NO.1: H11, H16, K30, H56, and / or, including a C-terminal domain obtained by mutating the C171 site based on the C-terminal domain of wild-type APOBEC3A as shown in SEQ ID NO.6.
[0065] In some embodiments of the present invention, the N-terminal domain amino acid sequence of the cytosine deaminase is shown in any one of SEQ ID NO. 2 to 5.
[0066] In some embodiments of the present invention, the amino acid sequence of the C-terminal domain of the cytosine deaminase is shown in SEQ ID NO.7.
[0067] In some embodiments of the present invention, the deamination domain of the engineered cytosine deaminase ProACD protein is derived from prokaryotes or eukaryotes. The eukaryotes are preferably mammals. The mammals are preferably rodents, even-toed ungulates, perissodactyls, lagomorphs, primates, etc. The primates are preferably Homo sapiens, orangutans, monkeys, or apes. In a preferred embodiment, the rodents are selected from rats and mice.
[0068] In some embodiments of the present invention, the deamination domain of the engineered cytosine deaminase ProACD protein is a natural deamination domain or a domain obtained by mutation based on a natural deamination domain.
[0069] In some embodiments of the present invention, the engineered cytosine deaminase comprises the following structure: NTD-deamination domain-CTD, where NTD represents the N-terminal domain and CTD represents the C-terminal domain. For example, the structure of the engineered cytosine deaminase is NTD-deamination domain-CTD-NTD-deamination domain-CTD, that is, NTD-deamination domain-CTD tandemly.
[0070] In some embodiments of the present invention, the domain donor of the engineered cytosine deaminase is selected from natural AID / APOBECs, and the engineered cytosine deaminase obtained is called ProACD. In one specific embodiment, the cytosine deaminase is selected from APOBEC1 or APOBEC3. In one specific embodiment, the cytosine deaminase is APOBEC3A. In one embodiment, the APOBEC3A is APOBEC3A after single-point or multi-point mutation. In one embodiment, the APOBEC3A is APOBEC3A obtained by mutating one or more of the following sites based on the wild-type APOBEC3A with the amino acid sequence shown in SEQ ID NO.8: H11, H16, C171, K30, H56.
[0071] The amino acid sequence of the engineered cytosine deaminase is shown in SEQ ID NO. 9-55.
[0072] The present invention also provides a method for preparing the engineered cytosine deaminase ProACD, the method comprising linking a nucleotide encoding the backbone of one cytosine deaminase protein with a nucleotide encoding the deamination domain of another cytosine deaminase protein to obtain a nucleotide encoding the engineered cytosine deaminase fragment, and then cloning the nucleotide encoding the engineered cytosine deaminase fragment into an expression vector and expressing it to obtain the engineered cytosine deaminase.
[0073] The specific experimental steps of the preparation method are conventional genetic engineering techniques in the field. In one embodiment, the experimental steps can be summarized as follows: assembling the DNA sequence encoding the deamination domain of the natural AID / APOBEC protein into the backbone of the human APOBEC3A protein to obtain the ProACD fusion protein fragment, then cloning the obtained ProACD fusion protein fragment into a vector, and then transfecting the vector into host cells to produce the engineered cytosine deaminase.
[0074] The preparation method utilizes AlphaFold2 for assistance, and the engineered cytosine deaminase ProACD obtained by the preparation method has cytosine-to-uracil activity.
[0075] The ProACDs obtained by the preparation method, which are "formed by fusing human APOBEC3A backbone domains (NTD and CTD) with various cytosine deaminase deamination domains from different species", exhibit editing specificity and efficiency completely different from that of natural APOBEC proteins when integrated into the CU-REWIRE system in multiple sequence contexts, including UC, AC, GC and CC.
[0076] The present invention also provides an enhanced PUF10 protein (ePUF10), wherein the ePUF10 protein is obtained by integrating leucine (L) and proline (P) into the fourth recognition unit (Repeat4, R4) of the PUF10 domain (PUF10).
[0077] PUF domain proteins are RNA recognition domains in the natural RNA-binding protein Pumilio and FBF (PUF). The ePUF10 protein is a redesigned RNA-binding protein with a similar structure to the natural PUF protein. It has 10 recognition unit sequences, each composed of 36 amino acids. The first, second, and fifth amino acids are responsible for recognizing target RNA bases. Different target bases can be recognized by regulating the composition of these three amino acids. By freely combining these recognition unit sequences that can recognize specific bases, the PUF10 protein can recognize 10 arbitrary RNA base sequences.
[0078] In one embodiment, the amino acid sequence of the PUF10 domain is shown in SEQ ID NO.56.
[0079] In some embodiments of the present invention, the amino acid sequence of the ePUF10 protein is shown in SEQ ID NO.57-64.
[0080] The present invention also provides the use of the engineered cytosine deaminase or the ePUF10 protein in the preparation of a base editing system.
[0081] The base editing system is selected from DNA base editing systems or RNA base editing systems. The DNA base editing system is a cytosine base editor (CBE) system, and / or a guanine base editor (GBE) system.
[0082] The present invention also provides a DNA base editing system, wherein the DNA base editing system includes the cytosine deaminase.
[0083] In some embodiments of the present invention, the DNA base editing system further includes a nuclease and / or guide RNA. The nuclease and guide RNA can be designed and selected according to existing technology. The nuclease may be a Cas nuclease.
[0084] The present invention also provides an RNA base editing system (CU-REWIRE, or CU-REWIRE fusion protein), wherein the RNA base editing system includes the engineered cytosine deaminase.
[0085] The RNA base editing system also includes the ePUF10 protein or Cas protein, which is directly linked to the engineered cytosine deaminase or linked to the engineered cytosine deaminase via a linker peptide.
[0086] In some embodiments of the present invention, the length of the linker peptide is 1 to 100 amino acids.
[0087] In some embodiments of the present invention, the linker peptide is selected from flexible peptides.
[0088] In some embodiments of the present invention, the ePUF10 protein and the engineered cytosine deaminase ProACD are fused at both ends by a flexible polypeptide to construct the CU-REWIRE fusion protein.
[0089] In some embodiments of the present invention, the flexible polypeptide is selected from XTEN linker peptides, GS linker peptides, NLS linker peptides, XTEN-NLS linker peptides, or XTEN-XTEN linker peptides. In some embodiments of the present invention, the amino acid sequences of the XTEN linker peptide, GS linker peptide, NLS linker peptide, XTEN-NLS linker peptide, or XTEN-XTEN linker peptide are shown in SEQ ID NO. 65-69, respectively.
[0090] In a preferred embodiment of the present invention, the RNA base editing system includes the engineered cytosine deaminase and the ePUF10 protein, wherein the ePUF10 protein is linked to the engineered cytosine deaminase via an XTEN linker peptide.
[0091] In some embodiments of the present invention, the amino acid sequence of the CU-REWIRE fusion protein is shown in SEQ ID NO. 70-93.
[0092] In the RNA base editing system, the ePUF10 protein recognizes and binds to RNA, and the engineered cytosine deaminase ProACD acts on single-stranded RNA to achieve C-to-U (cytosine to uracil) base editing. That is, the RNA base editing system can use the ePUF10 domain to target ProACD to the target RNA editing site, and ProACD will achieve C-to-U site-specific editing.
[0093] The CU-REWIRE RNA base editing system of the present invention can effectively perform C-to-U base editing in cultured cells without the use of gRNA. Moreover, all components of the CU-REWIRE system are derived from mammals, which can significantly reduce the immune response induced during in vivo editing.
[0094] The present invention also provides an isolated polynucleotide encoding any one or a fragment thereof: the engineered cytosine deaminase, the ePUF10 protein, or the base editing system.
[0095] The nucleotide sequences encoding the engineered cytosine deaminase, ePUF10 protein, or the isolated polynucleotides of the RNA base editing system can be inferred from their amino acid sequences using existing technologies, and are not specifically shown in this invention.
[0096] The amino acid sequence of the present invention may also be a sequence not shown in the present invention, but having 90% or more sequence identity with the sequence shown in the present invention, and having a similar function to the amino acid sequence shown in the present invention. Further, it may be obtained by substituting, deleting, or adding one or more (specifically 1-50, 1-30, 1-20, 1-10, 1-5, or 1-3) amino acids to the amino acid sequence shown in the present invention, and possessing the functional sequence of the amino acid sequence shown in the present invention. The amino acid sequence may have 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity with the nucleotide sequence shown in the present invention.
[0097] The present invention also provides a nucleic acid construct comprising the isolated polynucleotides described above.
[0098] In one embodiment, the nucleic acid construct is a nucleic acid construct comprising a separate polynucleotide encoding the engineered cytosine deaminase.
[0099] In one embodiment, the nucleic acid construct is a nucleic acid construct comprising a separate polynucleotide encoding the ePUF10 protein.
[0100] In one embodiment, the nucleic acid construct is a nucleic acid construct comprising a separate polynucleotide encoding the RNA base editing system. In one embodiment, the nucleotide sequence of the nucleic acid construct is shown in SEQ ID NO. 95.
[0101] The term "nucleic acid construct" refers to an artificially constructed nucleic acid segment that can be introduced into target cells or tissues. The nucleic acid construct can be various expression vectors, which include a vector backbone (empty vector) and an expression frame. The term "expression frame" refers to a sequence with the potential to encode a protein.
[0102] There is no specific limitation on the type of expression vector. An expression vector is a nucleic acid molecule that allows the insertion of foreign nucleotides without disrupting its ability to replicate and / or integrate into the host cell. Expression vectors may include nucleic acid sequences that allow them to replicate in the host cell, such as origins of replication. Expression vectors may also include one or more selective marker genes and other genetic factors. An expression vector is a vector containing the necessary regulatory sequences to enable the transcription and translation of one or more inserted genes. Expression vectors are selected from eukaryotic expression vectors or prokaryotic expression vectors.
[0103] The eukaryotic expression vector is selected from yeast expression vectors, insect expression vectors, or mammalian expression vectors. The mammalian expression vector is selected from retroviral expression vectors, lentiviral expression vectors, adenovirus expression vectors, and adeno-associated virus expression vectors.
[0104] The host cells are selected from eukaryotic or prokaryotic host cells. Eukaryotic host cells are selected from fungi such as yeast, insects, birds, plants, *C. elegans* or nematodes, or mammalian host cells. Non-limiting examples of insect cells are *Spodoptera frugiperda* (Sf) cells. Examples of yeast host cells are *Saccharomyces cerevisiae*, *Kluyveromyces lactis* (K. lactis), or *Yarrowialipolytica*. Examples of mammalian cells are COS cells, juvenile hamster kidney cells, mouse L cells, LNCaP cells, Chinese hamster ovary (CHO) cells, human embryonic kidney (HEK) cells, African green monkey cells, CV1 cells, Vero, or Hep-2 cells. Examples of prokaryotic host cells include bacterial cells such as Escherichia coli, Streptomyces, Bacillus subtilis, Salmonella typhi, or mycobacteria.
[0105] Those skilled in the art can transfect the expression vector into host cells using methods well-known in the art to obtain cells containing the coding gene of the cytosine deaminase, the ePUF10 protein, or the RNA base editing system. For example, the expression vector can be introduced into eukaryotic cells by calcium phosphate co-precipitation, electroporation, microinjection, liposome transfection, or transfection using polyamine transfection reagents.
[0106] The present invention also provides a viral vector system, the viral vector system comprising the nucleic acid construct. The viral vector system is an adeno-associated virus vector system.
[0107] The adeno-associated virus vector expression system is selected from Escherichia coli expression systems, yeast expression systems, insect expression systems, mammalian expression systems, and plant expression systems; preferably, the adeno-associated virus vector expression system is selected from any one of plasmid transient transfection expression systems, baculovirus expression systems, stable cell line expression systems, adenovirus expression systems, and poxvirus expression systems.
[0108] Furthermore, the adeno-associated virus vector system also includes a host cell. The host cell carries the adeno-associated virus vector. The host cell can be selected from various suitable host cells in the art, as long as it does not limit the purpose of the invention. Specifically, suitable cells can be cells that produce adeno-associated virus, such as 293 cells.
[0109] In some embodiments of the present invention, the adeno-associated virus vector system in the transient transfection expression system further includes a nucleic acid construct carrying the target gene and an auxiliary plasmid. In one embodiment of the present invention, the nucleic acid construct carrying the target gene includes a polynucleotide encoding the engineered cytosine deaminase, the ePUF10 protein, or the RNA base editing system. Specifically, the nucleic acid construct carrying the target gene is the aforementioned polynucleotide construct of the present invention that includes encoding the RNA base editing system.
[0110] This invention also provides an adeno-associated virus (AAV), which is packaged from the AAV vector system via viral packaging. The AAV can be used to treat various diseases, such as Duchenne muscular dystrophy and premature aging; the specific disease can be targeted based on the designed target sequence recognized by the ePUF10 protein.
[0111] The CU-REWIRE system of this invention can be successfully delivered to different tissues and organs of animals via AAV or LNP delivery systems to edit target RNA in the target organs. Specifically, the AAV-CU-REWIRE can be introduced into the Mef2c-L35P ASD mouse model via tail vein injection, and precisely edits and repairs target RNA in the brain of this mouse model with an efficiency of up to 50%. Moreover, the base-edited and repaired RNA can be translated normally, successfully restoring the expression of MEF2C protein and correcting autistic-like behaviors in Mef2c-L35PASD mice. It is evident that the CU-REWIRE system, through engineered ProACD, provides a powerful platform for in vivo base editing, paving the way for future clinical translation of gene therapy.
[0112] The present invention also provides a cell containing the aforementioned nucleic acid construct or a genome in which exogenous polynucleotides are integrated.
[0113] The cells described in this invention are obtained by converting the expression vector into host cells.
[0114] The present invention also provides a base editing method, wherein a target gene is brought into contact with the adeno-associated virus or the base editing system to achieve single base editing on the target gene.
[0115] In one embodiment, the single-base editing is C-to-U base editing. The target gene is a gene that binds to the ePUF10.
[0116] The base editing method is performed in cells or in vivo, in vitro, or in cell-free systems.
[0117] The present invention also provides a treatment method for a disease, the treatment method comprising administering the base editing system or the adeno-associated virus to a subject or isolated cells of a subject.
[0118] The specific type of disease is not limited; the disease type can be determined based on the pathogenic gene edited by the base editing system. For example, the disease could be a neurological disorder, specifically Alzheimer's disease.
[0119] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.
[0120] Before further describing specific embodiments of the present invention, it should be understood that the scope of protection of the present invention is not limited to the specific embodiments described below; it should also be understood that the terminology used in the embodiments of the present invention is for describing specific embodiments and not for limiting the scope of protection of the present invention; in the specification and claims of the present invention, unless otherwise expressly stated in the text, the singular forms "a", "an" and "this" include the plural forms.
[0121] When numerical ranges are given in the embodiments, it should be understood that, unless otherwise stated in the present invention, both endpoints of each numerical range and any value between the two endpoints may be selected. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. In addition to the specific methods, apparatus, and materials used in the embodiments, based on the knowledge of the prior art possessed by one of ordinary skill in the art and the description of this invention, any prior art methods, apparatus, and materials similar to or equivalent to those described, apparatus, and materials in the embodiments of this invention may be used to implement the present invention.
[0122] Materials and Methods
[0123] 1. Construction of engineered cytidine deaminase ProACD gene expression plasmid
[0124] To construct the ProACD fusion protein, the encoding DNA sequence of the deamination domain of the natural AID / APOBEC protein was first synthesized via gene synthesis. Then, the encoding DNA sequence of the deamination domain of the natural AID / APOBEC protein was assembled into the backbone of the human APOBEC3A protein to design a novel programmable ProACD fusion protein (amino acid sequence as shown in SEQ ID NO. 9-55). Subsequently, the encoding DNA sequence fragment of the obtained ProACD sequence was cloned into the pCI-Neo vector (purchased from Promega) to generate the pCI-ProACD-Flag vector.
[0125] 2. Construction of enhanced ePUF10 plasmid
[0126] The expression vector for RNA-binding protein PUF10 was constructed by first amplifying the ePUF10 sequence using PCR, then linearizing the pCl-neo plasmid by double digestion (NheI / NotI), and finally inserting the ePUF10 fragment into the pCl-neo vector using T4 DNA ligase to construct the pCl-ePUF10 plasmid (amino acid sequence as shown in SEQ ID NO.56-64).
[0127] 3. Construction of a novel CU-REWIRE5 protein expression plasmid
[0128] (1) Linearization of pCl-ProACD-Flag vector in 1.
[0129] (2) The ePUF10 sequence constructed in step 2 was inserted into the pCl-ProACD-Flag vector (see Material Method 1: purchased from Promega) by enzyme digestion and ligation to construct the pCl-ProACD-ePUF10(CU-REWIRE5) plasmid.
[0130] 4. Packaging and purification of adeno-associated virus
[0131] The adeno-associated virus (AAV) vector used in this study was the pAAV-hSyn-EGFP plasmid. First, the pAAV-hSyn-EGFP plasmid was linearized by enzyme digestion. Then, the ePUF10 and CU-REWIRE5 fragments obtained by PCR amplification were inserted into the pAAV-hSyn-EGFP plasmid through homologous recombination, yielding pAAV-hSyn-PUF10 (control group) and pAAV-hSyn-CU-REWIRE (experimental group), respectively. pAAV-hSyn-PUF10 and pAAV-hSyn-CU-REWIRE were then packaged into AAV-PHP.eB serotype adeno-associated virus, and the virus was purified by gradient centrifugation using iodixanol. The virus titer was then detected by qPCR.
[0132] 5. Mammalian cell culture and transfection
[0133] The cell line used in this study was HEK 293T (ATCC CRL-3216), which was cultured in DMEM high-glucose medium (HyClone) containing 10% fetal bovine serum (FBS) in a Thermo incubator at 37°C with 5% carbon dioxide.
[0134] Unless otherwise specified, all cell transfection experiments in this study used transient transfection. To investigate the editing effect of the CU-REWIRE tool on target sites in reporter genes, cells were co-transfected using the REWIRE system and the reporter gene, followed by cell collection for further analysis. All transfection reagents used in the above experiments were Lipofectamine 3000 (Thermo Fisher Scientific); please refer to the product manual for usage instructions.
[0135] 6. Mammal experiments
[0136] All mouse experiments conducted in this institute were approved and authorized by the Animal Welfare and Ethics Committee of Shanghai Jiao Tong University School of Medicine. The mice used were sourced from our laboratory's collection, and both female and male mice were used.
[0137] In this experiment, male rats were evenly divided into experimental and control groups. AAV-PHP.eB adeno-associated virus particles encoding the CU-REWIRE protein were injected into four-week-old Mef2c-L35P rats via tail vein injection. + / - In (C57BL / 6J) heterozygous mice, the injection dose per mouse was 1×10⁻⁶. 12vg (vector genomes). After injection, the genomes were carefully cultured until AAV-PHP.eB-CU-REWIRE was efficiently expressed (approximately 4 weeks). Then, RNA editing efficiency was assessed, and mouse social behavior was studied.
[0138] The experimental method for social behavior studies involves observing and analyzing the social behavior of mice. Under identical experimental conditions, the behavioral characteristics of mice in the RNA editing group and the control group were observed and compared. The differences in social behavior between the two groups were analyzed to investigate whether the autism-related pathological phenotypes in the RNA editing group mice were restored. Subsequently, different brain tissues were collected from the mice, RNA was extracted, and the editing efficiency of REWIRE on the target RNA in different brain tissues of mice was detected using RNA-seq or Sanger-seq.
[0139] Experimental methods not explicitly described in this invention, such as RNA sequencing, were performed using conventional techniques in the field. Example 1 demonstrates how structural optimization of the PUF10 protein improves the expression stability of the new CU-REWIRE base editing enzyme.
[0140] To design a more efficient and stable RNA recognition protein, this invention inserts an LP (L: leucine, P: proline) dipeptide into the fourth recognition unit (R4: recognition unit 4) of PUF10, as shown in SEQ ID NO. 56, to obtain an enhanced PUF10 (ePUF10) (amino acid sequence shown in SEQ ID NO. 57-64). The structure of ePUF10 is significantly different from that of PUF10, and the new structure is more likely to recognize and bind RNA. Figure 1 a).
[0141] Then, the ePUF10 module was fused with the cytidine deaminase APOBEC3A to develop a novel base editing enzyme, CU-REWIRE4.0. Figure 1 b. The amino acid sequence is shown in SEQ ID NO. 71. The results showed that the protein expression level of the ePUF10 version of CU-REWIRE4.0 with the LP insertion was significantly higher than that of the original CU-REWIRE3.0 in cells. Figure 1 c. The amino acid sequence is shown in SEQ ID NO. 70, indicating that the structure optimization of the ePUF10 protein with the insertion of the LP dipeptide can produce a more efficient and stable RNA base editor.
[0142] Example 2 verifies the editing efficacy of the new CU-REWIRE editing enzyme.
[0143] To verify whether the editing efficiency of the new version of CU-REWIRE has been significantly improved, we designed CU-REWIRE to target and edit the C32 of enhanced green fluorescent protein (EGFP) mRNA. 459 Target sites were used to evaluate the editing efficacy of CU-REWIRE 4.0. Figure 2 a-2d). mRNA high-throughput sequencing (RNA-Seq) results showed that CU-REWIRE4.0 achieved a significant improvement in C-to-U editing efficiency, reaching 82.3%, compared to 69.7% for CU-REWIRE3.0. Figure 2 a) indicates that the optimization of the ePUF10 structural domain can significantly improve the editing efficiency of CU-REWIRE4.0.
[0144] Off-target analysis at the transcriptome level (RNA-seq analysis covering 50X transcriptome) revealed 224 off-target base editing events in CU-REWIRE3.0 and 731 off-target base editing events in CU-REWIRE4.0. Figure 2 b) indicates that the editing power of the upgraded CU-REWIRE4.0 with the ePUF10 structure domain is significantly higher than that of the original CU-REWIRE3.0.
[0145] We further analyzed the edit sequence preference and edit precision of CU-REWIRE4.0. The results show that CU-REWIRE4.0 tends to edit the C( ) in the UC consensus motif. Figure 2 c). Furthermore, its editing precision is extremely high, further demonstrated by the fact that C-to-U base editing primarily occurs at the second position downstream of the ePUF10 binding site ( Figure 2 d). In summary, compared with CU-REWIRE3.0, CU-REWIRE4.0 significantly improved the C-to-U base editing efficiency at the target RNA site, indicating that the CU-REWIRE4.0 version has significant enhancements in both protein expression level and target base editing efficiency.
[0146] Example 3: Optimizing the amino acid composition of APOBEC3A to reduce off-target effects of the CU-REWIRE editing system
[0147] Natural APOBEC3A protein tends to form dimers to bind to single-stranded DNA or RNA, and the formation of these dimers increases off-target editing effects during base editing. In this patent, to reduce APOBEC3A dimer formation, we analyzed the crystal structure and molecular evolutionary characteristics of human APOBEC3A protein (amino acid sequence shown in SEQ ID NO. 8) and introduced single-point and multi-point mutations (H11, H16, C171, K30, H56, ...) into the APOBEC3A protein. Figure 3 a). Then, these mutated APOBEC3A versions were fused with the ePUF10 domain to construct different versions of the CU-REWIRE4 variant ( Figure 3 b-3c).
[0148] We designed CU-REWIRE4 to target and edit the C-terminus of EGFP mRNA. 459 Target sites were used to evaluate the editing efficacy of the improved CU-REWIRE4. Figure 3 c). mRNA-Seq results showed that point-mutated versions of CU-REWIRE4.X maintained high C-to-U editing efficiency while significantly reducing off-target effects. For example, CU-REWIRE4.1 with the C171A point mutation maintained high editing efficiency while significantly reducing off-target effects. Figure 3 c-3d). These data show that our optimization strategy of selectively introducing point mutations into the NTD and CTD backbone of APOBEC3A can significantly reduce the number of off-target edits in the CU-REWIRE system (reduced off-target numbers mean reduced tool side effects) and off-target edit efficiency (the lower the efficiency of off-target edits, the lower the side effects; the average efficiency of 11.4% was reduced to 7.9%, which is a very significant improvement). Figure 3 d). Further analysis of the data relationship between the editing efficiency and off-target effects of each CU-REWIRE 4.x version of the editing enzyme at the target site showed that these CU-REWIRE 4.x versions could effectively improve the editing efficiency of the editing enzyme. Figure 3 e) shows that our optimization results provide more diverse options for the application of base editing and greatly reduce side effects.
[0149] Example 4 uses artificial intelligence to split the APOBEC protein into three parts with independent functions.
[0150] Based on the experimental data of Example 3 ( Figure 3Our analysis of the APOBEC family of proteins revealed that APOBEC proteins consist of three relatively independent domains. While the deaminase domain is relatively conserved, the NTD and CTD domains are prone to evolutionary changes. This suggests that the deaminase domain is crucial for catalyzing C-to-U conversion, and that the NTD and CTD may specifically recognize different DNA / RNA targets. Furthermore, we identified a site (H70 / M153) for scientifically dissecting the APOBEC3A protein, creatively dividing it into three functionally independent domains: the NTD domain, the deaminase domain, and the CTD domain. These three domains each have relatively independent functions, but they cooperate to achieve C-to-U base editing. Figure 3 a).
[0151] In summary, we propose the following hypothesis: When the deaminase domain of the AID / APOBEC protein is fused with the NTD or CTD domains of APOBEC from different sources, it may create novel engineered cytosine base editing enzymes that can edit cytosine under different conditions.
[0152] To verify this hypothesis, we used AlphaFold2 to analyze the structures of APOBEC3 and APOBEC1 from human and rodent sources, and then dissected these proteins according to the dissection patterns we found (dissection points are shown in Figure 1). Figure 4 As shown in figures a-4e, the results reveal that the deaminase domains of the different APOBEC proteins separated from the original protein group exhibit strikingly similar structures. Figure 4 The results (a-4e) suggest that the functional modules of these proteins may be interchangeable, further demonstrating the feasibility of our designed splitting strategy and laying the foundation for our subsequent research.
[0153] Example 5: Artificial Intelligence-Assisted Development of Programmable Artificial Cytosine Deaminase (ProACD)
[0154] Based on the experimental data obtained in Examples 3 and 4, we dissociated natural cytidine deaminases from different sources to obtain thousands of NTD domains, deaminase domains, and CTD domain components. Figure 5 a).
[0155] Subsequently, this invention explored incorporating the deaminase domain of the dissected natural AID / APOBEC into our designed human APOBEC3A protein backbone, thus designing a novel engineered cytidine deaminase synthesis platform, which we call the Programmable Engineered Cytidine Deaminase Pre-ACD platform. The specific engineered cytidine deaminase produced by this platform is named ProACD. Figure 5 (b) can be used for base editing of DNA and RNA.
[0156] We analyzed the phylogenetic tree of the AID / APOBEC family using 1160 complete eukaryotic AID / APOBEC1-related genes recorded in the NCBI database, and found that APOBEC3A and APOBEC1 belong to closely related branches in this tree. Figure 5 c). Subsequently, with the assistance of artificial intelligence tools such as AlphaFold2, we successfully designed 1160 engineered ProACD proteins using these 1160 natural proteins as templates. Figure 5 d). Meanwhile, we used AlphaFold2 to predict... Figure 5 In section d, the crystal structure of the engineered ProACD protein synthesized using human and rodent deaminase domains as donors is shown. Figure 5 These ProACDs exhibit a structure remarkably similar to human APOBEC3A, suggesting they may possess base-editing capabilities similar to the human APOBEC3A protein. Figure 5 In summary, our method can design thousands of functional ProACD proteins (e-5f). Figure 5 ).
[0157] Example 6: Development of a new generation programmable RNA base editing system CU-REWIRE5 using ProACD
[0158] To design a more precise and efficient RNA base editing system, this invention integrates different types of ProACD and ePUF10 domain modules to develop a new generation of programmable RNA base editing system, CU-REWIRE5s. Figure 6 a). Specifically, we fused the ProACD and ePUF10 domain modules obtained in Example 5 to develop different versions of CU-REWIRE5.1-5.17 (amino acid sequences are shown in SEQ ID NO. 72-78) to further analyze the new functions of the CU-REWIRE5s editing system in detail. Figure 6 b).
[0159] Expression analysis in HEK293T cells showed that all CU-REWIRE5 variants could express the editing enzyme protein normally, and most CU-REWIRE5 protein expression was stable. Figure 6 c). We further analyzed whether these editing enzymes could edit RNA, and the results showed that most CU-REWIRE5 enzymes could edit the target RNA with high efficiency. Figure 6 d).
[0160] Example 7: The CU-REWIRE5 base editing system significantly enhances the targeting range and specificity of RNA base editing.
[0161] Next, to verify the efficacy and off-target effects of the new CU-REWIRE5 RNA editing technology, we designed CU-REWIRE5 to target and edit C-cells of EGFP mRNA carrying different mutation sites. 459 The target site is specifically determined by modifying a base upstream of the EGFP mRNA target site to a different base (where U... 458 C was changed to G 458 C, A 458 C and C 458 C), to test the sequence preference of CU-REWIRE5 for editing target RNA ( Figure 7 a).
[0162] This embodiment tested the editing specificity and activity of CU-REWIRE5 at target sites in different base environments. The results showed that multiple CU-REWIRE5 base editors containing ProACD proteins exhibited base editing sequence preferences and surprising editing activity that were distinctly different from those of the natural APOBEC protein, especially in target site environments where the natural APOBEC protein could not edit, such as in G... 458 C, A 458 C and C 458 It exhibits highly efficient C-to-U base editing activity under C environment ( Figure 7(b) For example, CU-REWIRE 5.1 and CU-REWIRE 5.8, whose ProACD protein fractions are donated by human APOBEC1 and APOBEC3H respectively, exhibit highly efficient C-to-U editing activity within the GC motif and lower C-to-U editing activity within the AC motif. Native APOBEC protein shows no editing activity in these GC and AC sequences. CU-REWIRE 5.15 exhibits highly efficient C-to-U editing activity within the CC and UC motifs. CU-REWIRE 5.16 exhibits highly efficient C-to-U editing activity within the AC and UC motifs. CU-REWIRE 5.3 and CU-REWIRE 5.17, these two editing enzymes, exhibit highly efficient C-to-U editing activity in all four motifs: GC, AC, CC, and UC. Native APOBEC protein only shows editing activity within the UC motif. In addition, some ProACDs also exhibit editing activity similar to that of natural APOBEC proteins, such as CU-REWIRE5.7 and CU-REWIRE5.13, which can only edit cytosine C in the UC motif. Figure 7 b).
[0163] Furthermore, we used transcriptome RNA-seq experiments to assess the overall off-target effects of CU-REWIRE5 at the whole transcriptome level. Figure 7 c). Our results indicate that CU-REWIRE5.1, 5.17, and 5.16, containing engineered ProACD proteins, exhibited higher off-target editing rates compared to CU-REWIRE4.1, which contains the natural APOBEC protein, and CU-REWIRE5.15, 5.7, 5.13, and 5.8 had fewer off-target editing sites. Figure 7 c).
[0164] There is a positive correlation between off-target editing rate and on-target editing efficiency (i.e., the higher the editing activity of CU-REWIRE5s, the more off-target editing events there are). This indicates that our newly designed ProACD protein has different editing sequence preferences and editing activities than the natural APOBEC protein. This further shows that the method used in this invention significantly expands the library of cytidine deaminase proteins. Our designed ProACD enzyme can serve as a supplement and replacement for natural cytidine deaminases, and has extremely high value for basic research and scientific translation.
[0165] Example 8: Analysis of Editing Sequence Preferences and Editing Accuracy of Editing GC Motifs in CU-REWIRE 5.1 and 5.8
[0166] The deamination domain donors of the ProACD protein fraction of the two editing enzymes, CU-REWIRE5.1 and CU-REWIRE5.8, are derived from human APOBEC1 and APOBEC3H, respectively. They exhibit highly efficient C-to-U editing activity within the GC motif, where the native APOBEC protein shows no editing activity; and lower C-to-U editing activity within the AC motif, where the native APOBEC protein shows no editing activity. Figure 8 (a-8e). Our results show that the editing efficiency of the ProACD enzyme is significantly different from that of the natural human APOBEC3A (edits only the UC motif) and human APOBEC1 (edits the UC motif with lower efficiency), indicating that a new cytidine editing activity has emerged in engineered ProACD. Figure 8 a).
[0167] We then evaluated the base editing properties of CU-REWIRE5.1 and CU-REWIRE 5.8 in detail within the EGFP sequence. The results showed that CU-REWIRE5.1 exhibited highly efficient C-to-U editing activity within the GC motif, lower C-to-U editing activity within the AC motif, and no editing activity within the UC and AC motifs. Figure 8 b). We further utilize Figure 7 RNA-seq data from C was used to analyze the base editing bias of off-target editing events at the transcriptome level. The analysis revealed that the editing motifs in off-target base editing events were associated with a single target site (C) in EGFP mRNA. 459 The sequence preferences observed when using C-to-U editing were largely consistent. Figure 8 c). Furthermore, the editing properties of CU-REWIRE 5.8 are basically the same as those of CU-REWIRE 5.8, only the efficiency is slightly lower. Figure 8 d-8e).
[0168] Subsequently, this embodiment further analyzed the editing precision of the CU-REWIRE5.1 base editing enzyme, particularly the selection of the editing window. Figure 9 a-9f). We found that CU-REWIRE5.1 (editing at the GC motif) showed the highest editing efficiency at 2- or 4-nucleotide (nt) positions downstream of the ePUF10 binding site in EGFP mRNA, as shown by different target reporter genes containing GC dinucleotide motifs at different positions downstream of the ePUF10 binding site. Figure 9 a).
[0169] Subsequently, we further evaluated the editing efficiency of CU-REWIRE5.1 in reporter genes carrying pathogenic mutations. For example, APOE4 allelic mutations at the Arg130Cys (c.T388C) and Arg173Cys (c.T526C) sites of the APOE4 protein are high-risk pathogenic mutations that can lead to Alzheimer's disease. We designed CU-REWIRE5.1 (CU5.1) enzymes with ePUF10 domains (amino acid sequences shown in SEQ ID NO. 87–90) to bind to the upstream sequences located at the target sites of APOE4 RNA c.T388C and c.T526C, respectively (binding to the target site C). 388 and C 526 At the upstream 2-nt or 3-nt position, it was found that CU-REWIRE5.1 bound 2-nt upstream of the mutation site can effectively convert the target base C to U, with an editing efficiency exceeding 90%. This editing can convert APOE4 into benign APOE3 ( Figure 9 b) and APOE1 ( Figure 9 c) Alleles. In another example, the c.68C>A (p.P23H) mutation in the human rhodopsin gene, known to cause retinitis pigmentosa (RP), can be effectively edited by CU-REWIRE5.1, which binds to 2-nt upstream of the mutation site. Figure 9 c) (Amino acid sequence as shown in SEQ ID NO. 91). We will use C 76 Edited as U 76 This is equivalent to introducing a stop codon into the mRNA carrying the mutation site, which can eliminate toxic protein mutants. Figure 9 c). Furthermore, CU-REWIRE5.1 can efficiently edit the T143C pathogenic mutation site of SOD1 mRNA associated with neurodegenerative diseases, especially when CU-REWIRE5.1 binds to the 2-nt upstream of the mutation site (amino acid sequence shown in SEQ ID NO. 92), it can efficiently convert pathogenic mRNA into normal functional mRNA, ultimately eliminating toxic protein mutants. Figure 9 e). In summary, based on the results of CU-REWIRE5.1 targeting and editing different mRNAs, we summarize the editing window of CU-REWIRE5.1 as 2–5 nt downstream of the ePUF10 binding site. Figure 9 f) provides more specific guidance for potential users of the base editing tool in this embodiment.
[0170] Example 9: Analysis of Edit Sequence Preferences and Edit Accuracy in CU-REWIRE 5.15 for Editing CC Motifs
[0171] Specifically, the deamination domain donors of the ProACD protein component of the CU-REWIRE5.15 editing enzyme are derived from mouse APOBEC3. It exhibits highly efficient C-to-U editing activity within the CC motif, while the native APOBEC protein shows no editing activity in this sequence. Figure 10 a-10c). Our results show that the editing efficacy of the ProACD enzyme is significantly different from that of the natural human APOBEC3A, indicating the emergence of novel cytidine editing activity in engineered ProACD (a-10c). Figure 10 a).
[0172] We then evaluated the base editing properties of CU-REWIRE5.15 in detail within the EGFP sequence. The results showed that CU-REWIRE5.15 exhibited highly efficient C-to-U editing activity within the CC and UC motifs, but no editing activity within the AC and GC motifs. Figure 8 b). We further utilize Figure 7 RNA-seq data from C was used to analyze the base editing bias of off-target editing events at the transcriptome level. The analysis revealed that the editing motifs in off-target base editing events were associated with a single target site (C) in EGFP mRNA. 459 The sequence preferences observed when using C-to-U editing were largely consistent. Figure 10 c).
[0173] Subsequently, this embodiment further analyzed the editing precision of the CU-REWIRE5.15 enzyme, particularly the selection of the editing window. Figure 11 a-11c). We found that CU-REWIRE5.15 (editing at the CC motif) had the highest editing efficiency 2-nt downstream of the ePUF10 binding site in EGFP mRNA. Figure 11 a) Subsequently, we used a similar approach to detect the editing window of CU-REWIRE5.15 in genes carrying pathogenic point mutations. We designed CU-REWIRE5.15 to target and edit the c.104T>C(p.L35P) mutation in the MEF2C gene, which is believed to cause severe autism spectrum disorder. We found that CU-REWIRE5.15 can effectively edit the target site C downstream of the ePUF10 binding site, thus editing C... 104 Convert to U 104 The editing efficiency is approximately 45%, and this editing can repair pathogenic point mutations in MEF2C. Figure 11 b).
[0174] In summary, based on the results of CU-REWIRE5.15 targeting and editing different mRNAs, we summarize the editing window of CU-REWIRE5.15 as 2–4 nt downstream of the ePUF10 binding site. Figure 11 c) provides more specific guidance for potential users of the base editing tools in this embodiment.
[0175] Example 10: Analysis of Edit Sequence Preferences and Edit Accuracy in CU-REWIRE 5.16 for Editing AC Motifs
[0176] Specifically, the deamination domain donors of the ProACD protein component of the CU-REWIRE5.16 editing enzyme are derived from mouse APOBEC3. It exhibits highly efficient C-to-U editing activity within the AC motif, while the native APOBEC protein shows no editing activity in this sequence. Figure 12 Our results show that the editing efficacy of the ProACD enzyme is significantly different from that of the natural human APOBEC3A, indicating the emergence of novel cytidine editing activity in engineered ProACD. Figure 12 a). We then evaluated the base editing properties of CU-REWIRE5.16 in detail within the EGFP sequence. The results showed that CU-REWIRE5.16 exhibited highly efficient C-to-U editing activity within the AC motif, but extremely low editing activity within the AC, CC, and GC motifs. Figure 12 b). We further utilize Figure 7 RNA-seq data from C was used to analyze the base editing bias of off-target editing events at the transcriptome level. The analysis revealed that the editing motifs in off-target base editing events were associated with a single target site (C) in EGFP mRNA. 459 The sequence preferences observed when using C-to-U editing were largely consistent. Figure 12 c).
[0177] Subsequently, this embodiment further analyzed the editing window of the CU-REWIRE5.16 enzyme ( Figure 12 We found that CU-REWIRE5.16 (editing at the AC motif) had the highest editing efficiency 3-nt downstream of the ePUF10 binding site in EGFP mRNA. Figure 12 d). In summary, based on the results of CU-REWIRE5.16 targeting and editing different mRNAs, we summarize the editing window of CU-REWIRE5.16 as 2–5 nt downstream of the ePUF10 binding site. Figure 12 e) provides more specific guidance for potential users of the base editing tools in this embodiment.
[0178] Example 11: Analysis of CU-REWIRE5s editing characteristics of arbitrary AC, CC, GC, and UC motifs.
[0179] The deamination domain donors of the ProACD protein fraction of the two editing enzymes, CU-REWIRE5.3 and CU-REWIRE5.17, are derived from mouse APOBEC1 and rat APOBEC1, respectively. They exhibit highly efficient C-to-U editing activity in four motifs: AC, CC, GC, and UC. Native APOBEC proteins lack RNA editing activity in these sequences. Figure 13 Our results show that the editing efficacy of the ProACD enzyme is significantly different from that of natural mouse APOBEC1 and rat APOBEC1, indicating the emergence of novel cytidine editing activity in engineered ProACD (a-13e). Figure 13 a).
[0180] We then evaluated the base editing properties of CU-REWIRE5.3 and CU-REWIRE 5.17 in the EGFP sequence in detail. The results showed that CU-REWIRE5.17 exhibited efficient C-to-U editing activity in the AC, CC, GC, and UC motifs, but lower editing activity in the AC and CC motifs. Figure 13 b). We further utilize Figure 7 RNA-seq data from C was used to analyze the base editing bias of off-target editing events at the transcriptome level. The analysis revealed that the editing motifs in off-target base editing events were associated with a single target site (C) in EGFP mRNA. 459 The sequence preferences observed when using C-to-U editing were largely consistent. Figure 13 c). Furthermore, the editing properties of CU-REWIRE 5.3 are basically the same as those of CU-REWIRE 5.17, only the efficiency is slightly lower. Figure 13 d-13e).
[0181] This embodiment further analyzes the editing window size of the CU-REWIRE5.17 base editing enzyme. Figure 14 a-14c). We found that CU-REWIRE5.17 exhibited high editing efficiency at the 2- or 4-nt position downstream of the ePUF10 binding site in EGFP mRNA. Figure 14 a). Moreover, CU-REWIRE5.17 can efficiently edit the T1541C site of DDX3X mRNA associated with neurodevelopmental disorders, especially when CU-REWIRE5.17 binds to the 2-nt or 3-nt upstream of the mutation site (amino acid sequence as shown in SEQ ID NO. 93), ultimately eliminating the toxic protein mutant. Figure 14b). In summary, based on the results of CU-REWIRE5.17 targeting and editing different mRNAs, we summarize the editing window of CU-REWIRE5.17 as 2–5 nt downstream of the ePUF10 binding site. Figure 14 c) provides more specific guidance for potential users of the base editing tools in this embodiment.
[0182] Example 12: Construction of a new generation base editing system CU-REWIRE5s using engineered ProACD enzymes derived from mammals.
[0183] This embodiment also utilizes the patterns discovered in Example 5 to explore the possibility of generating ProACD using deaminase domains from other mammals. We used AlphaFold2 to analyze APOBEC1s and APOBEC3s ( Figure 15 a-15c). We found that when the Z1 and Z2 domains of mammalian APOBEC3s bind to the NTD and CTD of human APOBEC3, they can form a functional ProACD (a-15c). Figure 15 a-15c). These engineered ProACDs were fused with the ePUF10 domain to construct different mammalian-derived versions of CU-REWIRE5s (amino acid sequences are shown in SEQ ID NO. 79–86). Figure 16 a), and then the editing activity of these CU-REWIRE5s containing engineered ProACD was tested on EGFP mRNA. Figure 16 (b) To further analyze the novel functionalities of the CU-REWIRE5s editing system in detail. Data shows that the newly synthesized CU-REWIRE5s exhibit highly efficient C-to-U editing activity in GC, AC, and CC motifs. Figure 16 c-16g).
[0184] These results highlight the diversity of engineered ProACD in C-to-U RNA base editing, and further demonstrate that the method used in this invention significantly expands the library of cytidine deaminase proteins. Our designed ProACD enzyme can serve as a supplement and replacement for natural cytidine deaminases, and has extremely high value for basic research and scientific translation.
[0185] Example 13: Evaluation of the in vivo RNA base editing efficacy of the CU-REWIRE5 system.
[0186] Efficient RNA base editing within the central nervous system has always been a challenging task. To explore the editing efficacy of the CU-REWIRE system in the central nervous system, we used CU-REWIRE5, which incorporates the ProACD enzyme, for in vivo RNA base editing. Using CU-REWIRE5.15 (CU5.15) as an example, we repaired an autism spectrum disorder (ASD) mouse model with the Mef2cL35P (c.104T>C) point mutation, which leads to severe ASD phenotypes.
[0187] We will use CU5.15 (which targets and binds to the A receptor of mouse Mef2c mRNA) 92 -A 101 The sequence was constructed into the AAV-PHP.eB vector, which can cross the blood-brain barrier (BBB). (This AAV vector is a commercially available vector. The amino acid sequence after inserting the promoter, CU-REWIRE5, and polyA of this invention is shown in SEQ ID NO. 94, and the nucleotide sequence is shown in SEQ ID NO. 95.) Figure 17 a) CU5.15 is driven by the neuron-specific promoter hSyn. The C-terminus of CU5.15 is linked to an EGFP reporter gene via P2A to track the expression profile of the CU5.15 protein. After packaging, AAV-CU5.15 was injected intravenously into 4-week-old Mef2c WT and L35P mice. + / - In mice ( Figure 17 a). Two months after AAV injection, the expression of AAV-CU5.15 in the mouse brain was observed. The results showed that AAV-CU5.15 (green fluorescent signal) was efficiently expressed in the target brain regions of mice, namely the cortex and hippocampus, indicating that CU5.15 can be applied to the treatment of related diseases of the nervous system. Figure 17 b).
[0188] Subsequently, we investigated the editing characteristics of the CU-REWIRE5 system in the mouse brain using Sanger sequencing and RNA sequencing of prefrontal cortex and hippocampal tissue samples. The results showed that in Mef2c L35P... + / - In the prefrontal cortex tissue of heterozygous mice, the control group carried normal T cells. 104 Approximately 50% of the Mef2C mRNA at this site carries the pathogenic C. 104 Mef2C mRNA at this site accounts for approximately 50%; in hippocampal tissue, the control group carries normal T... 104 Approximately 50% of the Mef2C mRNA at this site carries the pathogenic C. 104 Mef2C mRNA at this site accounts for approximately 50% ( Figure 17c-16d). In the experimental group injected with AAV-CU5.15, the prefrontal cortex carried normal T cells. 104 The proportion of Mef2C mRNA at this site increased significantly (averaging 72%, indicating a higher percentage of pathogenic C mRNA). 104 Mef2C mRNA levels at the site decreased to 28%; normal T cells were present in the hippocampus. 104 The proportion of Mef2C mRNA at this site also increased significantly (averaging 67%, meaning it carries pathogenic C). 104 Mef2C mRNA at this site decreased to 33%. Figure 17 (c-17d). Further detailed statistical analysis of AAV-CU5.15 in Mef2c L35P mice C 104 The specific editing efficiency at different sites showed that the average editing efficiency of AAV-CU5.15 in different brain regions ranged from 25% to 40%. Figure 17 e). Furthermore, at the Mef2cmRNA target site C 104 No editing of non-target sites was observed in nearby cytidine C, demonstrating the precision of CU-REWIRE5.15-mediated RNA editing. Figure 17 c). In summary, AAV-CU5.15 can efficiently edit target Mef2cmRNA in the target brain region of mice with extremely high editing precision, exhibiting single-base editing, meaning that it repairs point mutations without introducing new editing mutations. This data is significantly better than existing CRISPR-mediated base editing enzymes. Figure 17 c-17e).
[0189] Most notably, compared with the ePUF10-injected control group, the MEF2C protein expression level in the prefrontal cortex and hippocampus of the CU5.15-injected Mef2c L35P mice was significantly restored. Figure 17 f), especially in the hippocampus, the expression of MEF2C protein is almost completely restored. Figure 17 g), indicating that the CU5.15-edited Mef2cmRNA can be translated normally to produce a functional protein.
[0190] We further verified the recovery of MEF2C protein expression levels using immunostaining with anti-MEF2C antibody. Compared with the control group, MEF2C protein expression in the brains of mice in the CU5.15 administration group was significantly restored, with protein levels approaching those of Mef2c WT mice. Figure 18 (a-18b). Furthermore, the lifespan of Mef2c L35P mice in the CU5.15 injection group was significantly improved, returning to normal levels. Figure 18c) This fully demonstrates that CU-REWIRE can repair mutated RNA in the central nervous system, and the edited RNA can be translated normally to produce functional proteins, thereby improving the pathological phenotype of mice.
[0191] This invention also investigated the efficacy of CU5.15 in treating the autistic pathological phenotype in Mef2c L35P mice. First, we assessed the mice's social abilities using a three-box test. Figure 19 a) The results showed that mice in the CU5.15 injection group possessed basic social abilities. Figure 19 b), and at the same time, the mice's social novelty ability was significantly improved. Figure 19 c). Furthermore, behavioral experiments on social intruders showed that in the initial four trials with intruder partners, the abnormal social behavior of Mef2c L35P mice was completely rescued after treatment with CU5.15. Figure 19 These results indicate that CU5.15-mediated in vivo RNA base editing can fully restore aberrant MEF2C protein expression and improve social deficits associated with neurodevelopmental disorders related to the Mef2c-L35P mutation.
[0192] In summary, our research demonstrates that RNA base editing via CU-REWIRE5 provides a powerful tool for correcting genetic mutations in the brain and has significant clinical value for the treatment of ASD and related neurological disorders.
[0193] The above embodiments are for illustrating the implementation schemes disclosed in this invention and should not be construed as limiting the invention. Furthermore, various modifications and variations of the methods listed herein will be apparent to those skilled in the art without departing from the scope and spirit of the invention. Although the invention has been specifically described in conjunction with various specific preferred embodiments, it should be understood that the invention should not be limited to these specific embodiments. In fact, various modifications as described above that are obvious to those skilled in the art to obtain the invention should be included within the scope of this invention.
[0194] In this invention, the sequence numbers and sequence names of amino acids or nucleotides correspond as follows:
[0195]
Claims
1. An engineered cytosine deaminase, characterized in that, The amino acid sequence of the engineered cytosine deaminase is shown in SEQ ID NO. 9, 11, 25, 26, 29, 30, 32, 33, 35, 47, 53 or 55.
2. Use of the engineered cytosine deaminase according to claim 1 in the preparation of a base editing system.
3. The use according to claim 2, characterized in that, The base editing system is selected from RNA base editing systems.
4. An RNA base editing system, characterized in that, The RNA base editing system is the CU-REWIRE fusion protein or its encoding gene, and the amino acid sequence of the CU-REWIRE fusion protein is shown in SEQ ID NO. 72, 73, 78, 79, 80, 81, 86, 87, 88, 90, 91 or 92.
5. An isolated polynucleotide, characterized in that, The polynucleotide encodes any one or a fragment thereof: the engineered cytosine deaminase of claim 1 or the RNA base editing system of claim 4.
6. A nucleic acid construct, characterized in that, The nucleic acid construct comprises the polynucleotide as described in claim 5.
7. A viral vector system, characterized in that, The viral vector system includes the nucleic acid construct of claim 6.
8. The viral vector system as described in claim 7, characterized in that, The viral vector system is an adeno-associated virus vector system.
9. An adeno-associated virus, characterized in that, The adeno-associated virus is prepared by viral packaging of the viral vector system described in claim 7 or 8.
10. A cell, characterized in that, The cell contains the nucleic acid construct of claim 6 or the genome in which the exogenous polynucleotide of claim 5 is integrated.
11. A base editing method for purposes other than disease diagnosis and treatment, characterized in that, The target gene is contacted with the adeno-associated virus of claim 9 or the RNA base editing system of claim 4 to achieve single-base editing of the target gene.
12. The base editing method according to claim 11, characterized in that, The single-base editing is C-to-U base editing.
13. The base editing method according to claim 11, characterized in that, The target gene is a gene that binds to ePUF10, and the amino acid sequence of the ePUF10 protein is shown in SEQ ID NO.57.