Engineered cytosine deaminase, preparation method therefor and use thereof
By designing and engineering the fusion of cytosine deaminase ProACD and ePUF10 protein to construct the CU-REWIRE system, the problems of low efficiency and high off-target rate of existing RNA base editing tools are solved, and efficient and safe RNA base editing is achieved, which is suitable for gene therapy and basic research.
Patent Information
- Application Number
- PCT/CN2025/073468
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2025-01-20
- Publication Date
- 2025-10-16
AI Technical Summary
Existing RNA base editing tools have problems such as complex editing systems, low efficiency, significant interference with non-target RNA, and possible immune response. In particular, the APOBEC3A protein is extremely inefficient when editing specific sequences, limiting the application scope of C-to-U base editors.
An engineered cytosine deaminase ProACD was designed. By splitting the cytosine deaminase protein into N-terminal and C-terminal domains and fusing them with the deamination domain, the CU-REWIRE system was constructed in combination with the ePUF10 protein to achieve efficient and precise C-to-U base editing independent of gRNA.
It achieves stable expression in cells, reduces off-target rates, meets clinical application requirements, and provides a safe and efficient RNA base editing tool suitable for the treatment of genetic diseases such as Duchenne muscular dystrophy and progeria.
Smart Images

Figure PCTCN2025073468-APPB-I100001 
Figure 00000039_0000 
Figure 00000039_0001
Abstract
Description
An engineered cytosine deaminase and preparation method and use thereof TECHNICAL FIELD
[0001] The present application relates to the field of gene editing, in particular to an engineered cytosine deaminase and preparation method and use thereof. BACKGROUND
[0002] I. Scientific research on gene editing tools and value of gene therapy
[0003] Genes are the most important genetic material in organisms. How to correctly understand, predict and ultimately control the expression of genes so that cells can normally function is a major opportunity and challenge for the further development of biology in the present and future period.
[0004] Base editing, as a ubiquitous gene modification phenomenon in organisms, changes DNA or RNA sequence information to rewrite codons, create new RNA splicing sites, etc., to process a DNA or RNA into a mature mRNA that can translate into different functional proteins, diversify genetic information, and thus provide a better complexity evolution basis for organisms, playing an indispensable role in organisms. In addition, a large part of currently known human diseases are caused by gene mutations, of which about 58% of human genetic mutation diseases are caused by single base mutations in genes.
[0005] Therefore, it is of great research value to design and develop efficient and precise base editing enzymes to modify the sequence of a single base in DNA or RNA to manipulate the expression of target genes. At the same time, it is very meaningful to use base editing enzymes to precisely repair pathogenic single base mutations in target genes for the research and treatment of such genetic mutation diseases.
[0006] Current gene therapy strategies for single base mutation diseases mainly repair or replace mutant genes by directly rewriting the sequence information of target genes at the DNA or RNA level through base editing to treat diseases. Among them, DNA base editing, as a traditional gene manipulation strategy, can directly edit bases on DNA to rewrite the target DNA sequence and correct gene expression. Since changes in genomic DNA will accompany the life of a cell, long-term safety concerns about genomic DNA editing have always been a major problem.
[0007] II. Advantages and great application value of RNA base editing
[0008] Compared with DNA base editing therapy, RNA base editing directly edits the target base on RNA to rewrite the RNA sequence without changing the DNA sequence, which is a non-permanent and reversible gene expression regulation strategy, does not produce heritable editing products, is relatively easy to manipulate, and will not produce heritable editing products when used as a potential treatment in vivo, which is safer than DNA editing. Therefore, manipulating genes at the RNA level has better controllability and safety, making this type of gene therapy more conducive to the translation of basic research to clinical practice.
[0009] As a ubiquitous gene modification phenomenon in organisms, RNA base editing can directly regulate the alternative splicing, translation and degradation of RNA by changing the RNA sequence to rewrite the codon and create new RNA splicing sites, manipulate the expression of target genes, and ultimately achieve the diversification of genetic information, thereby enabling organisms to obtain a better complexity evolution foundation and play an indispensable role in organisms. RNA base editing in organisms can be divided into two categories according to different molecular targeting mechanisms: guide RNA (gRNA)-mediated RNA base editing and RNA binding protein (RBPs)-mediated base editing, and their target editing products can exert their respective physiological functions in different environments.
[0010] In eukaryotes, the existing base editing proteins mainly include ADARs and APOBECs. RNA adenosine deaminase ADAR (Adenosine deaminases acting on RNA) mainly acts on double-stranded RNA, and cytidine deaminase APOBEC (Apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like) mainly acts on single-stranded RNA. Among them, the N-terminal of ADAR protein contains an RNA binding protein domain, which is responsible for recognizing and positioning the target RNA, and the C-terminal contains an adenosine deamination catalytic domain, which is responsible for catalyzing the hydrolysis of target adenine and deaminating it to inosine. Inosine is recognized as guanosine during translation, realizing the base editing of adenosine to guanosine (A-to-G). APOBEC protein catalyzes the hydrolysis of RNA cytosine and deaminates it to uridine, realizing the base editing of cytidine to uridine (C-to-U).
[0011] Abnormal base editing process of RNA will affect the normal growth and development of organisms, and in higher organisms, it will lead to the occurrence of various diseases. Statistical data shows that about 58% of human genetic genetic mutation diseases are caused by single base mutation. Among them, 21593 guanosine to adenosine (G-to-A) and 6945 thymine to cytosine (T-to-C) single base pathogenic mutations have been found (as of March 2022). This type of disease can theoretically be repaired and corrected by DNA or RNA base editing.
[0012] Given that RNA base editing is directly or indirectly involved in most life activities, and its editing process is closely related to diseases. Using ADAR or APOBEC as an RNA base editing enzyme donor, and using RBP or gRNA to recruit ADAR or APOBEC to the target RNA editing site, precise base editing can be achieved at the target site. RNA base editing gene therapy is a strategy for directly repairing or replacing mutant genes at the RNA level, thereby correcting the abnormal processing and metabolism of RNA and essentially treating diseases. This strategy does not change the genome sequence and does not produce heritable byproducts, and can be used as a potential gene therapy method in vivo.
[0013] III. Research progress and design defects of existing RNA base editing tools
[0014] In recent years, some RNA base editing tools for adenosine to guanosine (A-to-I) have been developed, and the main editing strategy is to use gRNA to target the target RNA and recruit adenosine deaminase ADAR2 or Cas13-ADAR2 effector protein to form a target RNA / gRNA / ADAR ternary complex at the editing site, achieve A-to-I base editing, and have a good promotion foundation. However, there are still some deficiencies: (1) The editing system is complex. The gRNA, ADAR2 effector protein and target RNA in the system are co-folded and assembled into a ternary complex in the cell, which is the basis for base editing. (2) The editing efficiency is low. The assembly efficiency and proportion of the gRNA / ADAR effector protein / target RNA ternary complex is the rate-limiting step of efficient base editing. The whole process needs to go through the co-localization and folding assembly of gRNA and ADAR effector protein and gRNA and target RNA. (3) Disturb the expression of non-target RNA. Overexpression of gRNA components will contact and pair with non-target RNA, change the structure of non-target RNA, interfere with the normal expression of these RNAs, and thus disturb the normal expression of the transcriptome RNA. (4) Cas protein as the core protein of CRISRP system, derived from bacterial protein, will cause immune response and toxicity of drug administration object in clinical medical process.
[0015] In recent years, the development of cytidine to uridine (C-to-U) base editing tools has remained at a relatively early stage, and the main editing strategy is to use natural APOBEC cytidine deaminase. APOBEC protein is an effector domain of RNA binding protein (RNA binding proteins, RBPs). Natural APOBEC protein, as a large class of cytidine deaminases that act on single-stranded RNA in the human body, can achieve C-to-U base editing. Most APOBEC proteins have only 200 amino acids, which is much smaller than ADAR, and have more obvious advantages in biological delivery.
[0016] There are two types of C-to-U base editing tools: the first type is similar to A-to-I base editing strategy, which uses gRNA to target the target RNA and uses gRNA to recruit cytosine deaminase APOBEC3A-Cas13 effector protein to form a target RNA / gRNA / APOBEC3A-Cas13 ternary complex at the RNA editing site, and then perform C-to-U base editing. However, this system has a difficult to overcome drawback, APOBEC3A protein is a single-stranded RNA cytosine deaminase, gRNA will target and bind to target RNA and pair to form gRNA / target RNA double-stranded at the target RNA target site. The result of this targeted binding is to eliminate the single-stranded RNA environment required for APOBEC to function, thereby causing a decrease in editing efficiency and editing accuracy, making it difficult to apply in actual cases. The second type: In order to overcome the difficulty of CRISPR system in using APOBEC protein for C-to-U base editing, a gRNA-independent base editing system (CU-REWIRE) was developed. This system uses RNA-binding protein (PUF) to recognize and bind to target RNA and target APOBEC3A protein to the target RNA for C-to-U base editing. This process does not require the participation of gRNA, thereby solving the problem encountered by the first type of tool and achieving efficient and accurate base editing of the target RNA.
[0017] IV. The necessity and great application value of artificially designed engineered cytosine base editing enzymes
[0018] However, the activity of existing natural APOBEC proteins in RNA base editing systems is relatively low, and currently only human APOBEC3A can efficiently perform RNA level cytosine to uridine (C-to-U) site-specific base editing. However, the existing APOBEC3A protein has strong sequence preference when editing target bases, and it can only edit the base "C" in the "UC" sequence. APOBEC3A cannot edit the base "C" in the "AC" sequence, the base "C" in the "CC" sequence, and the base "C" in the "GC" sequence, and in these sequences, the editing efficiency is extremely low. This editing property of APOBEC3A limits the range of use of C-to-U base editors and becomes a rate-limiting step for the further application of C-to-U base editing tools.
[0019] In summary, the APOBEC3A protein with high activity can only edit the base "C" in the "UC" sequence; other natural APOBEC proteins have low editing activity and cannot meet the performance requirements for use, making it difficult to apply in C-to-U site-specific base editing tools.
[0020] Therefore, it is necessary to design a non-natural cytosine deaminase with high editing activity and capable of editing "C" in "UC", "AC", "CC" and "GC" sequences respectively.
[0021] It is necessary to design a new generation of DNA or RNA base editing tool using the artificially designed non-natural cytosine deaminase to improve the editing efficiency and accuracy, so as to realize C-to-U base editing more accurately and efficiently, as an alternative to the existing tool, and further applied to basic research and gene therapy. SUMMARY
[0022] In view of the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide an engineered cytosine deaminase and a preparation method and use thereof, to solve the problems in the prior art.
[0023] To achieve the above-mentioned objects and other related objects, the present application provides an engineered cytosine deaminase, which comprises a skeleton and a deamination domain arranged in the skeleton, the skeleton of the cytosine deaminase protein is composed of an N-terminal domain and a C-terminal domain of a cytosine deaminase protein, the N-terminal domain and the C-terminal domain are derived from a cytosine deaminase protein, and the deamination domain is derived from another cytosine deaminase protein.
[0024] The present application also provides a preparation method of the engineered cytosine deaminase, which comprises connecting a nucleotide encoding a skeleton of a cytosine deaminase protein with a nucleotide encoding a deamination domain of another cytosine deaminase protein to obtain a nucleotide encoding a fragment of the engineered cytosine deaminase, and then cloning the nucleotide encoding the fragment of the engineered cytosine deaminase into an expression vector and expressing to obtain the engineered cytosine deaminase.
[0025] The present application also provides an ePUF10 protein, which is obtained by integrating leucine and proline into the fourth recognition unit of the PUF10 domain.
[0026] The present application also provides the use of the engineered cytosine deaminase or the ePUF10 protein in preparing a base editing system.
[0027] The present application also provides a DNA base editing system, which comprises the cytosine deaminase.
[0028] The present application also provides an RNA base editing system, which comprises the engineered cytosine deaminase.
[0029] The present application also provides an isolated polynucleotide encoding any one of the following or a fragment thereof: the engineered cytosine deaminase, the ePUF10 protein, or the base editing system.
[0030] The present application also provides a nucleic acid construct comprising the isolated polynucleotide.
[0031] The present application also provides a viral vector system comprising the nucleic acid construct.
[0032] The present application also provides an adeno-associated virus (AAV) packaged by the viral vector system.
[0033] The present application also provides a cell containing the nucleic acid construct or the polynucleotide integrated into the genome of the cell.
[0034] The present application also provides a base editing method, which comprises contacting a target gene with the adeno-associated virus or the base editing system to edit a single base on the target gene.
[0035] As described above, the engineered cytosine deaminase, and the preparation method and use thereof, have the following beneficial effects: it is a gRNA-independent, simple and efficient, low off-target, and can be stably expressed in cells, meeting the efficient delivery requirements of later clinical applications. It can be used as a potential and relatively safe in vivo editing tool, and is expected to play an important role in some fields of gene therapy and basic research, such as treating Duchenne muscular dystrophy, premature aging and other genetic diseases. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 The optimization of ePUF10 protein based on crystal structure enhances the protein expression intensity of CU-REWIRE. In the left panel of a, the crystal structure of PUF10 and ePUF10 predicted based on AlphaFold2 tool is shown, and in the right panel, the schematic diagram of the recognition and binding of ePUF10 protein and substrate RNA is shown. In b, the schematic diagram of CU-REWIRE editing system designed using different versions of PUF10 is shown. The original version of CU-REWIRE 3.0 (upper panel) fused with PUF10 protein, and the upgraded version of CU-REWIRE 4.0 (lower panel) fused with ePUF10 protein, containing 10 RNA recognition units, targeting 10-nt RNA sequence. In c, the protein expression level analysis of different versions of CU-REWIRE editing system is shown.
[0037] Figure 2 The optimization of ePUF10 protein based on crystal structure enhances the editing efficiency of CU-REWIRE. In a, the CU-REWIRE 3.0 and CU-REWIRE 4.0 are used to edit the C459 PUF10 binding site in green fluorescent protein is underlined in blue, and the adjacent cytosine is marked in red (top panel). The heatmap shows the editing efficiency of each CU-REWIRE and control at the target site C 459 and all cytosines nearby. The editing efficiency of C 459 is shown on the right, determined by RNA-seq. b, Analysis of off-target effects of different versions of CU-REWIRE editing enzymes. Transcriptional C-to-U base editing events were determined by RNA-seq experiments. Orange diamonds represent the target site EGFP-C 459 , and the average efficiency of total off-target editing events is shown on the right. "n" represents the total number of cytosine off-target editing events detected. c, Analysis of editing activity of CU-REWIRE4 at the target site C 459 in different EGFP transcripts. d, Analysis of editing window of CU-REWIRE4 system. The efficiency of cytosines at different positions downstream of the binding site of CU-REWIRE4.0 was analyzed, and n-bp represents the number of bases between the editing site and the 3' end of the binding site of PUF10.
[0038] Figure 3 Optimization of the amino acid composition of APOBEC3A to reduce off-target effects of CU-REWIRE editing system. a, Schematic diagram of functional domains of different domains of human APOBEC3A. b, Schematic diagram of different versions of CU-REWIRE4.x recognizing and binding to the target EGFP mRNA. "X" represents different versions of APOBEC3A carrying different point mutations. c, Analysis of base editing efficiency of CU-REWIRE4.X on EGFP mRNA. The ePUF10 binding site in EGFP is underlined in blue, and the adjacent cytosine is marked in red. The heatmap shows the editing efficiency of each CU-REWIRE and ePUF10 control at the target site C 459 nearby. The editing efficiency of C 459 is shown on the right. Editing rates were determined by RNA-seq with three biological replicates. d, Analysis of transcriptional RNA base editing efficiency and off-target effects in samples treated with different CU-REWIRE4. Orange diamonds represent the target site EGFP-C 459 , and the average editing efficiency of off-target editing events is annotated next to the scatter plot. "n" represents the total number of cytosine base edits detected. e, Correlation analysis of on-target site editing efficiency (y-axis) and whole-transcriptome C-to-U off-target editing events (x-axis) of different CU-REWIRE4.X.
[0039] Figure 4. Splitting of different domains of APOBEC proteins. a. The crystal structure of human APOBEC3A (PDB:4XXO) and the splitting of different domains. Through sequence analysis, APOBEC3A protein was split into N-terminal domain (NTD), C-to-U deaminase domain and C-terminal domain (CTD) with independent functions. b-e. The crystal structure of native APOBEC proteins predicted by AlphaFold2 and the splitting of different domains. Mouse APOBEC3 (b), human APOBEC1 (c), mouse APOBEC1 (d) and rat APOBEC1 (e).
[0040] Figure 5. Development of programmable engineered cytosine base editor ProACD with the aid of artificial intelligence. a. The splitting of different domains of native AID / APOBEC proteins (the splitting strategy is derived from Figure 4). b. The assembly of programmable engineered cytosine base editor ProACD with the aid of artificial intelligence. c. Phylogenetic tree analysis of AID / APOBEC family proteins, with different branches marked in different colors. d. Artificial synthesis and production of engineered ProACD based on AI prediction. e-f. The crystal structure of native AID / APOBEC proteins and engineered ProACD proteins predicted by AlphaFold2. Top: the crystal structure of native AID / APOBECs predicted by artificial intelligence. Bottom: the crystal structure of ProACD predicted by artificial intelligence.
[0041] Figure 6. Development of new generation programmable RNA base editing system CU-REWIRE5 with ProACD. a. The design of programmable CU-REWIRE5 system with ProACD protein. Top, the schematic diagram of upgraded CU-REWIRE5 structure for C-to-U base editing. Middle, the programmable assembly of each module in CU-REWIRE5 system. The deamination domain of ProACD is provided by different sources of native APOBEC deaminase for specific C-to-U base editing. RNA recognition domain ePUF10, including 10 RNA recognition units, each of which is designed for specific RNA base recognition. Bottom, the precise binding of CU-REWIRE5 to the target EGFP mRNA, with the ePUF10 binding sequence of AACGUCUAUA and the editing target site of C 459 . b. The donor source table of deamination domain of ProACD in different versions of CU-REWIRE5. c. Protein expression level analysis of different versions of CU-REWIRE5 in HEK293T cells. d. Editing efficiency analysis of different versions of CU-REWIRE5 on the target site C 459 of EGFP mRNA.
[0042] Figure 7 Targeting scope and specificity analysis of CU-REWIRE5 base editors. a The diagram shows the targeting of CU-REWIRE5s to the C 459 Targeting recognition of target sites. Original U 458 C sequence was manually changed to A 458 C, G 458 C and C 458 C, to test the base editing sequence preference of CU-REWIRE5. b The statistical analysis of the editing efficiency of CU-REWIRE5s at the C 459 Editing efficiency in different base sequence contexts (A 458 C, C 458 C, U 458 C, G 458 C). e The binding sequence of PUF10 and the editing target sites are the same as in Figure 7a. c The global off-target effects of CU-REWIRE5s at the transcriptome level were assessed using RNA-seq experiments. (The samples were derived from Figure 7b). Orange diamonds represent the EGFPC 459 sites, and the right side shows the average editing efficiency of off-target editing events.
[0043] Figure 8 Editing sequence preference and editing accuracy analysis of CU-REWIRE5.1 and 5.8 editing GC motif. a The editing efficiency and accuracy analysis of different versions of CU-REWIRE5s at the C 459 on the EGFP mRNA target site. b The editing preference analysis of CU-REWIRE5.1 editing enzyme at on-target sites. c The editing preference analysis of CU-REWIRE5.1 editing enzyme in all off-target editing events in the transcriptome (data derived from Figure 7c). d The editing preference analysis of CU-REWIRE5.8 editing enzyme at on-target sites. e The editing preference analysis of CU-REWIRE5.8 editing enzyme in all off-target editing events in the transcriptome (data derived from Figure 7c).
[0044] Figure 9 Editing window analysis of CU-REWIRE5.1 (CU5.1) and gene therapy applications. a The editing window size analysis of CU-REWIRE5.1 editing the EGFPC 459 target site, and the figure shows the editing efficiency of CU-REWIRE5.1 at different distances from the PUF binding site. b The CU5.1 editing of the pathogenic target site C 459 on the APOE4 mRNA. c The CU5.1 editing of the pathogenic target site C 388, editing efficiency was determined by Sanger sequencing. The ePUF10 binding sites are marked in blue. c. CU5.1 editing pathogenic target site C 526 , editing efficiency was determined by Sanger sequencing. The ePUF10 binding sites are marked in blue. d. CU5.1 editing pathogenic target site C 76 of Rhodopsin mRNA. 143 of SOD1 mRNA. f. Summary of editing window size of CU-REWIRE5.1 editing target sites of EGFP, APOE4 Rhodopsin and SOD1 mRNA. The x-axis represents the distance from the editing site to the PUF binding site, and the y-axis represents the editing efficiency of C-to-U, which was determined by Sanger sequencing.
[0045] Figure 10 Editing sequence preference and editing accuracy analysis of CU-REWIRE5.15 editing CC motif. a. Editing efficiency and accuracy analysis of different versions of CU-REWIRE5s on the target site C 459 of EGFP mRNA. b. Editing preference analysis of CU-REWIRE5.15 editing enzyme at C 459 site. c. Editing preference analysis of CU-REWIRE5.15 editing enzyme in all off-target editing events in the transcriptome (data derived from Figure 7c).
[0046] Figure 11 Editing window analysis and gene therapy application of CU-REWIRE5.15 (CU5.15). a. Editing window size analysis of CU-REWIRE5.15 editing EGFP target site. b. Editing efficiency and accuracy analysis of different versions of CU-REWIRE5s on the target site C 459 of Mef2c mRNA. c. Editing efficiency and accuracy analysis of CU-REWIRE5.15 editing enzyme at C 104 site. e. The ePUF10 binding sites are marked in blue. c. Summary of editing window size of CU-REWIRE5.15 editing target sites of EGFP and Mef2c mRNA.
[0047] Figure 12 Editing sequence preference and editing accuracy analysis of CU-REWIRE5.16 editing AC motif. a. Editing efficiency and accuracy analysis of different versions of CU-REWIRE5s on the target site C 460 of EGFP mRNA. b. Editing preference analysis of CU-REWIRE5.16 editing enzyme at C 460Figure 13. CU-REWIRE5s editing enzyme properties analysis of cytosine C in any AC, CC, GC, UC motif. a. Editing efficiency and accuracy analysis of different versions of CU-REWIRE5s on EGFPC 460 target site. b. Editing window size analysis of CU-REWIRE5.17 editing EGFPC
[0048] Figure 13. CU-REWIRE5s editing enzyme properties analysis of cytosine C in any AC, CC, GC, UC motif. a. Editing efficiency and accuracy analysis of different versions of CU-REWIRE5s on EGFPC 459 target site. b. Editing window size analysis of CU-REWIRE5.17 editing EGFPC 459 target site. c. Editing preference analysis of CU-REWIRE5.17 editing enzyme on all off-target editing events in transcriptome (data derived from Figure 7c). d. Editing preference analysis of CU-REWIRE5.3 editing enzyme on target site C 459 target site. c. Editing preference analysis of CU-REWIRE5.17 editing enzyme on all off-target editing events in transcriptome (data derived from Figure 7c). d. Editing preference analysis of CU-REWIRE5.3 editing enzyme on target site C
[0049] Figure 14. Editing window analysis of CU-REWIRE5.17 (CU5.17) and gene therapy application. a. CU-REWIRE5.17 editing EGFPC 459 target site. b. Editing window size analysis of CU5.17 editing pathogenic target site C 1541 of DDX3X mRNA. Editing efficiency was determined by Sanger sequencing. e. PUF10 binding site is marked in blue. c. Editing window size summary of CU-REWIRE5.17 editing target site of different transcripts of EGFPC
[0050] Figure 15. Constructing a new generation of base editing system CU-REWIRE5s with engineered ProACD enzymes from mammalian origin. a-c. Crystal structure of engineered ProACD proteins predicted by AlphaFold2. The design and structure prediction method of ProACD proteins is the same as Figure 5.
[0051] Figure 16. Construction of a new generation of base editing system CU-REWIRE5s using mammalian-derived artificial ProACD enzyme. a, Schematic diagram of the design of programmable CU-REWIRE5 system using mammalian-derived ProACD protein. b, Schematic diagram of the precise binding of CU-REWIRE5 to the target EGFP mRNA, the ePUF10 binding sequence is AACGUCUAUA, and the editing target site is C 459 . c-f, Analysis of editing efficiency of different versions of CU-REWIRE5 at the target site C 459 on different transcripts of EGFP. g, Analysis of base editing preference of different versions of CU-REWIRE5 on different transcripts of EGFP.
[0052] Figure 17. Analysis of protein expression and editing efficiency of AAV-CU-REWIRE5 system in mouse brain. a, Schematic diagram of CU-REWIRE5.15 (CU5.15) packaged by AAV of PHP.eb serotype that can cross the blood-brain barrier and injected into mice through intravenous injection. b, Immunofluorescence protein expression analysis of CU5.15 (green) and DAPI (blue) in the cerebral cortex and hippocampal brain regions of mice 4 weeks after CU5.15 virus injection. Scale bar: 200 pm. c, Analysis of editing efficiency of CU5.15 at the target site C 104 in Mef2c mRNA of mice. The binding site of ePUF10 is marked in blue, and C 104 is marked in red (top). Editing efficiency was determined by three sequencing with three replicates. d, Analysis of the percentage of U + / in the prefrontal cortex and hippocampus tissues of Mef2c WT or L35P 104 mice edited by ePUF10 or CU5.15. e, Analysis of editing efficiency of CU5.15 at the C-to-U site C104 in Mef2c in each group. Editing efficiency formula: Cedited= (U CU5.15 -U ePUF10 ) / C total (1-U ePUF10 ). f, Analysis of MEF2C protein expression in the prefrontal cortex and hippocampus of MEF2C WT or L35P + / - mice in ePUF10 or CU5.15 administration groups. Antibody is mef2c antibody. GAPDH as internal control. g, Quantitative analysis of MEF2C protein expression using ImageJ (data from Figure 17f).
[0053] Figure 18. Analysis of MEF2C protein expression in the brain of Mef2c L35P mice. a, Analysis of MEF2C protein expression in the prefrontal cortex and hippocampus of MEF2C WT and L35P + / -In situ protein expression level analysis of MEF2C (red) and DAPI (blue) in the RSC and hippocampus of mice. Experimental results were determined by immunohistochemical staining. b. MEF2C protein fluorescence density quantification in RSC (ePUF10 vs. CU5.15, P = 0.0040) (left panel). MEF2C protein fluorescence density quantification in hippocampus (ePUF10 vs. CU5.15, P < 0.0001) (right panel). c. Mef2c L35P + / - Survival curve analysis of mice. WT (ePUF10) mice 23, L35P+ / -(ePUF10) mice 17, L35P + / - (CU5.15) mice 23).
[0054] Figure 19 Analysis of autism-related social behavior phenotypes in Mef2c L35P mice. a. Motion heat map analysis of Mef2c WT or L35P+ / - mice in social activity in the three-chamber test in different administration groups. b. Analysis of the time of social interaction of Mef2c mice with strange mice in the three-chamber test in different administration groups. c. Analysis of the cumulative social interaction time of Mef2c mice with strange or familiar mice in different administration groups. d. Diagram of the strange mouse intrusion test of mice. T0: experimental mice were isolated and fed for three days. T1-T4: data analysis of experimental mice social interaction test with the same mouse partner. T5: data analysis of experimental mice social interaction test with a new partner mouse. The recording time was the sniffing time between mice within 2 minutes. e. Analysis of the sniffing time of Mef2c WT+ePUF10 in the social intruder test. f. Analysis of the sniffing time of Mef2c L35P + / - +ePUF10 in the social intruder test. g. Analysis of the sniffing time of Mef2c L35P + / - +CU5.15 experimental group in the social intruder test. DETAILED DESCRIPTION
[0055] [Corrected according to Rule 91 on 03.04.2025] The present application provides a method for preparing and use of an artificially designed programmable cytosine deaminase, the method comprising splitting a cytosine deaminase protein into three segments, and fusing the N-terminal functional domain (NTD) and C-terminal functional domain (CTD) of the split cytosine deaminase protein with a deaminase domain of another cytosine deaminase protein to produce an engineered cytosine deaminase, named ProACD (Programmable Artificial Cytosine Deaminase). The cytosine deaminase can be used in RNA and DNA base editing systems.
[0056] The present application constructs a new RNA base editing tool independent of gRNA by fusing cytosine deaminase ProACD and programmable RNA binding protein ePUF10 (enhanced PUF10) protein, called CU-REWIRE (C-to-U RNA editing with individual RNA-binding enzyme). The CU-REWIRE of the present application targets RNA editing based on ePUF10, specifically fusing the RNA recognition domain (RNA recognition domain) that binds RNA, i.e. ePUF10 protein, and the effector domain (effector domain) that functions, i.e. ProACD protein, to construct a new base editing protein CU-REWIRE, which specifically targets the target RNA through the ePUF10 domain, and uses the cytidine deaminase ProACD domain to achieve C-to-U (Cytosine to Uridine) base editing at the target RNA site, thereby eliminating pathogenic RNA to correct pathogenic error point mutations; or editing expressed RNA molecules to specifically manipulate gene function.
[0057] Specifically, the present application first provides an engineered cytosine deaminase, which comprises a backbone and a deaminase domain arranged in the backbone, the backbone of the cytosine deaminase protein is composed of an N-terminal domain and a C-terminal domain of a cytosine deaminase protein, the N-terminal domain and the C-terminal domain are derived from a cytosine deaminase protein, and the deaminase domain is derived from another cytosine deaminase protein. That is, the N-terminal domain and the C-terminal domain are derived from the same cytosine deaminase protein, and the deaminase domain is derived from a cytosine deaminase protein different from the N-terminal domain and the C-terminal domain.
[0058] The "one cytosine deaminase protein" and "another cytosine deaminase protein" refer to cytosine deaminase proteins from different individuals. The different individuals can refer to individuals of different classes (e.g., mammalia), or individuals of the same species or genus.
[0059] In some embodiments of the present application, the skeleton of the cytosine deaminase protein is derived from the skeleton of a human cytosine deaminase protein. In one embodiment, the skeleton of the cytosine deaminase protein can be the skeleton of a natural cytosine deaminase protein, or a skeleton obtained by mutation based on a natural cytosine deaminase protein.
[0060] In some embodiments of the present application, the skeleton of the cytosine deaminase protein is the C-terminal domain (CTD) and the N-terminal domain (NTD) of the cytosine deaminase protein.
[0061] The N-terminal domain and the C-terminal domain are obtained by cutting a cytosine deaminase protein from the H70 site and the M153 site, and the obtained N-terminal and C-terminal fragments are the N-terminal domain and the C-terminal domain, respectively.
[0062] That is, the N-terminal domain is a fragment from the N-terminal of the cytosine deaminase protein to any of the sites within 5 amino acids above and below the H70 site (e.g., 65th, 66th, 67th, 68th, 69th, 71st, 72nd, 73rd, 74th, 75th), and the C-terminal domain is a fragment from any of the sites within 5 amino acids above and below the M153 site (e.g., 148th, 149th, 150th, 151st, 152nd, 154th, 155th, 156th, 157th, 158th) to the C-terminal of the cytosine deaminase protein.
[0063] In one embodiment, the N-terminal domain is a fragment from the first amino acid of the N-terminal of the cytosine deaminase protein to any of the sites within 5 amino acids above and below the H70 site, and the C-terminal domain is a fragment from any of the sites within 5 amino acids above and below the M153 site to the first amino acid of the C-terminal of the cytosine deaminase protein.
[0064] The deamination domain is a fragment between the H70 site or any of the sites within 5 amino acids above and below the H70 site, and the M153 site or any of the sites within 5 amino acids above and below the M153 site of the cytosine deaminase protein.
[0065] In some embodiments of the present application, the framework of APOBEC3A comprises: an N-terminal domain obtained by mutating any one or more of the following sites: H11, H16, K30, H56, on the basis of the wild-type APOBEC3A N-terminal domain with the amino acid sequence shown in SEQ ID NO. 1, and / or a C-terminal domain obtained by mutating site C171 on the basis of the wild-type APOBEC3A C-terminal domain with the amino acid sequence shown in SEQ ID NO. 6.
[0066] In some embodiments of the present application, the N-terminal domain of the cytosine deaminase has an amino acid sequence shown in any one of SEQ ID NO. 2-5.
[0067] In some embodiments of the present application, the C-terminal domain of the cytosine deaminase has an amino acid sequence shown in SEQ ID NO. 7.
[0068] In some embodiments of the present application, the deamination domain of the engineered cytosine deaminase ProACD protein is derived from a prokaryote or a eukaryote. The eukaryote is preferably a mammal. The mammal is preferably a rodent, an even-toed ungulate, an odd-toed ungulate, a lagomorph, a primate, etc. The primate is preferably Homo sapiens, chimpanzee, monkey, ape. In preferred embodiments, the rodent is selected from a rat and a mouse.
[0069] In some embodiments of the present application, the deamination domain of the engineered cytosine deaminase ProACD protein is a native deamination domain or a domain obtained by mutating a native deamination domain.
[0070] In some embodiments of the present application, the engineered cytosine deaminase comprises a structure of NTD-deamination domain-CTD, wherein NTD represents an N-terminal domain and CTD represents a C-terminal domain. For example, the structure of the engineered cytosine deaminase is NTD-deamination domain-CTD-NTD-deamination domain-CTD, which is obtained by connecting NTD-deamination domain-CTD in series.
[0071] In some embodiments of the present application, the domain donor of the engineered cytosine deaminase is selected from natural AID / APOBECs, and the engineered cytosine deaminase prepared therefrom is referred to as ProACD. In a specific embodiment, the cytosine deaminase is selected from APOBEC1 or APOBEC3. In a specific embodiment, the cytosine deaminase is APOBEC3A. In an embodiment, the APOBEC3A is APOBEC3A after single-point or multi-point mutation. In an embodiment, the APOBEC3A is APOBEC3A obtained after mutation at one or more of the following sites: H11, H16, C171, K30, H56, based on the wild-type APOBEC3A with the amino acid sequence shown in SEQ ID NO. 8.
[0072] The amino acid sequence of the engineered cytosine deaminase is shown in SEQ ID NO. 9-55.
[0073] The present application also provides a preparation method of the engineered cytosine deaminase ProACD, which comprises connecting a nucleotide encoding a cytosine deaminase protein skeleton with a nucleotide encoding a deamination domain of another cytosine deaminase protein, obtaining a nucleotide encoding a fragment of the engineered cytosine deaminase, and then cloning the nucleotide encoding the fragment of the engineered cytosine deaminase into an expression vector and expressing to obtain the engineered cytosine deaminase.
[0074] The specific experimental steps of the preparation method are conventional genetic engineering techniques in the art. In an embodiment, the experimental steps can be summarized as follows: assemble the deamination domain DNA sequence encoding the natural AID / APOBEC protein into the skeleton of the human APOBEC3A protein to obtain a ProACD fusion protein fragment, then clone the obtained ProACD fusion protein fragment into a vector, and then transfect the vector into a host cell to produce the engineered cytosine deaminase.
[0075] The preparation method utilizes AlphaFold2 for assistance, and the engineered cytosine deaminase ProACD obtained by the preparation method has cytosine to uracil activity.
[0076] The ProACDs obtained by the preparation method, which are formed by fusing the skeleton domain (NTD and CTD) of human APOBEC3A with various deamination domains of cytosine deaminases from different species, exhibit completely different editing specificity and efficiency in multiple sequence contexts including UC, AC, GC, and CC when integrated into the CU-REWIRE system.
[0077] The present application also provides an enhanced PUF10 protein (ePUF10), which is obtained by integrating leucine (L) and proline (P) into the fourth recognition unit (Repeat4, R4) of a PUF10 domain (PUF10).
[0078] The PUF domain protein is a RNA recognition domain of a natural RNA binding protein Pumilio and FBF (PUF). The ePUF10 protein is a RNA binding protein similar to the PUF protein in structure, which is redesigned on the basis of the natural PUF protein and has 10 recognition unit sequences, each of which is composed of 36 amino acids, wherein the first, second and fifth amino acids are responsible for recognizing target RNA bases, and different target bases are recognized by regulating the composition of the three amino acids; the PUF10 protein assembled by freely combining the recognition unit sequences capable of recognizing specific bases can recognize 10 RNA bases of any sequence.
[0079] In an embodiment, the amino acid sequence of the PUF10 domain is shown as SEQ ID NO. 56.
[0080] In some embodiments of the present application, the amino acid sequence of the ePUF10 protein is shown as SEQ ID NO. 57-64.
[0081] The present application also provides a use of the engineered cytosine deaminase or the ePUF10 protein in preparing a base editing system.
[0082] The base editing system is selected from a DNA base editing system or an RNA base editing system. The DNA base editing system is a cytosine base editor (CBE) system, and / or a guanine base editor (GBE) system.
[0083] The present application also provides a DNA base editing system comprising the cytosine deaminase.
[0084] In some embodiments of the present application, the DNA base editing system further comprises a nuclease and / or a guide RNA. The nuclease and the guide RNA can be designed and selected according to the prior art. The nuclease can be a Cas nuclease.
[0085] The present application also provides an RNA base editing system (CU-REWIRE, also referred to as CU-REWIRE fusion protein), which comprises the engineered cytosine deaminase.
[0086] The RNA base editing system further comprises the ePUF10 protein or the Cas protein, which is directly linked to the engineered cytosine deaminase or linked to the engineered cytosine deaminase through a linking peptide.
[0087] In some embodiments of the present application, the length of the linking peptide is 1-100 amino acids.
[0088] In some embodiments of the present application, the linking peptide is selected from a flexible polypeptide (linker).
[0089] In some embodiments of the present application, the ePUF10 protein and the engineered cytosine deaminase ProACD are fused at both ends by a flexible polypeptide to construct a CU-REWIRE fusion protein.
[0090] In some embodiments of the present application, the flexible polypeptide is selected from an XTEN linker, a GS linker, an NLS linker, an XTEN-NLS linker, or an XTEN-XTEN linker. In some embodiments of the present application, the amino acid sequences of the XTEN linker, the GS linker, the NLS linker, the XTEN-NLS linker, or the XTEN-XTEN linker are shown in SEQ ID NO. 65-69, respectively.
[0091] In a preferred embodiment of the present application, the RNA base editing system comprises the engineered cytosine deaminase and the ePUF10 protein, which is linked to the engineered cytosine deaminase through an XTEN linker.
[0092] In some embodiments of the present application, the amino acid sequence of the CU-REWIRE fusion protein is shown in SEQ ID NO. 70-93.
[0093] In the RNA base editing system, the ePUF10 protein recognizes and binds to RNA, and the engineered cytosine deaminase ProACD acting on single-stranded RNA realizes C-to-U (cytosine to uracil) base editing, i.e., the RNA base editing system can target ProACD to the target RNA editing site through the ePUF10 domain, and realize C-to-U site-directed editing by ProACD.
[0094] The CU-REWIRE RNA base editing system of the present application does not require the use of gRNA, i.e., it can effectively perform C-to-U base editing in cultured cells. Moreover, each component of the CU-REWIRE system is derived from mammals, which can significantly reduce the immune response induced by the system when editing in vivo.
[0095] The present application also provides an isolated polynucleotide encoding any one of the following or a fragment thereof: the engineered cytosine deaminase, the ePUF10 protein, or the base editing system.
[0096] The nucleotide sequence of the isolated polynucleotide encoding the engineered cytosine deaminase, ePUF10 protein, or the RNA base editing system can be obtained by using the prior art to deduce from its amino acid sequence, which is not specifically shown in the present application.
[0097] The amino acid sequence of the present application can also be a sequence that is not shown in the present application and has 90% or more sequence identity with the amino acid sequence shown in the present application, and has a function similar to the amino acid sequence shown in the present application. Further, it can be a sequence obtained by substituting, deleting, or adding one or more (specifically 1-50, 1-30, 1-20, 1-10, 1-5, or 1-3) amino acids to the amino acid sequence shown in SEQ ID Nos. 9-55, and has a function sequence of the amino acid sequence shown in the present application. The amino acid sequence can have 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity with the nucleotide sequence shown in the present application.
[0098] The present application also provides a nucleic acid construct comprising the isolated polynucleotide.
[0099] In one embodiment, the nucleic acid construct is a nucleic acid construct comprising an isolated polynucleotide encoding the engineered cytosine deaminase.
[0100] In one embodiment, the nucleic acid construct is a nucleic acid construct comprising an isolated polynucleotide encoding the ePUF10 protein.
[0101] In one embodiment, the nucleic acid construct is a nucleic acid construct comprising an isolated polynucleotide encoding the RNA base editing system. In one embodiment, the nucleotide sequence of the nucleic acid construct is shown in SEQ ID No. 95.
[0102] The "nucleic acid construct" refers to an artificially constructed nucleic acid segment that can be introduced into target cells or tissues, and can be various expression vectors including a vector backbone, i.e., an empty vector, and an expression frame. The term "expression frame" refers to a sequence having the potential to encode a protein.
[0103] The expression vector is not particularly limited. The expression vector refers to a nucleic acid molecule that allows insertion of foreign nucleotides without destroying the ability of the vector to replicate and / or integrate in a host cell. The expression vector can include nucleic acid sequences that allow it to replicate in a host cell, such as an origin of replication. The expression vector can also include one or more selectable marker genes and other genetic factors. The expression vector is a vector that includes the necessary regulatory sequences to transcribe and translate an inserted gene or genes. The expression vector is selected from a eukaryotic expression vector or a prokaryotic expression vector.
[0104] The eukaryotic expression vector is selected from a yeast expression vector, an insect expression vector or a mammalian expression vector. The mammalian expression vector is selected from a retroviral expression vector, a lentiviral expression vector, an adenoviral expression vector, an adeno-associated viral expression vector.
[0105] The host cell is selected from a eukaryotic host cell or a prokaryotic host cell. The eukaryotic host cell is selected from a fungus such as a yeast, an insect, an avian, a plant, a C. elegans or a nematode or a mammalian host cell. A non-limiting example of an insect cell is a Spodoptera frugiperda (Sf) cell. Examples of yeast host cells are S. cerevisiae, Kluyveromyces lactis (K. lactis) or Yarrowia lipolytica. Examples of mammalian cells are COS cells, baby hamster kidney cells, mouse L cells, LNCaP cells, Chinese hamster ovary (CHO) cells, human embryonic kidney (HEK) cells, African green monkey cells, CV1 cells, Vero or Hep-2 cells. Examples of prokaryotic host cells include bacterial cells such as E. coli, Streptomyces, B. subtilis, Salmonella typhi or mycobacteria.
[0106] The skilled person can transfect the expression vector into a host cell to obtain a cell comprising the coding gene of the cytosine deaminase or the ePUF10 protein or the RNA base editing system according to methods well known in the art. For example, introducing the expression vector into a eukaryotic cell can be performed by calcium phosphate co-precipitation, electroporation, microinjection, lipofection or transfection with a polyamine transfection reagent.
[0107] The present application also provides a viral vector system comprising the nucleic acid construct. The viral vector system is an adeno-associated viral vector system.
[0108] The adeno-associated viral vector expression system is selected from the group consisting of an E. coli expression system, a yeast expression system, an insect expression system, a mammalian expression system, a plant expression system; preferably, the adeno-associated viral vector expression system is selected from the group consisting of any one of a plasmid transient transfection expression system, a baculovirus expression system, a stable cell line expression system, an adenovirus expression system, a poxvirus expression system.
[0109] Further, the adeno-associated viral vector system further comprises a host cell. The host cell carries the adeno-associated viral vector. The host cell can be selected from any suitable host cell in the art as long as it does not limit the inventive purpose of the present application. Specifically, the suitable cell can be a cell for producing adeno-associated virus, for example, 293 cell.
[0110] In some embodiments of the present application, in the plasmid transient transfection expression system, the adeno-associated viral vector system further comprises a nucleic acid construct carrying a target gene and a helper plasmid. In an embodiment of the present application, the nucleic acid construct carrying a target gene comprises a polynucleotide encoding the engineered cytosine deaminase or the ePUF10 protein or the RNA base editing system. Specifically, the nucleic acid construct carrying a target gene is the aforementioned polynucleotide construct comprising a polynucleotide encoding the RNA base editing system of the present application.
[0111] The present application also provides an adeno-associated virus (AAV) which is packaged by the adeno-associated viral vector system. The adeno-associated virus can be used to treat various diseases, such as Duchenne muscular dystrophy, premature aging, and the specific disease category can be achieved according to the target sequence recognized by the designed ePUF10 protein.
[0112] The CU-REWIRE of the present application can be successfully delivered into different tissues and organs of animals by AAV or LNP delivery system to edit the target RNA of the target organ. Specifically, AAV-CU-REWIRE can be introduced into Mef2c-L35P ASD mouse models by tail vein injection, and can precisely edit and repair the target RNA in the brain of the mouse model, with an efficiency of 50%, and the RNA repaired by base editing can be normally translated, successfully repairing the expression of MEF2C protein and correcting the autism-like behavior of Mef2c-L35P ASD mice. It can be seen that the CU-REWIRE system provides a powerful platform for in vivo base editing by engineering ProACD, paving the way for future clinical translation of gene therapy.
[0113] The present application also provides a cell containing the nucleic acid construct or having the exogenous polynucleotide integrated into the genome.
[0114] The cell of the present application is obtained by transforming the expression vector into a host cell.
[0115] The present application also provides a base editing method, which comprises contacting a target gene with the adeno-associated virus or the base editing system to edit a single base on the target gene.
[0116] In one embodiment, the single base editing is C-to-U base editing. The target gene is a gene that binds to the ePUF10.
[0117] The base editing method is performed in a cell, or in vivo, an ex vivo cell, or a cell-free system.
[0118] The present application also provides a method for treating a disease, which comprises administering the base editing system or the adeno-associated virus to a subject or an ex vivo cell of the subject.
[0119] The type of the disease is not particularly limited, and the type of the disease can be determined according to the pathogenic gene edited by the base editing system. The disease is, for example, a disease of the nervous system, and is, for example, Alzheimer's disease.
[0120] The above embodiments of the present application are described in detail by specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the disclosure. The present application can also be implemented or applied by other different specific embodiments, and various modifications or changes can be made to the details in the specification based on different views and applications without departing from the spirit of the present application.
[0121] Before further describing the specific embodiments of the present application, it should be understood that the scope of protection of the present application is not limited to the following specific embodiments; it should also be understood that the terms used in the embodiments of the present application are used to describe the specific embodiments, but not to limit the scope of protection of the present application; in the specification and claims of the present application, the singular forms "a", "an" and "the" include the plural forms unless otherwise explicitly stated in the text.
[0122] When the embodiments give a numerical range, it should be understood that, unless otherwise stated by the present application, each numerical range and any numerical value between the two endpoints can be selected. Unless otherwise defined, all technical and scientific terms used in the present application have the same meaning as generally understood by those skilled in the art. In addition to the specific methods, devices, materials used in the embodiments, any method, device and material of the prior art similar or equivalent to those described in the embodiments of the present application can also be used to implement the present application according to the master of the prior art and the description of the present application by those skilled in the art.
[0123] Materials and methods
[0124] 1. Construction of engineered cytidine deaminase ProACD gene expression plasmid
[0125] To construct the ProACD fusion protein, first, the coding DNA sequence of the amino acid sequence of the deamination domain of the natural AID / APOBEC protein was synthesized by gene synthesis, and then the coding DNA sequence of the amino acid sequence of the deamination domain of the natural AID / APOBEC protein was assembled into the skeleton of the human APOBEC3A protein to design a new programmable ProACD fusion protein (amino acid sequence as shown in SEQ ID NO. 9-55), and then the coding DNA sequence fragment of the obtained ProACD sequence was cloned into the pCI-Neo vector (purchased from Promega Corporation) to generate the pCI-ProACD-Flag vector.
[0126] 2. Construction of enhanced ePUF10 plasmid
[0127] The RNA binding protein PUF10 expression vector was constructed by first amplifying the ePUF10 sequence by PCR, linearizing the pCl-neo plasmid by double digestion (NheI / NotI), and then using T4 DNA ligase to insert the ePUF10 fragment into the pCl-neo vector to construct the pCl-ePUF10 plasmid (amino acid sequence as shown in SEQ ID NO. 56-64).
[0128] 3. Construction of novel CU-REWIRE5 protein expression plasmid
[0129] (1) Linearization of the pCl-ProACD-Flag vector in 1.
[0130] (2) The ePUF10 sequence constructed in 2 was inserted into the pCl-ProACD-Flag vector by enzyme digestion and ligation (see Materials and Methods 1: purchased from Promega Corporation), to construct the pCl-ProACD-ePUF10 (CU-REWIRE5) plasmid.
[0131] 4. Packaging and purification of adeno-associated virus
[0132] The vector of adeno-associated virus (AAV) used in this study is pAAV-hSyn-EGFP plasmid. First, linearize the pAAV-hSyn-EGFP plasmid by enzyme digestion, then insert the ePUF10 and CU-REWIRE5 fragments amplified by PCR into the pAAV-hSyn-EGFP plasmid by homologous recombination, respectively, to obtain pAAV-hSyn-PUF10 (control group) and pAAV-hSyn-CU-REWIRE (experimental group). Then, package pAAV-hSyn-PUF10 and pAAV-hSyn-CU-REWIRE into AAV-PHP.eB serotype adeno-associated virus, respectively, then use iodixanol gradient centrifugation for virus purification, then collect the virus, and detect the virus titer by qPCR.
[0133] 5. Mammalian cell culture and transfection
[0134] The cell line used in this study is HEK 293T (ATCC CRL-3216), which is cultured in DMEM high glucose medium (HyClone) containing 10% fetal bovine serum (FBS), and the culture environment is a 37°C constant temperature incubator (Thermo) with a 5% concentration of carbon dioxide.
[0135] The cell transfection experiment involved in this study uses the transient transfection method, except for special annotation. To study the editing effect of the editing tool CU-REWIRE on the target site in the reporter gene, the cells are co-transfected with the REWIRE system and the reporter gene, and then the cells are collected for further analysis and detection. The transfection reagent used in the above experiment is Lipofectamine 3000 (Thermo Fisher Scientific), and the use method is described in the product manual.
[0136] 6. Mammalian experiment
[0137] All mouse experiments conducted in this study are approved and authorized by the Animal Welfare and Ethics Committee of the School of Medicine, Shanghai Jiao Tong University. The mice used in this study are from the laboratory preservation, and both female and male mice are used.
[0138] In this experiment, male mice were evenly divided into experimental and control groups, and AAV-PHP.eB adeno-associated virus particles encoding CU-REWIRE protein were injected into four-week-old Mef2c-L35P + / - (C57BL / 6J) heterozygous mice, and the injection dose of each mouse was 1×10 12vg (vector genomes). After careful cultivation after injection, when AAV-PHP.eB-CU-REWIRE is highly expressed (about 4 weeks), the RNA editing efficiency is detected, and the social behavior of mice is studied.
[0139] The experimental method of the social behavior experiment is to observe and analyze the social behavior of mice. By observing and recording the behavioral characteristics of the RNA editing group mice and the control group mice under the same experimental conditions, the differences in social behavior between the two groups of mice are analyzed and compared to study whether the autism-related pathological phenotype of the RNA editing group mice is restored. Then different brain tissues of the mice are collected, RNA is extracted, and RNA-seq or Sanger-seq is used to detect the editing efficiency of REWIRE on the target RNA in different brain tissues of the mice.
[0140] The experimental methods not explicitly described in the present application, such as RNA sequencing methods, are performed using conventional techniques in the art.
[0141] Example 1: Improve the expression stability of the new version of CU-REWIRE base editor by optimizing the structure of PUF10 protein
[0142] In order to design a more efficient and stable RNA recognition protein, the present application inserts an LP (L: leucine, P: proline) dipeptide into the fourth recognition unit (R4: recognition unit 4) of PUF10 with an amino acid sequence as shown in SEQ ID NO. 56 to obtain an enhanced PUF10 (ePUF10) (the amino acid sequence is shown as SEQ ID NO. 57-64). The structure of ePUF10 is significantly different from that of PUF10, and the new structure is more easily recognized and combined with RNA (Figure 1a).
[0143] Then the ePUF10 module and the cytidine deaminase APOBEC3A are fused to develop a new base editor CU-REWIRE4.0 (Figure 1b, the amino acid sequence is shown as SEQ ID NO. 71). The results show that the protein expression amount of the ePUF10 version of CU-REWIRE4.0 with LP insertion in cells is significantly higher than that of the original version of CU-REWIRE3.0 (Figure 1c, the amino acid sequence is shown as SEQ ID NO. 70), indicating that the structure optimization of ePUF10 protein with LP dipeptide insertion can produce a more efficient and stable RNA base editor with higher expression.
[0144] Example 2: Verify the editing efficiency of the new version of CU-REWIRE editing enzyme
[0145] To verify whether the editing efficiency of the new version of CU-REWIRE is significantly improved, we designed CU-REWIRE to target and edit the C of enhanced green fluorescent protein (EGFP) mRNA. 459 The target sites were then analyzed to evaluate the editing efficacy of CU-REWIRE4.0 (Figures 2a-2d). mRNA high-throughput sequencing (RNA-Seq) results showed that CU-REWIRE4.0 achieved a significant improvement in C-to-U editing efficiency, reaching 82.3%, compared to 69.7% for CU-REWIRE3.0 (Figure 2a), indicating that optimizing the ePUF10 domain can significantly enhance the editing efficacy of CU-REWIRE4.0.
[0146] The results of off-target analysis at the transcriptome level (RNA-seq analysis with 50X transcriptome coverage) showed that CU-REWIRE3.0 had 224 off-target base editing events and CU-REWIRE4.0 had 731 off-target base editing events (Figure 2b), indicating that the editing efficiency of the upgraded version of CU-REWIRE4.0 with the ePUF10 domain is significantly higher than that of the original CU-REWIRE3.0.
[0147] We further analyzed the editing sequence preference and editing accuracy of CU-REWIRE4.0, and the results showed that CU-REWIRE4.0 tends to edit C in the UC consensus motif (Figure 2c). At the same time, its editing accuracy is extremely high, which is further reflected in the fact that C-to-U base editing mainly occurs at the second position downstream of the ePUF10 binding site (Figure 2d). In summary, compared with CU-REWIRE3.0, the C-to-U base editing efficiency of CU-REWIRE4.0 at the target RNA target site is significantly improved, indicating that the CU-REWIRE4.0 version has significant enhancements in protein expression levels and target base editing efficiency.
[0148] Example 3 Optimizing the amino acid composition of APOBEC3A to reduce off-target effects of the CU-REWIRE editing system
[0149] The natural APOBEC3A protein tends to form dimers to bind single-stranded DNA or single-stranded RNA, and the formation of such dimers increases off-target editing effects in the base editing process. In this patent, in order to reduce the formation of APOBEC3A dimers, we analyzed the crystal structure and molecular evolution characteristics of human APOBEC3A (amino acid sequence as shown in SEQ ID NO. 8) protein, and introduced single-point and multi-point mutations (H11, H16, C171, K30, H56, Figure 3a) in the APOBEC3A protein. Then, these mutated APOBEC3A versions were fused with ePUF10 domains to construct different versions of CU-REWIRE4 variants (Figures 3b-3c).
[0150] We designed CU-REWIRE4 to target edit the C 459 target site to evaluate the editing efficacy of improved CU-REWIRE4 (Figure 3c). The mRNA-Seq results show that the point mutation version of CU-REWIRE4.X significantly reduces off-target effects while maintaining high C-to-U editing efficiency. For example, CU-REWIRE 4.1 with C171A point mutation significantly reduces off-target effects while maintaining high editing efficiency (Figures 3c-3d). These data show that our optimization strategy of selectively introducing point mutations in the NTD and CTD framework of APOBEC3A can significantly reduce the number of off-target edits (reduced number of off-targets, reduced side effects of the tool) and the efficiency of off-target editing (lower efficiency of off-target editing, lower side effects, average efficiency 11.4%, reduced to 7.9%, a very significant effect) in the CU-REWIRE system (Figure 3d). Further analysis of the data relationship between the editing efficiency of the editing enzyme at the target site and the off-target effect of each CU-REWIRE4.x version shows that these CU-REWIRE4.x versions can effectively improve the editing efficiency of the editing enzyme (Figure 3e), showing that our optimization results provide more diverse choices for base editing applications and greatly reduce side effects.
[0151] Example 4 Splitting APOBEC proteins into three independently functional parts with the aid of artificial intelligence
[0152] According to the experimental data of Example 3 (Figure 3), and our analysis of APOBEC family proteins, we found that APOBEC proteins are composed of three relatively independent domains, the deaminase domain is relatively conserved, but the NTD and CTD are prone to evolutionary changes, which means that the deaminase domain is essential for catalyzing C-to-U conversion, while the NTD and CTD may specifically recognize different DNA / RNA targets. At the same time, we found the site of scientific splitting of APOBEC3A protein (H70 / M153 site), which creatively splits the APOBEC protein into three domains with independent functions, namely NTD domain, deaminase domain (Deaminase domain), and CTD domain. The three domains have relatively independent functions and cooperate with each other to achieve C-to-U base editing (Figure 3a).
[0153] Based on the above, we propose the following hypothesis: when the deaminase domain of AID / APOBEC protein is fused with APOBEC NTD or CTD domain from different sources, it may create new types of engineered cytosine base editors that can edit cytosine in different environments.
[0154] To verify this hypothesis, we analyzed the structure of APOBEC3 and APOBEC1 from human and rodent sources using AlphaFold2, and split these proteins according to the splitting rule we found (splitting points are shown in Figures 4a-4e), and the results are shown in the figures. The deaminase domain (Deaminase domain) of different APOBEC proteins split out has a surprising similar structure (Figures 4a-4e), suggesting that the functional modules of these proteins may be interchangeable, further indicating that our designed splitting strategy is feasible, laying a foundation for our subsequent research.
[0155] Example 5 Artificial intelligence assisted development of programmable engineered cytosine base editor ProACD (Programmable Artificial Cytosine Deaminase)
[0156] According to the experimental data obtained in Examples 3 and 4, we split different sources of natural cytosine deaminases to obtain thousands of NTD domains, deaminase domains (Deaminase domain), and CTD domain components (Figure 5a).
[0157] [Corrected according to Rule 91 on 03.04.2025] Subsequently, the present invention explores the placement of the deaminase domain of the split natural AID / APOBEC into the human APOBEC3A protein scaffold we designed, designing a new type of engineered cytidine deaminase synthesis platform, which we call the programmable engineered cytidine deaminase Pre-ACD synthesis platform, the specific engineered cytidine deaminase produced by the platform is named ProACD (Figure 5b), which can be used for base editing of DNA and RNA.
[0158] We analyzed the phylogenetic tree of the AID / APOBEC family using 1160 complete eukaryotic AID / APOBEC1-related genes recorded in the NCBI database, and found that APOBEC3A and APOBEC1 belong to closely related branches in the tree (Figure 5c). Subsequently, we used artificial intelligence tools such as AlphaFold2 to assist in designing 1160 engineered ProACD proteins using the 1160 natural proteins as templates (Figure 5d). At the same time, we used AlphaFold2 to predict the crystal structure of the engineered ProACD proteins in Figure 5d, using human and rodent deaminase domains as donors for synthesis (Figures 5e-5f). These ProACDs show extremely similar structures to human APOBEC3A, suggesting that they may have similar base editing efficacy to human APOBEC3A protein (Figures 5e-5f). In summary, our method can design thousands of functional ProACD proteins (Figure 5).
[0159] Example 6 Development of a new generation of programmable RNA base editing system CU-REWIRE5 using ProACD
[0160] In order to design a more accurate and efficient RNA base editing system, the present invention fuses different types of ProACD and ePUF10 domain modules to develop a new generation of programmable RNA base editing system CU-REWIRE5s (Figure 6a). Specifically, we fused the ProACD and ePUF10 domain modules obtained in Example 5 to develop different versions of CU-REWIRE5.1-5.17 (amino acid sequences are shown in SEQ ID NO. 72-78, respectively), to further analyze the new functions of the CU-REWIRE5s editing system (Figure 6b).
[0161] Expression analysis in HEK293T cells showed that all CU-REWIRE5 variants could express editing enzyme protein normally, and most CU-REWIRE5 proteins were stable (Figure 6c). We further analyzed whether these editing enzymes could edit RNA, and the results showed that most CU-REWIRE5 could edit target RNA, and the editing efficiency was very high (Figure 6d).
[0162] Example 7 CU-REWIRE5 base editor system can significantly enhance the targeting range and specificity of RNA base editing
[0163] Next, in order to verify the efficiency and off-target effect of CU-REWIRE5 in editing RNA, we designed CU-REWIRE5 to target and edit the C 459 Target site, specifically, one base upstream of the EGFP mRNA target site was changed to different bases (where U 458 C was changed to G 458 C, A 458 C and C 458 C), in order to test the sequence preference of CU-REWIRE5 in editing target RNA (Figure 7a).
[0164] This example tested the editing specificity and activity of CU-REWIRE5 in different base environments of the target site. The results showed that multiple CU-REWIRE5 base editors containing ProACD protein exhibited completely different base editing sequence preferences and surprising editing activity from natural APOBEC protein, especially in some target site environments that natural APOBEC protein cannot edit, such as in G 458 C, A 458 C and C 458C environment (Fig. 7b). For example, CU-REWIRE 5.1 and CU-REWIRE 5.8, whose ProACD protein components’ deamination domain donors are derived from human APOBEC1 and APOBEC3H, respectively, show high C-to-U editing activity in GC motifs and lower C-to-U editing activity in AC motifs, while native APOBEC proteins have no editing activity in these GC and AC sequences. CU-REWIRE 5.15 shows high C-to-U editing activity in CC and UC motifs. CU-REWIRE 5.16 shows high C-to-U editing activity in AC and UC motifs. CU-REWIRE 5.3 and CU-REWIRE 5.17 show high C-to-U editing activity in GC, AC, CC, and UC motifs, while native APOBEC proteins only have editing activity in UC motifs. In addition, some ProACDs also show similar editing activity to native APOBEC proteins, such as CU-REWIRE 5.7 and CU-REWIRE 5.13, which can only edit cytosine C in UC motifs (Fig. 7b).
[0165] In addition, we used transcriptome RNA-seq experiments to evaluate the overall off-target effects of CU-REWIRE5 at the whole transcriptome level (Fig. 7c). Our results show that, compared with CU-REWIRE 4.1 containing native APOBEC proteins, CU-REWIRE 5.1, 5.17, and 5.16 containing engineered ProACD proteins show higher off-target editing rates, and CU-REWIRE 5.15, 5.7, 5.13, and 5.8 have fewer off-target editing sites (Fig. 7c).
[0166] There is a positive correlation between off-target editing rates and on-target editing efficiency (i.e., the higher the editing activity of CU-REWIRE5s, the more off-target editing events), which indicates that our newly designed ProACD proteins have different editing sequence preferences and editing activities from native APOBEC proteins, and further indicates that the method used in the present application significantly expands the library of cytidine deaminase proteins, and the ProACD enzymes we designed can serve as a supplement and replacement for native cytidine deaminase, with extremely high basic research and scientific transformation value.
[0167] Example 8 Editing sequence preference and editing precision analysis of CU-REWIRE 5.1 and 5.8 editing GC motifs
[0168] The deamination domain donors of the ProACD protein components of the two editing enzymes, CU-REWIRE5.1 and CU-REWIRE 5.8, are derived from human APOBEC1 and APOBEC3H, respectively. They showed efficient C-to-U editing activity within the GC motif, where the natural APOBEC protein has no editing activity; and showed lower C-to-U editing activity within the AC motif, where the natural APOBEC protein has no editing activity (Figures 8a-8e). Our results showed that the editing efficiency of the ProACD enzyme was significantly different from that of the natural human APOBEC3A (which only edits the UC motif) and human APOBEC1 (which has low efficiency in editing the UC motif), indicating that new cytidine editing activity has emerged in the engineered ProACD (Figure 8a).
[0169] We then evaluated the base editing properties of CU-REWIRE5.1 and CU-REWIRE 5.8 in detail in the EGFP sequence. The results showed that CU-REWIRE5.1 showed efficient C-to-U editing activity within the GC motif, lower C-to-U editing activity within the AC motif, and no editing activity in the UC and AC motifs (Figure 8b). We further used the RNA-seq data in Figure 7c to analyze the base editing preferences of off-target editing events at the transcriptome level and found that the editing motifs in the off-target base editing events were consistent with the editing motifs in the single target site (C 459 ) using C-to-U editing (Figure 8c). Furthermore, the editing properties of CU-REWIRE5.8 were largely consistent with those of CU-REWIRE5.8, albeit with slightly lower efficiency (Figures 8d-8e).
[0170] Subsequently, this example further analyzed the editing accuracy of the CU-REWIRE5.1 base editing enzyme, especially the screening of the editing window (Figures 9a-9f). We found that CU-REWIRE5.1 (editing at the GC motif) showed the highest editing efficiency at the 2- or 4-nucleotide (nt) position downstream of the ePUF10 binding site in EGFP mRNA, as shown by different target reporter genes containing GC dinucleotide motifs at different positions downstream of the ePUF10 binding site (Figure 9a).
[0171] Subsequently, we further evaluated the editing efficiency of CU-REWIRE5.1 in reporter genes carrying pathogenic mutations. For example, APOE4 allele mutations at Arg130Cys (c.T388C) and Arg173Cys (c.T526C) sites of APOE4 protein are high-risk pathogenic mutations of Alzheimer's disease, which can cause the occurrence of Alzheimer's disease. We designed CU-REWIRE5.1 (CU5.1) enzymes with ePUF10 domain (amino acid sequences are shown in SEQ ID NO. 87-90) respectively, which are combined with the upstream sequences of APOE4 RNA c.T388C and c.T526C target sites (combined at 2-nt or 3-nt upstream of the target site C 388 and C 526 respectively), and found that CU-REWIRE5.1 combined 2-nt upstream of the mutation site can effectively convert the target base C to U, with an editing efficiency of more than 90%, which can convert APOE4 to benign APOE3 (Figure 9b) and APOE1 (Figure 9c) alleles, respectively. In another example, the human rhodopsin gene c.68C>A (p.P23H) mutation known to cause retinitis pigmentosa (RP) can be effectively edited by CU-REWIRE5.1 combined with 2-nt upstream of the mutation site (Figure 9c) (amino acid sequence is shown in SEQ ID NO. 91). We edited C 76 to U 76 , which is equivalent to introducing a stop codon in the mRNA carrying the mutation site, which can eliminate toxic protein mutants (Figure 9c). Moreover, CU-REWIRE5.1 can efficiently edit the T143C pathogenic mutation site of SOD1 mRNA associated with neurodegenerative diseases, especially when CU-REWIRE5.1 is combined with 2-nt upstream of the mutation site (amino acid sequence is shown in SEQ ID NO. 92), which can efficiently convert the pathogenic mRNA to normal functional mRNA, and ultimately eliminate toxic protein mutants (Figure 9e). In summary, based on the results of CU-REWIRE5.1 targeting editing different mRNAs, we summarized the editing window of CU-REWIRE5.1 as 2-5 nt downstream of the ePUF10 binding site (Figure 9f), which provides more specific guidance for potential users of base editing tools in this embodiment.
[0172] Example 9 Analysis of editing sequence preference and editing accuracy of CU-REWIRE5.15 editing CC motif
[0173] Specifically, the deaminase domain of the ProACD protein component of CU-REWIRE5.15 editors was derived from murine APOBEC3. It showed high efficient C-to-U editing activity in CC motif, which was not edited by native APOBEC proteins (Fig. 10a-10c). Our results showed that the editing efficiency of ProACD enzyme was significantly different from that of native human APOBEC3A, indicating that new cytidine editing activity appeared in the engineered ProACD (Fig. 10a).
[0174] We next evaluated the base editing properties of CU-REWIRE5.15 in detail in EGFP sequence, and the results showed that CU-REWIRE5.15 showed high efficient C-to-U editing activity in CC motif and UC motif, and no editing activity in AC and GC motif (Fig. 8b). We further analyzed the base editing preference of off-target editing events at the transcriptome level using the RNA-seq data in Fig. 7c, and found that the editing motif in off-target base editing events was basically consistent with the sequence preference observed when using C-to-U editing on the single target site (C 459 ) of EGFP mRNA (Fig. 10c).
[0175] Subsequently, this embodiment further analyzed the editing accuracy of CU-REWIRE5.15 enzyme, especially the screening of editing window (Fig. 11a-11c). We found that CU-REWIRE5.15 (editing in CC motif) had the highest editing efficiency in EGFP mRNA 2-nt downstream of ePUF10 binding site (Fig. 11a). Subsequently, we used a similar method to detect the editing window of CU-REWIRE5.15 in the gene carrying pathogenic point mutation, and we designed CU-REWIRE5.15 to target edit the MEF2C gene c.104T>C (p.L35P) mutation, which is considered to cause severe autism spectrum disorder. We found that CU-REWIRE5.15 could effectively edit the target site C 104 to U 104 downstream of ePUF10 binding site with an editing efficiency of about 45%, which could repair the pathogenic point mutation of MEF2C (Fig. 11b).
[0176] In summary, combined with the results of CU-REWIRE5.15 targeted editing of different mRNAs, we summarized the editing window of CU-REWIRE5.15 as 2-4 nt downstream of ePUF10 binding site (Fig. 11c), which provided more specific guidance for potential users of the base editing tool in this embodiment.
[0177] Example 10 Editing sequence preference and editing accuracy analysis of CU-REWIRE5.16 editing AC motif
[0178] Specifically, the deaminase domain of the ProACD protein component of CU-REWIRE5.16 editing enzyme was derived from murine APOBEC3. It showed high efficient C-to-U editing activity within AC motif, while the native APOBEC protein showed no editing activity in this sequence (Fig. 12a-12c). Our results showed that the editing efficiency of ProACD enzyme was significantly different from the native human APOBEC3A, indicating that new cytidine editing activity appeared in the engineered ProACD (Fig. 12a). We next evaluated the base editing properties of CU-REWIRE5.16 in detail in EGFP sequence, and the results showed that CU-REWIRE5.16 showed high efficient C-to-U editing activity within AC motif, and very low editing activity in AC, CC and GC motifs (Fig. 12b). We further analyzed the base editing preference of off-target editing events at the transcriptome level using the RNA-seq data of Fig. 7c, and found that the editing motifs in off-target base editing events were basically consistent with the sequence preference observed when using C-to-U editing on a single target site (C 459 ) of EGFP mRNA (Fig. 12c).
[0179] Subsequently, this embodiment further analyzed the editing window of CU-REWIRE5.16 enzyme (Fig. 12d-12e). We found that CU-REWIRE5.16 (editing in AC motif) had the highest editing efficiency 3-nt downstream of the ePUF10 binding site in EGFP mRNA (Fig. 12d). In summary, combined with the results of CU-REWIRE5.16 targeting editing different mRNAs, we summarized the editing window of CU-REWIRE5.16 as 2~5nt downstream of the ePUF10 binding site (Fig. 12e), which provides more specific guidance for potential users of base editing tools in this embodiment.
[0180] Example 11 Editing sequence preference and editing accuracy analysis of CU-REWIRE5.16 editing AC motif
[0181] The deaminase domain donors of the ProACD protein component of CU-REWIRE 5.3 and CU-REWIRE 5.17 are derived from mouse APOBEC1 and rat APOBEC1, respectively. They show high efficiency of C-to-U editing activity in AC, CC, GC, UC four motifs, while the native APOBEC proteins have no RNA editing activity in these sequences (Figures 13a-13e). Our results show that the editing efficiency of ProACD enzymes is significantly different from the native mouse APOBEC1 and rat APOBEC1, indicating that new cytidine editing activity has emerged in the engineered ProACD (Figure 13a).
[0182] We next evaluated the base editing properties of CU-REWIRE 5.3 and CU-REWIRE 5.17 in detail in the EGFP sequence, and the results showed that CU-REWIRE 5.17 showed high efficiency of C-to-U editing activity in AC, CC, GC, UC four motifs, and lower editing activity in AC and CC motifs (Figure 13b). We further analyzed the base editing preference of off-target editing events at the transcriptome level using the RNA-seq data in Figure 7c, and found that the editing motifs in off-target base editing events were basically consistent with the sequence preference observed when using C-to-U editing on a single target site (C 459 ) of EGFP mRNA (Figure 13c). In addition, the editing properties of CU-REWIRE 5.3 were basically consistent with CU-REWIRE 5.17, except that the efficiency was slightly lower (Figures 13d-13e).
[0183] This embodiment further analyzes the editing window size of CU-REWIRE 5.17 base editing enzyme (Figures 14a-14c). We found that CU-REWIRE 5.17 showed higher editing efficiency at the 2- or 4-nt position downstream of the ePUF10 binding site in the EGFP mRNA (Figure 14a). Moreover, CU-REWIRE 5.17 can efficiently edit the T1541C site of DDX3X mRNA associated with neurodevelopmental disorders, especially when CU-REWIRE 5.17 binds 2-nt or 3-nt upstream of the mutant site (the amino acid sequence is shown in SEQ ID NO. 93), ultimately eliminating the toxic protein mutant (Figure 14b). In summary, combined with the results of CU-REWIRE 5.17 targeting editing different mRNAs, we summarize the editing window of CU-REWIRE 5.17 as 2-5 nt downstream of the ePUF10 binding site (Figure 14c), which provides more specific guidance for potential users of the base editing tool in this embodiment.
[0184] Example 12. Constructing a new generation of base editing system CU-REWIRE5s with engineered ProACD enzymes from mammalian sources
[0185] This example also explored the possibility of generating ProACD using deaminase domains from other mammals based on the rules discovered in Example 5. We analyzed APOBEC1s and APOBEC3s from mammalian species using AlphaFold2 (Figures 15a-15c). We found that the Z1 and Z2 domains of mammalian APOBEC3s can form functional ProACD when combined with the NTD and CTD of human APOBEC3 (Figures 15a-15c). We fused these engineered ProACD with ePUF10 domains to construct different versions of CU-REWIRE5s from mammalian sources (amino acid sequences are shown in SEQ ID NO. 79-86, respectively) (Figure 16a), and then tested the editing activities of these CU-REWIRE5s containing engineered ProACD on EGFP mRNA (Figure 16b) to further analyze the new functions of CU-REWIRE5s editing system in detail. The data showed that the newly synthesized CU-REWIRE5s displayed high efficiency of C-to-U editing activity in GC, AC and CC motifs (Figures 16c-16g).
[0186] These results highlight the diversity of engineered ProACD in C-to-U RNA base editing, further demonstrating that the method employed by the present application significantly expands the library of cytidine deaminase proteins, and the ProACD enzymes designed by us can serve as a supplement and replacement for natural cytidine deaminases, with extremely high basic research and scientific transformation value.
[0187] Example 13. Evaluating the efficacy of CU-REWIRE5 system in in vivo RNA base editing
[0188] Efficient RNA base editing within the central nervous system has been a highly challenging task. To explore the editing efficacy of CU-REWIRE system within the central nervous system, we used CU-REWIRE5 fused with ProACD enzymes for in vivo RNA base editing. We took CU-REWIRE5.15 (CU5.15) as an example, and used CU5.15 to repair the Mef2c L35P (c.104T>C) point mutation in an autism spectrum disorder (ASD) mouse model, which causes severe ASD phenotypes.
[0189] We used CU5.15 (which targets the binding of Mef2c mRNA A 92 -A 101The amino acid sequence of CU5.15 after insertion of the promoter, CU-REWIRE5, and polyA of the present application is shown as SEQ ID NO. 94, and the nucleotide sequence is shown as SEQ ID NO. 95 (Figure 17a). CU5.15 is driven by the neuron-specific promoter hSyn, and the C-terminus of CU5.15 is connected to the EGFP reporter gene via P2A, which is used to track the expression profile of CU5.15 protein. After AAV-CU5.15 is packaged, it is injected into 4-week-old Mef2c WT and L35P + / - mice via the tail vein (Figure 17a). Two months after AAV injection, the expression of AAV-CU5.15 in the mouse brain was observed, and the results showed that AAV-CU5.15 (green fluorescent signal) can be efficiently expressed in the target brain regions of mice, i.e., the cortex and hippocampus, indicating that CU5.15 can be applied to the treatment of related diseases of the nervous system (Figure 17b).
[0190] Subsequently, we detected the editing characteristics of the CU-REWIRE5 system in the mouse brain by performing Sanger sequencing and RNA sequencing on the prefrontal cortex and hippocampal tissue samples of mice. The experimental results showed that in the Mef2c L35P + / - hybrid mice, the proportion of Mef2C mRNA carrying the normal T 104 site in the control group was about 50%, i.e., the proportion of Mef2C mRNA carrying the pathogenic C 104 site was about 50%; in the hippocampal tissue, the proportion of Mef2C mRNA carrying the normal T 104 site in the control group was about 50%, i.e., the proportion of Mef2C mRNA carrying the pathogenic C 104 site was about 50% (Figures 17c-16d). In the experimental group injected with AAV-CU5.15, the proportion of Mef2C mRNA carrying the normal T 104 site in the prefrontal cortex increased significantly (the average value reached 72%, i.e., the proportion of Mef2C mRNA carrying the pathogenic C 104 site decreased to 28%); the proportion of Mef2C mRNA carrying the normal T 104 site in the hippocampus also increased significantly (the average value increased to 67%, i.e., the proportion of Mef2C mRNA carrying the pathogenic C 104 site decreased to 33%) (Figures 17c-17d). Further detailed statistics of AAV-CU5.15 in the C 104The specific editing efficiency of the site, the results showed that the average editing efficiency of AAV-CU5.15 in different brain regions was between 25% and 40% (Figure 17e). In addition, no non-target editing was observed in the cytosine C near the Mef2c mRNA target site C 104 target site (Figure 17c), demonstrating the accuracy of CU-REWIRE5.15-mediated RNA editing. In summary, AAV-CU5.15 can efficiently edit the target Mef2c mRNA in the target brain region of mice, and the editing accuracy is extremely high, which is manifested as single-base editing, that is, while repairing point mutations, no new editing mutations are introduced, which is significantly better than the existing CRISPR system-mediated base editing enzyme (Figures 17c-17e).
[0191] Most notably, compared with Mef2c L35P mice injected with ePUF10 control, the expression level of MEF2C protein in the prefrontal cortex and hippocampus brain regions of Mef2c L35P mice injected with CU5.15 was significantly restored (Figure 17f), especially in the hippocampus, the expression of MEF2C protein was close to complete restoration (Figure 17g), which indicated that the edited Mef2c mRNA by CU5.15 could be normally translated to produce functional protein.
[0192] We further verified the restoration of MEF2C protein expression level by immunostaining using anti-MEF2C antibody. Compared with the control group, the expression of MEF2C protein in the brain of mice in the CU5.15 administration group was significantly restored, and the protein was restored to near the level of Mef2c WT mice (Figures 18a-18b). In addition, the lifespan of Mef2c L35P mice injected with CU5.15 was also significantly improved, and the lifespan was restored to normal level (Figure 18c), which fully demonstrated that CU-REWIRE can repair mutant RNA in the central nervous system, and the edited RNA can be normally translated to produce functional protein to improve the pathological phenotype of mice.
[0193] The efficacy of CU5.15 in treating the autism pathological phenotype of Mef2c L35P mice was also studied. First, we detected the social ability of mice by three-box test (Figure 19a), and the results showed that the mice in the CU5.15 injection group had basic social ability (Figure 19b), and the social novelty ability of the mice was significantly improved (Figure 19c). In addition, the results of social intruder behavioral experiments showed that the abnormal social behavior of Mef2c L35P mice was completely rescued after CU5.15 treatment in the first four tests with intruder partners (Figures 19d-19g). These results show that RNA base editing mediated by CU5.15 can completely restore abnormal MEF2C protein expression and improve social deficits associated with Mef2c-L35P mutation-related neurodevelopmental disorders in vivo.
[0194] In summary, our research shows that RNA base editing by CU-REWIRE5 provides a powerful tool for correcting genetic mutations in the brain, and has high clinical treatment value for the treatment of ASD disease and related nervous system diseases.
[0195] The above examples are intended to illustrate the embodiments disclosed in the present application and should not be construed as limiting the present application. In addition, various modifications listed herein and changes in the method of the invention are obvious to those skilled in the art without departing from the scope and spirit of the invention. Although the present application has been specifically described in conjunction with various preferred embodiments thereof, it should be understood that the present application should not be limited to these specific embodiments. In fact, various modifications as described above to obtain the invention that are obvious to those skilled in the art should be included within the scope of the present application.
[0196] The sequence numbers and sequence names of amino acids or nucleotides in the present application correspond as follows:
Claims
1. An engineered cytosine deaminase, characterized in that: The engineered cytosine deaminase comprises a skeleton and a deamination domain disposed within the skeleton. The skeleton of the cytosine deaminase protein is composed of the N-terminal domain and the C-terminal domain of the cytosine deaminase protein. The N-terminal domain and the C-terminal domain are derived from one cytosine deaminase protein, and the deamination domain is derived from another cytosine deaminase protein.
2. The engineered cytosine deaminase according to claim 1, characterized in that The skeleton of the cytosine deaminase protein is the skeleton of a cytosine deaminase protein of mammalian origin; preferably, the skeleton of the cytosine deaminase protein is derived from the skeleton of a human cytosine deaminase protein; preferably, the skeleton of the cytosine deaminase protein is the skeleton of a natural cytosine deaminase protein, or a skeleton obtained by mutation on the basis of a natural cytosine deaminase protein.
3. The engineered cytosine deaminase according to claim 1, characterized in that The deamination domain of the cytosine deaminase protein is derived from mammals; preferably, the deamination domain of the cytosine deaminase protein is a natural deamination domain or a domain obtained by mutation based on the natural deamination domain.
4. The engineered cytosine deaminase according to claim 1, characterized in that The N-terminal domain is a fragment from the N-terminal end of the cytosine deaminase protein to the H70 site or any site within 5 amino acids above and below the H70 site, and / or the C-terminal domain is a fragment from the M153 site of the cytosine deaminase protein or any site within 5 amino acids above and below the M153 site to the C-terminal end; and / or the deamination domain is a fragment between the H70 site of the cytosine deaminase protein or any site within 5 amino acids above and below it, to the M153 site or any site within 5 amino acids above and below the M153 site.
5. The engineered cytosine deaminase according to claim 1, characterized in that The engineered cytosine deaminase comprises the following structure: NTD-deamination domain-CTD; NTD stands for N-terminal domain, and CTD stands for C-terminal domain.
6. The engineered cytosine deaminase according to claim 1, characterized in that The cytosine deaminase is AID / APOBECs; preferably, the cytosine deaminase is selected from the APOBEC1 family or the APOBEC3 family.
7. The engineered cytosine deaminase according to claim 1, characterized in that The cytosine deaminase is APOBEC3A; preferably, the APOBEC3A skeleton is an APOBEC3A skeleton that has undergone single-point or multi-point mutations; More preferably, the APOBEC3A backbone comprises: an N-terminal domain of the wild-type APOBEC3A having an amino acid sequence as shown in SEQ ID NO. 1, or an N-terminal domain obtained by mutation of any one or more of the following sites: H11, H16, K30, H56, and / or a C-terminal domain of the wild-type APOBEC3A having an amino acid sequence as shown in SEQ ID NO. 6, or a C-terminal domain obtained by mutation of the C171 site thereof.
8. The engineered cytosine deaminase according to claim 1, characterized in that The amino acid sequence of the N-terminal domain of the cytosine deaminase is shown in any one of SEQ ID NOs. 2 to 5; and / or the amino acid sequence of the C-terminal domain of the cytosine deaminase is shown in SEQ ID NO.
7.
9. The engineered cytosine deaminase according to claim 1, characterized in that The amino acid sequences of the engineered cytosine deaminase are shown in SEQ ID NOs. 9 to 55.
10. The method for preparing the engineered cytosine deaminase according to any one of claims 1 to 9, characterized in that: The preparation method comprises linking nucleotides encoding a cytosine deaminase protein skeleton with nucleotides encoding a deamination domain of another cytosine deaminase protein to obtain nucleotides encoding the engineered cytosine deaminase fragment, and then cloning the nucleotides encoding the engineered cytosine deaminase fragment into an expression vector and expressing the nucleotides to obtain the engineered cytosine deaminase.
11. An ePUF10 protein, characterized in that The ePUF10 protein is obtained by integrating leucine and proline into the fourth recognition unit of the PUF10 domain.
12. The ePUF10 protein according to claim 11, characterized in that The amino acid sequence of the PUF10 domain is shown in SEQ ID NO.
56.
13. The ePUF10 protein according to claim 11, characterized in that The amino acid sequence of the ePUF10 protein is shown in SEQ ID NO.
57.
14. Use of the engineered cytosine deaminase according to any one of claims 1 to 9 or the ePUF10 protein according to any one of claims 11 to 13 in preparing a base editing system.
15. The use according to claim 14, characterized in that The base editing system is selected from a DNA base editing system or an RNA base editing system; preferably, the DNA base editing system is a cytosine base editor system, and / or a guanine base editor system.
16. A DNA base editing system, characterized in that The DNA base editing system comprises the engineered cytosine deaminase according to any one of claims 1 to 9.
17. An RNA base editing system, characterized in that The RNA base editing system comprises the engineered cytosine deaminase or its encoding gene according to any one of claims 1 to 9.
18. The RNA base editing system according to claim 17, wherein The RNA base editing system is a CU-REWIRE fusion protein or its encoding gene, wherein the CU-REWIRE fusion protein includes the engineered cytosine deaminase described in any one of claims 1 to 9, and further includes the ePUF10 protein or Cas protein described in any one of claims 11 to 13, wherein the ePUF10 protein or Cas protein is directly connected to the engineered cytosine deaminase, or is connected to the engineered cytosine deaminase through a connecting peptide.
19. The RNA base editing system according to claim 18, characterized in that The connecting peptide is selected from a flexible polypeptide, preferably, the length of the connecting peptide is 1 to 100 amino acids; preferably, the flexible polypeptide is selected from an XTEN connecting peptide, preferably, the amino acid sequence of the connecting peptide is shown in any one of SEQ ID NOs. 65 to 69.
20. The RNA base editing system according to claim 18, wherein The amino acid sequence of the CU-REWIRE fusion protein is shown in SEQ ID NOs. 71-93.
21. An isolated polynucleotide, characterized in that The polynucleotide encodes any of the following or a fragment thereof: the engineered cytosine deaminase according to any one of claims 1 to 9, the ePUF10 protein according to any one of claims 11 to 13, or the RNA base editing system according to any one of claims 17 to 20.
22. A nucleic acid construct, characterized in that The nucleic acid construct comprises the polynucleotide of claim 21.
23. A viral vector system, characterized in that The viral vector system comprises the nucleic acid construct according to claim 22; preferably, the viral vector system is an adeno-associated viral vector system.
24. An adeno-associated virus, characterized in that The adeno-associated virus is formed by viral packaging of the viral vector system according to claim 23.
25. A cell, characterized in that The cell contains the nucleic acid construct of claim 22 or the exogenous polynucleotide of claim 21 integrated into its genome.
26. A base editing method, characterized in that The target gene is contacted with the adeno-associated virus according to claim 24 or the RNA base editing system according to any one of claims 17 to 20 to achieve single base editing on the target gene.
27. The base editing method according to claim 26, wherein The single base editing is C-to-U base editing; preferably, the target gene is a gene that binds to the ePUF10.
28. A method for treating a disease, characterized in that: The treatment method comprises administering the RNA base editing system of any one of claims 17 to 20 or the adeno-associated virus of claim 24 to the subject or the subject's ex vivo cells.
29. The method of claim 28, wherein: The disease is a disease of the nervous system, preferably Alzheimer's disease.
Citation Information
Patent Citations
Base editors with improved precision and specificity
CN110959040A
Using split deaminases to limit unwanted off-target base editor deamination
CN111093714A
RNA fixed-point editing by artificially constructing RNA editing enzyme and related application
CN111793627A
Bipartite base editor (BBE) architectures and type-ii-c-cas9 zinc finger editing
US20200140842A1