Fusion protein, RNA editing level detection method and application thereof

By using fusion proteins and deaminases to edit RNA at the initiation or termination of translation, and combining this with high-throughput sequencing technology, the problem of large sample sizes and high costs required for detecting RNA translation levels in existing technologies has been solved, enabling simple and low-cost detection of RNA editing levels.

CN121991936APending Publication Date: 2026-05-08SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411550870.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies require large amounts of biological samples and high-throughput RNA sequencing to detect RNA translation levels. The experimental procedures are complex and the reagents are expensive, which limits their application.

Method used

By employing fusion proteins, including combinations of translation initiation or termination factors and deaminases, RNA editing is achieved at the initiation or termination of translation. Combined with high-throughput sequencing technology, this simplifies the detection method for RNA editing levels.

Benefits of technology

It enables the acquisition of RNA editing levels using a small number of biological samples and a single test. The experimental procedure is simple and inexpensive, making it suitable for general molecular biology laboratories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121991936A_ABST
    Figure CN121991936A_ABST
Patent Text Reader

Abstract

The invention discloses a fusion protein, an RNA editing level detection method and application thereof, and relates to the technical field of nucleic acid editing.The fusion protein comprises a first protein and a second protein, and the first protein comprises any one of a translation initiation factor and a translation termination factor; the second protein comprises deaminase; the fusion protein is used for editing RNA at the beginning or ending of translation. Wherein the translation initiation factor and the translation termination factor are respectively combined with the RNA chain at the translation initiation and the translation termination, so that the fusion protein is in contact with the RNA, and all the RNA for translation initiation and translation termination are more accurately positioned; when the fusion protein is in contact with the RNA, the deaminase performs deamination editing on nucleotides near the contact point, so that a marker for RNA editing is left, and the RNA translation level can be detected by using the marker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of nucleic acid editing technology, and in particular to a method for detecting fusion proteins and RNA editing levels, and its application. Background Technology

[0002] According to the central dogma, genetic information is stored in deoxyribonucleic acid (DNA). DNA is transcribed to produce ribonucleic acid (RNA), which is then translated to produce proteins, ultimately exhibiting specific biological functions. The regulation of translation plays a crucial role in many physiological and pathological processes, such as organ development, cancer, and the occurrence and development of neurodegenerative diseases. Therefore, systematically detecting RNA translation levels is essential for in-depth exploration of the molecular mechanisms of gene expression regulation, development, and disease.

[0003] While existing technologies such as ribosomal imprinting can detect RNA translation levels, they present several challenges: 1) they require a large number of biological samples; 2) high-throughput RNA sequencing must be performed on each sample to estimate the translation level; and 3) the experimental procedures are complex and reagents are expensive. These factors limit RNA translation level identification to a small number of biological samples. Summary of the Invention

[0004] The main objective of this invention is to propose a method for detecting fusion proteins and RNA editing levels, and its application, aiming to provide a method for calculating RNA translation levels that requires less biological sample volume, fewer sequencing attempts, and a simpler experimental procedure.

[0005] To achieve the above objectives, the present invention proposes a fusion protein comprising a first protein and a second protein.

[0006] The first protein includes either a translation initiation factor or a translation termination factor;

[0007] The second protein includes a deaminase;

[0008] The fusion protein is used to edit RNA at the start or end of translation.

[0009] In one embodiment, the deaminase includes any one of ADAR, APOBEC, TadA, ADAR variants, APOBEC variants, and TadA variants.

[0010] In one embodiment, the translation initiation factor includes any one of EIF1AX, EIF3G, EIF3D, and EIF4E.

[0011] In one embodiment, the translation termination factor includes GSPT1.

[0012] In one embodiment, the fusion protein is obtained by gene cloning.

[0013] This invention also proposes a method for detecting RNA editing levels, comprising the following steps:

[0014] S10. Obtain the cell sample to be tested;

[0015] S20. Express the fusion protein in the cell sample to be tested, screen, and obtain a cell sample to be tested enriched with the fusion protein, wherein the fusion protein is the aforementioned fusion protein;

[0016] S30. Extract RNA from the cell sample containing the enriched fusion protein;

[0017] S40. The editing level of the RNA is calculated based on the RNA of the cell sample to be tested containing the enriched fusion protein.

[0018] In one embodiment, S40 includes:

[0019] S401. Sequencing the RNA to obtain the number of RNA reads (R) covering trusted editing sites and the number of RNA reads (E) in the test cell sample enriched with the fusion protein that have undergone trusted editing.

[0020] S402. Calculate the RNA editing level of each gene in the test cell sample containing the enriched fusion protein according to the formula Z = E / R, where Z is the RNA editing level of each gene in the test cell sample containing the enriched fusion protein, E is the number of RNA reads in the test cell sample containing the enriched fusion protein that have undergone reliable editing, and R is the number of RNA reads in the test cell sample containing the enriched fusion protein that cover the reliable editing site.

[0021] In one embodiment, in step S20, the expression of the fusion protein in the cell sample to be tested includes transient transfection or stable transfection.

[0022] In one embodiment, in step S401, the method for sequencing the RNA includes high-throughput sequencing and Sanger sequencing.

[0023] The present invention also provides an application of the method for detecting RNA editing level in detecting RNA translation level.

[0024] In the technical solution of this invention, the fusion protein includes a deaminase and a translation initiation factor / translation termination factor. The translation initiation factor and the translation termination factor bind to the RNA strand at the initiation and termination of translation, respectively, so that the fusion protein contacts the RNA and more accurately locates all RNAs that have started or terminated translation. When the fusion protein contacts the RNA, the deaminase deaminates and edits nucleotides near the contact point, thereby leaving a marker for RNA editing. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0026] Figure 1 This is a flowchart of the RNA editing level detection method in Example 1 of the present invention;

[0027] Figure 2 This is a graph showing the results of Western blot detection of fusion protein expression in Example 1 of the present invention;

[0028] Figure 3 This is a graph showing the translation level results of two different genes in HEK293T cells in Example 2 of the present invention;

[0029] Figure 4 The graph shows the results of the base mutation level of the fusion protein detected at the whole transcriptome level using high-throughput sequencing technology in Example 3 of the present invention.

[0030] Figure 5 A comparison diagram of RNA editing levels of protein-coding genes and non-protein-coding genes in Example 3 provided by the present invention;

[0031] Figure 6 The graph shows the correlation analysis results between the RNA editing level and the RNA translation level obtained by ribosomal imprinting technology in Example 3 of this invention, where the fusion protein is introduced.

[0032] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer shall apply. Where the manufacturers of reagents or instruments are not specified, they are all conventional products that can be purchased commercially. Furthermore, the meaning of "and / or" throughout the text includes three parallel solutions; for example, "A and / or B" includes solution A, or solution B, or a solution where both A and B are satisfied simultaneously. In addition, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] According to the central dogma, genetic information is stored in deoxyribonucleic acid (DNA). DNA is transcribed to produce ribonucleic acid (RNA), which is then translated to produce proteins, ultimately exhibiting specific biological functions. The regulation of translation plays a crucial role in many physiological and pathological processes, such as organ development, cancer, and the occurrence and development of neurodegenerative diseases. Therefore, systematically detecting RNA translation levels is essential for in-depth exploration of the molecular mechanisms of gene expression regulation, development, and disease.

[0035] Detecting RNA translation levels using high-throughput sequencing technologies, such as ribosomal imprinting, requires first collecting a large number of cells (approximately 10 million cells per sample). These cells are then treated with actinomycete ketone to inhibit ribosome movement on RNA. Subsequent, complex molecular biology procedures, including cell lysis, RNase digestion, recovery of small RNA fragments, and construction of small RNA sequencing libraries, take approximately 4 to 5 days to obtain a sequencing library characterizing ribosomal imprinting. Simultaneously, a total RNA sequencing library from the same cell sample must be constructed to obtain RNA content information for each gene. Finally, bioinformatics analysis correlates the ribosomal imprinting and RNA content of each gene in the two sequencing libraries to ultimately calculate the RNA translation level.

[0036] While existing technologies such as ribosomal imprinting can detect RNA translation levels, several challenges exist: 1) a large number of biological samples are required; 2) high-throughput RNA sequencing must be performed on each sample to estimate the translation level; and 3) the experimental procedures are complex and reagents are expensive. These factors limit RNA translation level identification to a relatively small number of biological samples.

[0037] In view of this, the present invention provides a fusion protein comprising a first protein and a second protein.

[0038] The first protein includes either a translation initiation factor or a translation termination factor;

[0039] The second protein includes a deaminase;

[0040] The fusion protein is used to edit RNA at the start or end of translation.

[0041] The constructed fusion protein is either a translation initiation factor-deaminase or a translation termination factor-deaminase.

[0042] RNA editing refers to the insertion, deletion, or substitution of bases in the coding region of transcribed RNA, resulting in differences between the mature RNA and the DNA template that originally encoded it.

[0043] Nucleoside deaminases are a wide range of enzymes that catalyze the removal of amino groups from various nucleosides in living organisms. For example, adenosine deaminase (ADAR) can catalyze the removal of amino groups from adenosine (A) to convert it into hypoxanthine nucleoside (I). Therefore, by utilizing the deamination properties of nucleoside deaminases, the conversion of nucleotides on the RNA chain can be achieved, that is, RNA editing can be realized.

[0044] Translation initiation factors and translation termination factors play crucial roles in protein biosynthesis, participating in the initiation and termination phases of translation, respectively, ensuring accuracy and efficiency. Translation initiation factors are primarily responsible for initiating the translation process, enabling mRNA to bind to ribosomes; translation termination factors are responsible for recognizing stop codons at the end of the translation process and promoting polypeptide chain release and ribosome dissociation.

[0045] In the technical solution of this invention, the fusion protein includes a deaminase and a translation initiation factor / translation termination factor. The translation initiation factor and the translation termination factor bind to the RNA strand at the initiation and termination of translation, respectively, so that the fusion protein contacts the RNA and more accurately locates all RNAs that have started or terminated translation. When the fusion protein contacts the RNA, the deaminase deaminates and edits nucleotides near the contact point, thereby leaving a marker for RNA editing.

[0046] In some embodiments of the present invention, the deaminase includes any one of ADAR, APOBEC, TadA, ADAR variants, APOBEC variants, and TadA variants. The ADAR is an adenosine deaminase, a class of proteins containing a deaminase-binding domain that catalyzes the conversion of adenosine residues (A) in RNA molecules into creatinine (I), i.e., the conversion from A to I. The ADAR family mainly includes three types of ADAR proteins: ADAR1 (isomers p110 and p150), ADAR2, and ADAR3. The APOBEC is an apolipoprotein B mRNA editing catalytic polypeptide. APOBEC is a class of cytosine deaminases. APOBEC proteins utilize their deaminase activity to catalyze the conversion of cytosine nucleotides (C) in mRNA or DNA to uracil (U), or cytosine nucleotides (C) to thymine nucleotides (T), by binding to RNA or DNA. The APOBEC family includes APOBEC1, APOBEC2, APOBEC3, and APOBEC4, among others. APOBEC3 includes multiple members such as APOBEC3A and APOBEC3B. The main function of the TadA deaminase is to deaminate adenine (A) and convert it to hypoxanthine (I). Specific variants of ADAR, APOBEC, and TadA may include different subtypes, isoforms, or members with specific functions or expression patterns, such as the deaminase domain of the human ADARB1 protein carrying the E488Q mutation, the full-length rat-derived APOBEC1 protein, and TadA.8e. However, variants of ADAR, APOBEC, and TadA that have the above-mentioned deamination or RNA conversion functions are all within the scope of protection of this invention.

[0047] In some embodiments of the present invention, the translation initiation factor includes any one of EIF1AX, EIF3G, EIF3D, and EIF4E. Using a fusion protein containing the above translation factors, the fusion protein binds to RNA more efficiently during the translation initiation phase, allowing the deaminase in the fusion protein to react more quickly with the bases on the RNA to edit the RNA, thereby improving the level of RNA editing.

[0048] In some embodiments of the present invention, the translation termination factor includes GSPT1. Using GSPT1, the fusion protein binds to RNA more efficiently during the translation termination phase, allowing the deaminase in the fusion protein to react more quickly with the bases on the RNA to edit the RNA, thereby improving the RNA editing level.

[0049] In some embodiments of the present invention, the fusion protein is obtained by gene cloning. The fusion protein of the present invention is a single polypeptide chain formed by linking two or more proteins together through genetic engineering. For example, if a translation initiation factor and a deaminase form a fusion protein, it means that during the construction of the expression vector, the coding gene for the translation initiation factor and the coding gene for the deaminase are tandemly linked, so that the protein product obtained after transcription and translation simultaneously contains the functional regions of both proteins.

[0050] It should be noted that the fusion order of the translation initiation / termination factor and the deaminase can be interchanged. For example, the translation initiation / termination factor may be at the N-terminus of the fusion protein, and the deaminase may be at the C-terminus of the fusion protein; or the deaminase may be at the N-terminus of the fusion protein, and the translation initiation / termination factor may be at the C-terminus of the fusion protein.

[0051] This invention also proposes a method for detecting RNA editing levels, comprising the following steps:

[0052] S10. Obtain the cell sample to be tested;

[0053] S20. Express the fusion protein in the cell sample to be tested, screen, and obtain a cell sample to be tested enriched with the fusion protein, wherein the fusion protein is the aforementioned fusion protein;

[0054] S30. Extract RNA from the cell sample containing the enriched fusion protein;

[0055] S40. The editing level of the RNA is calculated based on the RNA of the cell sample to be tested containing the enriched fusion protein.

[0056] It should be noted that in step S20, after the fusion protein is expressed in the cell sample to be tested, it will bind to naturally occurring endogenous factors in the cell to form a functional complex (such as the translation initiation complex).

[0057] When a translation initiation factor is fused with a deaminase, the newly generated fusion protein is expressed in the cell. During the early stages of translation, the translation initiation factor interacts with mRNA and ribosomes to help guide the initiation of new protein chain synthesis. At the same time, the deaminase portion retains its original function and binds to its corresponding substrate after translation is completed, playing a deamination modification role.

[0058] The advantages of the above methods for detecting RNA editing levels are: 1) they can be used to identify RNA editing levels with a small number of biological samples; 2) each biological sample only needs to be tested once to obtain the RNA editing level; 3) the experimental procedure is simple, suitable for general molecular biology laboratories, and the reagent cost is low.

[0059] In some embodiments of the present invention, S40 includes:

[0060] S401. Sequencing the RNA to obtain the number of RNA reads (R) covering trusted editing sites and the number of RNA reads (E) in the test cell sample enriched with the fusion protein that have undergone trusted editing.

[0061] S402. Calculate the RNA editing level of each gene in the test cell sample containing the enriched fusion protein according to the formula Z = E / R, where Z is the RNA editing level of each gene in the test cell sample containing the enriched fusion protein, E is the number of RNA reads in the test cell sample containing the enriched fusion protein that have undergone reliable editing, and R is the number of RNA reads in the test cell sample containing the enriched fusion protein that cover the reliable editing site.

[0062] It should be noted that in step S401, the sequencing method for the RNA can be either Sanger sequencing or high-throughput sequencing; there is no limitation, as long as the RNA sequence can be obtained. When using Sanger sequencing, the EditR software is used to calculate the base editing ratio at specific positions in the Sanger sequencing results. When using high-throughput sequencing, reliable editing sites are detected using Sailor (1.2.0) software, and the number of RNA reads with reliable editing sites for each gene in the cell sample enriched with the fusion protein is obtained using Biostar214299 software. Furthermore, FeatureCounts (v2.0.1) software is used to calculate the number of edited RNA sequencing reads for each gene and the total number of RNA sequencing reads in the cell sample enriched with the fusion protein. Using the formula Z = E / R, the RNA editing level corresponding to each gene in the cell sample can be calculated relatively accurately.

[0063] In some embodiments of the present invention, step S20, expressing the fusion protein in the cell sample to be tested, includes transient transfection or stable transfection. It should be noted that transient transfection does not integrate the foreign gene into the cell genome, and the foreign gene is not expressed in passaged cells; stable transfection can transfect the foreign gene into the cell genome, thereby expressing the foreign gene in passaged cells.

[0064] In the technical solution of the present invention, either transient transfection or stable transfection can be used to transfect the fusion protein sequence into the cell sample to be tested, as long as the fusion protein can be expressed in the sample to be tested.

[0065] Transient transfection can quickly produce cells expressing high levels of fusion proteins for detection, while stable transfection has a longer experimental cycle.

[0066] In some embodiments of the present invention, in step S401, the method for sequencing the RNA includes high-throughput sequencing and Sanger sequencing.

[0067] It should be noted that high-throughput RNA sequencing refers to a sequencing method that indirectly obtains RNA sequences by randomly fragmenting RNA, reverse transcribing it into cDNA, and then performing high-throughput sequencing on the cDNA.

[0068] By employing high-throughput sequencing technology, RNA can be sequenced and analyzed to determine the sequence of each transcript fragment with single nucleotide resolution accuracy. At the same time, there are no cross-reaction and background noise problems caused by fluorescence simulation signals in traditional microarray hybridization. Furthermore, there is no need to pre-design specific probes, and transcriptome analysis can be performed directly on the cell samples to be tested.

[0069] Meanwhile, each biological sample only needs to undergo high-throughput RNA sequencing once, and the RNA editing level is detected by editing RNA through fusion proteins, while the RNA expression level is also detected.

[0070] The present invention also provides an application of the method for detecting RNA editing level in detecting RNA translation level.

[0071] Given that RNA editing is irreversible, and that translation initiation / termination factor fusion proteins only come into contact with RNA and edit it when translation occurs / terminates, RNA editing levels can be used to indicate the degree to which translation initiation / termination factors bind to RNA, and then combined with RNA expression levels to quantitatively detect RNA translation levels.

[0072] The technical solution of the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings. It should be understood that the following embodiments are only used to explain the present invention and are not intended to limit the present invention.

[0073] Experimental methods:

[0074] Transient expression (plasmid): Taking HEK293T cells as an example, HEK293T cells were seeded in 6-well cell culture plates one day in advance, with 400,000 cells per well. On the day of the experiment, 2 μg of the pcDNA6 transient expression vector plasmid carrying the fusion protein coding sequence or deaminase coding sequence was transfected into HEK293T cells using PEI or Lipo transfection reagents (PEI: Yisheng Biotechnology (Shanghai) Co., Ltd., catalog number 40816ES02; Lipo: Yisheng Biotechnology (Shanghai) Co., Ltd., catalog number 40802ES01. Refer to the transfection reagent instructions for specific operations). Forty-eight hours after transfection, cells that successfully expressed the fusion protein / isolated deaminase protein were sorted and enriched using a flow cytometer based on the expression level of the fluorescent protein. These cells, with significantly higher fluorescent protein expression levels than the negative control (cells not transfected with mRNA / plasmid), were used for RNA extraction and high-throughput RNA sequencing library construction (for specific procedures, refer to the RNA extraction reagent (Invitrogen, catalog number 15596018CN) and the high-throughput RNA sequencing library construction kit instructions (Yisheng Biotechnology (Shanghai) Co., Ltd., catalog number 12309ES08).

[0075] Transient expression (mRNA): Taking HEK293T cells as an example, before transient expression of mRNA, a transient expression vector is needed as a template. Capped and tailed mRNAs are obtained using an in vitro transcription kit (Invitrogen, catalog number AM1345). Using HEK293T cells as an example, HEK293T cells are seeded in 6-well cell culture plates one day in advance, with 400,000 cells per well. On the day of the experiment, mRNA is transfected into HEK293T cells using the Lipo transfection reagent (refer to the transfection reagent instructions for specific procedures). Forty-eight hours after transfection, cells successfully expressing the fusion protein / isotropic deaminase protein are selected and enriched using flow cytometry based on the expression level of the fluorescent protein, i.e., cells with significantly higher fluorescent protein expression levels than the negative control (cells without mRNA / plasmid transfection). These cells are then used for RNA extraction and high-throughput RNA sequencing library construction (refer to the RNA extraction reagent and high-throughput RNA sequencing library construction kit instructions for specific procedures).

[0076] Stable Expression (Lentivirus): Taking HEK293T cells as an example, before stable expression using lentivirus, the stable expression vector and lentiviral helper packaging plasmids (pMD2.G (Addgene, #12259) and psPAX2 (Addgene, #12260)) need to be co-transfected into HEK293T cells. 48 hours after transfection, the culture medium supernatant is collected and mixed with polybrene (refer to the polybrene instructions, Yisheng Biotechnology (Shanghai) Co., Ltd., catalog number 40804ES76) and incubated with the cells to be infected for 48 hours. After 48 hours, the culture medium is replaced with one containing cyhalofop-p-ethyl (the final concentration of cyhalofop-p-ethyl in the medium is 10 μg / mL) for drug screening until all uninfected wild-type cells die. Cells successfully infected with the stable expression vector are obtained approximately 4-5 days later.

[0077] On the day of the stable expression experiment, HEK293T cells were seeded in 6-well cell culture plates, 400,000 cells per well. Doxycycline at a concentration of 1 μg / mL was added to the culture medium to induce the expression of the fusion protein / isotropic deaminase protein. After 48 hours of induction, cells were collected for RNA extraction and high-throughput RNA sequencing library construction (refer to the instructions for RNA extraction reagents and high-throughput RNA sequencing library construction kits for specific procedures).

[0078] Example 1: Construction of a fusion protein expression vector

[0079] (1) Construction of deaminase plasmid:

[0080] The full-length amino acid sequence of Human ADARB1 (hADAR2) was obtained by searching NCBI, and its deaminase domain was optimized based on reported literature to enhance RNA editing efficiency and reduce its preference for dsRNA.

[0081] The full-length amino acid sequence of Rat APOBEC1 (APOBEC1) was obtained by querying NCBI.

[0082] After the DNA sequences encoding the optimized ADAR2dd (hADAR2 deaminase domain) and the full-length amino acid sequence of APOBEC1 were synthesized by Qingke Biotechnology, primers were designed for PCR amplification.

[0083] The amino acid sequence of Human ADAR2dd is shown in SEQ ID NO.1:

[0084] MDIEDEENMSSSSTDVKENRNLDNVSPKDGSTPGPGEGSQLSNGGGGGPGRKRPLEEGSNGHSKYRLKKRRKTPGPVLPKNALMQLNEIKPGLQYTLLSQTGPVHAPLFVMSVEVNGQVFEGSGPTKKKAKLHAAEKALRSFVQFPNASEAHLAMGRTLSVNTDFTSDQADFPDTLFNGFETPDKAEPPFYVGSNGDDSFSSSGDLSLSASPVPASLAQPPLPVLPPFPPPSGKNPVMILNELRPGLKYDFLSESGESHAKSFVMSVVVDGQFFEGSGRNKKLAKARAAQSALAAIFNLHLDQTPSRQPIPSEGLQLHLPQVLADAVSRLVLGKFGDLTDNFSSPHARRKVLAGVVMTTGTDVKDAKVISVSTGTKCINGEYMSDRGLALNDCHAEIISRRSLLRFLYTQLELYLNNKDDQKRSIFQKSERGGFRLKENVQFHLYISTSPCGDARIFSPHEPILEGSRSYTQAGVQWCNHGSLQPRPPGLLSDPSTSTFQGAGTTEPADRHPNRKARGQLRTKIESGEGTIPVRSNASIQTWDGVLQGERLLTMSCSDKIARWNVVGIQGSLLSIFVEPIYFSSIILGSLYHGDHLSRAMYQRISNIEDLPPLYTLNKPLLSGISNAEARQPGKAPNFSVNWTVGDSAIEVINATTGKDELGRASRLCKHALYCRWMRVHGKVPSHLLRSKITKPNVYHESKLAAKEYQAAKARLFTAFIKAGLGAWVEKPTEQDQFSLTP。

[0085] The amino acid sequence of Rat APOBEC1 is shown in SEQ ID NO.2:

[0086] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK。

[0087] The optimized amino acid sequence of Human ADAR2dd is shown in SEQ ID NO.3:

[0088] MLHLPQVLADAVSRLVLGKFGDLTDNFSSPHARRKVLAGVVMTTGTDVKDAKVISVSTGTKCINGEYMSDRGLALNDCHAEIISRRSLLRFLYTQLELYLNNKDDQKRSIFQKSERGGFRLKENVQFHLYISTSPCGDARIFSPHEPILEEPADRHPNRKARGQLRTKIESGQGTIPVRSNASIQTWDGVLQGERLLTMSCSDKIARWNVVGIQGSLLSIFVEPIYFSSIILGSLYHGDHLSRAMYQRISNIEDLPPLYTLNKPLLSGISNAEARQPGKAPNFSVNWTVGDSAIEVINATTGKDELGRASRLCKHALYCRWMRVHGKVPSHLLRSKITKPNVYHESKLAAKEYQAAKARLFTAFIKAGLGAWVEKPTEQDQFSLT。

[0089] The optimized DNA sequence of Human ADAR2dd is shown in SEQ ID NO.4:

[0090]

[0091] The DNA sequence of Rat APOBEC1 is shown in SEQ ID NO.5:

[0092] ATGAGCTCAGAGACTGGCCCAGTGGCTGTGGACCCCACATTGAGACGGCGGATCGAGCCCCATGAGTTTGAGGTATTCTTCGATCCGAGAGAGCTCCGCAAGGAGACCTGCCTGCTTTACGAAATTAATTGGGGGGGCCGGCACTCCATTTGGCGACATACATCACAGAACACTAACAAGCACGTCGAAGTCAACTTCATCGAGAAGTTCACGACAGAAAGATATTTCTGTCCGAACACAAGGTGCAGCATTACCTGGTTTCTCAGCTGGAGCCCATGCGGCGAATGTAGTAGGGCCATCACTGAATTCCTGTCAAGGTATCCCCACGTCACTCTGTTTATTTACATCGCAAGGCTGTACCACCACGCTGACCCCCGCAATCGACAAGGCCTGCGGGATTTGATCTCTTCAGGTGTGACTATCCAAATTATGACTGAGCAGGAGTCAGGATACTGCTGGAGAAACTTTGTGAATTATAGCCCGAGTAATGAAGCCCACTGGCCTAGGTATCCCCATCTGTGGGTACGACTGTACGTTCTTGAACTGTACTGCATCATACTGGGCCTGCCTCCTTGTCTCAACATTCTGAGAAGGAAGCAGCCACAGCTGACATTCTTTACCATCGCTCTTCAGTCTTGTCATTACCAGCGACTGCCCCCACACATTCTCTGGGCCACCGGGTTGAAA。

[0093] The primer sequences of Human ADAR2dd are:

[0094] Forward primer F1: ATGGGCCAGCTGCATTTACCGCAGG;

[0095] Reverse primer R1: CGTGAGTGAGAACTGGTCC.

[0096] The primer sequences of Rat APOBEC1 are:

[0097] Upstream primer F2: ATGAGCTCAGAGACTGGC;

[0098] Downstream primer R2: TTTCAACCCGGTGGCCCA.

[0099] Amplification system: 25 μL of 2×Phanta Flash Master Mix, 2 μL of upstream primer (10 μM), 2 μL of downstream primer (10 μM), 10 ng of template DNA (plasmid template), and bring the volume to 50 μL with ddH2O.

[0100] Amplification program: 98℃ for 30s; 98℃ for 10s, 60℃ for 5s, 72℃ (5s / kb, adjusted according to PCR product length), 30 cycles; 72℃ for 1min.

[0101] (2) Construction of tag peptide-isolated peptide-fluorescent protein particles:

[0102] HA (hemagglutinin) was used as the tag peptide, T2A as the separation peptide, and mCherry as the fluorescent protein. The DNA sequence of HA-T2A-mCherry was designed and ligated into the TetON plasmid. Primer pairs were designed to amplify the DNA sequence of HA-T2A-mCherry, resulting in the DNA sequence of tag peptide-separation peptide-fluorescent protein.

[0103] The DNA sequence of HA-T2A-mCherry is shown in SEQ ID NO.6:

[0104] TATCCGTATGATGTTCCGGATTATGCAGGATCCGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTGGAGGAGAATCCCGGCCCTTTAATTAACATGGTGAGCAAGGGCGAGGAGGATAACATGGCCATCATCAAGGAGTTCATGCGCTTCAAGGTGCACATGGAGGGCTCCGTGAACGGCCACGAGTTCGAGATCGAGGGCGAGGGCGAGGGCCGCCCCTACGAGGGCACCCAGACCGCCAAGCTGAAGGTGACCAAGGGTGGCCCCCTGCCCTTCGCCTGGGACATCCTGTCCCCTCAGTTCATGTACGGCTCCAAGGCCTACGTGAAGCACCCCGCCGACATCCCCGACTACTTGAAGCTGTCCTTCCCCGAGGGCTTCAAGTGGGAGCGCGTGATGAACTTCGAGGACGGCGGCGTGGTGACCGTGACCCAGGACTCCTCCCTGCAGGACGGCGAGTTCATCTACAAGGTGAAGCTGCGCGGCACCAACTTCCCCTCAGACGGCCCCGTAATGCAGAAGAAAACCATGGGCTGGGAGGCCTCCTCCGAGCGGATGTACCCCGAGGACGGCGCCCTGAAGGGCGAGATCAAGCAGAGGCTGAAGCTGAAGGACGGCGGCCACTACGACGCTGAGGTCAAGACCACCTACAAGGCCAAGAAGCCCGTGCAGCTGCCCGGCGCCTACAACGTCAACATCAAGTTGGACATCACCTCCCACAACGAGGACTACACCATCGTGGAACAGTACGAACGCGCCGAGGGCCGCCACTCCACCGGCGGCATGGACGAGCTGTACAAGTAA。

[0105] The primer sequences of HA-T2A-mCherry are as follows:

[0106] Forward primer F3: TATCCGTATGATGTTCCGGATT;

[0107] Reverse primer R3: TTACTTGTACAGCTCGTCCA.

[0108] (3) Synthesis of translation initiation / termination factor DNA

[0109] Taking translation initiation factor EIF3G and translation termination factor GSPT1 as examples, the DNA sequences of translation initiation factor EIF3G and translation termination factor GSPT1 were found through NCBI, and the above DNA sequences were synthesized. Primers were designed based on the above DNA sequences and amplified.

[0110] The DNA sequence of EIF3G is shown in SEQ ID NO.7:

[0111] ATGCCTACTGGAGACTTTGATTCGAAGCCCAGTTGGGCCGACCAGGTGGAGGAGGAGGGGGAGGACGACAAATGTGTCACCAGCGAGCTCCTCAAGGGGATCCCTCTGGCCACAGGTGACACCAGCCCAGAGCCAGAGCTACTGCCGGGAGCTCCACTGCCGCCTCCCAAGGAGGTCATCAACGGAAACATAAAGACAGTGACAGAGTACAAGATAGATGAGGATGGCAAGAAGTTCAAGATTGTCCGCACCTTCAGGATTGAGACCCGGAAGGCTTCAAAGGCTGTCGCAAGGAGGAAGAACTGGAAGAAGTTCGGGAACTCAGAGTTTGACCCCCCCGGACCCAATGTGGCCACCACCACTGTCAGTGACGATGTCTCTATGACGTTCATCACCAGCAAAGAGGACCTGAACTGCCAGGAGGAGGAGGACCCTATGAACAAACTCAAGGGCCAGAAGATCGTGTCCTGCCGCATCTGCAAGGGCGACCACTGGACCACCCGCTGCCCCTACAAGGATACGCTGGGGCCCATGCAGAAGGAGCTGGCCGAGCAGCTGGGCCTGTCTACTGGCGAGAAGGAGAAGCTGCCGGGAGAGCTAGAGCCGGTGCAGGCCACGCAGAACAAGACAGGGAAGTATGTGCCGCCGAGCCTGCGCGACGGGGCCAGCCGCCGCGGGGAGTCCATGCAGCCCAACCGCAGAGCCGACGACAACGCCACCATCCGTGTCACCAACTTGTCAGAGGACACGCGTGAGACCGACCTGCAGGAGCTCTTCCGGCCTTTCGGCTCCATCTCCCGCATCTACCTGGCTAAGGACAAGACCACTGGCCAATCCAAGGGCTTTGCCTTCATCAGCTTCCACCGCCGCGAGGATGCTGCGCGTGCCATTGCCGGGGTGTCCGGCTTTGGCTACGACCACCTCATCCTCAACGTCGAGTGGGCCAAGCCGTCCACCAAC。

[0112] The primer sequence of EIF3G is:

[0113] Upstream primer F4: ATGCCTACTGGAGACTTTGATTCGA;

[0114] Downstream primer R4: GTTGGTGGACGGCTTGGC.

[0115] The DNA sequence of GSPT1 is shown in SEQ ID NO.8:

[0116]

[0117] The primer sequence for GSPT1 is as follows:

[0118] Upstream primer F5: ATGGATCCGGGCAGTGGC;

[0119] Downstream primer R5: GTCTTTCTCTGGAACCAGTT.

[0120] Amplification system: 25 μL of 2×Phanta Flash Master Mix, 2 μL of upstream primer (10 μM), 2 μL of downstream primer (10 μM), 10 ng of template DNA (plasmid template), and bring the volume to 50 μL with ddH2O.

[0121] The amplification program was 98℃ for 30 seconds; 98℃ for 10 seconds, 60℃ for 5 seconds, 72℃ (5 seconds / kb, adjusted according to the length of the PCR product), 30 cycles; 72℃ for 1 minute.

[0122] (4) Constructing a stable expression vector:

[0123] The TetON plasmid was double-digested with restriction endonucleases BamHI and PmeI to obtain a linearized vector. The DNA sequences of the translation initiation factor EIF3G / translation termination factor GSPT1, the deaminase APOBEC1 / HumanADAR2dd from the deaminase plasmid, the DNA sequence of the tag peptide-separation peptide-fluorescent protein HA-T2A-mCherry, and the linearized vector were ligated using a homologous recombination kit (purchased from Yisheng Biotechnology (Shanghai) Co., Ltd., catalog number 10923ES20) to obtain a stable expression vector.

[0124] (5) Detection of fusion protein expression

[0125] Taking the translation termination factor GSPT1 and a transient expression system as an example, cells with transient expression of the fusion protein for 48 hours were collected, and the expression level of the fusion protein was detected by Western blot. Figure 2 As shown, after transient expression for 48 hours, the expression level of the GSPT1 fusion protein was significantly higher than that of the endogenous GSPT1 protein, indicating that the fusion protein can be normally expressed in cells.

[0126] Example 2: Sanger sequencing method for detecting RNA editing level

[0127] Upstream and downstream primers were designed to target and amplify the reporter gene. Taking HEK293T cells as an example, before stable expression using lentivirus, the GSPT1 fusion protein stable expression vector and lentiviral helper packaging plasmids (pMD2.G (Addgene, #12259) and psPAX2 (Addgene, #12260)) were co-transfected into HEK293T cells. Forty-eight hours after transfection, the culture medium supernatant was collected and mixed with polybrene (refer to the polybrene instructions, Yisheng Biotechnology (Shanghai) Co., Ltd., catalog number 40804ES76). This mixture was then incubated with HEK293T cells expressing either a normally translatable reporter gene or a reporter gene expressing a translationally blocked reporter gene for 48 hours. After 48 hours, the culture medium was replaced with one containing cyhalofop-p-ethyl (the final concentration of cyhalofop-p-ethyl in the medium was 10 μg / mL) for drug screening until all uninfected wild-type cells died. Cells successfully infected with the stable expression vector were obtained approximately 4–5 days later. After obtaining cells with stable expression, the fusion protein was induced to express for 48 hours with doxycycline at a final concentration of 1 μg / mL. RNA was then extracted and reverse transcribed (Yisheng Biotechnology (Shanghai) Co., Ltd., 11141ES10) to obtain cDNA. Using the cDNA as a template, the reporter gene cDNA was amplified by PCR using designed pre- and post-primers. The PCR products were then subjected to Sanger sequencing to detect the RNA editing level. Figure 3 As shown in the diagram, different colored peaks represent corresponding bases: A is green, T is red, C is blue, and G is black; peak area represents base content. It is clear that in the control group cells expressing ADAR2dd, no significant base editing was observed even when the reporter gene was translated normally. However, in the experimental group cells expressing GSPT1-ADAR2dd, significant base editing was observed only when the reporter gene was translated normally. This result indicates that base editing mediated by the GSPT1-ADAR2dd fusion protein accurately reflects mRNA translation.

[0128] Upstream primer sequence F6: CACATGGTCCTGCTGGAGTT;

[0129] Downstream primer sequence R6: CCAGAGAGACCCAGTACAAGC.

[0130] Reverse transcription pretreatment system: 1 μg total RNA, 3 μL 5×g DNA digester mix, and RNase-free H2O to a final volume of 15 μL.

[0131] Reverse transcription pretreatment procedure: 42℃ for 2 min.

[0132] Reverse transcription reaction system: 15 μL of reverse transcription pretreatment system and 5 μL of 4×Hifair III SuperMix plus.

[0133] Reverse transcription program: 25℃ for 5 min; 55℃ for 30 min; 85℃ for 5 min.

[0134] PCR amplification system: 25 μL of 2×Phanta Flash Master Mix, 2 μL of upstream primer (10 μM), 2 μL of downstream primer (10 μM), 5 μL of template cDNA (1:10 dilution), and bring the volume to 50 μL with ddH2O.

[0135] PCR amplification program: 98℃ for 30s; 98℃ for 10s, 60℃ for 5s, 72℃ for 5s, 30 cycles; 72℃ for 1min.

[0136] Example 3: Detection of RNA Editing Levels Using High-Throughput Sequencing

[0137] Taking GSPT1 and the transient expression system as an example, the transient expression vector prepared in Example 1 was transformed into HEK293T cells. After 48 hours of expression, the cells were collected for RNA extraction and high-throughput RNA sequencing library construction (for specific operations, refer to the instructions of the RNA extraction reagent and the high-throughput RNA sequencing library construction kit). The constructed library was sequenced using an Illumina high-throughput sequencer to obtain the raw FASTQ file. The Fastq files were processed using the following software to obtain the RNA editing level of each gene: STAR (v2.7.9a) (parameters: --alignEndsType EndToEnd --outBAMcompression 10 --outSAMtype BAMSortedByCoordinate --outSAMmapqUnique 60), Sailor (1.2.0) (parameter: reverse_stranded_library:true)), Biostar214299 (default parameters), and FeatureCounts (v2.0.1) (parameter: -pBCQ 30-s 2).

[0138] like Figure 4 As shown, in the comparison of base mutation levels, it can be seen that the base mutation level is significantly higher in cells expressing the GSPT1 fusion protein than in cells expressing only deaminase. Secondly, after classifying the detected genes into protein-coding genes and non-protein-coding genes, as shown... Figure 5As shown, it is easy to see that the RNA editing level of protein-coding genes is significantly higher than that of non-protein-coding genes, effectively reflecting the RNA translation level. Finally, taking the GSPT1-ADAR2dd fusion protein as an example, a correlation analysis was performed between the RNA editing level generated by the fusion protein and the RNA translation level obtained by ribosomal imprinting technology, such as... Figure 6 As shown, it is easy to see that the RNA editing level produced by the GSPT1-ADAR2dd fusion protein and the RNA translation level detected by ribosomal imprinting technology are significantly positively correlated, and the correlation is much higher than that between the ADAR2dd protein and ribosomal imprinting technology.

[0139] In summary, the translation initiation or termination factors in the fusion protein provided by this invention can more accurately and rapidly locate all RNAs that have begun or terminated translation in the cell, enabling the deaminase in the fusion protein to contact the RNA, achieving deamination editing, and thus leaving a marker of RNA editing. This functional fusion protein can be used to edit RNA in cell samples, and the RNA translation level of the cell sample can be quickly assessed by detecting the level of RNA editing.

[0140] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the patent protection scope of the present invention.

Claims

1. A fusion protein, characterized in that, Including the first and second proteins, The first protein includes either a translation initiation factor or a translation termination factor; The second protein includes a deaminase; The fusion protein is used to edit RNA at the start or end of translation.

2. The fusion protein as described in claim 1, characterized in that, The deaminase includes any one of ADAR, APOBEC, TadA, ADAR variant, APOBEC variant, and TadA variant.

3. The fusion protein as described in claim 1, characterized in that, The translation initiation factor includes any one of EIF1AX, EIF3G, EIF3D, and EIF4E.

4. The fusion protein as described in claim 1, characterized in that, The translation termination factor includes GSPT1.

5. The fusion protein as described in claim 1, characterized in that, The fusion protein was obtained through gene cloning.

6. A method for detecting RNA editing levels, characterized in that, Includes the following steps: S10. Obtain the cell sample to be tested; S20. Express the fusion protein in the cell sample to be tested, screen it, and obtain a cell sample to be tested enriched with the fusion protein, wherein the fusion protein is the fusion protein according to any one of claims 1 to 5; S30. Extract RNA from the cell sample containing the enriched fusion protein; S40. The editing level of the RNA is calculated based on the RNA of the cell sample to be tested containing the enriched fusion protein.

7. The method for detecting RNA editing levels as described in claim 6, characterized in that, S40 includes: S401. Sequencing the RNA to obtain the number of RNA reads (R) covering trusted editing sites and the number of RNA reads (E) in the test cell sample enriched with the fusion protein that have undergone trusted editing. S402. Calculate the RNA editing level of each gene in the test cell sample containing the enriched fusion protein according to the formula Z = E / R, where Z is the RNA editing level of each gene in the test cell sample containing the enriched fusion protein, E is the number of RNA reads in the test cell sample containing the enriched fusion protein that have undergone reliable editing, and R is the number of RNA reads in the test cell sample containing the enriched fusion protein that cover the reliable editing site.

8. The method for detecting RNA editing levels as described in claim 6, characterized in that, In step S20, the expression of the fusion protein in the cell sample to be tested includes transient transfection or stable transfection.

9. The method for detecting RNA editing levels as described in claim 7, characterized in that, In step S401, the methods for sequencing the RNA include high-throughput sequencing and Sanger sequencing.

10. The application of the method for detecting RNA editing level as described in any one of claims 6 to 9 in detecting RNA translation level.