Method for screening DNA library of intracellular mRNA sensor switch and application thereof
By designing a DNA library with degenerate base switch sequences and a multi-step screening method, the problem of limited mRNA sensor sequence selection was solved, achieving high-throughput screening and high-success-rate screening of mRNA sensor switch sequences, thus improving the sensitivity and specificity of detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING CANCER HOSPITAL PEKING UNIV CANCER HOSPITAL
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-28
AI Technical Summary
In existing technologies, the selection of sensor sequences targeting a single site on mRNA is limited, resulting in low throughput and poor performance, making it difficult to achieve high-throughput screening.
A DNA library containing a switch sequence with degenerate bases was designed. Through steps such as chemical synthesis, preparation of circular expression vector library, construction of minimal linear expression vector library, seamless cloning, RCA rolling circle amplification, telomerase digestion and cell transfection, high-throughput mRNA sensor switch sequences were screened.
It enables high-throughput screening of single mRNA sites, with a high success rate in identifying sensor switch sequences, providing more sensor switch sequence options and improving detection sensitivity and specificity.
Smart Images

Figure CN121344095B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of molecular biology and relates to a DNA library screening method for intracellular mRNA sensor switches and its application. Background Technology
[0002] Adenosine Deaminase Acting on RNA (ADAR) is a type of RNA editing enzyme that deaminates adenosine (A) to inosine (I) in double-stranded RNA, which is then recognized by the cell as guanine (G). This A-to-I editing plays a role in various physiological and pathological processes and can "rewrite" genetic information at the transcriptional level. In recent years, researchers have begun to explore the use of ADAR's RNA editing function for precise RNA manipulation. For example, some studies have used endogenous ADAR to correct and edit disease-related mRNAs for the treatment of genetic diseases such as Duchenne muscular dystrophy.
[0003] With the development of technologies such as single-cell sequencing, researchers have discovered that different cell types and states possess their own unique transcriptional profiles. However, the lack of tools to directly utilize these RNA differences for monitoring and manipulating cells has long been a bottleneck in biology and medicine. ADAR-mediated A-to-I editing, due to its precise RNA modification characteristics, is considered to have great potential in live-cell biosensing, and can be used to monitor intracellular transcriptional levels in real time. The concept of RNA sensors thus arises: by designing specific RNA elements and leveraging the editing activity of ADAR, the presence of specific RNA molecules within the cell can be converted into easily detectable signals, thereby enabling the detection and recording of cell states and specific regulation.
[0004] In recent years, a series of ADAR-based programmable RNA sensor technologies have been proposed. Their core principle typically involves "coupling" the signal to be detected (such as the presence of a certain RNA) with the expression of a reporter gene: the sensor RNA contains a guide sequence that can bind complementary to the target RNA, but it also contains a stop codon UAG switch (corresponding to TAG in the DNA sequence), and a downstream payload sequence encoding a reporter protein. In the absence of target RNA, the sensor RNA remains in a state of closed downstream translation (switch off), containing the stop codon, and does not produce a reporter protein. Once the target RNA is present and binds to the sensor RNA to form a double-stranded structure, the ADAR enzyme edits the mismatch site A (usually designed so that A in the stop codon UAG forms a mismatch with C in the target CCA, increasing the probability of this site being edited), transforming it into a sequence that encodes an amino acid (UAG becomes UIG), thereby removing the original stop codon's inhibition of translation (switch on), and generating the downstream payload protein signal. This design makes payload protein production directly related to the presence of the target RNA, completing the transduction process from RNA detection to signal amplification.
[0005] Currently, there are two representative ADAR sensor systems:
[0006] 1. RADAR (live RNA sensing using adenosine deaminases acting on RNA) technology:
[0007] Live RNA sensing using adenosine deaminases acting on RNA (RADAR): Reported by Gao et al. from Stanford University around 2022 (a study in *Nature Biotechnology* 2022, https: / / www.nature.com / articles / s41587-022-01493-x, referred to as "RADAR technology"), this technology emphasizes design improvements that remove sequence restrictions and enhance applicability. Traditional linear sensors require the sensor RNA to form an AC mismatch (UAG–CCA) center with the target RNA for efficient ADAR editing, which limits the range of selectable target sequences. RADAR, by optimizing the RNA element structure, removes the dependence on CCA complementarity for specific sequences, allowing for GCA, UCA, and CAA, providing greater flexibility in target sequence selection. Design software is provided based on its established rules for sensor design, where, except for the mismatch between the target site and the UAG, all other regions are fully complementary. UAG is edited into UIG only when specific RNA signals are present in a particular cell type / state, thereby removing translational repression of downstream proteins and thus achieving higher specificity and sensitivity. This can be used for applications such as precisely eliminating diseased cells or regulating immune cells.
[0008] However, RADAR technology can only design one sensor sequence for a single site on mRNA, and the availability of this sequence needs to be validated. The detection method relies on flow cytometry to detect differences. Software designed using this rule has a limited variety of sensor sequences, low throughput, and poor performance.
[0009] 2. Hairpin RNA sensor technology:
[0010] In 2025, Qin et al. from Zhejiang University of Technology reported a novel hairpin RNA sensor. Inspired by the hairpin structure of natural pre-miRNA, this sensor initially exhibits an "inactive" conformation, lacking a complete double-stranded editing site. Only in the presence of a target molecule does the hairpin sensor undergo a conformational change to form a stable hairpin double-stranded region, containing a central UAG terminator site that can be edited into a UIG by ADAR. After editing, the translational repression of downstream reporter proteins (such as fluorescent proteins) is relieved, thereby indicating the presence of the target. This hairpin design increases the stability of the double-stranded region through the pairing or mismatch of adjacent bases, significantly improving ADAR recruitment and editing efficiency. However, this technology does not provide a universal design and screening method and has low throughput.
[0011] Therefore, developing a high-throughput screening method for RNA sensors that target and sense a single site on mRNA has become an urgent problem to be solved. Summary of the Invention
[0012] To address the issue that only one sequence can be selected for a single sensor site on mRNA, the present invention aims to provide a DNA library screening method for intracellular mRNA sensor switches, which can provide more sensor switch sequences for a single site on intracellular mRNA.
[0013] Another object of the present invention is to provide an application of the above-described screening method.
[0014] To achieve the above objectives, the present invention provides a DNA library screening method for intracellular mRNA sensor switching, comprising the following steps:
[0015] 1) Design a DNA library containing a switch sequence with degenerate bases.
[0016] The switch sequence consists of a 5' end sequence, a middle region sequence, and a 3' end sequence; the 5' end sequence and the 3' end sequence are both 84 nt in length, and the 5' end sequence and the 3' end sequence are complementary to the target mRNA sequence, respectively.
[0017] The intermediate region sequence is a variable sequence of 27 nt in length. This intermediate region sequence is represented by 9 triplet codons from the 5' end to the 3' end as 5'-FOZ FOZ FOZ FOZ TAG FOZ FOZ FOZ FOZ FOZ-3', where F is the first base of the codon, O is the second base of the codon, and Z is the third base of the codon.
[0018] After the 5' end sequence of the 84nt and the 3' end sequence of the 84nt are complementary to the coding strand DNA sequence of the target mRNA, the sequence position of the middle region is matched with the sequence on the coding strand DNA of the target mRNA from the 3' end to the 5' end, which is represented in the form of 9 triplets as 3'-foz foz foz foz acc foz foz foz foz-5'.
[0019] Where f is A, F is T; f is T, F is A; f is C, F is G; f is G, F is C;
[0020] When o is A, O is K; when o is T, O is N; when o is C, O is B; when o is G, O is N.
[0021] When z is A, Z is K; when z is T, Z is N; when z is C, Z is B; when z is G, Z is N.
[0022] Among them, the degenerate base K is G or T; the degenerate base S is G or C; the degenerate base N is A, T, C or G; and the degenerate base B is G, T or C.
[0023] The DNA library containing the switch sequence of degenerate bases was chemically synthesized.
[0024] 2) Prepare a circular expression vector library;
[0025] 3) Prepare a minimal linear expression vector library;
[0026] 4) Seamless cloning;
[0027] 5) RCA rolling circle amplification;
[0028] 6) Telomerase digestion;
[0029] 7) Cell transfection;
[0030] 8) RNA extraction and reverse transcription;
[0031] 9) Library construction for next-generation sequencing using cDNA as a template;
[0032] 10) Next-generation sequencing;
[0033] 11) Analyze the data and compare them to obtain the DNA library of the mRNA sensor switch.
[0034] Furthermore, in step 1), if the 5' end sequence and the 3' end sequence of the 84nt switch sequence contain a stop codon TAG, TGA, or TAA that belongs to the same reading frame as the TAG stop codon of the 27nt middle region sequence, then the corresponding TAG in the 5' end sequence and the 3' end sequence is modified to TCG, TGA to GGA, and TAA to TCA.
[0035] Furthermore, the method for preparing a circular expression vector library in step 2) is as follows: a first reporter gene, a first 2A peptide sequence, a second 2A peptide sequence, and a second reporter gene are inserted sequentially in the 5'-3' direction between the promoter and PolyA of the initial mammalian expression plasmid vector backbone to obtain a circular plasmid vector; the circular plasmid vector is amplified and purified; the switch sequence DNA library containing degenerate bases is amplified and purified by PCR, and then the purified switch sequence DNA library containing degenerate bases is inserted between the first 2A peptide sequence and the second 2A peptide sequence of the circular plasmid vector to obtain a circular expression vector library.
[0036] Furthermore, the first reporter gene is mCherry, eGFP, or SEAP; the second reporter gene is mCherry, eGFP, or SEAP; and the first reporter gene and the second reporter gene located on the same initial mammalian expression plasmid vector backbone are different; wherein, the nucleotide sequence of mCherry is shown in Seq ID No.1, the nucleotide sequence of eGFP is shown in Seq ID No.2, and the nucleotide sequence of SEAP is shown in Seq ID No.3. When mCherry, eGFP, or SEAP is selected as the first reporter gene, the stop codons TAG, TAA, or TGA at the end of its sequence are removed. When mCherry, eGFP, or SEAP is selected as the second reporter gene, the start codon ATG at the beginning of the sequence is removed.
[0037] Furthermore, the first 2A peptide sequence is selected from the P2A sequence, T2A sequence, E2A sequence, or F2A sequence; the second 2A peptide sequence is selected from the P2A sequence, T2A sequence, E2A sequence, or F2A sequence; wherein the nucleotide sequence of the P2A sequence is shown in Seq ID No. 4, the nucleotide sequence of the T2A sequence is shown in Seq ID No. 5, the nucleotide sequence of the E2A sequence is shown in Seq ID No. 6, and the nucleotide sequence of the F2A sequence is shown in Seq ID No. 7.
[0038] Furthermore, the method for preparing the minimized linear expression vector library in step 3) is as follows: design a forward primer and a reverse primer, wherein the forward primer introduces a telomerase restriction site and is located upstream of the 5' end of the promoter sequence; the reverse primer introduces a telomerase restriction site and is located downstream of the 3' end of the PolyA sequence; use the forward primer and the reverse primer to perform PCR on the circular expression vector library to obtain a minimized linear expression vector library.
[0039] Furthermore, the seamless cloning method in step 4) is to perform seamless cloning on the minimized linear expression vector library to obtain a minimized circular expression vector library with connected ends; the RCA rolling circle amplification method in step 5) is to perform RCA rolling circle amplification on the minimized circular expression vector library to obtain a minimized linear expression vector library with ultra-long periodic repetition.
[0040] Furthermore, step 6) involves telomerase digestion: the ultra-long periodically repeated minimized linear expression vector library is digested with telomerase to obtain a minimized linear expression vector library with covalently closed ends; step 7) involves cell transfection: the minimized linear expression vector library with covalently closed ends is transfected into cells expressing the target mRNA and cells not expressing the target mRNA for 72 hours.
[0041] Furthermore, the method for extracting RNA and reverse transcription in step 8) is as follows: RNA is extracted from cells expressing the target mRNA and cells not expressing the target mRNA, and the corresponding cDNA is obtained by reverse transcription.
[0042] The present invention also provides the application of the above-described screening method in screening DNA libraries of sensor switches targeting MAGEA4 mRNA.
[0043] The beneficial effects of this invention are as follows:
[0044] This invention provides a method and application for screening DNA library sequences of intracellular sensor switches targeting mRNA. This screening method can be a general method for screening DNA library sequences of single-site sensor switches of intracellular mRNA. It can construct corresponding sensor switch DNA library sequences based on the single-site sequence of the target mRNA. The screened sensor switch DNA sequences have high throughput and a high success rate of candidate libraries. Attached Figure Description
[0045] Figure 1 This is a schematic flowchart of the DNA library screening method for the intracellular mRNA sensor switch provided by the present invention.
[0046] Figure 2 The plasmid map of the circular plasmid vector prepared in Example 1 of this invention.
[0047] Figure 3 The structural diagram of the minimized linear expression vector prepared in Example 1 of the present invention.
[0048] Figure 4 Agarose gel electrophoresis image of the minimized linear carrier prepared in Example 1 of the present invention.
[0049] Figure 5 The image shows the RCA reaction and the agarose gel electrophoresis image of the product after telomerase digestion in Example 1 of this invention.
[0050] Figure 6 Agarose gel electrophoresis image of the purified minimal linear expression vector with terminal covalent closure prepared in Example 1 of this invention.
[0051] Figure 7 Agarose gel electrophoresis image of the sequencing sample prepared in Example 1 of this invention.
[0052] Figure 8 Fluorescence results of cells after transfection with expression vectors constructed from specific sequences and control sequences in the DNA library prepared in Example 1 of this invention. Detailed Implementation
[0053] The embodiments of the present invention will now be described in detail and comprehensively so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby providing a clearer and more explicit definition of the scope of protection of the present invention.
[0054] like Figure 1 The present invention illustrates a DNA library screening method for an intracellular mRNA sensor switch, comprising the following steps:
[0055] 1) Design a DNA library containing a switch sequence with degenerate bases.
[0056] The switch sequence consists of a 5' end sequence, a middle region sequence, and a 3' end sequence; the 5' end sequence and the 3' end sequence are both 84 nt in length, and the 5' end sequence and the 3' end sequence are complementary to the target mRNA sequence, respectively.
[0057] The intermediate region sequence is a variable sequence of 27 nt in length. This intermediate region sequence is represented by 9 triplet codons from the 5' end to the 3' end as 5'-FOZ FOZ FOZ FOZ TAG FOZ FOZ FOZ FOZ FOZ-3', where F is the first base of the codon, O is the second base of the codon, and Z is the third base of the codon.
[0058] FOZ represents the common form of triplet codons in the same reading frame as the TAG stop codon. The middle region sequence contains a total of 4 FOZs, 1 TAG stop codon, and 4 FOZs, for a total of 9 triplets, with a total length of 27nt.
[0059] After the 5' end sequence and 3' end sequence of the 84nt mRNA are complementary to the coding strand DNA sequence of the target mRNA, the sequence position of the intermediate region is matched with the corresponding sequence on the coding strand DNA of the target mRNA, which is represented in the form of 9 triplets from the 3' end to the 5' end as 3'-foz foz foz foz acc foz foz foz foz-5'. Here, matching refers to a one-to-one correspondence at position. The matched bases can be complementary sequences or mismatched sequences.
[0060] Where f is A, F is T; f is T, F is A; f is C, F is G; f is G, F is C; o is A, O is K; o is T, O is N; o is C, O is B; o is G, O is N; z is A, Z is K; z is T, Z is N; z is C, Z is B; z is G, Z is N.
[0061] Among them, the degenerate base K is G or T; the degenerate base S is G or C; the degenerate base N is A, T, C or G; and the degenerate base B is G, T or C.
[0062] When other stop codons TAG, TGA, or TAA appear in the same reading frame as the TAG stop codon in the 5' end sequence and 3' end sequence of the 84nt switch sequence, respectively, the corresponding TAG in the 5' end sequence and the 3' end sequence is modified to TCG, TGA to GGA, and TAA to TCA.
[0063] Here, f, o, and z correspond to the bases complementary to the F, O, and Z positions in the switch sequence, respectively, and acc (3'-5' direction) is the complementary sequence of the TAG (5'-3' direction) stop codon.
[0064] The DNA library containing the switch sequence of degenerate bases was chemically synthesized.
[0065] 2) Preparation of a circular expression vector library
[0066] A circular plasmid vector was obtained by inserting the first reporter gene, the first 2A peptide sequence, the second 2A peptide sequence, and the second reporter gene sequentially in the 5'-3' direction between the promoter and PolyA of the initial mammalian expression plasmid vector backbone; the circular plasmid vector was then amplified and purified.
[0067] The DNA library containing the switch sequence with degenerate bases was amplified and purified by PCR. Then, the purified DNA library containing the switch sequence with degenerate bases was inserted between the first 2A peptide sequence and the second 2A peptide sequence of the circular plasmid vector to obtain a circular expression vector library.
[0068] The first reporter gene is mCherry, eGFP, or SEAP; the second reporter gene is mCherry, eGFP, or SEAP; and the first reporter gene and the second reporter gene located on the same initial mammalian expression plasmid vector backbone are different; the nucleotide sequence of mCherry is shown in Seq ID No. 1, the nucleotide sequence of eGFP is shown in Seq ID No. 2, and the nucleotide sequence of SEAP is shown in Seq ID No. 3. When mCherry, eGFP, or SEAP is selected as the first reporter gene, the stop codons TAG, TAA, or TGA at the end of the sequence are removed; when mCherry, eGFP, or SEAP is selected as the second reporter gene, the start codon ATG at the beginning of the sequence is removed.
[0069] The first 2A peptide sequence is selected from the P2A, T2A, E2A, or F2A sequences; the second 2A peptide sequence is selected from the P2A, T2A, E2A, or F2A sequences; the nucleotide sequence of the P2A sequence is shown in Seq ID No. 4, the nucleotide sequence of the T2A sequence is shown in Seq ID No. 5, the nucleotide sequence of the E2A sequence is shown in Seq ID No. 6, and the nucleotide sequence of the F2A sequence is shown in Seq ID No. 7.
[0070] 3) Preparation of a minimal linear expression vector library
[0071] Design forward and reverse primers, wherein the forward primer introduces a telomerase cleavage site and is located upstream of the 5' end of the promoter sequence; the reverse primer introduces a telomerase cleavage site and is located downstream of the 3' end of the PolyA sequence.
[0072] The forward primer is located within 500 bp upstream of the 5' end of the promoter sequence, and the reverse primer is located within 500 bp downstream of the 3' end of the PolyA sequence. If there is an enhancer sequence preceding the promoter sequence, it is located within 500 bp upstream of the 5' end of the enhancer.
[0073] Using the forward and reverse primers, PCR was performed on the circular expression vector library to obtain a minimized linear expression vector library, such as... Figure 3 The structure shown.
[0074] The minimized linear expression vector library retains only the portion from the promoter to the end of the PolyA terminus on the original mammalian expression plasmid vector backbone, removing redundant elements such as resistance genes from the original mammalian expression plasmid vector backbone. If there is an enhancer sequence preceding the promoter sequence, the enhancer can be retained to improve the transcription efficiency of the minimized linear expression vector library.
[0075] 4) Seamless cloning
[0076] Seamless cloning of the minimal linear expression vector library yields a minimal circular expression vector library with connected ends.
[0077] 5) RCA rolling circle amplification
[0078] RCA rolling circle amplification was performed on the minimized circular expression vector library to obtain a minimized linear expression vector library with ultra-long periodic repetitions.
[0079] 6) Telomerase digestion
[0080] The minimal linear expression vector library with ultra-long periodic repetitions was digested with telomerase to obtain a minimal linear expression vector library with covalently closed ends.
[0081] 7) Cell transfection
[0082] The minimized linear expression vector library with covalently closed ends was transfected into cells expressing the target mRNA and cells not expressing the target mRNA for 72 hours.
[0083] 8) RNA extraction and reverse transcription
[0084] RNA was extracted from cells expressing the target mRNA and cells not expressing the target mRNA, and the corresponding cDNA was obtained by reverse transcription.
[0085] 9) Library construction using cDNA as a template for next-generation sequencing.
[0086] 10) Next-generation sequencing.
[0087] 11) Analyze the data and compare them to obtain the DNA library of the mRNA sensor switch.
[0088] This invention designs and synthesizes 84nt sequences complementary to the target mRNA at both ends of the switch sequence, with a 12nt sequence containing degenerate bases in the middle, plus a TAG and another 12nt sequence containing degenerate bases (totaling 27nt). By introducing degenerate bases, the library size of the sensor is significantly increased.
[0089] The sensor sequence described above is a DNA sequence. The corresponding RNA sequence after in vitro transcription or intracellular transcription will have all Ts replaced with Us.
[0090] Materials and equipment:
[0091] 1. Q5 High-Fidelity Enzyme (Q5) ® The Hot Start High-Fidelity 2X Master Mix was purchased from New England Biolabs, catalog number M0494S. This Q5 high-fidelity enzyme is used for short fragment library amplification.
[0092] 2. Q5 High-Fidelity Enzyme (NEBNext) ® The High-Fidelity 2X PCR Master Mix was purchased from New England Biolabs, catalog number M0541S. This Q5 high-fidelity enzyme is used for amplification of long-fragment vectors.
[0093] 3. The PCR instrument was an Eppendorf Mastercycler. ® X50-PCR thermal cycler.
[0094] 4. Agarose was purchased from Beijing TransGen Biotechnology Co., Ltd., item number: GS201-01.
[0095] 5. Electrophoresis buffer (50X Tris acetate EDTA buffer (TAE)) was purchased from Shanghai Beyotime Biotechnology Co., Ltd., product number: ST716.
[0096] 6. The electrophoresis apparatus (PowerPac™ HC Power Supply) was purchased from Bio-Rad, product number: 1645052.
[0097] 7. The agarose gel DNA rapid extraction kit was purchased from Jiangsu Kangwei Century Biotechnology Co., Ltd., catalog number: CW2303M.
[0098] 8. NEBuilder ® High-fidelity DNA assembly premix (NEBuilder) ® The HiFi DNA Assembly MasterMix was purchased from New England Biolabs, catalog number: E2621S.
[0099] 9. The Phi29 DNA polymerase kit was purchased from Beijing TransGen Biotech Co., Ltd., catalog number: LP101-01.
[0100] 10. Telomerase was purchased from Yisheng Biotechnology (Shanghai) Co., Ltd., TelN Protelomerase 5U / μL, catalog number: 14540ES72.
[0101] 11. The F-Vector, R-Vector, F-Library, and R-Library sequences were synthesized by Suzhou Genewiz Biotechnology Co., Ltd.
[0102] 12. The DNA molecular weight standard Trans5K Marker was purchased from Beijing TransGen Biotech Co., Ltd., catalog number: BM141-01.
[0103] 13. Cell culture dish (Corning® BioCoat™ Type I collagen-coated 100 mm TC-treated culture dish) purchased from Corning, catalog number: 354450.
[0104] 14. DMEM medium (Gibco™ DMEM (high sugar)) was purchased from Gibco, catalog number: 11965092.
[0105] 15. FBS: Fetal Bovine Serum, Premium, purchased from Thermo Fisher, item number: A5670201.
[0106] 16. A375 cells and 293T cells were purchased from the Cell Bank of the Chinese Academy of Sciences.
[0107] 17. The transfection reagent jetPRIME® DNA / siRNA was purchased from Sartorius Leper (Shanghai) Trading Co., Ltd., item number: 101000015.
[0108] 18. The rapid total RNA extraction kit for cells / tissues was purchased from Suzhou NCMBiotech Co., Ltd. (product number: M5105).
[0109] 19. RNase Inhibitor was purchased from Shanghai Beyotime Biotechnology Co., Ltd., product number: R0102.
[0110] 20. DNase I was purchased from Shanghai Beyotime Biotechnology Co., Ltd., product number: D7073.
[0111] 21. SuperScript™ II reverse transcriptase was purchased from Thermo Fisher, catalog number: 18064014.
[0112] 22. The DNA Marker I molecular weight standard was purchased from Tiangen Biotech (Beijing) Co., Ltd., catalog number: MD101-01.
[0113] 23. The vector pcDNA3.1-CMV-mCherry-T2A-P2A-eGFP was synthesized by Beijing Maijin Biotechnology Co., Ltd.
[0114] Example 1
[0115] MAGEA4, belonging to the melanoma-associated antigen family A4 and the cancer-testis antigen, is an important target in tumor research. In normal tissues, it is expressed only in testicular germ cells and silenced in somatic cells, but it is aberrantly activated in tumor tissues. MAGEA4 is highly expressed in various malignant tumors, making it an ideal target for cancer therapy. Therefore, screening sensor libraries targeting MAGEA4 mRNA is of significant value for designing novel targeted cancer therapies.
[0116] 1) Design a DNA library sequence targeting the sensing switch of the MAGEA4 gene's intracellular mRNA:
[0117] The target coding chain sequence of MAGEA4 is shown in Seq ID No. 8: 5'-TGA GTT GCA GCC AGG GCTGTG GGG AAG GGG CAG GGC TGG GCC AGT GCA TCT AAC AGC CCT GTG CAG CAG CTT CCCTTG CCT CGT GTA ACA TGA GGC CCA TTC TTC ACT CTG TTT GAA GAA AAT AGTCAG TGT TCT TAG TAG TGG GTT TCT ATT TTG TTG GAT GAC TTG GAG ATT TAT CTC TGTTTC CTT TTA CAA-3'.
[0118] The complementary strand of the target coding chain sequence of MAGEA4 mentioned above is shown in Seq ID No. 9:
[0119] 5'-TTG TAA AAG GAA ACA GAG ATA AAT CTC CAA GTC ATC CAA CAA AAT AGAAAC CCA CTA CTA AGA ACA CTG ACT ATT TTC TTC AAA CAG AGT GAA GAA TGG GCC TCA TGT TAC ACG AGG CAA GGG AAG CTG CTG CAC AGG GCT GTT AGA TGC ACT GGCCCA GCC CTG CCC CTT CCC CAC AGC CCT GGC TGC AAC TCA-3'.
[0120] The switch sequence designed for the MAGEA4 gene is shown in Seq ID No. 10:
[0121] 5'-TTG TCA AAG GAA ACA GAG ATA AAT CTC CAA GTC ATC CAA CAA AAT AGAAAC CCA CTA CTA AGA ACA CTG ACT ATT TTC TTC AAA CNB ABK GNN GNN TAG GNN TNN TNN TNN ACG AGG CAA GGG AAG CTG CTG CAC AGG GCT GTT AGA TGC ACT GGCCCA GCC CTG CCC CTT CCC CAC AGC CCT GGC TGC AAC TCA-3'.
[0122] In the target coding sequence of MAGEA4 shown in Seq ID No. 8, the bolded italicized CCA is a mismatch site corresponding to the TAG in the 27nt middle region of the switch sequence. This site corresponds to the bolded italicized TGG in the complementary strand of the target coding sequence of MAGEA4 shown in Seq ID No. 9. In the switch sequence designed for the MAGEA4 gene shown in this application (Seq ID No. 10), this TGG site is mutated to the stop codon TAG (bold italicized) to achieve sensor shutdown. The underlined TTA in the target coding sequence of MAGEA4 shown in Seq ID No. 8 corresponds to the underlined TAA in the complementary strand of the target coding sequence of MAGEA4 shown in Seq ID No. 9. TAA is another stop codon generated by the 5' end sequence in the same reading frame as the TAG in the 27nt middle region sequence. This results in multiple stop codons in the sequence, leading to premature translation termination and directly blocking the switching effect regulated by the TAG in the 27nt middle region sequence. Therefore, in the switch sequence designed for the MAGEA4 gene as shown in Seq ID No. 10 of this application, TAA is mutated to TCA.
[0123] Suzhou Genewiz Biotechnology Co., Ltd. was commissioned to chemically synthesize the above-mentioned 195bp switch sequence DNA library containing degenerate bases as shown in Seq ID No. 10, resulting in a total switch sequence DNA library of 1 OD.
[0124] 2) Preparation of a circular plasmid expression vector library
[0125] A circular plasmid vector was obtained by inserting the first reporter gene, the first 2A peptide sequence, the second 2A peptide sequence, and the second reporter gene sequentially in a 5'-3' direction between the promoter and PolyA of the initial mammalian expression plasmid vector backbone.
[0126] In this embodiment, the first reporter gene is selected as the DNA sequence encoding mCherry as shown in Seq ID No.1, and the expressed protein mCherry can emit red fluorescence. The second reporter gene is selected as the DNA sequence encoding eGFP as shown in Seq ID No.2, and the expressed protein eGFP can emit green fluorescence. The first 2A peptide sequence is selected as the coding sequence of T2A as shown in Seq ID No.5, and the second 2A peptide sequence is selected as the coding sequence of P2A as shown in Seq ID No.4.
[0127] Specifically, the initial mammalian expression plasmid vector backbone sequence for constructing the vector is selected from the pcDNA3.1-CMV sequence, mainly utilizing the promoter CMV and PolyA tail. Other mammalian expression plasmid vector backbones with promoters and PolyA tails can also be used.
[0128] The sequence shown in Seq ID No. 11 was constructed using pcDNA3.1-CMV and named pcDNA3.1-CMV-mCherry-T2A-P2A-eGFP. The plasmid map of the vector is shown below. Figure 2 As shown, the vector was synthesized by Beijing Maijin Biotechnology Co., Ltd.
[0129] The switch sequence is inserted between T2A and P2A in the vector pcDNA3.1-CMV-mCherry-T2A-P2A-eGFP. This can be achieved using any existing technology, but in this embodiment, it is achieved using seamless cloning technology.
[0130] The seamless cloning process was implemented as follows: PCR amplification of linearized pcDNA3.1-CMV-mCherry-T2A-P2A-eGFP was performed using Q5 high-fidelity enzyme (M0494S) to obtain the linearized vector 5'-P2A-eGFP-Vector-mCherry-T2A-3', as shown in Seq ID No. 12. The amplification primers were F-Vector and R-Vector. The F-Vector sequence is shown in Seq ID No. 13: 5'-GGTTCAGGTTCAAGCTCAGGATCCGGAGCAACAAACTTC-3', and the R-Vector sequence is shown in Seq ID No. 14: 5'-ACTGCTTCTAGACCCACCAGGGCCGGGATTCTCCTC-3'. The reaction conditions were set according to the Q5 high-fidelity enzyme (M0494S) manufacturer's instructions.
[0131] PCR amplification of the switch sequence DNA library was performed using Q5 high-fidelity enzyme (M0541S). The amplification primers were as follows: F-Library sequence (Seq ID No. 15): 5'-GGTGGGTCTAGAAGCAGTTTGTCAAAGGAAACAGAGATAAATCTC-3'; R-Library sequence (Seq ID No. 16): 5'-TGAGCTTGAACCTGAACCTGAGTTGCAGCCAGGGCTG-3'. The reaction conditions were set according to the Q5 high-fidelity enzyme (M0541S) manufacturer's instructions.
[0132] After PCR, the 5'-P2A-eGFP-Vector-mCherry-T2A-3' and switch sequence DNA libraries were subjected to 4% agarose gel electrophoresis at 130V. The 5'-P2A-eGFP-Vector-mCherry-T2A-3' and switch sequence DNA libraries were purified by gel extraction using an agarose gel DNA rapid recovery kit.
[0133] The purified 5'-P2A-eGFP-Vector-mCherry-T2A-3' and switch sequence DNA library were mixed at a molar ratio of 1:1 and then processed according to NEBuilder. ® Following the instructions for the high-fidelity DNA assembly premix, seamless cloning was performed to form a complete circular expression vector library 5'-Vector-mCherry-T2A-(switch sequence DNA library)-P2A-eGFP-Vector-3', with the sequence shown in Seq ID No. 17.
[0134] 3) Preparation of a minimal linear expression vector library
[0135] PCR amplification of the circular expression vector library 5'-Vector-mCherry-T2A-(switch sequence DNA library)-P2A-eGFP-Vector-3' was performed using Q5 high-fidelity enzyme (M0494S). Forward and reverse primers were designed using 5'Vector-mCherry-T2A-(switch sequence DNA library)-P2A-eGFP-Vector-3' as a template. The forward primer F-TEL introduced a telomerase restriction site and was located upstream of the 5' end of the CMV enhancer sequence upstream of the CMV promoter. Since the CMV enhancer and CMV promoter jointly play a role in transcription initiation, the CMV promoter is assumed to include the CMV enhancer region. The reverse primer R-TEL introduced a telomerase restriction site and was located downstream of the 3' end of the PolyA sequence. The sequence of the forward primer F-TEL is shown in Seq ID. Seq ID No. 18 shows the sequence: 5'-CATTATACGCGCGTATAATGGACTATTGTGTGCTGATATCTCCCGATCCCCTATGGT-3', and the reverse primer R-TEL sequence is shown in Seq ID No. 19: 5'-CATTATACGCGCGTATAATGGGCAATTGTGTGCTGATAAGCTGGTTCTTTCCGCCTC-3'.
[0136] Using the forward primer F-TEL and the reverse primer R-TEL, PCR was performed on the circular expression vector library to obtain a minimized linear expression vector library 5'-CMV promoter-mCherry-T2A-(switch sequence DNA library)-P2A-eGFP-PolyA-3', the sequence of which is shown in Seq ID No. 20. Agarose gel electrophoresis was performed on the minimized linear expression vector library 5'-CMV promoter-mCherry-T2A-(switch sequence DNA library)-P2A-eGFP-PolyA-3', and the results are shown below. Figure 4 As shown, lane 1 is the Trans5K Marker, and lane 2 is the 5'-CMV promoter-mCherry-T2A-(switch sequence DNA library)-P2A-eGFP-PolyA-3'.
[0137] The 3756 bp band in lane 2 was purified using an agarose gel DNA rapid recovery kit, using the 5'-CMV promoter-mCherry-T2A-(switch sequence DNA library)-P2A-eGFP-PolyA-3'.
[0138] The minimized linear expression vector library retains only the promoter (including enhancer) to PolyA portion of the original mammalian expression plasmid vector backbone, and removes redundant elements such as resistance genes from the original mammalian expression plasmid vector backbone, so that this DNA library can be transcribed into only complete mRNA sensors in mammalian cells.
[0139] 4) Seamless cloning
[0140] After purification, the 5'-CMV promoter-mCherry-T2A-(switch sequence DNA library)-P2A-eGFP-PolyA-3' of the minimized linear expression vector library was seamlessly cloned to obtain the minimized circular expression vector library with the ends linked together.
[0141] 5) RCA rolling circle amplification
[0142] The minimized circular expression vector library was subjected to RCA rolling circle amplification using the Phi29 DNA polymerase kit and the kit's included random primers. The DNA was then purified using an agarose gel DNA rapid recovery kit, followed by agarose gel electrophoresis. The results are shown below. Figure 5 As shown, lane 1 is the Trans5K Marker, lane 2 is the RCA reaction product, and the band represents a minimal linear expression vector library with ultra-long periodic repetitions.
[0143] 6) Telomerase digestion
[0144] The ultra-long periodic repetition minimized linear expression vector library was digested with TELN telomerase and then subjected to agarose gel electrophoresis. The results are as follows: Figure 5 As shown, lane 1 band represents the Trans5K marker, lane 2 band represents the RCA reaction product, and lane 3 band represents the product after TELN telomerase digestion. From... Figure 5 It can be seen that the RCA reaction product is greater than 5000 bp, and after telomerase digestion, a homogeneous product of 3736 bp is obtained. The difference between this 3736 bp product and the 3756 bp 5'-CMV promoter-mCherry-T2A-(switch sequence DNA library)-P2A-eGFP-PolyA-3' shown in Seq ID No. 20 is that, during seamless cloning to obtain the minimized circular expression vector library with connected ends, a total of 20 bp of repeated telomerase sequences at both ends were removed, leaving only one telomerase site. The reduced 20 bp does not affect the sequence from the CMV promoter to PolyA in the minimized linear expression vector library, and telomerase digestion simultaneously causes both ends to covalently close, thus obtaining a minimized linear expression vector library with covalently closed ends.
[0145] right Figure 5 The 3736bp minimal linear expression vector library with covalently closed ends in lane 3 of the middle agarose gel was purified and then subjected to agarose gel electrophoresis again. The results are as follows. Figure 6 As shown, lane 1 contains the Trans5K Marker, and lane 2 contains a 3736 bp minimal linear expression vector library with covalently closed ends. From... Figure 6 It can be seen that successful agarose gel recovery improves purity, and improved purity is beneficial to improving subsequent transfection efficiency.
[0146] 7) Cell transfection
[0147] A375 cells that highly express MAGEA4 mRNA and 293T cells that do not express MAGEA4 mRNA were selected, with 293T cells serving as a negative control.
[0148] Using cell culture dishes, DMEM (10% FBS), 37℃, 5% CO2 culture environment, A375 cells and 293T cells were cultured until the cell confluence was about 80%.
[0149] The minimized linear expression vector library with covalently closed ends was transfected into 293T and A375 cells for 72 hours;
[0150] The procedure was as follows: A375 cells and 293T cells were transfected with 10 μg of purified 3736 bp minimal linear expression vector library with covalently closed ends using the transfection reagent jetPRIME® DNA / siRNA. The cells were cultured for 72 hours in a 37°C, 5% CO2 environment after 24 hours of transfection by changing the medium.
[0151] 8) RNA was extracted from 293T and A375 cells and reverse transcribed to obtain cDNA.
[0152] A375 cells and 293T cells were collected separately. Genomic DNA was removed according to the instructions of the rapid total RNA extraction kit for cells / tissues. After reacting with RNase inhibitor and DNase I according to the instructions, RNA was purified again according to the instructions of the rapid total RNA extraction kit for cells / tissues to obtain total RNA from A375 cells and total RNA from 293T cells. Then, RNase inhibitor was added to the total RNA from A375 cells and total RNA from 293T cells again.
[0153] Following the instructions for SuperScript™ II reverse transcriptase, 5 μg of total RNA was taken from each of the A375 and 293T total RNA samples, and gene-specific primers GSP were added to a final concentration of 10 pmol / μL for reverse transcription, yielding 400 μL of cDNA product from A375 and 400 μL of cDNA product from 293T after reverse transcription.
[0154] The gene-specific primer GSP sequence is shown in Seq ID No. 21: 5'-CGTCGCCGTCCAGCTCGACCAG-3' is the reverse primer sequence starting from eGFP, and it was chemically synthesized.
[0155] 9) Library construction for next-generation sequencing using cDNA as a template
[0156] Library construction and pre-amplification were performed using Q5 high-fidelity enzyme (M0541S). The procedures for A375 cDNA and 293T cDNA products were the same, specifically, the total system volume was 800 μL, of which 400 μL was A375 cDNA or 293T cDNA, 40 μL was the forward primer F-Library (10 μM) as shown in Seq ID No. 15, 40 μL was the reverse primer R-Library (10 μM) as shown in Seq ID No. 16, and 320 μL was 2×Mix. Since the switch sequence DNA library contains degenerate bases, while the rest of the sequence is known, only the fragment containing the degenerate region in the switch sequence DNA library needed to be sequenced. The forward primer F-Library and reverse primer R-Library used for PCR were also the primers used in the previous PCR amplification of the switch sequence DNA library.
[0157] PCR program: 98℃ for 3 min, 25 thermal cycles (98℃ for 15 s, 63℃ for 30 s, 72℃ for 30 s), 72℃ for 2 min, and incubation at 4℃.
[0158] After performing agarose gel electrophoresis on the PCR products, a 200bp-300bp gel block was cut off, and the pre-amplified products were recovered using an agarose gel DNA rapid recovery kit.
[0159] To add sequencing adapters to the pre-amplified product, the PCR reaction system was as follows: 400 μL of pre-amplified product, 40 μL of the forward primer final-F (10 μM) of the sequencing adapter primers, 40 μL of the reverse primer final-R (10 μM) of the sequencing adapter primers, and 2 × Mix 320 μL.
[0160] The sequence of the forward primer final-F is shown in Seq ID No. 22: 5'-AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTCCGATCTACCCACTACTAAGAACACTGACT-3', and the sequence of the reverse primer final-R is shown in Seq ID No. 23: 5'-CAAGCAGAAGACGGCATACGAGATTCGCCTTAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTATCTAACAGCCCTGTGCAGC-3'.
[0161] PCR program: 98℃ for 3 min, 30 thermal cycles (98℃ for 15 s, 66℃ for 30 s, 72℃ for 30 s), 72℃ for 2 min, and incubation at 4℃. The PCR products were subjected to agarose gel electrophoresis, and the results are as follows: Figure 7 As shown, lane 1 represents DNA Marker I, lane 2 represents the sample to be sequenced from A375 cells, and lane 3 represents the sequencing sample from 293T cells. From... Figure 7 As can be seen, lanes 2 and 3 contain sequences of uniform size (223 bp), as shown in Seq ID No. 24. This indicates successful library construction. The final product containing sequencing adapters was obtained by gel extraction and recovery.
[0162] 10) Next-generation sequencing
[0163] The final product with sequencing adapters was sequenced by Suzhou Genewiz Biotechnology Co., Ltd. using Illumina PE150, with a sequencing depth of 40G.
[0164] 11) Analyze the data and compare them to obtain the DNA library of the mRNA sensor switch.
[0165] Using the software AptaSUITE: A Full-Featured Bioinformatics Framework for the Comprehensive Analysis of Aptamers from HT-SELEX Experiments. Hoinka, J., Backofen, R. and Przytycka, TM (2018). Molecular Therapy - Nucleic Acids, 11, 515–517. https: / / doi.org / 10.1016 / j.omtn.2018.04.006, FASTQ data filtering was performed, retaining only reads with a length of 56 bp between sequencing primer sequences. Sequence alignment was performed, and switch sequences were extracted. After removing common sequences in target cells (with target sequence mRNA) and control cells (without target sequence mRNA), sequences that were identical in A375 cells and 293T cells except for the A position of TAG were found. Sequences that changed TAG to TGG in target cells and did not show the TAG to TGG phenomenon in control cells were selected.
[0166] Sensor sequences were screened to identify MAGEA4 mRNA, which is highly expressed in A375 cells but not in 293T cells, resulting in 1296 sensor sequences.
[0167] 13. Cell-level validation of sequence expression results
[0168] Although the 1296 sensor sequences obtained from sequencing demonstrate the success of the library, it is still necessary to test whether they can function properly in cells. Since testing all sensor sequences would be too costly, one sequence was randomly selected for cellular expression validation. This sequence was named switch sequence 1, and its sequence is shown in Seq ID No. 25: 5'-TTGTCAAAGGAAACAGAGATAAATCTCCAAGTCATCCAACAAAATAGAAACCCACTACTAAGAACACTGACTATTTTCTTCAAACCCACGGTCGTGTAGGGGTTGTCTTTCACGAGGCAAGGGAAGCTGCTGCACAGGGCTGTTAGATGCACTGGCCCAGCCCTGCCCCTTCCCCACAGCCCTGGCTGCAACTCA-3'.
[0169] The switch sequence 1 was chemically synthesized by Suzhou Genewiz Biotechnology Co., Ltd. Switch sequence 1 was amplified using Q5 high-fidelity enzyme (M0541S) (using F-Library and R-Library primers). After PCR, the DNA was purified by electrophoresis on a 4% agarose gel at 130V using an agarose gel DNA rapid recovery kit. The purified DNA was then mixed with the linearized vector 5'-P2A-eGFP-Vector-mCherry-T2A-3' prepared in the previous steps and processed according to NEBuilder. ® The high-fidelity DNA assembly premix was used to perform seamless cloning to form a circular structure pcDNA3.1-CMV-mCherry-T2A-switch sequence 1-P2A-eGFP, as shown in Seq ID No. 26.
[0170] Meanwhile, using traditional RADAR technology to design candidate sequences for MAGEA4 at https: / / github.com / kristjaneerik / radar-rna-sensing, only one candidate sequence was obtained, named the RADAR sequence, as shown in Seq ID No. 27: 5'-CCACTACTAAGAACACTGACTATTTTCTTCAAACAGAGTGAAGAATAGGCCTCATGTTACACGAGGCAAGGGAAGCTGCTGCACAGGGCT-3'. The pcDNA3.1-CMV-mCherry-T2A-RADAR-P2A-eGFP sequence shown in Seq ID No. 28 was synthesized by Beijing Maijin Biotechnology Co., Ltd.
[0171] The pcDNA3.1-CMV-mCherry-T2A-switch sequence 1-P2A-eGFP and pcDNA3.1-CMV-mCherry-T2A-RADAR-P2A-eGFP were transformed to extract plasmids for large-scale production. Transformation was performed using DMT E. coli competent cells (DMTChemically Competent Cell, catalog number: CD511-01), with incubation at 37℃, 220 RPM, for 16 h (LB medium (containing ampicillin) was purchased from Shanghai Beyotime Biotechnology Co., Ltd., product number: ST165-500ml). Plasmid extraction was performed using EndoFree Plasmid Midi Kit, catalog number: CW2105S, Jiangsu Kangwei Century Biotechnology Co., Ltd.
[0172] Since the verification section of the example only requires the construction of expression vectors for switch sequence 1 and RADAR, the plasmid expression vectors can be constructed directly without minimizing the expression vector. Therefore, the plasmids pcDNA3.1-CMV-mCherry-T2A-switch sequence 1-P2A-eGFP and pcDNA3.1-CMV-mCherry-T2A-RADAR-P2A-eGFP can be directly transfected into cells for expression, and the switching effect can be observed.
[0173] The pcDNA3.1-CMV-mCherry-T2A-switch sequence 1-P2A-eGFP and the pcDNA3.1-CMV-mCherry-T2A-RADAR-P2A-eGFP sequence were transfected into A375 cells and 293T cells, respectively, using the jetPRIME® DNA / siRNA transfection reagent. Fluorescence microscopy imaging was performed 24 hours later, and the results are as follows: Figure 8 As shown. From Figure 8As can be seen, fields 1 and 3 are photographs taken at the same location but with different fluorescence channels under a fluorescence microscope, representing A375 cells transfected with pcDNA3.1-CMV-mCherry-T2A-switch sequence 1-P2A-eGFP (referred to as "switch sequence 1 A375"); fields 2 and 4 are photographs taken at the same location but with different fluorescence channels under a fluorescence microscope, representing A375 cells transfected with pcDNA3.1-CMV-mCherry-T2A-RADAR-P2A-eGFP (referred to as "RADAR A375"); it can be seen that in successfully transfected cells (emitting red light), the number of cells emitting green light in "switch sequence 1 A375" is significantly higher than that in "RADAR A375". Fields 5 and 7 are photographs taken under a fluorescence microscope at the same location using different fluorescence channels. These are 293T cells transfected with pcDNA3.1-CMV-mCherry-T2A-switch sequence 1-P2A-eGFP (referred to as "switch sequence 1 293T"). Fields 6 and 8 are photographs taken under a fluorescence microscope at the same location using different fluorescence channels. These are 293T cells transfected with pcDNA3.1-CMV-mCherry-T2A-RADAR-P2A-eGFP (referred to as "RADAR 293T"). It can be seen that in successfully transfected cells (emitting red light), the number of "switch sequence 1 293T" cells emitting green light is similar to that of "RADAR 293T", but significantly lower than that of "switch sequence 1". A375” demonstrates that after the switch sequence 1 transcribes the sensor mRNA sequence, which is complementary to the MAGEA4 mRNA, the UAG stop codon in the 27nt middle region generates a large number of A to I edits in A375 cells, stopping the repression of downstream eGFP gene translation and translating a large amount of eGFP, which shows green fluorescence under a fluorescence microscope.
[0174] As can be seen from the above embodiments, the DNA sequences screened by the present invention have high throughput and high success rate of candidate libraries. Moreover, the screened sequences are all sequences that have been edited within 72 hours, which can play a "guiding role" in the targeted therapy of tumor cells that specifically express the MAGEA4 gene in the future.
[0175] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A method for screening DNA libraries for intracellular mRNA sensor switches, characterized in that, Includes the following steps: 1) Design a DNA library containing a switch sequence with degenerate bases. The switch sequence consists of a 5' end sequence, a middle region sequence, and a 3' end sequence; the 5' end sequence and the 3' end sequence are both 84 nt in length, and the 5' end sequence and the 3' end sequence are complementary to the target mRNA sequence, respectively. The intermediate region sequence is a variable sequence of 27 nt in length. This intermediate region sequence is represented by 9 triplet codons from the 5' end to the 3' end as 5'-FOZ FOZ FOZ FOZ TAG FOZ FOZ FOZ FOZ FOZ-3', where F is the first base of the codon, O is the second base of the codon, and Z is the third base of the codon. After the 5' end sequence of the 84nt and the 3' end sequence of the 84nt are complementary to the coding strand DNA sequence of the target mRNA, the sequence position of the middle region is matched with the sequence on the coding strand DNA of the target mRNA from the 3' end to the 5' end, which is represented in the form of 9 triplets as 3'-foz foz foz foz acc foz foz foz foz-5'. Where f is A, F is T; f is T, F is A; f is C, F is G; f is G, F is C; When o is A, O is K; when o is T, O is N; when o is C, O is B; when o is G, O is N. When z is A, Z is K; when z is T, Z is N; when z is C, Z is B; when z is G, Z is N. Among them, the degenerate base K is G or T; the degenerate base S is G or C; the degenerate base N is A, T, C or G; and the degenerate base B is G, T or C. The DNA library containing the switch sequence of degenerate bases was chemically synthesized. If the 5' end sequence and the 3' end sequence of the 84nt switch sequence contain a stop codon TAG, TGA, or TAA that belongs to the same reading frame as the TAG stop codon of the 27nt middle region sequence, then the corresponding TAG in the 5' end sequence and the 3' end sequence is modified to TCG, TGA to GGA, and TAA to TCA. 2) Prepare a circular expression vector library; 3) Prepare a minimal linear expression vector library; 4) Seamless cloning; 5) RCA rolling circle amplification; 6) Telomerase digestion; 7) Cell transfection; 8) RNA extraction and reverse transcription; 9) Library construction for next-generation sequencing using cDNA as a template; 10) Next-generation sequencing; 11) Analyze the data and compare them to obtain the DNA library of the mRNA sensor switch; Sequence alignment was performed based on the results of next-generation sequencing. After removing common sequences from the switch sequences in cells expressing the target mRNA and cells not expressing the target mRNA, sequences that were identical in both cells except for the A position of the middle TAG were identified. Sequences in which the TAG changed to TGG in cells expressing the target mRNA and sequences in which the TAG changed to TGG did not exist in cells not expressing the target mRNA were selected. This yielded the DNA library of the mRNA sensor switch.
2. The DNA library screening method for intracellular mRNA sensor switching as described in claim 1, characterized in that, The method for preparing a circular expression vector library in step 2) is as follows: insert the first reporter gene, the first 2A peptide sequence, the second 2A peptide sequence, and the second reporter gene sequentially in the 5'-3' direction between the promoter and PolyA of the initial mammalian expression plasmid vector backbone to obtain a circular plasmid vector; amplify and purify the circular plasmid vector; The DNA library containing the switch sequence with degenerate bases was amplified and purified by PCR. Then, the purified DNA library containing the switch sequence with degenerate bases was inserted between the first 2A peptide sequence and the second 2A peptide sequence of the circular plasmid vector to obtain a circular expression vector library.
3. The DNA library screening method for intracellular mRNA sensor switching as described in claim 2, characterized in that, The first reporter gene is mCherry, eGFP, or SEAP; the second reporter gene is mCherry, eGFP, or SEAP; and the first reporter gene and the second reporter gene located on the same initial mammalian expression plasmid vector backbone are different. The nucleotide sequence of mCherry is shown in Seq ID No. 1, the nucleotide sequence of eGFP is shown in Seq ID No. 2, and the nucleotide sequence of SEAP is shown in Seq ID No.
3. When mCherry, eGFP, or SEAP is selected as the first reporter gene, the stop codons TAG, TAA, or TGA at the end of the sequence are removed. When mCherry, eGFP, or SEAP is selected as the second reporter gene, the start codon ATG at the beginning of the sequence is removed.
4. The DNA library screening method for intracellular mRNA sensor switching as described in claim 2, characterized in that, The first 2A peptide sequence is selected from the P2A sequence, T2A sequence, E2A sequence, or F2A sequence; the second 2A peptide sequence is selected from the P2A sequence, T2A sequence, E2A sequence, or F2A sequence. The nucleotide sequence of P2A is shown in Seq ID No. 4, the nucleotide sequence of T2A is shown in Seq ID No. 5, the nucleotide sequence of E2A is shown in Seq ID No. 6, and the nucleotide sequence of F2A is shown in Seq ID No.
7.
5. The DNA library screening method for intracellular mRNA sensor switching as described in claim 1, characterized in that, The method for preparing the minimized linear expression vector library in step 3) is as follows: design a forward primer and a reverse primer, wherein the forward primer introduces a telomerase restriction site and is located upstream of the 5' end of the promoter sequence; the reverse primer introduces a telomerase restriction site and is located downstream of the 3' end of the PolyA sequence; use the forward primer and the reverse primer to perform PCR on the circular expression vector library to obtain a minimized linear expression vector library.
6. The DNA library screening method for intracellular mRNA sensor switching as described in claim 1, characterized in that, The seamless cloning method in step 4) is to perform seamless cloning on the minimized linear expression vector library to obtain a minimized circular expression vector library with connected ends; the RCA rolling circle amplification method in step 5) is to perform RCA rolling circle amplification on the minimized circular expression vector library to obtain a minimized linear expression vector library with ultra-long periodic repetition.
7. The DNA library screening method for intracellular mRNA sensor switching as described in claim 6, characterized in that, The method for telomerase digestion in step 6) is as follows: the ultra-long periodically repeated minimized linear expression vector library is digested with telomerase to obtain a minimized linear expression vector library with covalently closed ends; the method for cell transfection in step 7) is as follows: the minimized linear expression vector library with covalently closed ends is transfected into cells expressing the target mRNA and cells not expressing the target mRNA for 72 hours.
8. The DNA library screening method for intracellular mRNA sensor switching as described in claim 1, characterized in that, The method for extracting RNA and reverse transcription in step 8) is as follows: RNA is extracted from cells expressing the target mRNA and cells not expressing the target mRNA, and the corresponding cDNA is obtained by reverse transcription.
9. Application of the screening method as described in any one of claims 1 to 8 in screening DNA libraries for sensor switches targeting MAGEA4 mRNA.
Citation Information
Patent Citations
Screening method for identification of efficient pre-trans-splicing molecules
US20060154257A1
Deaminase-based RNA sensors
US20230123513A1