Method for screening and identifying ribosome entry sites in short chain

By constructing a short-chain IRES sequence library and a dual reporter gene system, combined with RNC-seq sequencing technology, the problems of low efficiency and poor reliability in existing IRES screening methods were solved, and IRES sequences with high translational activity were screened out, thereby improving the expression level and application potential of circRNA.

CN121674529APending Publication Date: 2026-03-17ZHEJIANG CANCER HOSPITAL +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511858217.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing IRES screening methods are inefficient, have poor reliability, and have a high false positive rate. They are difficult to accurately locate functional IRES sequences and cannot effectively verify their translation efficiency and physiological function in cells.

Method used

A short-chain IRES sequence library was constructed, and the short-chain IRES were cloned into a screening vector through homologous recombination. Fluorescent proteins and peptides were expressed in cells using a dual reporter gene system, and ribosome-bound IRES sequences were captured by RNC-seq sequencing technology to achieve efficient screening and validation.

Benefits of technology

This improved the screening efficiency and reliability of IRES sequences, identified sequences with high translational activity, enhanced the expression level of circRNA, reduced the false positive rate, and expanded the application scenarios of IRES elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121674529A_ABST
    Figure CN121674529A_ABST
Patent Text Reader

Abstract

The invention discloses a method for screening and identifying ribosome entry sites in a short chain, which comprises the following steps of: (1) screening a short chain IRES sequence with the length of less than 100nt from an IRES database, constructing a screening library and synthesizing; (2) constructing a screening vector containing double reporter genes, cloning the short-chain IRES into a linearized screening vector, and generating circular RNA which only depends on the short-chain IRES sequence to recruit ribosome expression of the reporter genes in cells by the screening vector cloned with the short-chain IRES sequence; and (3) carrying out cell transfection on the screening vector cloned with the short-chain IRES sequence, screening out cells expressing reporter genes through cell sorting, and sequencing through a ribosome-newborn peptide chain compound to obtain the short-chain IRES sequence capable of recruiting ribosome to start circular RNA translation in the cells. According to the invention, a short-chain IRES sequence library is constructed, and short IRES capable of starting circular RNA translation is screened, so that transfection limitation caused by overlarge molecular weight of the current IRES sequence is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biochemistry, and in particular to a method for screening and identifying short-chain internal ribosome entry sites. Background Technology

[0002] The internal ribosome entry site (IRES) sequence, a major translation initiation site for circular RNA (circRNA), is a non-coding sequence widely found in eukaryotes and plays a crucial role in many diseases, mediating the expression of various cancer-related genes. IRES sequences can drive the translation of nucleic acid drugs in vivo and serve as therapeutic targets for many diseases, possessing enormous development potential and application prospects.

[0003] The commonly used IRES sequence: The internal ribosome entry site (IRES) of Coxsackievirus B3 (CVB3) is located in the 5' untranslated region (5'UTR) of the viral genome. Belonging to type I enteroviruses, the specific region of the IRES spans from 1 to 740 nt and is responsible for mediating cap-independent translation of viral RNA. CVB3-IRES is a primary research target and has been applied in retroviral vector technology to facilitate the introduction of more genes into somatic cells for stable multi-gene expression. However, the primary structure of CVB3-IRES is lengthy, which is unfavorable for the expression of short peptides and liposome encapsulation. IRES sequences typically exceed 600 nucleotides, limiting their application in viral vectors with limited cloning capabilities.

[0004] Current IRES screening methods suffer from systemic flaws, resulting in low screening efficiency, poor result reliability, high false positive rates, and difficulties in standardized comparisons and subsequent functional validation. This poses a significant challenge to the identification and understanding of endogenous IRES within cells.

[0005] Traditional IRES screening typically uses reporter gene vectors containing the complete 5' UTR sequence. These UTRs can be hundreds or even thousands of bases long. These ultra-long UTRs may contain randomly formed secondary structural or sequence elements with weak IRES-like activity, which are not true functional IRES. In full-length UTR-based screening, these "noise" signals can mask the true, strong IRES signal, leading to false positives.

[0006] In cell-based high-throughput screening, transfection or infection with vectors containing very long sequences reduces efficiency and limits library complexity. Functional IRES are typically driven by relatively compact core components. Full-length UTR screening cannot accurately locate these core regions, resulting in high subsequent resolution costs. Most IRES screenings based on reporter genes or direct RNA sequencing detect the presence, abundance, or expression of reporter proteins. This cannot directly and specifically reflect whether the sequence truly and efficiently mediates the initiation of ribosomes at internal sites within the cell. RNA-seq measures total RNA abundance, which is greatly affected by transcription and degradation and has a weak correlation with translation efficiency. Reporter gene methods measure the final protein product, which is affected by multiple processes including transcription, translation, and protein stability.

[0007] Current screening methods lack specific probes or markers to distinguish different translation initiation mechanisms. Many IRES screening studies stop at preliminary positive results (such as increased reporter gene activity or sequencing enrichment), lacking rigorous, systematic, and multi-faceted functional validation to confirm that the screened sequences are truly functional and physiologically relevant IRES. A large number of "IRES candidates" generated by screening remain in databases or low-confidence lists, with low credibility, unable to be widely accepted by the field and used for subsequent in-depth research, resulting in wasted resources and cognitive confusion. Summary of the Invention

[0008] This invention provides a method for screening and identifying short-chain internal ribosome entry sites, constructing a short-chain IRES sequence library, and screening for short IRES that can initiate circular RNA translation, thereby overcoming the transfection limitations caused by the excessively large molecular weight of current IRES sequences.

[0009] The technical solution of the present invention is as follows: A method for screening and identifying short-chain internal ribosome entry sites, comprising the following steps: (1) Short IRES sequences with a length of less than 100 nt were selected from the IRES database, and a screening library was constructed and synthesized; (2) Construct a selection vector containing two reporter genes, and clone the short IRES into the linearized selection vector. The selection vector with the cloned short IRES sequence generates a circular RNA in the cell that recruits ribosomes to express the reporter gene in dependence only on the short IRES sequence. (3) Cell transfection was performed on the selection vector cloned with the short IRES sequence. Cells expressing the reporter gene were selected by cell sorting, and the short IRES sequence that can recruit ribosomes to initiate the translation of intracellular circular RNA was obtained by sequencing the ribosome-newborn peptide chain complex.

[0010] This invention addresses the three major bottlenecks in long IRES sequences—limited nucleic acid drug loading capacity, low transfection efficiency, and insufficient duration of drug efficacy—by establishing a novel short-chain IRES screening system that integrates screening and validation.

[0011] In step (1), the IRES database includes the IRESite database and / or the IRESbase database.

[0012] Preferably, step (1) further includes: removing the following sequences from the initial screening targets to construct a screening library and synthesize them: ① IRES sequences containing high GC content regions, ② IRES sequences containing long repeat motifs.

[0013] High GC content means GC% ≥ 60% and the length of long repeating motifs n ≥ 3.

[0014] IRES sequences containing high GC content regions can form overly stable stem structures, restricting dynamic conformational changes and reducing translation efficiency. Sequences containing long repetitive motifs may form unconventional structures, interfering with the binding of IRES to ribosomes or translation factors. Removing these two types of IRES sequences can improve the quality of sequence libraries.

[0015] To screen for IRES sequences that efficiently initiate circRNA translation within cells, this invention constructs a self-splicing circularization screening system.

[0016] Preferably, the screening vector comprises a fluorescent protein gene, a functional protein gene, and an EcoRV restriction site; the functional protein gene includes a functional protein head fragment and a functional protein tail fragment, wherein the EcoRV restriction site is located between the functional protein head fragment and the functional protein tail fragment; the functional protein head fragment is located downstream of the functional protein tail fragment.

[0017] Furthermore, the screening vector comprises a fluorescent protein-5' intron-functional protein tail fragment-stop codon-EcoRV restriction site-functional protein head fragment-3' intron.

[0018] The fluorescent protein is mRuby fluorescent protein; the functional protein is EGFP fluorescent protein or SIINFEKL polypeptide.

[0019] The intron mentioned is the ZKSCAN1 intron.

[0020] Preferably, step (2) includes: (2-1) Construct a screening vector containing two reporter genes; (2-2) Design homologous recombination sequences at both ends of the short IRES and clone the short IRES into the enzyme-digested linearized selection vector by means of homologous recombination using the homologous recombination sequences at both ends.

[0021] The homologous recombination product contains a fluorescent protein-5' intron-functional protein tail fragment-stop codon-short IRES-functional protein head fragment-3' intron. After in vitro backsplicing and circularization, the functional protein head fragment and the functional protein tail fragment are linked to form a complete functional protein gene that expresses the functional protein. This head-to-tail backsplicing mechanism ensures that: (1) the translation of the functional protein gene is initiated by the IRES element; (2) the functional protein gene is translated only in circRNA form, and if a complete gene fragment is not formed, the functional protein will not be expressed. The expression of the functional protein gene is strictly driven by the short IRES sequence, which serves as the initiation site for circRNA to recruit ribosomes and initiate translation.

[0022] This invention also discloses the application of short-chain IRES sequences screened using the aforementioned method for screening and identifying short-chain internal ribosome entry sites in the expression of target genes.

[0023] This invention also discloses the application of short-chain IRES sequences screened using the aforementioned method for screening and identifying short-chain internal ribosome entry sites in the preparation of nucleic acid drugs.

[0024] The present invention also discloses an expression vector comprising a short IRES sequence selected by the screening and identification method and a target gene, wherein the short IRES sequence is located upstream of the target gene.

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) Screening for highly stable IRES sequences with a length of <100nt enhances the recruitment ability of pharmaceutical IRES elements to ribosomes and increases the expression level of circRNA, which is expected to achieve a breakthrough in reducing clinical dosage.

[0026] (2) Screening out highly efficient IRES sequence elements is expected to improve the potency and therapeutic effect of nucleic acid vaccines, reduce the immune rejection response of nucleic acid vaccines, reduce the side effects of circRNA, and expand the application scenarios of the screened IRES elements.

[0027] (3) The novel short-chain IRES elements with high translational activity screened by the method of the present invention can promote the clinical application breakthrough of "enhanced efficacy and reduced toxicity" of nucleic acid vaccines in the future, and provide a core regulatory element library for circRNA drug development. Attached Figure Description

[0028] Figure 1 This invention relates to a method for screening and identifying ribosome entry sites within short chains.

[0029] Figure 2For the construction of IRES libraries and screening systems and cell validation, (a) the source of IRES sequences, (b) the number and proportion of each component, (c) the composition of the screening system, (d) cell transfection imaging, and (e) flow cytometry results.

[0030] Figure 3 For the analysis of lentivirus transfection binding RNC-seq results, (a) RNC-seq sequencing experimental procedure, (b) summary of gene categories of ribosome-binding positive sequences from RNC-seq sequencing, (c) percentage of positive sequences from RNC-seq sequencing in each category of the library, (d) CircRNA verification of cell viability of positive IRES sequences, and (e) prediction of short IRES sequence structure.

[0031] Figure 4 Flow cytometry peaks were used to validate cell viability using short IRES sequences.

[0032] Figure 5 This is a schematic diagram of the design and preparation process of the screening system.

[0033] Figure 6 This is a schematic diagram of the IRES library recombination and recombinant plasmid library preparation process.

[0034] Figure 7 The image shows the results of high-throughput sequencing to detect the quality of recombinant plasmid libraries.

[0035] Figure 8 The images are agarose gel electrophoresis diagrams. (a) IRES library amplification, (b) plasmid linearization of the screening system, (c) plasmid homologous recombination, and (d) PCR amplification and verification of the recombinant fragment. Detailed Implementation

[0036] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate the understanding of the present invention and do not limit it in any way.

[0037] Intra-reactive proteins (IRES) are core regulatory elements in the translation of circRNA into nucleic acid drugs. Their activity directly affects the translation efficiency of circRNA and is an inevitable direction for the development of circRNA nucleic acid vaccines. Miniaturization design has become a technological high ground in nucleic acid vaccine development, namely, shortening the length of the IRES sequence while increasing the drug loading density of the vaccine. This invention addresses the three major bottlenecks in long IRES sequences—limited nucleic acid drug loading, low transfection efficiency, and insufficient duration of efficacy—and aims to establish a novel short-acting IRES screening system. Using a high-throughput technology platform, this invention will systematically analyze the sequence characteristics, functional motifs, and three-dimensional structure-activity relationships of short IRES, aiming to screen miniature elements with high initiation protein translation capabilities, providing key theoretical support for optimizing circRNA drug design. Figure 1As shown, firstly, a short IRES sequence screening library was constructed. The IRES library was then cloned into a linearized screening system via homologous recombination using homologous arms designed at both ends of the library for cell transfection. Secondly, the constructed screening system can generate circular RNA within cells that recruits ribosomes to express the SIINFEKL peptide or green fluorescent protein solely based on the IRES sequence. Cells expressing the SIINFEKL peptide were screened by cell sorting, and the IRES sequence capable of recruiting ribosomes to initiate intracellular circular RNA translation was obtained through ribosome-nascent peptide chain sequencing. This validation system was used to demonstrate the ability of the IRES sequence to recruit ribosomes, thus forming a complete system integrating screening and validation.

[0038] IRES screening library establishment This invention primarily references two major IRES databases: IRESite and IRESbase. These two databases contain the most comprehensive IRES sequence information to date. This invention selects sequences shorter than 100 nt from these two main databases as initial screening targets, and refines them through structural optimization based on ribosome binding region fragments documented in existing IRES libraries. To improve the quality of the sequence library, the following sequences are systematically excluded: (1) Sequences containing high GC content (GC%≥60%) regions will form overly stable stem structures, limiting the dynamic conformational changes of IRES and reducing translation efficiency. (2) Sequences containing long repetitive motifs (n≥3) may form unconventional structures, interfering with the binding of IRES to ribosomes or translation factors. Finally, this invention establishes a screening library containing 7500 short-chain IRES sequences, such as Figure 2 The IRES sequences shown in (a) originate from human genomes, viral genomes, and other categories, with human genomes accounting for 39.12%, viral genomes accounting for 40.03%, and other categories accounting for 20.85%. Figure 2 (b) and the library synthesis was completed on the Agilent SurePrint platform. The construction of this library provides an important tool for the high-throughput screening and validation of the subsequent IRES function.

[0039] II. Design and Preparation of the Screening System To screen for IRES sequences that efficiently initiate circRNA translation within cells, this invention constructs a self-splicing circularization screening system. Figure 2In the middle (c) section, a dual reporter gene design was employed, utilizing a self-splicing circularization design to achieve intracellular circularization. Its structure was mRuby-5'Intron-FEKL-STOP-EcoRV(IRES)-SIIN-3'Intron. Red fluorescent proteins mRuby and SIINFEKL were used for fluorescent labeling and functional verification, respectively. SIINFEKL is a model antigen peptide widely used in immunological research, often employed as a model antigen for testing novel vaccine vectors. As the 257-264th amino acid fragment of chicken ovalbumin (OVA) (Ser-Ile-Ile-Asn-Phe-Glu-Lys-Leu), it exhibits high affinity for mouse MHC class I molecule H-2Kb. The SIINFEKL gene was divided into two equal fragments and constructed into a selection vector. After transcription into mRNA, the peptide was back-spliced ​​by intracellular endonuclease recognizing the flanking ZKSCAN1 introns, reconnecting the split SIIN and FEKL gene fragments to form the complete SIINFEKL gene. This reverse splicing mechanism ensures that: (1) the translation of the SIINFEKL polypeptide gene is initiated by the IRES element; and (2) the SIINFEKL polypeptide gene is translated only in circRNA form, and the polypeptide will not be expressed if a complete gene fragment is not formed.

[0040] This invention first verified the feasibility of the screening system through preliminary experiments. The CVB3 WO Ex_pcDNA3.1 (mRuby-5'intron-FP-STOP-CVB3-EG-3'intron) vector plasmid's core design utilizes a fissuring enhanced green fluorescent protein (EGFP) gene as a reporter gene, whose expression depends on the self-circularization of intracellular RNA. Green fluorescent protein is only produced when the RNA successfully circularizes, forming the complete EGFP target gene. Figure 2 (d) Figure 2 As shown in (e), Figure 2 As shown in (d), normal expression of two fluorescent proteins was observed under the microscope. EGFP expression is strictly driven by the CVB3-IRES sequence, which serves as the initiation site for circRNA to recruit ribosomes and initiate translation.

[0041] III. RNC-seq sequencing captures full-length IRES-circRNA in positive cells. The complete workflow of RNC-seq-based sequencing analysis is as follows: Figure 3As shown in (a), the process includes cell processing, nucleic acid processing, library construction, and sequencing. RNC-seq is commonly used to study mRNAs undergoing translation and their associated nascent peptide chains. Its core principle is to reveal the real-time translation status of genes by extracting mRNA fragments bound to ribosomes in cells, constructing libraries, and sequencing. In the sample processing stage before RNC-seq, an innovative step of delinearizing RNA is added, significantly reducing the interference of native cell mRNA on sequencing results and enriching circRNAs with IRES library sequences as ribosome recruitment sites as much as possible. Through RNC-seq, this invention captured a total of 682 sequences with IRES activity. The captured IRES sequences were aligned to a short-chain IRES library, and the results are shown below. Figure 3 (b) Figure 3 As shown in (c). The specific analysis is as follows: (1) Human genome source: 496 sequences (accounting for 72.73% of positive sequences and 16.9% of human genome sequences). (2) Viral genome source: 61 sequences (accounting for 8.94% of positive sequences and 2% of viral genome sequences). (3) Other sources: 125 sequences (accounting for 18.33% of positive sequences and 8% of viral genome sequences). The positive sequences captured this time fully demonstrate that IRES sequences are species-selective. More than 70% of the positive sequences in this time are from human genomes, while less than 10% are from viruses. RNC-seq can capture full-length mRNA in translation, thus reflecting the complete translation process more comprehensively. This technology avoids the result bias that may be introduced after non-ribosomal protected sequences are degraded by RNA nucleases.

[0042] IV. Short-chain IRES exhibit high translational activity in cells. Twenty IRES sequences with the highest enrichment levels, bound to ribosomes and captured by RNC-seq, were selected for functional validation, and a unique translation plasmid was constructed. The synthesis of the validation plasmid was commissioned to Suzhou Levofloxacin Biotechnology Co., Ltd. Corresponding IRES-circRNAs were prepared; these IRES sequences exhibited high cell activity and cell tropism (e.g., ...). Figure 4 (As shown). To ensure the accuracy and reliability of the experimental results, these IRES-circRNAs were independently transfected to eliminate the potential influence of interactions between different IRES on the experimental results. Figure 3 As shown in (d), the expression levels of CircRNA synthesized from IRES sequences that exhibit high activity in transcriptome sequencing are as follows: the average fluorescence intensity of sequences such as CircIREStyl-1675 and CircIREStyl-7328 are higher than those of the double-positive control group (CVB3-IRES, crTMV-IRES).

[0043] We hypothesize that longer IRES sequences can affect ribosome recognition of short target genes, while shorter IRES sequences are more suitable for the translation and expression of short peptides or antigens. This result indicates that the screened highly active IRES sequences have important application value in the construction of circRNAs, providing strong support for further optimization of circRNA function and regulatory mechanisms. RNA secondary structure prediction of highly active sequences was performed using RNAstructure software (e.g., ...). Figure 3 As shown in (e), the software can quickly identify the core functional domains in IRES sequences, including key regions such as pseudoknots, stem-loop structures, and GNRA tetranucleotide loops. The results show that all sequences exhibit significant stem-loop structure features. Such structures are ubiquitous in IRES functional elements and may maintain their translational activity through the following mechanisms: (1) Spatial conformational stability: Stem-loop structures can fix the spatial folding of the IRES core region, forming a specific interface for interaction with the 40S ribosomal subunit or translation initiation factors. (2) Dynamic regulation of site exposure: The flexibility of the loop may promote the exposure of key nucleotide residues, thereby mediating ribosome recognition and recruitment. This data suggests that preserving such domains when designing exogenous circRNAs is crucial for maintaining their translational ability. The functional necessity of specific stem-loops can be further verified through point mutation or truncation experiments.

[0044] This technological system will drive the transformation of IRES research from an "empirical screening" to a "rational design" paradigm, providing a standardized platform for the development of circRNA synthetic biology tools. As a unique class of RNA functional elements, IRES sequences demonstrate multidimensional application value in viral replication, precise gene regulation, and nucleic acid drug development through their mediated non-classical translation mechanisms. Addressing the systemic methodological gaps in current IRES research, this invention successfully constructs a groundbreaking three-in-one technological system of "screening-validation-optimization." With the rapid rise of circRNA as a novel therapeutic vector, the functional analysis and engineering modification of IRES elements have become core components for overcoming the bottleneck in nucleic acid drug delivery efficiency.

[0045] Example 1 A method for screening and identifying short-chain internal ribosome entry sites, comprising the following steps: (a) Establishment of IRES screening library Through systematic statistical analysis and summarization of target sequences for IRES screening, this invention primarily references the two most authoritative IRES databases: IRESite and IRESbase. This invention focuses on IRES sequences shorter than 100 nt from these two databases as the main screening targets. Combining known initiation translation core fragments in IRES sequences with online RNAfold prediction tools, the minimum free energy (MFE) secondary structure and centroid structure of IRES sequences proven to initiate mRNA translation were analyzed. This structural information provides important evidence for functional studies of IRES sequences, further optimizing and combining them to generate new IRES sequences for screening. This invention also incorporates research findings from a Stanford University research group on circRNA translation initiation elements, ultimately constructing a shorter and more efficient IRES screening library. Figure 2 As shown in (a), the IRES sequences originated from human genomes, viral genomes, and other sources. Human genomes accounted for 39.12%, viral genomes for 40.03%, and other sources for 20.85% (e.g., human genomes). Figure 2 (as shown in (b)). Finally, a screening library containing 7,500 short-chain IRES sequences was established, and efficient synthesis was performed on the Agilent SurePrint platform.

[0046] (II) Design and preparation of the screening system The process is as follows Figure 5 As shown, this includes screening system amplification, linearization, and IRES library amplification.

[0047] To screen for IRES sequences that efficiently initiate circRNA translation within cells, this invention constructs a highly efficient screening system (e.g., Figure 2As shown in (c), its structure is mRuby-5'Intron-FEKL-STOP-EcoRV(IRES)-SIIN-3'Intron. In this system, the EcoRV restriction site is used for subsequent homologous recombination insertion of the IRES library. The screening system contains two key target genes: mRuby (Red Fluorescent Protein) and SIIN-FEKL peptide, used for fluorescent labeling and functional verification, respectively. mRuby is excited at 558 nm and emitted at 605 nm, exhibiting bright red fluorescence. mRuby serves as an indicator gene for plasmid transfection and expression, used to monitor the function of the screening system in real time. To study the ability of short IRES elements to initiate circRNA translation, this invention specifically selected SIINFEKL peptide as the target gene because of its short sequence and mature detection method, which can effectively reduce the translation difficulty. SIINFEKL is a peptide composed of 8 amino acids, widely used in cell receptor recognition and antigen presentation. This polypeptide is easily synthesized and modified, and often binds to major histocompatibility complex class I molecules, serving as a model antigen for studying cellular recognition and response to specific antigens. The expression level of the SIINFEKL polypeptide was detected by flow cytometry to assess the ability of the IRES sequence to recruit ribosomes and initiate circRNA translation within cells. After transcription into mRNA, the SIIN-FEKL polypeptide gene undergoes backsplicing via intracellular endonuclease recognition of the flanking ZKSCAN1 introns, reconnecting the split SIIN and FEKL polypeptide genes to form the complete SIINFEKL polypeptide gene. After extensive amplification, it is linearized and used as a vector for subsequent homologous recombination of the IRES library.

[0048] (III) IRES library recombination and preparation of recombinant plasmid libraries The process is as follows Figure 6 As shown, this includes plasmid homologous recombination to competent cells, amplification, and cell transfection and sorting.

[0049] The IRES library consists of 7500 single-stranded DNA molecules, with homologous recombination sequences designed at both ends. Using polymerase chain reaction (PCR), we converted the single-stranded DNA library into double-stranded DNA and cloned the IRES library into a selection vector (pcDNA3.1+) using the principle of homologous recombination. This library construction provides an important tool for subsequent high-throughput screening and validation of IRES function. For IRES library cloning, we previously designed the Fw homologous arm -AGCTGTACAAGTAATAG and the Rv homologous arm -TCGCCCTTGCTCACC, which can be used for homologous recombination. Homologous recombination was performed using NEB's NEBuilder HiFi DNA Assembly Master Mix. The homologous recombination system (20 µL) consisted of the following components: vector: 1 µg linearized plasmid, insert: 0.1 µg PCR library, 10 µL NEBuilder HiFi DNA Assembly Master Mix, and ddH2O to a final volume of 20 µL. The mixture was incubated at 50 °C for 1 hour, and the recombination product was stored at -40 °C. The recombinant IRES library plasmid could be directly used to transform competent NEB 10-beta Competent E. coli: 100 µL of NEB 10-beta Competent E. coli competent cells (thawed on ice) were added, and the mixture was gently tapped against the tube wall several times and incubated on ice for 30 minutes. The mixture was then heat-shocked in a 42 °C water bath for 45-60 seconds, quickly transferred to an ice bath, and incubated for 2 minutes. Add 700 µL of antibiotic-free sterile Stable Outgrowth Medium (SOM) liquid medium to a centrifuge tube, mix well, and incubate at 200 rpm for 60 minutes at 37°C in a shaker. Centrifuge the resuscitation solution at 2,000 rpm for 5 minutes, discard 700 µL of supernatant, and spread the remaining liquid evenly onto LB agar plates containing the appropriate antibiotic. Invert the plates and incubate overnight at 37°C. To ensure the diversity of the IRES library recombinant plasmids as much as possible, scrape all colonies from the LB agar plates and transfer them to 250 mL Erlenmeyer flasks, add an appropriate amount of LB liquid medium, and incubate with vigorous shaking. After incubation, extract the recombinant IRES library plasmids. Outsourced sequencing showed that the quality of the recombinant IRES library plasmids was excellent, meeting the requirements for subsequent experiments (e.g., sequencing). Figure 7 (As shown).

[0050] (iv) Agarose gel electrophoresis Select a plastic gel casting tank with a sufficient number of samples. Prepare 50 ml of 1x Tris-Acetate-EDTA buffer in an Erlenmeyer flask, add 1 g of agarose powder, microwave on medium heat for three minutes, cool to approximately 60 degrees Celsius, and add 1 µL of staining agent. After mixing well, quickly pour the mixture into the gel casting tank, being careful not to create air bubbles. Quickly insert a suitable comb to form the sample wells. After the gel has completely solidified, carefully remove the comb and place the gel into the electrophoresis tank. Loading system (6 µL): DNA loading buff (6x) 1 µL; digested DNA sample 200-300 ng; ddH2O to 6 µL. Incubate at 95 degrees Celsius for 1 minute, then incubate on ice for 2-3 minutes. Select the corresponding sample wells for loading; one well is for the DNA ladder, one well for the undigested plasmid, and the other well for the digested plasmid. Perform electrophoresis in constant voltage mode at 140 V for 40 minutes. After electrophoresis, images were taken using a gel imaging system.

[0051] The results of IRES library amplification (a), plasmid linearization of the screening system (b), plasmid homologous recombination (c), PCR amplification and verification of recombinant fragments (d) are shown in the agarose gel electrophoresis results. Figure 8 As shown.

[0052] (v) Transient cell transfection Seed cells into 10cm culture plates 24 hours in advance, and transfect when cell confluence reaches 70-90%. The transfection system is as follows: 500µLopti-MEM. TM Lipofectamine diluted in culture medium TM 43.4 µg of reagent 3000 was mixed thoroughly. A 500 µL Lopti-MEM was then used. TM Dilute 28 µg of the IRES recombinant plasmid library with culture medium to prepare a DNA premix, then add P3000. TM Mix 56 µL of reagent thoroughly. Add the diluted DNA premix to the diluted Lipofectamine at a 1:1 ratio. TM In a 3000 mL solution, gently mix. Incubate the mixture at room temperature for 15 minutes to allow the DNA and liposomes to fully bind and form a complex. Add the incubated DNA-liposome complex dropwise to the plated cells, gently shaking the plate to ensure even distribution. Cell culture and analysis: Incubate the cells at 37°C for 2-4 days, then analyze the transfected cells. Fluorescence microscopy reveals that the cells appear entirely red, indicating successful plasmid transfection and normal expression.

[0053] (vi) RNA Nucleic Acid Extraction After sorting, cells were seeded in 10cm culture dishes at a density of approximately 5-50 million / mL. The translation elongation inhibitor actinomycin was added to the culture medium, gently mixed, and incubated at 37°C for 10-15 minutes. The culture medium was removed, and the cells were placed on ice and washed twice with pre-chilled PBS. An appropriate amount of cell lysis buffer was added, and the cells were lysed on ice for 15-20 minutes. After lysis, the cells were centrifuged at 12,000 g for 10 minutes at 4°C, and the supernatant was transferred to a new centrifuge tube. The obtained cell lysate was immediately used for subsequent experiments. The cell lysate was ultracentrifuged at 100,000 g through a 1 M sucrose pad at 4°C for 2 hours to separate and purify the ribosome-RNA complex. Total RNA from the ribosome complex was extracted using Takara RNAiso Plus reagent, with the following steps: 1 / 5 volume of chloroform was added to the homogenate lysis buffer, the centrifuge tube was tightly capped, and the mixture was vigorously shaken by hand for 15 seconds. After the solution was fully emulsified, it was allowed to stand at room temperature for 5 minutes. The mixture was then centrifuged at 12,000 g for 15 minutes at 4°C. Carefully remove the centrifuge tube. At this point, the solution has separated into three layers: a colorless supernatant (RNA phase), a middle white protein layer, and a colored lower organic phase. Transfer the supernatant to a new centrifuge tube. Add an equal volume of isopropanol to the supernatant, invert to mix, and let stand at room temperature for 10 minutes. Centrifuge at 12,000 g for 10 minutes at 4°C, discarding the supernatant. RNA precipitate will appear at the bottom of the tube. Carefully discard the supernatant, and slowly add 1 mL of 75% ethanol along the wall of the centrifuge tube, gently inverting to wash the tube wall. Centrifuge at 12,000 g for 5 minutes at 4°C, carefully discarding the ethanol. Dry the RNA precipitate at room temperature for 2-5 minutes. Add an appropriate amount of RNase-free water to dissolve the precipitate, gently pipetting if necessary to aid dissolution. Once the RNA is completely dissolved, aliquot and store at -80°C.

[0054] (vii) Establishing RNC-seq sequencing libraries After the RNA samples passed quality control, the library was constructed using the GENESEED® RNA-seq Library Prep Kit for Illumina. The designed DNA probes were hybridized to the RNA samples, and rRNA was removed from the total RNA through specific binding. The purified RNA was then fragmented to obtain RNA fragments suitable for sequencing. Using the fragmented RNA as a template, first-strand cDNA was synthesized using random primers and reverse transcriptase. A strand-specific method was used to incorporate dUTPs for labeling during second-strand cDNA synthesis, simultaneously performing end repair. An A-tail was added to the 3' end of the double-stranded cDNA for subsequent adapter ligation. Illumina sequencing adapters were ligated to both ends of the cDNA fragments. The ligation products were purified using magnetic beads to remove unligated adapters and other impurities. Fragments of the target size were selected by gel electrophoresis or magnetic bead sorting. Before PCR amplification, the second-strand cDNA template labeled with dUTP was digested with uracil DNA glycosylase to ensure the chain specificity of the library. Library amplification was performed using high-fidelity DNA polymerase. The PCR reaction conditions were set as follows: (1) 98℃, 1:00 min; (2) 98℃, 0:30 min; (3) 65℃, 0.30 min; (4) 72℃, 1:00 min; (5) Repeat step (2), 15x; (6) 72℃, 5:00; (7) 4℃, ∞. After amplification, the library was purified and recovered using magnetic beads to remove primer dimers and other impurities. The RNC-seq sequencing library was then constructed.

[0055] (viii) CircRNA Synthesis For IRES sequence verification, a validation plasmid (pET-28a) was designed. The plasmid needed to be linearized. The linearization system (100 µL) consisted of: CircRNA validation plasmid (20 µg), XbaI endonuclease (3 µL), 10x NEBuff (10 µL), and ddH2O to 100 µL. The reaction system was incubated at 37°C for 3 hours, followed by heating to 65°C and holding for 30 minutes to ensure complete inactivation of the endonuclease. After the enzyme digestion reaction, the product was purified, and the concentration of the purified product was determined using a spectrophotometer or quantitative fluorescence instrument. For in vitro transcription, we used the NEB T7 transcription kit. The in vitro transcription system (20 µL) consisted of: plasmid DNA (1 µg), T7 RNA Polymerase Mix (1 µL), NTP mixing buffer (10 µL), and ddH2O to 20 µL. The reaction system was incubated at 37°C for 3 hours, followed by the addition of 1 µL of Dnase I and incubation for another 30 minutes to remove DNA contamination. 11 µL of LiCl was added, and the mixture was thoroughly mixed. The sample was then transferred to a -40°C freezer for overnight storage to promote RNA precipitation. The next day, the sample was centrifuged at 13,300 rpm for 1 hour at 4°C, and the supernatant was discarded. The precipitate was washed with 200 µL of 70% ethanol and centrifuged again at 13,300 rpm for 15 minutes. After centrifugation, the supernatant was carefully aspirated, and the precipitate was air-dried at room temperature. The precipitate was dissolved in 40 µL of nuclease-free water, and the RNA concentration was determined using a spectrophotometer or quantitative PCR instrument. The solution was then stored for later use. For linear mRNA, an in vitro circularization strategy was used. The in vitro circularization system (50 µL) consisted of: RNA transcribed in the previous step (90 µg), T4 RNA Ligase 2 buff (10 µL), GTP (1 µL), and ddH2O to a final volume of 50 µL. The system was incubated at 56°C for 30 minutes. Purification was performed after incubation. A New England RNA purification kit was used for purification. After RNA circularization, RNase R digestion was performed to remove linear RNA. The digestion system (100 µL) consisted of: RNA (80 µg), RNase R enzyme (2 µL), RNase R Reaction Buffer (6 µL), and ddH2O to a final volume of 100 µL. The mixture was incubated at 37°C for 45 min and then at 65°C for 15 min. After incubation, the above purification steps were repeated to obtain pure, single IRES-circRNA.

[0056] We used nuclease digestion or restriction enzyme digestion combined with PCR to fragment genomic or transcriptome-derived UTRs and clone them into reporter vectors. Fragment sizes were optimized to 87 nt to enrich regions potentially containing core IRES elements. Flow cytometry sorting (FACS) and integrated Ribosome-Nascent Chain Complex Sequencing (RNC-seq) technology enabled parallel and rapid screening of nearly ten thousand UTR fragments.

[0057] This invention constructs a small-scale, fragmented IRES library and performs standardized screening using high-throughput technologies such as FACS / NGS. During the screening process or on the screening results, RNC-seq analysis is performed to specifically capture and locate ribosome initiation events, effectively distinguishing IRES from other translation initiation mechanisms and significantly reducing false positives. From rigorous biochemical mechanism verification to functional physiological verification, orthogonal methods are used for cross-validation. Only by passing this system can a sequence be confirmed as a truly functional and physiologically relevant IRES.

[0058] The present invention has the following advantages: First, based on literature mining (PubMed / EMBL-EBI) and database resources (IRESbase v2.1, CircAtlas), candidate IRES sequences were screened, and a potentially unique short-chain IRES functional library (less than 100 bp) was constructed. Compared to traditional long-chain IRES or chemically modified base research, this strategy significantly increases the capacity of circRNA vectors by shortening the length of functional elements. By establishing an IRES sequence library and screening system, we successfully screened IRES sequences with higher functional activity than positive controls. Compared to traditional mRNA drugs that rely on the 5' cap structure to initiate translation, IRES sequence-mediated cap-independent translation also reduces the risk of degradation by decapping enzymes in vivo, making it suitable for nucleic acid vaccines requiring long-term expression.

[0059] Second, innovative genomic-translatome combined analysis: (1) At the genomic level, FastNGS high-throughput sequencing technology was used to locate IRES sequences. (2) At the translatome level, RNC-seq sequencing technology was combined to capture IRES-driven dynamic translation events. By integrating dual-omics data, the deep relationship between IRES activity and cells was revealed. Using RNC-seq technology, we can understand the transcription and translation of circRNAs with IRES library sequences in cells. The IRES sequences detected by this technology all have the ability to recruit ribosomes. By comparing and combining genomics and transcriptomics, we understand that applying RNC-seq sequencing technology to the IRES sequence screening system can minimize the bias of screening results brought about by genomic sequencing, enabling us to obtain the most accurate screening results.

[0060] Third, establish a closed-loop research system for screening and validation: Primary screening: assessing the basic activity of IRES based on a dual reporter system; Secondary validation: quantifying single-molecule translation events through a combination of cell translation systems and live-cell imaging. The IRES sequence is a core component of circRNA drug development. By establishing an IRES sequence validation system and synthesizing unique IRES-circRNAs, we can precisely locate highly functional IRES sequences within cells, eliminating functional enhancements caused by interactions between nucleic acid elements.

[0061] Fourth, our experiments have demonstrated that short IRES sequences can recruit ribosomes to initiate circRNA translation. Furthermore, in the design of nucleic acid drugs, this can reduce the mutational impact of redundant sequences and the difficulty of nucleic acid drug encapsulation. Establishing an IRES sequence screening-validation-optimization platform will provide insights for the development of nucleic acid drugs.

[0062] The embodiments described above provide a detailed explanation of the technical solutions and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for screening and identifying short internal ribosome entry sites, characterized in that, The method comprises the following steps: (1) screening short-chain IRES sequences with a length less than 100 nt from an IRES database, constructing a screening library and synthesizing; (2) constructing a screening vector containing double reporter genes, and cloning the short-chain IRES into the linearized screening vector, so that the screening vector containing the short-chain IRES sequence generates a circular RNA in cells, which expresses the reporter genes only by recruiting ribosomes of the short-chain IRES sequence; (3) transfecting the screening vector containing the short-chain IRES sequence into cells, screening cells expressing the reporter genes by cell sorting, and obtaining the short-chain IRES sequence capable of recruiting ribosomes to initiate translation of the circular RNA in cells by sequencing of ribosome-new peptide chain complexes.

2. The method of screening and identifying short internal ribosome entry sites according to claim 1, wherein, The IRES database comprises the IRESite database and / or the IRESbase database.

3. The method of screening and identifying short internal ribosome entry sites according to claim 1, wherein, Step (1) further comprises: eliminating the following sequences from the preliminary screening target to construct a screening library and synthesize: ① IRES sequences containing a high GC content region, and ② IRES sequences containing a long repeat motif.

4. The method of screening and identifying short internal ribosome entry sites according to claim 1, wherein, The screening vector comprises a fluorescent protein gene, a functional protein gene, and an EcoRV enzyme cutting site; the functional protein gene comprises a functional protein first fragment and a functional protein tail fragment, and the EcoRV enzyme cutting site is located between the functional protein first fragment and the functional protein tail fragment; the functional protein first fragment is located downstream of the functional protein tail fragment.

5. The method of screening and identifying short internal ribosome entry sites according to claim 4, wherein, The screening system comprises a fluorescent protein-5' intron-functional protein tail fragment-EcoRV enzyme cutting site-functional protein first fragment-3' intron.

6. The method of screening and identifying short internal ribosome entry sites according to claim 4, wherein, The fluorescent protein is an mRuby fluorescent protein; and the functional protein is an EGFP fluorescent protein or a SIINFEKL polypeptide.

7. The method of screening and identifying short internal ribosome entry sites according to claim 1, wherein, Step (2) comprises: (2-1) constructing a screening vector containing double reporter genes; (2-2) designing homologous recombination sequences at both ends of the short-chain IRES, and cloning the short-chain IRES into the enzyme-cut linearized screening vector by homologous recombination at both ends.

8. Application of a short-chain IRES sequence screened by the screening and identification method of the short-chain internal ribosome entry site according to any one of claims 1-7 in expression of a target gene.

9. Application of a short-chain IRES sequence screened by the screening and identification method of the short-chain internal ribosome entry site according to any one of claims 1-7 in preparation of a nucleic acid drug.

10. An expression vector, characterized by, A nucleic acid drug comprising a short-chain IRES sequence screened by the screening and identification method of the short-chain internal ribosome entry site according to any one of claims 1-7 and a target gene, wherein the short-chain IRES sequence is located upstream of the target gene.