A method for identifying and screening coding orfs in long non-coding rnas of skeletal muscle
By using a multi-omics approach to screen for encoding ORFs in long non-coding RNAs of skeletal muscle, and combining ribosome binding and proteomics validation, the problem of high false positives in existing technologies has been solved, achieving efficient screening and improved accuracy.
Patent Information
- Application Number
- CN202310031964.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-05
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2043-01-05
AI Technical Summary
Current technologies lack effective means to screen for coding open reading frames (ORFs) in long non-coding RNAs of skeletal muscle, and traditional methods have a high false positive rate.
A multi-omics approach was adopted, including transcriptomics, translatomics, and proteomics analysis, combined with molecular biology validation. By screening for encoding ORFs through ribosome binding and proteomics validation, the false positive rate was reduced.
This method effectively screens out ORFs encoded by long non-coding RNAs in skeletal muscle, reduces the probability of false positives, improves the accuracy of screening, and facilitates the efficient screening of key molecules regulating muscle development.
Smart Images

Figure CN115966256B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics technology, specifically to a method for analyzing, identifying, and screening encoding ORFs in long non-coding RNAs of skeletal muscle. Background Technology
[0002] Long non-coding RNAs (lncRNAs) are linear RNA molecules with a length between 200 bp and 100 kbp that have essentially no coding ability and mainly originate from non-coding regions. Some lncRNAs are highly expressed in muscle and play an important regulatory role in muscle development. LncRNAs participate in the transcriptional regulation of muscle development through trans and cis mechanisms. The cis and trans regulatory effects of lncRNAs depend on the location of the target gene. Cis regulation refers to the expression of lncRNAs at or near the same chromosomal locus as the target gene, i.e., the transcriptional regulation of neighboring target genes. Trans regulation, on the other hand, can inhibit or activate gene transcription at independent chromosomal loci, i.e., the transcriptional regulation of distant target genes.
[0003] Short-chain amino acid peptides play important roles in the growth, development, and stress resistance of plants and animals. Most of these functional peptides are obtained through processing precursor proteins or direct translation of open reading frames (ORFs) present in the genome, sometimes located in the untranslated region of messenger RNA (mRNA). In addition to conventional peptides derived from precursor processing, recent plant and animal studies have increasingly revealed the existence and function of non-conventional peptides (NCPs). These peptides are translated from the 5'UTR or 3'UTR of transcripts and are currently defined as peptides encoded by non-coding RNAs (ncRNAs). LncRNAs, circRNAs, and pre-miRNAs can all encode non-conventional peptides. LncRNAs can have one or more small ORFs (sORFs) with fewer than 300 bases, which can be translated into small peptides with a length of less than 100 amino acids.
[0004] LncRNAs were previously defined as linear RNAs that do not encode proteins. However, with advancements in research, it has been discovered that some lncRNAs can not only encode small peptides but also use these peptides to perform biological functions. For example, skeletal muscle-related lncRNAs (LINC00948 in humans and AK009351 in mice) contain a short open reading frame of 138 nucleotides and potentially encode 46 highly conserved amino acids. Small peptides encoded by lncRNAs are conserved in vertebrates and are named myoregulatory proteins (MLNs). MLNs are expressed in almost all skeletal muscles and encode transmembrane α-helices. MLNs share strong structural similarities with phosphoproteins and myolipin, whose transmembrane regions contain many of the same conserved residues. Recently, researchers discovered a 90-amino acid peptide encoded by lncRNA LINC00961, which is highly expressed in the skeletal muscle of both humans and mice and promotes skeletal muscle development.
[0005] With the advent of the omics era, transcriptomics, translatomics, and proteomics technologies can reveal gene expression patterns at different levels. Currently, there is a lack of effective methods for identifying coding open reading frames (ORFs) in lncRNAs. Existing techniques are based on genomic analysis to predict potential coding ORFs (such as patent CN202110996256.2), but this method only analyzes single omics data, resulting in a high false positive rate, and the predicted ORFs may not necessarily encode small peptides. Therefore, there is a need to develop a multi-omics method based on transcriptomics, translatomics, and proteomics for screening coding ORFs in skeletal muscle long non-coding RNAs, combined with molecular biology validation methods, to achieve both batch screening of coding ORFs in skeletal muscle lncRNAs and a significant reduction in the number of false positives for screened LncRNA coding ORFs. Summary of the Invention
[0006] To address the above problems, this invention proposes a method for screening encoding ORFs in long non-coding RNAs of skeletal muscle.
[0007] The method for screening ORFs in long non-coding RNAs of skeletal muscle provided by this invention includes the following steps:
[0008] Step 1) Take skeletal muscle samples and perform library construction and sequencing on the transcriptome, translatome, and proteome respectively. The library construction and sequencing methods are consistent with the conventional procedures, and transcriptome sequencing data, translatome sequencing data, and proteome mass spectrometry data are obtained respectively.
[0009] Step 2) Analyze the transcriptome data obtained in Step 1) and screen out ORFs with potential coding capabilities within long non-coding RNAs as set 1;
[0010] Step 3) Analyze the transcriptome and translatome data obtained in Step 1) and calculate the ribosome release score (RRS) of ORFs of long non-coding RNAs and mRNAs respectively. ORFs with an RRS score greater than 1.24 are selected as candidate ORFs. The intersection of these candidate ORFs with the above set 1 is taken to form set 2.
[0011] The formula for calculating RRS in step 3) is: RRS = (RPKMCDS / RPKM3′UTR) Ribo-seq / (RPKMCDS / RPKM3′UTR) RNA-seq In the formula, RPKMCDS represents the expression level of transcripts located in the coding region, RPKM3′UTR represents the expression level of transcripts located 3′ upstream of the coding region in the non-coding region, Ribo-seq indicates that the data analyzed in this part comes from translatome sequencing data, and RNA-seq indicates that the data analyzed in this part comes from transcriptome sequencing data.
[0012] Step 4) Analyze all small peptides in skeletal muscle using the proteomic data obtained in Step 1), compare them with the theoretically translatable small peptides that can be formed by the ORFs screened in Set 2, and screen out the ORFs that can be uniquely matched as Set 3. The ORFs in Set 3 are the ORFs that can be encoded in the long non-coding RNA of skeletal muscle.
[0013] Preferably, the step 2) of screening ORFs with potential coding ability within long non-coding RNAs as set 1 is as follows: First, use the ORFfinder tool to predict all potential ORFs within long non-coding RNAs in the transcriptome data, and extract ORFs with the start codon ATG and a length range of 60-450 bp; then, use the Fickett and Hexamer algorithms to select ORFs with potential coding ability from the extracted ORFs, using the top 95% of mRNA coding ability scores as the criteria for screening coding ORFs, and select ORFs with potential coding ability as set 1.
[0014] The beneficial effects of this invention are as follows:
[0015] 1. The method of this invention greatly reduces the probability of false positive ORFs: Steps 3) and 4) of this invention are key steps in solving false positives. Their basic principle follows the central dogma of gene expression, from DNA transcription to RNA, RNA binding to ribosomes for translation, and finally, the identification of the translated peptide (protein) at the protein level to obtain the corresponding product. Traditional methods are based on DNA transcription into RNA, and then predicting potential ORFs within the transcribed RNA. Step 2) of this invention includes this prediction and calculation. Based on this, the present invention continues to track which parts of the transcript RNA bind to ribosomes for translation, which is the scheme described in step 3). In this step, LncRNA-ORFs that do not bind to ribosome transcripts are removed. LncRNA-ORFs that do not bind to ribosome transcripts are false positive ORFs. Step 4) detects ribosome-bound LncRNA-ORFs at the protein level, which is to further remove LncRNA-ORFs without small peptide products based on step 3). That is, by checking whether small peptide products can be detected, we can infer whether these LncRNA-ORFs can encode, and further remove LncRNA-ORFs that cannot encode in set 2, which greatly reduces the probability of false positive molecules and finally obtains the encoding ORFs in the long non-coding RNA of skeletal muscle.
[0016] 2. This invention can effectively screen for encoding ORFs in long non-coding RNAs of skeletal muscle, which is beneficial for efficiently screening key molecules that regulate muscle development. Attached Figure Description
[0017] Figure 1 The results of screening LncRNA-ORFs using the Fickett and Hexamer algorithms in Example 2;
[0018] Figure 2 This is a length distribution diagram of the translated genome sequencing fragments in Example 2;
[0019] Figure 3 This is a map showing the location annotations of the translated genome sequence fragments from Example 2 within the reference genome.
[0020] Figure 4 This is a diagram showing the abundance of translatome sequencing fragments and their relative positions to start codons in Example 2.
[0021] Figure 5 Analysis of the reading frame composed of three consecutive codons in Example 2;
[0022] Figure 6 Venn diagram of the prediction results of translatable LncRNA-ORFs in buffalo skeletal muscle and cattle skeletal muscle using three algorithms in Example 2;
[0023] Figure 7 This was used to verify the lncRNA-ORFs of the longissimus dorsi muscle of buffalo and cattle in Example 3 by qPCR.
[0024] Figure 8 This is a schematic diagram of the construction of the pEGFP-N1-lncORFs reporter vector in Example 3;
[0025] Figure 9 Fluorescence observation of 293T cells transfected with the LncORF reporter vector in Example 3. Detailed Implementation
[0026] The present invention will be further described below with reference to the embodiments.
[0027] Example 1: Transcriptome, translatome, and proteome analyses were performed on skeletal muscle samples from buffalo and cattle to build libraries.
[0028] 1.1. Experimental Materials
[0029] Longissimus dorsi muscle (n=30) of 24-month-old river buffalo and crossbred yellow cattle was collected from a local slaughterhouse in Xixiangtang District, Nanning City. After removing non-muscle tissue with instruments, the muscle samples were immediately flash-frozen in liquid nitrogen. The samples were stored at -80℃ for transcriptome, translatome sequencing, proteomic analysis, and validation experiments.
[0030] 1.2. Transcriptional library
[0031] Samples of the longissimus dorsi muscle from 30 water buffalo and 30 yellow cattle were pooled separately. Total RNA was extracted using the Trizol method (ref), and ribosomal RNA was removed. The resulting RNA was randomly fragmented into short fragments. Using the fragmented RNA as a template, cDNA was synthesized into the first strand using six-base random hexamers. Then, buffer, dNTPs, RNase H, and DNA polymerase I were added to synthesize the second strand of cDNA. The second strand was purified using a QiaQuick PCR kit, eluted with EB buffer, and underwent end repair, addition of base A, and the addition of sequencing adapters. The second strand was then degraded by UNG (Uracil-N-Glycosylase). Fragment size selection was performed using agarose gel electrophoresis, followed by PCR amplification, lncRNA library construction, and sequencing. Finally, the sequencing library was sequenced using an Illumina HiSeq™ 4000. The buffalo reference genome version used in this study is GCF_003121395.1_UOA_WB_1; the yellow cattle reference genome version is GCF_002263795.1_ARS-UCD1.2.
[0032] 1.3. Translation Group Repository
[0033] The longissimus dorsi muscle samples from 30 buffalo and 30 cattle were separately mixed, ground into powder in liquid nitrogen, and then dissolved in 400 μL of lysis buffer. The mixture was incubated on ice for 10 min, followed by centrifugation at 20,000 g for 10 min at 4°C, and the supernatant was collected. To prepare ribosomal footprints (RFs), 7.5 μL of ribonuclease I and 5 μL of deoxyribonuclease I were added to 300 μL of lysis buffer and gently mixed and incubated on a shaker at room temperature for 45 min. Nuclease digestion was terminated by adding 10 μL of SUPERase·In RNase inhibitor to the lysate. Ribosomes were then recovered, and the liquid was centrifuged at 600 g for 4 min at room temperature using a size exclusion column (illustra Microsin S-400HR column; GE Healthcare; catalog number 27-5140-01). 100 μL of the digested RFs were added to the column and centrifuged at 600 g for 2 min. Next, 10 μL of 10% (wt / vol) SDS was added to the elution buffer, and RFs larger than 17 bp were isolated using the RNA Clean and Concentrator-25 kit (ZymoResearch; R1017). rRNA was removed using a DNA probe complementary to the rRNA sequence. Then, RNase H and DNase I were used to digest the probe, and the RFs were purified using magnetic beads (Vazyme). After obtaining the above ribosomal footprint, [the following steps were performed] using [the appropriate method / method]. Multiple Small RNA Library Prep Set for (catalogno. E7300S, E7300L) Ribo-seq libraries were constructed. In short, adaptors were added to both ends of the RF, followed by reverse transcription and PCR amplification. The PCR products, ranging from 140-160 bp, were enriched, and cDNA libraries were constructed and sequenced using a Denovo Illumina HiSeq™ 2500.
[0034] 1.4. Protein Assembly Library
[0035] Samples of the longissimus dorsi muscle from 30 buffalo and 30 cattle were mixed separately, and protein buffer (500 mM Tris-HCl, 50 mM EDTA, 700 mM sucrose, 100 mM KCl, 2% β-mercaptoethanol, and 1 mM benzyl sulfonyl fluoride, pH 8.0) was added. After grinding, an equal volume of Tris-saturated phenol was added to extract the protein. The mixture was centrifuged at 5500 g at 4°C for 10 min, and the supernatant containing protein was collected. Pre-cooled acetone was then added, and the protein was precipitated at -20°C. The total protein precipitate was dried and dissolved in a solution containing 4% SDS for mass spectrometry analysis. 100 μg of protein from each sample was used for enzymatic digestion using the FASP method, and the digested peptides were then analyzed by LC-MS-MS using Q-EXACTIVE (Thermo, USA). The collected sample data were analyzed for protein identification using ProteomeDiscoverer 2.1.0182 (Thermo Fisher Scientific, Rockford, IL, USA). The relevant data processing parameters were set as follows: Search engine: Sequest HT; Protein database: Uniprot (https: / / www.uniprot.org / uniprot and transcriptome-predicted ncRNAs-encoded peptide library); Enzyme: Trypsin; MissCleavages: 2; Peptide Mass Tolerance: ±10ppm; Fragment Mass Tolerance: ±0.02Da; Peptide FDR: Less than 1%; Protein Q Value: Less than 1%.
[0036] Example 2: Screening for ORFs in long non-coding RNAs from bovine skeletal muscle using multi-omics methods
[0037] 2.1 Obtaining Set 1 using Transcriptome Data: First, the long non-coding RNAs (lncRNAs) expressed in the skeletal muscle of cattle and buffalo were identified using the transcriptome sequencing data obtained in Example 1. 2,263 lncRNAs were identified in cattle skeletal muscle, and 2,343 lncRNAs were identified in buffalo skeletal muscle. To obtain the potential open reading frames (OPGs) of all linear non-coding transcripts, the ORFfinder tool was used to predict all potential OPGs for the above-mentioned lncRNA transcripts. According to the ORFfinder prediction results, the 2,263 lncRNAs expressed in cattle muscle contained a total of 145,459 possible OPGs; the 2,343 lncRNAs expressed in buffalo muscle contained a total of 181,009 possible OPGs. To further evaluate the coding potential of these lncRNA-derived open reading frames (lncRNA-ORFs), we used mRNA coding capacity scores as a reference for coding potential, and the top 95% of mRNAs by coding capacity scores as the criteria for screening potentially coding open reading frames. We then used the Fickett and Hexamer algorithms to calculate the coding potential of the screened potential open reading frames. Finally, as follows... Figure 1 As shown, Figure 1 A represents the screening results for LncRNA-ORFs encoded by bovine skeletal muscle. Figure 1 B represents the screening results of buffalo skeletal muscle encoding LncRNA-ORFs. We screened ORFs with similar mRNA coding capabilities. In the figure, most of the dots representing mRNAs are located in the upper right quadrant, i.e., the quadrant where Hexamer is greater than 0 and Fickett is greater than 0.72. LncRNA-ORFs located in the same quadrant can be regarded as ORFs from non-coding RNAs with higher coding potential. In this way, LncRNA-ORFs falling in other quadrants were excluded, and these screened LncRNA-ORFs formed LncRNA-ORF set 1.
[0038] 2.2. Set 2 was obtained through transcriptome and translatome data analysis: Ribosome-binding transcripts were identified through transcriptome and translatome analysis, and ribosome-binding LncRNA-ORFs were searched. The RRS score of mRNA was used as a reference to evaluate ribosome-binding LncRNA-ORFs. LncRNA-ORFs with an RRS greater than 1.24 were selected, and their intersection with Set 1 was used to form Set 2. Specifically, in this embodiment, the length of ribosome-binding RNA sequences with translatome sequencing characteristics is typically 20-35 bp, such as... Figure 2 As shown, Figure 2 This is a length distribution map of the translated genome sequencing fragments, including all translated ORFs, encompassing both mRNA-derived and lncRNA-derived ORFs. Figure 2 A represents the read length distribution of ribosome-captured transcripts from cattle. Figure 2 B shows the read length distribution of buffalo ribosome-captured transcripts. It can be seen that the transcript read lengths obtained from translatome sequencing of buffalo and cattle skeletal muscle samples are distributed between 20 and 40 bases, with the highest abundance between 25 and 30 bases. This conforms to the basic rules governing the space occupied by ribosome-bound transcripts, indicating that the translatomics data is reliable and can be used for subsequent analysis to identify LncRNA-ORFs.
[0039] To locate the genomic regions of ribosome-binding transcripts, the genomic distribution regions of two bovine skeletal muscle ribosome-captured transcripts were statistically analyzed, such as... Figure 3 As shown, the proportions of cattle transcripts and buffalo transcripts distributed in the coding region were 79.68% and 79.68%, respectively. Figure 3 A) and 85.83% ( Figure 3 B) indicates that the DNA coding region is the main translation region. However, if... Figure 3 As shown in the figure, there are still ribosome-binding transcripts in about 15-20% of the non-coding regions (intron regions, 3' and 5' non-coding regions). These transcripts are usually distinguished from mRNA and are defined as non-coding RNA. Most of the LncRNAs in this example come from this region, and the following analysis will focus on these transcripts.
[0040] In this embodiment, to analyze and determine the initiation and termination locations of LncRNA-ORF translation, the ribosome dwell time was statistically analyzed. For example... Figure 4 As shown in Figure A, the sequence of bovine skeletal muscle translatomes shows a peak at position -12 in the coding region, indicating that the ribosome lingers for a relatively long time 12 bp upstream of the coding start base. A similar pattern is observed in buffalo skeletal muscle translatome data, such as... Figure 4 As shown in B, the sequenced fragment of buffalo skeletal muscle translatome shows a peak at -16 in the coding region, and the ribosomes remain for a relatively long time at a position 16 bp upstream of the coding start base. Figure 4 This indicates that the ribosome lingers for a relatively long period 10-20 bp before the initiation base, and a similar situation exists before the termination base, which will not be elaborated further here. The region between the initiation and termination of translation is the ribosome-binding translation region. Identifying the translation initiation and termination sites provides strong evidence for the search for translatable LncRNA-ORF molecules.
[0041] In addition, to determine the correct LncRNA-ORFs that can encode small peptides, such as Figure 5As shown, three consecutive bases in a ribosomal sequence can form three completely different open reading frames (ORFs). The three different colored bars in the figure represent the relative proportions of the three types of reading frames. This invention evaluated the probability of ORFs formed by three consecutive codons for each ORF and selected the longest and most probable coding reading frame for subsequent analysis. Figure 5 A and Figure 5 In section B, for a 27-base-length segment, the three columns from left to right represent the ORFs encoded by the first to third bases, respectively. The segment represented by the middle column, which encodes the second base and has the highest relative proportion and the longest possible length, is the analyzable reading frame. This method ensures that the ORF analysis results are as long as possible, thus guaranteeing the integrity of the ORFs.
[0042] To further reduce the error rate of transcriptome bioinformatics analysis prediction results, such as Figure 6 A and Figure 6 As shown in Figure B, a Venn diagram is plotted using the analysis results of the three methods for analyzing the codeability of LncRNA-ORFs, thus obtaining LncRNA-ORFs that simultaneously meet the screening criteria of the three algorithms. Figure 6 The Fickett score and ORF score are known and conventional bioinformatics analysis methods for assessing coding ability, and are algorithms already known to those skilled in the art. They serve as reference methods for evaluating LncRNA-ORFs, and their specific algorithms will not be elaborated further. The ribosome release score (RRS), as the core analysis algorithm of this invention, combines translatomics and transcriptomics data, and ultimately, by integrating them with the Fickett score and ORF score, obtains the encoding LncRNA-ORFs.
[0043] Specifically, the formula for calculating RRS is RRS = (RPKMCDS / RPKM3′UTR) Ribo-seq / (RPKMCDS / RPKM3′UTR) RNA-seq In the formula, RPKMCDS represents the expression level of transcripts located in the coding region, RPKM3′UTR represents the expression level of transcripts located in the 3' non-coding region upstream of the coding region, Ribo-seq indicates that the data analyzed in this part comes from translatome sequencing data, and RNA-seq indicates that the data analyzed in this part comes from transcriptome sequencing data. This invention calculates the RRS of all LncRNA-ORFs and mRNAs in the transcriptome and translatome data, and uses the top 95% of the mRNA RRS as a reference. Finally, LncRNA-ORFs with RRS values greater than 1.24 are identified as candidate LncRNA-ORFs. These candidate LncRNA-ORFs are then intersected with set 1 to form set 2. For example... Figure 6In set A, these LncRNA-ORFs are also present in set 1, specifically in the 932 buffalo muscle-derived LncRNA-ORFs at the central intersection. Figure 6 The 747 bovine muscle-derived LncRNA-ORFs in the central intersection region of B simultaneously met the screening criteria of the Fickett score, ORF score, and RRS algorithm, and were all identified as encoding LncRNA-ORFs. These LncRNA-ORFs were used as set 2.
[0044] 2.3 Obtaining Set 3 through Proteomic Data: Proteomic analysis of all small peptides in muscle is performed, and these are compared with theoretically translatable small peptides formed by the ORFs screened in Set 2. Only ORFs that can be uniquely matched are selected as Set 3. The ORFs in Set 3 are the ORFs that can be encoded by the long non-coding RNAs of skeletal muscle. Specifically, Set 2 is first converted into theoretically translatable small peptide sequences. Then, the mass spectra of skeletal muscle expressed proteins and small peptides are analyzed using liquid chromatography-mass spectrometry (LC-MS / MS) and compared with the aforementioned theoretically translatable small peptide sequences and the uniprot public protein database. Results comparing with known protein sequences in the uniprot public protein database are removed. Non-coding polypeptides with a uniquely matched peptide number greater than or equal to 1 are listed as candidate potential coding RNAs, thereby identifying new polypeptides encoded by non-coding RNAs. Finally, the identified new polypeptides encoded by non-coding RNAs are used as the basis for determining the non-coding RNA open reading frames as Set 3.
[0045] This embodiment confirms the existence of small peptides encoded by lncRNA-ORFs in the total protein expressed in muscle tissue by analyzing mass spectrometry data and protein-encoding data generated from proteomics. This step completes the screening process in the bioinformatics step. The specific analysis software and parameters are as follows: The collected sample data were analyzed for protein identification using Proteome Discoverer 2.1.0182 software. The main data analysis parameters were set as follows: search engine: Sequest HT; protein databases: Uniprot and lncRNA-ORFs encoded peptide library; digestive enzyme: Trypsin; peptide mass tolerance: ±10ppm; peptide molecular weight tolerance: ±0.02Da; FDR: less than 1%.
[0046] By searching and comparing the database, information on the peptides encoded by LncRNA-ORFs in Table 1 was obtained. Among them, 17 small peptides of cattle LncRNA-ORFs and 11 small peptides of buffalo LncRNA-ORFs were found to have unique corresponding peptide segments in the mass spectrometry data. This indicates that these LncRNA-ORFs have identified specific encoded small peptides in the total protein expressed in muscle tissue, which further strengthens the evidence that they can encode. These LncRNA-ORFs are used as set 3. The ORFs in set 3 are the ORFs that can be encoded in the long non-coding RNA of skeletal muscle obtained by multi-omics joint screening.
[0047] Table 1. Mass spectrometry identification results of small peptides encoded by LncRNA-ORFs
[0048]
[0049]
[0050] Example 3: Verification of whether the screened LncRNA-ORFs can initiate the encoding of small peptides in cells.
[0051] 3.1 Detecting whether the screened LncRNA-ORF encoding LncRNAs are truly expressed using qPCR assays.
[0052] The LncRNA-ORFs in set 3 obtained in Example 2 were transcripts that were most likely encoding polypeptides after final screening. To verify whether these LncRNA-ORFs were truly expressed at the transcriptional level, we randomly selected 28 LncRNA-ORFs (11 from buffalo and 17 from yellow cattle) from set 3 for qPCR expression verification. The LncRNA-ORF verification primers used in this example are shown in Table 2, and the qPCR method was the same as the conventional procedure.
[0053] like Figure 7 A and Figure 7 The LncRNA-ORF transcripts used for verification, as shown in B, were all tested for expression levels using qPCR. The results showed that the expression abundance of different LncRNA-ORF transcripts varied, and all of them had expression levels greater than 0. This experiment proved that the LncRNA-ORFs screened in Example 2 were all truly expressed transcripts, and the results obtained from sequencing screening were accurate.
[0054] Table 2 Primer sequences used for LncRNA-ORF detection
[0055]
[0056]
[0057]
[0058] In the table, B is the abbreviation for buffalo, C is the abbreviation for cattle, F is the upstream primer, and R is the downstream primer.
[0059] 3.2 Constructing the Report Carrier
[0060] Two LncRNA-ORFs were randomly selected from the above 28 ORFs. The cattle LncORF-39162 was named Lnc253-459 according to its positional relationship, and the buffalo LncORF-804 was named Lnc1168-926 according to its positional relationship.
[0061] The four types of carriers constructed are as follows Figure 8 As shown: EGFP is the green fluorescent protein gene used to indicate whether the vector has been translated correctly. Figure 8 A is the modified vector pEGFP-N1, which underwent point mutation via PCR, changing the EGFP start codon ATG to ATT. Figure 8 B's pEGFP-N1, without EGFP start codon modification, was used as a control vector. Subsequently... Figure 8 C represents the LncRNA-ORFs to be validated. This invention validates Lnc253-459 and Lnc1168-926. Four vectors for each of these two LncRNA-ORFs were constructed, resulting in two sets of vectors and a total of eight vectors. The start codon ATG was mutated to ATT using PCR point mutation, and the corresponding... Figure 8 The D start codon ATG was not mutated, serving as a control LncRNA-ORF fragment.
[0062] Figure 8Fragments A, B, C, and D were ligated as follows to form four vectors. The expression of these vectors was used to determine whether the lncRNA-ORF could encode a peptide. The lncRNA-ORF fragments were ligated into the multiple cloning site (MCS) of the pEGFP-N1 vector via Xho I and BamHI restriction sites. A was ligated to C and A to D to form the pEGFPmut-LncORF and pEGFPmut-LncORFmut vectors, respectively. Similarly, B was ligated to C and B to D to form the pEGFP-LncORF and pEGFP-LncORFmut vectors. The pEGFPmut-LncORF vector served as the experimental group, used to detect whether the ligated LncORF translated into a small peptide. pEGFPmut-LncORFmut was the negative control group with a double mutation, indicating that neither the vector nor LncORF possessed translational ability when both start codons were mutated. pEGFP-LncORF was the positive control group, indicating that both EGFP and LncORF could be translated to display green fluorescent protein, demonstrating normal translatability of the vector. pEGFP-LncORFmut was the single mutation control group, indicating that even when LncORF could not be translated, the EGFP protein had the ability to independently translate the protein, demonstrating that the two parts were independent. All four vectors ensured that LncORF and EGFP sequences within the same reading frame could be translated into fusion-expressed proteins. After successful ligation, the vector construction was completed after verification by Sanger sequencing and BLAST alignment.
[0063] 3.3 Transfect cells, observe fluorescence, and interpret the results.
[0064] The four reporter vectors (pEGFPmut-LncORF, pEGFPmut-LncORFmut, pEGFP-LncORF, and pEGFP-LncORFmut) were transfected into 293T cells. The transfection method was the same as conventional molecular biology methods and will not be described in detail here. After culturing the transfected cells for 24 hours, the presence of green fluorescence signals in the 293T cells was observed to visually determine whether LncRNA-ORF could encode a small peptide. In this invention, EGFP and LncRNA-ORF are within the same reading frame. Therefore, when LncRNA-ORF is translated, the mutated EGPF start codon does not stop translation. The small peptide encoded by LncRNA-ORF will form a fusion protein with EGFP, thus determining whether LncRNA-ORF is translated. Furthermore, EGPF without a mutated start codon can independently translate to produce green fluorescent protein and can be used as a positive control in this experiment.
[0065] Using this LncRNA-ORF reporter vector combination, this invention detected that two LncRNA-derived reading frames, Lnc253-459 and Lnc1168-926, encode small peptides. Here, Lnc253-459 (… Figure 9 Taking A as an example, please explain how to interpret the results. Figure 9 The observation method for the results in group B, Lnc1168-926, is the same. (For example...) Figure 9 As shown in Figure A, the start codon of the Lnc253-459 sequence in the pEGFP-Lnc253-459mut vector was mutated, preventing the translation of the small peptide. However, EGFP fluorescence signal was still observed in the figure, indicating that the EGFP start codon of the reporter gene was not mutated, and its translation was not affected by the inability to translate Lnc253-459. This vector serves as a control vector, demonstrating that EGFP expression function is independent when the reporter gene is not mutated. The observation of EGFP signal in the pEGFPmut-Lnc253-459 vector indicates that normal translation of the preceding Lnc253-459 gene drives the expression of the EGFP gene within the same reading frame with a mutated start codon. This directly proves that Lnc253-459 has translational capability, making this vector a key vector for indicating whether LncRNA-ORF can be translated. When the start codons of both Lnc253-459 and EGFP were mutated to ATT, the 293T cells transfected with the vector showed no green fluorescence signal. This indicates that the constructed mutant reporter vector can effectively indicate the lack of translation ability of LncORF, and this vector served as the negative control in the experiment. The last group, pEGFP-Lnc253-459, had neither Lnc253-459 nor EGFP mutated. It normally expressed small peptides and EGFP protein, and showed a strong green fluorescence signal, serving as the positive control in this experiment. The reporter vector combination constructed in this invention uses pEGFPmut-LncORF as the key reporter vector to indicate whether LncRNA-ORF can encode peptides. When the pEGFPmut-LncORF vector expresses green fluorescent protein in 293T cells, it indicates that LncRNA-ORF can encode small peptides. Conversely, when the pEGFPmut-LncORF vector does not express green fluorescent protein in 293T cells, it indicates that LncRNA-ORF cannot encode small peptides. The other three vectors can effectively indicate whether this report system is functioning properly. When the other three vectors are expressed normally in 293T cells, the accuracy of the experimental results is guaranteed.
[0066] Therefore, the scheme described in Example 3 can effectively identify whether LncRNA-ORF can encode small peptides in cells, further verifying that the LncRNA-ORFs obtained by the multi-omics joint screening of the present invention truly encode small peptide products.
[0067] The verification in this embodiment demonstrates that the skeletal muscle long non-coding RNAs obtained by the screening method of the present invention generally have the ability to encode small peptides within the cell. The method of the present invention can greatly reduce the false positive problem of the screened LncRNAs encoding ORFs by combining multiple omics approaches.
[0068] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of equivalents of the claims be included within the present invention.
[0069] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for analyzing, identifying, and screening encoding ORFs in long non-coding RNAs of skeletal muscle, characterized in that, Includes the following steps: Step 1) Take skeletal muscle samples and perform library construction and sequencing on the transcriptome, translatome, and proteome respectively. The library construction and sequencing methods are consistent with the conventional procedures, and transcriptome data, translatome data, and proteome data are obtained respectively. Step 2) Analyze the transcriptome data obtained in Step 1) and screen out ORFs with potential coding capabilities within long non-coding RNAs as set 1; Step 3) Analyze the transcriptome and translatome data obtained in Step 1) and calculate the ribosome release score (RRS) of ORFs of long non-coding RNAs and mRNAs respectively. ORFs with an RRS score greater than 1.24 are selected as candidate ORFs. The intersection of these candidate ORFs with the above set 1 is taken to form set 2. The formula for calculating RRS is: , in the formula Indicates the expression level of transcripts located in the coding region, This indicates that it is located 3 upstream of the coding area. ’ Expression levels of non-coding transcripts This indicates that the data analyzed in this section comes from translatome sequencing data. This indicates that the data analyzed in this section comes from transcriptome sequencing data; Step 4) Analyze all small peptides in skeletal muscle using the proteomic data obtained in Step 1), compare them with the theoretically translatable small peptides that can be formed by the ORFs screened in Set 2, and screen out the ORFs that can be uniquely matched as Set 3. The ORFs in Set 3 are the ORFs that can be encoded in the long non-coding RNA of skeletal muscle. The step 2) of screening ORFs with potential coding ability within long non-coding RNAs as set 1 is as follows: First, the ORFfinder tool is used to predict all potential ORFs within long non-coding RNAs in the transcriptome data, and ORFs with the start codon ATG and a length range of 60~450 bp are extracted; then, the fickett and hexamer algorithms are used to select ORFs with potential coding ability from the extracted ORFs, with the top 95% of mRNA coding ability scores as the criteria for screening.
Citation Information
Patent Citations
Analysis method of long-chain non-coding RNA translation small peptide based on translation group
CN110556163A
Method for identifying non-coding RNA polypeptide
CN114038500A