Method for high-throughput identification of lncrna-encoded peptides and application thereof
By constructing an index of lncRNA-encoded peptides using high-throughput identification methods, the problems of incomplete identification and indexing in existing technologies have been solved. This enables accurate subtyping and prognostic assessment of lncRNA-encoded peptides in cancer patients, providing an effective basis for clinical diagnosis and treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENYANG PHARMA UNIV
- Filing Date
- 2022-11-08
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies cannot effectively identify lncRNA-encoded peptides, resulting in incomplete indexes. Furthermore, algorithmic flaws cause small peptides to be discarded, leading to a lack of reliable annotations and indexes, which makes them unsuitable for accurate subtyping and prognostic assessment of clinical patients.
High-throughput identification methods were employed, and an open reading frame library was constructed using the GENCODE and NONCODE databases. Combined with proteogenomics indexing and mass spectrometry data analysis, computational proteomics analysis was performed using MaxQuant software to screen for lncRNA-encoded peptides. Prognostic models were then constructed using LASSO regression and nonnegative matrix factorization algorithms.
A comprehensive index of human lncRNA-encoded peptides has been established, enabling high-throughput identification and quantitative detection of lncRNA-encoded peptides. This information can be used for precise subtyping and prognosis of cancer patients, providing a basis for clinical diagnosis and treatment.
Smart Images

Figure CN115985397B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to a high-throughput identification method for lncRNA-encoded peptides and its application in precise subtyping of clinical patients and construction of prognostic models. Background Technology
[0002] Long non-coding RNAs (lncRNAs) are traditionally defined as transcripts longer than 200 nucleotides that lack coding potential, playing crucial regulatory roles in normal physiological and pathological processes. Recent studies have shown that some lncRNAs also possess the potential to encode bioactive small peptides. For example, lncRNA HOXB-AS3 encodes a 53-amino acid peptide that can inhibit the progression of rectal cancer; rectal cancer patients with low expression of this lncRNA exhibit poorer prognosis (Huang et al., MolCell, 2017). The lncRNA gene LINC-00266-1 encodes a 71-amino acid peptide RBRP, which is significantly upregulated in colorectal cancer tissues; colorectal cancer patients with high RBRP expression levels show even worse prognosis (Zhu et al., Nat Commun, 2020). These studies suggest that small peptides encoded by lncRNAs are potential targets for tumor therapy and could be developed as biomarkers for clinical diagnosis and prognosis in cancer patients.
[0003] However, achieving this goal currently faces the following challenges: First, there is a lack of reliable annotations and indexes, making it impossible to identify lncRNA-encoded peptides (lncPEPs) through routine analysis of mass spectrometry data; second, the existing annotated lncRNAs represent only a small fraction of the actual lncRNAs, resulting in an incomplete index built upon them; and third, algorithmic defects lead to the discarding of a large number of small peptides that are not ATG translation initiation sites.
[0004] In recent years, the development of ribosomal blotting technology and proteogenomics tools has provided powerful data resources and analytical methods for solving the aforementioned challenges. Therefore, there is an urgent need to develop a high-throughput method and workflow for identifying lncRNA-encoded peptides. Based on this, lncRNA-encoded peptides can be further developed as biomarkers for tumor molecular subtyping and prognosis. Summary of the Invention
[0005] The purpose of this invention is to develop a method for high-throughput detection of lncRNA-encoded peptides and its application in precise subtyping and prognostic model construction for clinical patients.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A high-throughput identification method for lncRNA-encoded peptides is performed according to the following steps:
[0008] Step 1: Identify potential lncRNA open reading frames;
[0009] A candidate open reading frame (OPG) library was constructed using human gene annotations from the GENCODE and NONCODE databases. Human Ribo-seq data was collected to evaluate the tribase periodicity of the candidate OPGs.
[0010] Step 2: Construct a proteogenomics index;
[0011] The obtained open reading frame nucleotide sequences were translated into peptide sequences using the Biostrings toolkit, and the peptide index was supplemented by integrating the SmProt and cncRNAdb databases. The blastp algorithm was used to align the peptide sequences to known human protein data, and peptide index entries that could be completely aligned to known proteins were removed.
[0012] Step 3: Quantitative proteomics analysis;
[0013] The obtained protein genomic index was used for computational proteomics analysis of lncRNA-encoded peptides using computational proteomics methods. The specific steps are as follows:
[0014] Quantitative proteomics mass spectrometry data of cancer tissues and their corresponding non-cancerous tissues from cancer patients were collected from public protein mass spectrometry databases. The obtained lncRNA-encoded peptide index and the human proteome reference sequence provided by Uniprot (https: / / www.uniprot.org / ) were used as common indexes, and computational proteomics analysis was performed using MaxQuant (https: / / www.maxquant.org / ) software.
[0015] Step 4: Filtering lncRNA-encoded peptide fragments;
[0016] The lncRNA-encoded peptide fragments obtained by the above methods are screened, specifically including:
[0017] Remove all peptide fragments that can be completely aligned to known proteins;
[0018] Remove all peptide fragments that can be completely aligned to known nonsynonymous single nucleotide polymorphism events;
[0019] Use a random sample consistency regression model to remove all peptide fragments whose corrected retention time deviates from the predicted retention time.
[0020] Step 5: Determine the strength of the lncRNA-encoded peptide;
[0021] The final peptide fragment intensities are merged, and the open reading frame groups are merged.
[0022] A class of lncRNA-encoded peptides, serving as tumor prognostic biomarkers, are obtained using a high-throughput identification method for these lncRNA-encoded peptides. The application of these lncRNA-encoded peptides as tumor biomarkers in a risk score prediction model for predicting the prognosis of hepatocellular carcinoma patients specifically includes the following steps:
[0023] S1. Based on patient survival and relapse information, lncRNA-encoded peptide strength information (lncRNA-encoded peptide strength obtained by the high-throughput identification method), the LASSO regression model is used to obtain the several lncRNA-encoded peptides that have the greatest impact on patient survival and relapse.
[0024] S2. Using the Cox regression model, calculate the Coef value and P value of the correlation between each lncRNA-encoded peptide and patient survival and relapse.
[0025] S3. Based on the correlation coefficient between lncRNA-encoded peptides and patient survival and relapse, a risk score is calculated for patient survival and relapse. The formula is: Risk score = Coef(lncPEP 1) × Exp(lncPEP 1) + Coef(lncPEP 2) × Exp(lncPEP 2) + ... + Coef(lncPEP N) × Exp(lncPEP N).
[0026] For the differentiation of cancer patients, based on the intensity of lncRNA-encoded peptides obtained by the high-throughput identification method, a non-negative matrix factorization algorithm is used to perform precise molecular subtyping of cancer patients.
[0027] The application of the lncRNA-encoded peptide as a tumor biomarker in the precise molecular subtyping of cancer patients specifically includes the following steps:
[0028] Based on the lncRNA-encoded peptide intensity obtained from the high-throughput identification method, a non-negative matrix factorization algorithm was used to perform precise molecular typing of cancer patients. Specifically, lncPEPs with an absolute median difference in mass spectrometry intensity greater than 0 in cancer tissue samples were used as input. The nmf function of the NMF package (https: / / cran.r-project.org / web / packages / NMF / index.html) was used to perform non-negative matrix factorization analysis on the cancer patients.
[0029] Compared with the prior art, the present invention has the following beneficial effects.
[0030] This invention establishes a comprehensive and complete index of human lncRNA-encoded peptides. The method described in this invention can identify lncRNA-encoded peptides in proteomic samples and quantify their expression levels. This invention also provides a predictive model based on a lncRNA-encoded peptide network to predict the prognosis of cancer patients, and validates the feasibility of this model in a dataset, providing a basis for the precise diagnosis and treatment of clinical liver cancer patients. Attached Figure Description
[0031] Figure 1 This is the technical solution process adopted in this invention.
[0032] Figure 2 AB is a representative lncRNA-encoded peptide, along with its peptide fragment intensity heatmap and mass spectrometry.
[0033] Figure 3 AB is a peptide intensity heatmap of all 16 significantly upregulated lncRNA-encoded peptides and 6 significantly downregulated lncRNA-encoded peptides in liver cancer samples, as well as box plots of some representative lncRNA-encoded peptides.
[0034] Figure 4 AB represents the results of a univariate Cox regression analysis of the intensity of the ten lncRNA-encoded peptides that have the greatest impact on survival (A) and recurrence (B) in liver cancer patients.
[0035] Figure 5 AB represents the results of Cox regression analysis of the survival score and recurrence score of liver cancer patients.
[0036] Figure 6 A and B are a comparison of survival rates between patients with high survival risk scores and patients with low survival risk scores (A), and a comparison of recurrence rates between patients with high recurrence risk scores and patients with low recurrence risk scores (B).
[0037] Figure 7 AC represents the molecular grouping of liver cancer patients based on the intensity of lncRNA-encoded peptides. A is a heatmap of consistency between the two subgroups of patients, B is a survival analysis of the two subgroups of patients, and C is the chi-square test result of the subgroup division and whether the patients have cirrhosis. Detailed Implementation
[0038] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The following description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and additions without departing from the method of the present invention, and these improvements and additions should also be considered within the scope of protection of the present invention.
[0039] Example
[0040] A high-throughput identification method for lncRNA-encoded peptides, such as Figure 1 As shown, the specific operation steps are as follows:
[0041] Step 1: Identify potential lncRNA open reading frames;
[0042] Human Ribo-seq data were retrieved and downloaded from the GEO database (https: / / www.ncbi.nlm.nih.gov / geo / ). The following analysis was performed on the Ribo-seq data: First, quality control was performed on the sequencing data. Adapters were removed using Trim Galore software (https: / / www.bioinformatics.babraham.ac.uk / projects / trim_galore / ). The sequencing data were then aligned to the human reference genome using STAR software (https: / / github.com / alexdobin / STAR). The tribase periodicity of the sequencing fragments aligned to the human genome was assessed using Ribotricer software (https: / / github.com / smithlabcode / ribotricer). A candidate open reading frame (ORF) library was constructed using human gene annotations provided by the GENCODE (https: / / www.gencodegenes.org / pages / biotypes.html) and NONCODE (http: / / noncode.org / ) databases. Phase scoring was performed, and a scoring threshold of 0.44 was used to obtain lncRNAs with coding potential and their ORF information.
[0043] Step 2: Construct a proteogenomics index;
[0044] Open reading frames (OLFs) of lncRNAs from different samples were collected, and redundant ORFs were removed based on the following criteria: if an ORF has the same stop codon and is located on the same transcript, the longest ORF is retained; if an ORF has the same stop codon but is located on different transcripts, each ORF is retained separately. The nucleotide sequences were translated into peptide sequences using the translate function of the Biostrings toolkit (https: / / bioconductor.org / packages / release / bioc / html / Biostrings.html). The SmProt (http: / / bigdata.ibp.ac.cn / SmProt / index.html) and cncRNADB (http: / / www.rna-society.org / cncrnadb / ) databases were integrated as supplementary databases. The BLASTP algorithm was used to align peptide sequences to known human protein data, and peptide entries that could be completely aligned to known proteins were removed.
[0045] Step 3: Quantitative proteomics analysis;
[0046] Label-free quantitative proteomics mass spectrometry data of cancer tissues and corresponding adjacent non-cancerous tissues from 96 patients with early-stage hepatocellular carcinoma were collected from the iProX proteomics integrated resource library (iProX, http: / / 111.198.139.98 / page / home.html). Using the obtained lncRNA-encoded peptide index and the human proteome reference sequence provided by Uniprot (https: / / www.uniprot.org / ) as common indexes, computational proteomics analysis was performed using MaxQuant (https: / / www.maxquant.org / ) software.
[0047] Step 4: Filtering lncRNA-encoded peptide fragments;
[0048] The possible lncRNA-encoded peptide fragments obtained from the above process are further filtered. Human nonsynonymous mononucleotide diversity data are downloaded from dbSNP (https: / / www.ncbi.nlm.nih.gov / snp / ), and diversity protein data are constructed based on Uniprot sequences. The blastp algorithm is used to align the obtained lncRNA peptide fragments to the diversity protein data and the known protein data, respectively, and the lncRNA peptide fragments that can be completely aligned to both the diversity protein data and the known protein data are filtered out.
[0049] The R language's `sample` function was used to randomly select 2,000 known protein peptide fragments from the top 20% of Andromeda scores and the bottom 20% of posterior error probabilities (PEPs). These fragments were then used as input to correct the DeepLC model (https: / / github.com / compomics / DeepLC). The corrected DeepLC model was then used to predict the retention time of lncRNA peptide fragments. Using the predicted and measured retention times, a robust linear regression model with Random Sample Consistency (RANSAC) regression algorithm was applied to filter out peptide fragments with excessively large discrepancies between the predicted and measured retention times.
[0050] Step 5: Determine the strength of the lncRNA-encoded peptide;
[0051] The final peptide fragment intensities are combined to merge open reading frame (ORF) groups based on the following criteria: if two ORFs share evidence of a common peptide fragment and are located on different transcripts of the same gene, they are merged and named after the gene they share; if two ORFs share evidence of a common peptide fragment but are located on different genes, they are discarded because their specific origin cannot be determined.
[0052] The lncRNA-encoded peptides detected above were used as tumor biomarkers:
[0053] Using the intensity of lncRNA-encoded peptides with an absolute median difference greater than 0 as input, LASSO regression analysis was performed on the impact of lncRNA-encoded peptide intensity on patient survival and recurrence based on survival and prognostic information of hepatocellular carcinoma patients. Cox regression analysis was conducted on the top ten lncRNA-encoded peptides with the greatest impact on patient survival and recurrence, and the correlation coefficients (Coef value and P-value) were calculated. For lncRNA-encoded peptides with a P-value less than 0.05, patient survival and recurrence risk scores were calculated using the formula: Risk score = Coef(lncPEP 1) × Exp(lncPEP 1) + Coef(lncPEP 2) × Exp(lncPEP 2) + ... + Coef(lncPEP N) × Exp(lncPEP N).
[0054] Nonnegative matrix factorization algorithm:
[0055] The lncPEP values with a median absolute difference in mass spectrometry intensity greater than 0 were used as inputs (12 values) across all 96 cancer tissue samples. The NMF package (https: / / cran.r-project.org / web / packages / NMF / index.html) nmf function was used to perform nonnegative matrix factorization (NMF) analysis on liver cancer patients, calculating the score metric in 50 iterations and clustering in 400 iterations.
[0056] Test results:
[0057] 1. Identification and quantitative detection of lncRNA-encoded peptides in hepatocellular carcinoma tissues
[0058] From the proteomic data of all 96 pairs of patients with early-stage hepatocellular carcinoma, we ultimately obtained 1094 open reading frames (ORFs). 53% were from non-coded lncRNA genes, 23% from genetically encoded lncRNA genes, and 24% were genetically encoded pseudogenes. The mass spectra and expression heatmaps of representative lncRNA-encoded peptides are shown in the figure. Figure 2 (A-2B). Differential analysis was performed on 99 lncRNA-encoded peptides that appeared in more than 5% of the samples and multiple times in both cancer and non-cancer samples. A p-value less than 0.05 and an absolute logFC value greater than 0.5 were used as thresholds. Of these, 16 showed significant upregulation and 6 showed significant downregulation. Figure 3 A-3B).
[0059] 2. Using lncRNA-encoded peptides as biomarkers has good prognostic value.
[0060] Cox regression analysis was performed on the ten lncPEPs that had the greatest impact on survival time and disease-free survival, with a p-value less than 0.05 as the threshold. These included five lncRNA-encoded peptides: NONHSAG004137.2_ORF1, ENSG00000213513_ORF1, ENSG00000235748_ORF1, ENSG00000229000_ORF1, and ENSG00000261553_ORF1, whose intensity significantly affected patient survival. Three lncRNA-encoded peptides, including NONHSAG060836.1_ORF1, ENSG00000213513_ORF1, and ENSG00000237757_ORF1, had a significant impact on patient relapse. Figure 4(A-4B). We used the intensity of several significantly influential lncRNA-encoded peptides to score patient survival and recurrence, assessing survival score and recurrence score for 96 patients. Univariate Cox regression analysis was performed on both scores based on patient survival and recurrence information. The survival score had a significant impact on patient survival, and the recurrence score had a more significant impact on patient recurrence than other clinical characteristics. Figure 5 A-5B).
[0061] In survival analysis, there were significant differences in survival and recurrence probability between patients with the two scores. Figure 6 (A-6B) Based on two scores, the survival and recurrence status of patients in the second year were predicted (at this time, among all 96 patients, 42 relapsed, 54 did not relapse, 26 died, and 70 survived). Time ROC curves were plotted. The AUC for predicting patients' survival in the second year based on the survival score was 0.846, and the AUC for predicting patients' recurrence in the second year based on the recurrence score was 0.749.
[0062] For all 96 patients with hepatocellular carcinoma, the tissues were well divided into two subgroups ( Figure 7 A), of which Group 1 had 28 patients and Group 2 had 68 patients. Survival analysis showed that the survival time of patients in Group 1 was significantly shorter than that of patients in Group 2. Figure 7 B), disease-free survival time was not significantly affected. This subgroup classification was significantly correlated with whether a patient had cirrhosis in their clinical characteristics (chi-square test P < 0.01). Figure 7 C).
[0063] In summary, the above studies demonstrate that lncRNA-encoded peptides have good potential as biomarkers for predicting patient prognosis.
Claims
1. A method for high-throughput identification of lncRNA-encoded peptides, characterized in that, Follow these steps: Step 1: Identify potential lncRNA open reading frames; A candidate open reading frame library was constructed using human gene annotations from the GENCODE and NONCODE databases. Human Ribo-seq data was collected, and the tribase periodicity of the candidate open reading frames was evaluated. First, the sequencing data underwent quality control. Trim Galore software was used to remove adapters. STAR software was used to align the sequencing data to the human reference genome. Ribotricer software was used to evaluate the tribase periodicity of the sequencing fragments aligned to the human genome. A candidate open reading frame library was constructed using GENCODE and NONCODE database annotations. Phase scoring was performed, and 0.44 was used as the scoring threshold to obtain lncRNAs with coding potential and their open reading frame information. Step 2: Construct a proteogenomics index; Open reading frames (OPFs) of lncRNAs from different samples were collected, and redundant ORFs were removed. The criteria were as follows: if an ORF has the same stop codon and is located on the same transcript, the longest ORF is retained; if an ORF has the same stop codon but is located on different transcripts, each ORF is retained separately. The nucleotide sequences were translated into peptide sequences using the translate function of the Biostrings toolkit. The SmProt and cncRNADB databases were integrated as supplements, and the blastp algorithm was used to align the peptide sequences to known human protein data. Peptide entries that could be completely aligned to known proteins were removed. Step 3: Quantitative proteomics analysis; Quantitative proteomics mass spectrometry data of cancer tissues and corresponding non-cancerous tissues from cancer patients were collected from public protein mass spectrometry databases. The obtained lncRNA-encoded peptide index and the human proteome reference sequence provided by Uniprot were used as common indexes, and computational proteomics analysis was performed using MaxQuant software. Step 4: Filtering lncRNA-encoded peptide fragments; The obtained possible lncRNA-encoded peptide fragments were further filtered; human nonsynonymous mononucleotide diversity data were downloaded from dbSNP, and diverse protein data were constructed based on Uniprot sequences; the lncRNA peptide fragments were aligned to the diverse protein data and known protein data respectively using the blastp algorithm, and the lncRNA peptide fragments that could be completely aligned to the diverse protein data and known protein data were filtered. The R language `sample` function was used to randomly select 2,000 known protein peptide fragments from the top 20% with the highest Andromeda scores and the bottom 20% with the lowest error probabilities. Based on their amino acid sequences, modifications, and retention times, the DeepLC model was corrected. The corrected DeepLC model was then used to predict the retention times of lncRNA peptide fragments. Using the predicted and measured retention times, a robust linear regression model with random sample consistency regression algorithm was used to filter and remove peptide fragments with excessively large differences between the predicted and measured retention times. Step 5: Determine the strength of the lncRNA-encoded peptide; The final peptide fragment intensities are merged, and the open reading frame groups are merged.
2. Application of lncRNA encoding peptide as a tumor biomarker in constructing a risk score prediction model for predicting the prognosis of cancer patients, characterized in that, The lncRNA encodes peptides NONHSAG004137.2_ORF1, ENSG00000213513_ORF1, ENSG00000235748_ORF1, ENSG00000229000_ORF1, ENSG00000261553_ORF1, NONHSAG060836.1_ORF1, and ENSG00000237757_ORF1; the application method specifically includes the following steps: S1. Based on patient survival and relapse information, lncRNA-encoded peptide strength information, LASSO regression model was used to screen out several lncRNA-encoded peptides that have the greatest impact on patient survival and relapse. S2. Using the Cox regression model, calculate the Coef value and P value of the correlation between each screened lncRNA-encoded peptide and patient survival and relapse. S3. Based on the correlation coefficient Coef value obtained in step S2, construct the risk score formula: Risk score = Coef(lncPEP 1)×Exp (lncPEP 1) + Coef (lncPEP 2)×Exp (lncPEP 2) + ··· +Coef (lncPEP N)×Exp (lncPEP N).
3. The use of lncRNA-encoded peptides as tumor biomarkers in the construction of a molecular typing model for precision medicine for cancer patients, characterized by, The lncRNA encodes peptides NONHSAG004137.2_ORF1, ENSG00000213513_ORF1, ENSG00000235748_ORF1, ENSG00000229000_ORF1, ENSG00000261553_ORF1, NONHSAG060836.1_ORF1, and ENSG00000237757_ORF1; the application method specifically includes the following steps: Using lncPEP values where the absolute median difference in mass spectrometry intensity is greater than 0 in cancer tissue samples as input, the NMF package's nmf function is used to perform non-negative matrix factorization algorithm analysis on cancer patients.