A method for establishing a biomarker recognition model for oral squamous cell carcinoma

By processing and analyzing transcriptomic data from the OSCC tumor database, using Perl and R languages ​​to mine differentially expressed RNA factors, and combining this with a Cox regression model, an OSCC biomarker identification model was established. This model addresses the problem of limited identification methods in existing technologies, improving the accuracy of identification and treatment efficacy.

CN115762644BActive Publication Date: 2026-04-03CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-24
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing methods for identifying OSCC biomarkers are limited and lack specificity, resulting in unsatisfactory treatment outcomes, especially in patients with advanced clinical stages.

Method used

We used Perl and R to process and analyze transcriptome data from the OSCC tumor database, mine differentially expressed RNA factors in OSCC transcriptomes, analyze their relationship with immune infiltration, screen RNA factors significantly associated with survival based on Cox regression model, and establish a biomarker recognition model.

Benefits of technology

This improves the accuracy of OSCC biomarker identification, enabling a better understanding of the tumor immune microenvironment and helping to improve treatment outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762644B_ABST
    Figure CN115762644B_ABST
Patent Text Reader

Abstract

This application discloses a method for establishing a biomarker recognition model for oral squamous cell carcinoma (OSCC). The method includes processing and analyzing transcriptomic data from an OSCC tumor database to identify differentially expressed RNA factors in the OSCC transcriptome; analyzing the immune infiltration relationship between OSCC cell abundance and differentially expressed factors; evaluating the actual impact of the identified differentially expressed factors on tumorigenesis and development mechanisms; screening and retaining RNA factors significantly related to survival; and identifying biomarkers that can be used as OSCC. This invention provides a method for establishing a biomarker recognition model for oral squamous cell carcinoma, which can improve the sensitivity and accuracy of biomarker identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical big data technology, and in particular to a method for establishing a biomarker identification model for oral squamous cell carcinoma. Background Technology

[0002] Transcriptomics is a discipline that studies gene transcription and its regulatory mechanisms at the cellular level. It studies gene expression at the RNA (ribonucleic acid) level. The transcriptome is the sum of all RNA that a living cell can transcribe. Research on biomarkers and drug targets in transcriptomics is a crucial tool for studying cell phenotypes and functions, providing new insights for identifying disease biomarkers, exploring pathophysiological mechanisms, and investigating related biological pathways.

[0003] Oral squamous cell carcinoma (OSCC) is the most common cancer of the oral cavity and maxillofacial region. It is a highly aggressive cancer prone to metastasis and local recurrence, and conventional treatments such as surgery, radiotherapy, and chemotherapy are often ineffective. Due to the unique physiological and anatomical location of the oral cavity, the disease often significantly impacts patients' chewing, swallowing, speech, and breathing functions, even threatening their lives. Most OSCC patients are diagnosed at an advanced stage, resulting in poor treatment outcomes. Overall, the treatment outcomes for OSCC patients remain unsatisfactory, and there is a lack of targeted research on OSCC biomarker identification models.

[0004] Conventional biomarker identification methods involve converting genomic or transcriptomic sequencing or expression information into a matrix, analyzing differential expression patterns, and then identifying one or more differentially expressed factors related to biological pathways as tumor biomarkers. Invasion is a major biological characteristic of malignant tumors, primarily due to the invasive and penetrating capabilities of tumor cells. Analyzing cell abundance in tumor samples is beneficial for exploring the relationship between biomarkers and the tumor immune microenvironment. Summary of the Invention

[0005] The purpose of this invention is to provide a method for establishing a biomarker identification model for oral squamous cell carcinoma (OSCC). This method fully utilizes the programming language capabilities of Perl and R to process and analyze transcriptome data from OSCC tumor databases, the immune infiltration relationship between OSCC sample cell abundance and differentially expressed factors, and evaluates the actual impact of the mined differentially expressed factors on the tumor development mechanism, thereby making the identification of OSCC biomarkers more accurate.

[0006] This invention discloses a method for establishing a biomarker recognition model for oral squamous cell carcinoma, comprising:

[0007] Transcriptome data from the OSCC tumor database were processed and analyzed to identify differentially expressed RNA factors in the OSCC transcriptome.

[0008] Analysis revealed the relationship between cell abundance and differentially expressed factors in OSCC samples, and their relationship to immune infiltration.

[0009] The study aimed to evaluate the actual impact of differentially expressed factors on the mechanisms of tumor development and progression, screen and retain RNA factors that are significantly related to survival, and identify biomarkers that can be used as OSCC.

[0010] The transcriptome data from the OSCC tumor database were processed and analyzed. Further analysis and mining of differentially expressed RNA factors in the OSCC transcriptome included:

[0011] Transcriptomic components of OSCC patient tumor samples were analyzed using RNA sequencing.

[0012] Transcriptome data from the OSCC tumor database were processed and analyzed to identify differentially expressed RNA factors in the OSCC transcriptome, including:

[0013] The transcriptome data from the OSCC tumor database was processed and analyzed using Perl and R languages ​​to obtain the expression matrix of OSCC transcriptomics factors and to analyze and mine differentially expressed factors in the OSCC transcriptome.

[0014] Further analysis revealed the immune infiltration relationship between cell abundance and differentially expressed factors in OSCC samples, including:

[0015] The target regulatory relationships among various RNAs were obtained using the Circ2Traits online database tool and the FunRich tool for predictive intersection analysis.

[0016] Further analysis revealed the immune infiltration relationship between cell abundance and differentially expressed factors in OSCC samples, including:

[0017] Biological pathway mapping results of differentially expressed factors in the OSCC transcriptome were obtained through gene ontology and pathway enrichment analysis, matching the involved cellular components, molecular functions and biological pathways.

[0018] Analysis revealed the following relationships between cell abundance and differentially expressed factors in OSCC samples, indicating immune infiltration:

[0019] Based on the linear support vector regression principle of CI BERSORT operation in R language, the expression matrix of human immune cell subtypes is deconvolved, and tumor-infiltrating immune cells are quantified from RNA sequencing data to obtain the immune infiltration relationship between cell abundance and differentially expressed factors in OSCC samples.

[0020] The actual impact of differentially expressed factors on tumorigenesis and development mechanisms was evaluated, and RNA factors significantly associated with survival were screened and retained. Biomarkers that could serve as OSCC biomarkers were identified, including:

[0021] Based on univariate Cox regression model analysis, RNA factors that are significantly associated with survival were screened and retained, and biomarkers that can be used as OSCC were identified.

[0022] This invention primarily addresses the issue of limited methods for identifying OSCC biomarkers. Based on the mining of biomarkers and drug targets in OSCC circular RNA (circular RNA), microRNA (miRNA), and messenger RNA (mRNA), this invention explores the relationship between transcriptome expression and tumor immune invasion in OSCC, and establishes an OSCC biomarker identification model based on the Cox regression model (Cox probability-hazards model 1).

[0023] Differentially expressed RNA factors exhibit significantly different expression patterns under varying environmental conditions. Therefore, differentially expressed RNA factors can serve as cellular markers for identification. This invention processes and analyzes transcriptomic data from an OSCC tumor database to identify differentially expressed RNA factors in the OSCC transcriptome. Since the main biological characteristic of malignant tumors is invasion, and the biological feature of malignant tumors is that tumor cells have a certain ability to invade, analyzing the cell abundance in tumor samples is beneficial for exploring the relationship between biomarkers and the tumor immune microenvironment. Therefore, this invention analyzes and obtains the immune infiltration relationship between cell abundance and differentially expressed factors in OSCC samples. By utilizing differentially expressed RNA factors as cellular markers for identification and analyzing the differentially expressed RNA factors related to immune infiltration, RNA markers that can distinguish OSCC can be obtained. Attached Figure Description

[0024] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1This is a flowchart of a method for establishing a biomarker recognition model for oral squamous cell carcinoma proposed in this invention;

[0027] Figure 2 This is a flowchart of a method for establishing a biomarker recognition model for oral squamous cell carcinoma proposed in this invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0029] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.

[0030] Furthermore, the use of terms such as "first" and "second" in this invention is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" and "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but only on the basis of being achievable by those skilled in the art. When the combination of technical solutions is contradictory or impossible to implement, such a combination of technical solutions should be considered non-existent and not within the scope of protection claimed by this invention.

[0031] This invention discloses a method for establishing a biomarker recognition model for oral squamous cell carcinoma, comprising:

[0032] Step 100: Process and analyze the transcriptome data from the OSCC tumor database to analyze and mine differentially expressed RNA factors in the OSCC transcriptome.

[0033] Step 200: Analyze the immune infiltration relationship between cell abundance and differentially expressed factors in OSCC samples;

[0034] Step 300: Evaluate the actual impact of the obtained differentially expressed factors on the mechanisms of tumorigenesis and development, screen and retain RNA factors significantly related to survival, and identify biomarkers that can serve as OSCC. Step 100: Process and analyze the transcriptome data from the OSCC tumor database, and analyze and mine OSCC transcriptome RNA.

[0035] Differential expression factors further include:

[0036] Transcriptomic components of OSCC patient tumor samples were analyzed using RNA sequencing.

[0037] Differentially expressed RNA factors exhibit significantly different expression patterns under varying environmental conditions. Therefore, differentially expressed RNA factors can serve as cellular markers for identification. This invention processes and analyzes transcriptomic data from an OSCC tumor database to identify differentially expressed RNA factors in the OSCC transcriptome. Since the main biological characteristic of malignant tumors is invasion, and the biological feature of malignant tumors is that tumor cells have a certain ability to invade, analyzing the cell abundance in tumor samples is beneficial for exploring the relationship between biomarkers and the tumor immune microenvironment. Therefore, this invention analyzes and obtains the immune infiltration relationship between cell abundance and differentially expressed factors in OSCC samples. By utilizing differentially expressed RNA factors as cellular markers for identification and analyzing the differentially expressed RNA factors related to immune infiltration, RNA markers that can distinguish OSCC can be obtained.

[0038] Step 100 involves processing and analyzing the transcriptome data from the OSCC tumor database, including analyzing and mining differentially expressed RNA factors in the OSCC transcriptome:

[0039] The transcriptome data from the OSCC tumor database was processed and analyzed using Perl and R languages ​​to obtain the expression matrix of OSCC transcriptomics factors and to analyze and mine differentially expressed factors in the OSCC transcriptome.

[0040] Using the `limma` package in R, data from TCGA was analyzed, yielding 482 differentially expressed circRNAs and 92 miRNAs. The GSE118750 dataset was obtained from the GEO (Gene Expression Omnibus) database, and using hash functions in Perl and the `limma` package in R, 961 differentially expressed mRNAs were analyzed. All differential expression thresholds required a Pvalue value less than 0.05 and a log2 shift change value greater than 1.2.

[0041] In biological research, more features are often transformed into mathematical models. These gene expressions often exhibit correlations. Organizing the transcriptome data from the OSCC tumor database into a sample matrix facilitates data comparison, processing, and screening. Perl, with its dynamic language capabilities, offers powerful flexibility, while R, easy to program from a computer science perspective, allows for high-quality analysis with high freedom. Therefore, using Perl and R to process data enables faster and more flexible generation of matrices and preliminary screening of differentially expressed factors that meet the criteria, reducing the workload of subsequent analysis. Step 200, analyzing the immune infiltration relationship between OSCC sample cell abundance and differentially expressed factors, further includes:

[0042] The target regulatory relationships among various RNAs were obtained using the Circ2Traits online database tool and the FunRich tool for predictive intersection analysis.

[0043] The target miRNAs of differentially expressed circRNAs were predicted using the Circ2Traits database, and the intersections of these with the differentially expressed miRNAs were obtained. The target mRNAs of the intersecting miRNAs were predicted using the FunRich tool, and then the intersections of these with the differentially expressed mRNAs in OSCC were obtained, yielding target regulatory relationships among 89 circRNAs, 43 miRNAs, and 223 mRNAs.

[0044] The Circular RNA 2Traits database is a database of diseases or specific traits associated with circular RNAs. It establishes a link between circular RNAs and diseases by analyzing the interactions between circular RNAs and disease-related miRNAs. Furthermore, to further analyze the regulatory pathways of circular RNAs in diseases, a ceRNA regulatory network of circular RNAs has been constructed. FunRich is a standalone software tool primarily used for gene and protein functional enrichment and interaction network analysis. By using the FunRich toolkit to predict RNAs and find their intersections with differentially expressed RNAs, the targeted regulatory relationships between various RNAs can be determined, effectively narrowing down the range of RNA biomarkers to be searched and reducing workload.

[0045] Step 200 further analyzed the immune infiltration relationship between cell abundance and differentially expressed factors in OSCC samples, including:

[0046] Biological pathway mapping results of differentially expressed factors in the OSCC transcriptome were obtained through gene ontology and pathway enrichment analysis, matching the involved cellular components, molecular functions and biological pathways.

[0047] Using the STRING database and Cytoscape software (MCODE plugin set to default parameters), the interactions between mRNA expression products were plotted as a protein-protein interaction network, and key subnetworks were identified. The ClusterProfiler, ggplot2, colorspace, strini, pathview, and org.Hs.eg.db packages in R were used to obtain the GO and KEGG mapping results of differentially expressed factors in these submodules.

[0048] Since the expression products of tumor cells differ significantly from those of normal cells, software prediction of the expression products of differentially expressed factors can more accurately determine whether the expression factor is the desired OSCC biomarker, thereby improving the recognition rate.

[0049] Step 200 analysis revealed the immune infiltration relationship between cell abundance and differentially expressed factors in OSCC samples, including:

[0050] Based on the linear support vector regression principle of CI BERSORT operation in R language, the expression matrix of human immune cell subtypes is deconvolved, and tumor-infiltrating immune cells are quantified from RNA sequencing data to obtain the immune infiltration relationship between cell abundance and differentially expressed factors in OSCC samples.

[0051] Based on the principle of linear support vector regression (LiNOR) using the CI BERSORT operation in R language, the expression matrix of human immune cell subtypes is deconvolved. Tumor-infiltrating immune cells are quantified from RNA sequencing data to obtain the immune infiltration relationship between cell abundance and differentially expressed factors in OSCC samples, and to verify the actual impact of the mined differentially expressed factors on the tumor development mechanism.

[0052] Tumor-infiltrating immune cells can influence tumor progression and the success of anticancer treatments by exerting both pro- and anti-tumor effects. Quantifying tumor-infiltrating immune cells helps reveal the multifaceted role of the immune system in human cancers and its involvement in tumor escape mechanisms and treatment responses. Using CI BERSORT software, we can quantify tumor-infiltrating immune cells from tumor RNA sequencing data.

[0053] CI BERSORT is a tool for deconvolution of expression matrices of human immune cell subtypes based on the principle of linear support vector regression. It is superior to other methods for deconvolution analysis of expression matrices in microarrays and sequencing matrices, particularly for unknown mixtures and expression matrices containing similar cell types.

[0054] Step 300 evaluates the actual impact of differentially expressed factors on the mechanisms of tumor development and progression, screens and retains RNA factors significantly associated with survival, and identifies biomarkers that can serve as OSCC, including:

[0055] Based on univariate Cox regression model analysis, RNA factors that are significantly associated with survival were screened and retained, and biomarkers that can be used as OSCC were identified.

[0056] Univariate Cox regression analysis was used, with log-1 ambda from the Lasso algorithm as the penalty coefficient for regression analysis. Based on the obtained partial likelihood bias values, 43 mRNAs and 6 miRNAs closely related to OSCC survival were selected as biomarkers for oral squamous cell carcinoma.

[0057] The univariate Cox regression model can perform regression analysis on each feature to determine whether the feature is significantly related to survival. The RNA factors that are significantly related to survival are the OSCC biomarkers described in this invention.

[0058] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for establishing a biomarker recognition model for oral squamous cell carcinoma, characterized in that, include: Transcriptome data from the OSCC tumor database were processed and analyzed to identify differentially expressed RNA factors in the OSCC transcriptome. Analysis revealed the relationship between cell abundance and differentially expressed factors in OSCC samples, and their relationship to immune infiltration. The actual impact of the differentially expressed factors on the tumorigenesis and development mechanism was evaluated, and RNA factors that were significantly related to survival were screened and retained to identify biomarkers that could be used as OSCC. The analysis revealed the immune infiltration relationship between OSCC sample cell abundance and differentially expressed factors, which further included: The target regulatory relationships among various RNAs were obtained using the Circ2Traits online database tool and the FunRich tool for predictive intersection analysis. The analysis revealed the immune infiltration relationship between OSCC sample cell abundance and differentially expressed factors, which further included: Biological pathway mapping results of differentially expressed factors in the OSCC transcriptome were obtained through gene ontology and pathway enrichment analysis, matching the involved cellular components, molecular functions and biological pathways. The analysis revealed the following relationships between cell abundance and differentially expressed factors in OSCC samples, including: Based on the linear support vector regression principle of CIBERSORT operation in R language, the expression matrix of human immune cell subtypes is deconvolved, and tumor-infiltrating immune cells are quantified from RNA sequencing data to obtain the immune infiltration relationship between cell abundance and differentially expressed factors in OSCC samples.

2. The method for establishing a biomarker recognition model for oral squamous cell carcinoma according to claim 1, characterized in that, The processing and analysis of transcriptome data from the OSCC tumor database, and the analysis and mining of differentially expressed RNA factors in the OSCC transcriptome, further include: Transcriptomic components of OSCC patient tumor samples were analyzed using RNA sequencing.

3. The method for establishing a biomarker recognition model for oral squamous cell carcinoma according to claim 1, characterized in that, The processing and analysis of transcriptome data from the OSCC tumor database, including the analysis and mining of differentially expressed RNA factors in the OSCC transcriptome, includes: The transcriptome data from the OSCC tumor database was processed and analyzed using Perl and R languages ​​to obtain the expression matrix of OSCC transcriptomics factors, and to analyze and mine differentially expressed factors in the OSCC transcriptome.

4. The method for establishing a biomarker recognition model for oral squamous cell carcinoma according to claim 1, characterized in that, The evaluation revealed the actual impact of differentially expressed factors on the mechanisms of tumor development and progression. RNA factors significantly associated with survival were screened and retained, and biomarkers that could serve as OSCC were identified, including: Based on univariate Cox regression model analysis, RNA factors that are significantly associated with survival were screened and retained, and biomarkers that can be used as OSCC were identified.

Citation Information

Patent Citations

  • Application of combined STAT signal channel related gene in colorectal cancer prognosis model

    CN114334147A

  • Hybrid model for the classification of carcinoma subtypes

    US20130172203A1