A scoring system established based on single-cell analysis for evaluating the prognosis of head and neck squamous cell carcinoma

By analyzing single-cell data and using deconvolution algorithm to calculate the cell composition structure of HNSCC patients, a scoring system based on single-cell analysis was constructed, which solved the problem that the existing system failed to fully explore the relationship between metabolic reprogramming and patient heterogeneity, and achieved an effective evaluation of the prognosis of HNSCC.

CN117116346BActive Publication Date: 2025-05-30SUN YAT SEN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310988594.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-05-22
Filing Date
2023-08-07
Publication Date
2025-05-30
Estimated Expiration
2043-08-07

AI Technical Summary

Technical Problem

The existing scoring system constructed by single-cell sequencing analysis related to head and neck squamous cell carcinoma (HNSCC) has been limited in data, and the existing metabolomics scoring system has failed to describe the impact of metabolic reprogramming on patient prognosis from the cellular level.

Method used

By analyzing single-cell data, the composition of high-metabolic active cells and low-metabolic active cells was obtained, and the cell composition structure of 486 HNSCC patients was calculated using the deconvolution algorithm, and the proportion of high-metabolic active cell data to all cells was defined as a scoring system, and a scoring system based on single-cell analysis was constructed.

Benefits of technology

An effective scoring system that can evaluate the prognosis of HNSCC is established, providing a cellular level to describe the impact of metabolic reprogramming on patient prognosis and help improve the prognosis of HNSCC patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117116346B_ABST
    Figure CN117116346B_ABST
Patent Text Reader

Abstract

The present invention provides a scoring system for evaluating the prognosis of head and neck squamous cell carcinoma established based on single-cell analysis, including: obtaining sample data, including single-cell data of head and neck squamous cell carcinoma and large-cohort RNA data, and the large-cohort RNA data including RNA sequence data and clinical information of head and neck squamous cell carcinoma patients; performing quality control on the single-cell data and the large-cohort RNA data; clustering and annotating the cells in the single-cell data; defining malignant cells and metabolic activity genes; clustering the cells according to metabolic activity; establishing the relationship between the proportion of cell populations and the prognosis of head and neck squamous cell carcinoma according to the correlation between the metabolic activity of the clustered cells and head and neck squamous cell carcinoma; screening metabolic biomarkers according to the scoring system. Combining single-cell sequencing and RNA-Seq technologies to construct a scoring system from the cellular and clinical prognosis levels, and providing a reference for analyzing metabolic characteristics and the impact of reprogramming on tumor malignancy and drug resistance by using single-cell sequencing and tumor metabolic reprogramming technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and particularly to a scoring system for evaluating the prognosis of head and neck squamous cell carcinoma established based on single-cell analysis. Background Art

[0002] Head and neck squamous cell carcinoma (HNSCC) is the most common malignant tumor in the head and neck and the sixth most common malignant tumor in the body, with approximately 900,000 new cases and 500,000 death cases each year. Despite various intervention methods such as surgery, radiotherapy, chemotherapy, and immunotherapy, the prognosis remains poor. Further exploring the development mechanism of HNSCC and identifying new molecular biomarkers as more effective treatment targets are of great significance for improving the prognosis of HNSCC patients.

[0003] The reprogramming of energy metabolism has been recognized as a common phenomenon in tumors and a necessary means for tumor cells to maintain rapid proliferation. The metabolic reprogramming of malignant tumor cells is a complex and multi-step process and is related to many molecular alterations. In recent years, the research on metabolic reprogramming during the occurrence and development of head and neck squamous cell carcinoma has been a research hotspot in the field, and alterations in metabolic pathways such as glycolysis and amino acid metabolism have been proven to play important roles in HNSCC.

[0004] Tumor molecular typing refers to the systematic molecular characterization of tumors through omics technologies such as genomics, transcriptomics, proteomics, and metabolomics, exploring the occurrence and development mechanisms and intrinsic characteristics of tumors from multiple levels, classifying tumors according to the expression profiles of tumor molecular characteristics, and further subclassifying tumors on this basis, which is beneficial to improving tumor diagnosis, prognosis stratification, treatment guidance, recurrence monitoring, and drug research and development. Some researchers have conducted various omics molecular typings on pancreatic cancer, cervical cancer, pediatric brain cancer, breast cancer, etc.

[0005] Single-cell sequencing analysis refers to omics research conducted at the single-cell level. In recent years, as a cutting-edge technology, single-cell sequencing analysis has been widely applied to the research of tumor cell heterogeneity, tumor cell metastasis mechanisms, tumor stem cells, etc., playing an important role in understanding the occurrence, development, metastasis, and drug resistance mechanisms of tumors, providing many new molecular targets and strategies for tumor diagnosis and treatment, and research applying this technology has been frequently published in high-level journals.

[0006] In the existing scoring systems constructed by single-cell sequencing analysis related to HNSCC, due to the limited number of patients in the scRNA-seq dataset, few studies have comprehensively explored the relationship between metabolic reprogramming reflected in cell composition and patient heterogeneity. It often depends on a specific cell ratio without considering the patient's prognosis, which limits its clinical application. The existing scoring systems constructed based on HNSCC metabolomics often rely on a specific gene expression pattern or the modeling of a group of gene expressions, without describing the impact of metabolic reprogramming on patient prognosis at the cellular level, nor the scoring system related to metabolic reprogramming. Summary of the Invention

[0007] In the present invention, by analyzing single-cell data, a composition including highly metabolically active cells, low metabolically active cells, and integrated non-malignant cells is obtained. Then, using the deconvolution algorithm, the cell composition structure of 486 HNSCC patients is estimated through gene information on transcriptome data. Among them, highly metabolically active cells are malignant cells with high-frequency cell metabolic reprogramming, mainly distributed in metastatic samples in scRNA data, while low metabolically active cells are malignant cells with relatively low metabolic activity, mainly distributed in primary samples in scRNA data. Defining the proportion of highly metabolically active cell data in all cells as the scoring system, a scoring system based on single-cell analysis for evaluating the prognosis of HNSCC is thus constructed.

[0008] Among them, the scoring method based on single-cell analysis for evaluating the prognosis of head and neck squamous cell carcinoma includes:

[0009] Obtain sample data, where the sample data includes single-cell data of head and neck squamous cell carcinoma and large cohort RNA data, and the large cohort RNA data includes RNA sequence data and clinical information of head and neck squamous cell carcinoma patients;

[0010] Perform quality control on the single-cell data and the large cohort RNA data;

[0011] Cluster and annotate the cells in the single-cell data;

[0012] Define malignant cells and metabolically active genes;

[0013] Cluster the cells according to their metabolic activity;

[0014] Establish the relationship between the proportion of different cell clusters and the prognosis of head and neck squamous cell carcinoma based on the correlation between the metabolic activity of the clustered cells and head and neck squamous cell carcinoma;

[0015] Screen metabolic biomarkers according to the scoring system.

[0016] Furthermore, the quality control of the single-cell data specifically includes:

[0017] Remove cells with gene expression ≤ 200 and cells with gene expression ≥ 2500;

[0018] Remove doublet cells;

[0019] Remove genes expressed in fewer than 3 cells;

[0020] Remove mitochondrial genes;

[0021] Merge the datasets.

[0022] Quality control of large cohort RNA data may include: (1) retaining samples not soaked in formalin; (2) retaining samples with complete clinical data; (3) retaining samples with clear clinical outcomes and survival times.

[0023] Furthermore, clustering and annotating the cells in the single-cell data specifically includes:

[0024] Perform dimensionality reduction clustering on all cells, and filter out cells with ribosomal gene expression exceeding 35% in each subpopulation;

[0025] Find highly expressed genes in each subpopulation, and annotate each cell subpopulation based on the identified highly expressed genes, specifically expressed genes, and reported cell marker genes;

[0026] The annotated cells are divided into subpopulations such as epithelial cells, T cells, B cells, plasma cells, monocytes, mast cells, fibroblasts, and malignant cells.

[0027] Furthermore, the definition of malignant cells includes:

[0028] Take the malignant cells in the sample data for dimensionality reduction clustering;

[0029] Classify according to the origin of malignant cells in each group:

[0030] The cell group with the largest number of metastatic cancer sample sources is defined as the metastatic group, the cell group with the largest number of primary cancer sample sources is defined as the primary group, and the remaining malignant cell groups with comparable metastatic cancer and primary cancer sample sources are defined as the mixed group.

[0031] Furthermore, the grouping of cells according to metabolic activity includes:

[0032] Calculate the cut-off value of metabolic activity;

[0033] Define malignant cells with a metabolic activity higher than the cut-off value as high-metabolic-activity cells, and malignant cells with a metabolic activity lower than the cut-off value as low-metabolic-activity cells;

[0034] Define cells other than malignant cells in the dimensionality reduction clustering as non-malignant cells.

[0035] Further, the defined metabolic activity genes include:

[0036] Screen for differential genes between metastatic population cells and primary population cells;

[0037] Perform enrichment analysis on the differential genes;

[0038] Define the differential genes enriched in the metabolic pathway as metabolism-related genes;

[0039] Perform GSEA analysis on the metabolism-related genes, and define the metabolism-related genes enriched in the metastatic cell population as metabolic activity genes.

[0040] Further, the scoring system specifically includes:

[0041] Using the deconvolution algorithm, according to the gene expression profiles of high-metabolic-activity cells and low-metabolic-activity cells, calculate the proportion of high-metabolic-activity cells corresponding to all RNA sequence data of patients, and define this proportion as the scoring system.

[0042] On the other hand, the present invention provides a scoring system for evaluating the prognosis of head and neck squamous cell carcinoma established based on single-cell analysis, including:

[0043] A sample acquisition module for acquiring sample data, where the sample data includes single-cell data of head and neck squamous cell carcinoma and large cohort RNA data, and the large cohort RNA data includes RNA sequence data and clinical information of head and neck squamous cell carcinoma patients;

[0044] A quality control module for performing quality control on the single-cell data and the large cohort RNA data;

[0045] A clustering and annotation module for clustering and annotating the cells in the single-cell data;

[0046] A definition module for defining malignant cells and metabolic activity genes;

[0047] A cell population division module for dividing cells according to metabolic activity;

[0048] A scoring module for establishing the relationship between the proportion of different cell populations and the prognosis of head and neck squamous cell carcinoma according to the correlation between the metabolic activity of the cells after population division and head and neck squamous cell carcinoma as the scoring system;

[0049] A screening module for screening metabolic biomarkers according to the scoring system;

[0050] Performing quality control on the single-cell data specifically includes:

[0051] Removing cells with ≤200 expressed genes and cells with ≥2500 gene expression counts;

[0052] Removing doublet cells;

[0053] Remove genes expressed in fewer than 3 cells;

[0054] Remove mitochondrial genes;

[0055] Merge the datasets;

[0056] Clustering and annotating the cells in the single-cell data specifically includes:

[0057] Perform dimensionality reduction clustering on all cells, and filter out cells with ribosomal gene expression exceeding 35% in each subpopulation;

[0058] Find highly expressed genes in each subpopulation, and annotate each cell subpopulation according to the found highly expressed genes, specifically expressed genes, and reported cell marker genes;

[0059] The annotated cells are divided into subpopulations such as epithelial cells, T cells, B cells, plasma cells, monocytes, mast cells, fibroblasts, and malignant cells.

[0060] Furthermore, the definition module includes a malignant cell definition module and a metabolic activity gene definition module;

[0061] The malignant cell definition module is used for:

[0062] Perform dimensionality reduction clustering on the malignant cells in the sample data, and classify them according to the origin of the malignant cells in each group: the cell group with the most metastatic cancer sample sources is defined as the metastatic group, the cell group with the most primary cancer sample sources is defined as the primary group, and the remaining malignant cell groups with comparable metastatic and primary cancer sample sources are defined as the mixed group;

[0063] The metabolic activity gene definition module is used for:

[0064] Screen for differential genes between metastatic group cells and primary group cells;

[0065] Perform enrichment analysis on the differential genes;

[0066] Define the differential genes enriched in the metabolic pathway as metabolism-related genes;

[0067] Perform GSEA analysis on the metabolism-related genes, and define the metabolism-related genes enriched in the metastatic cell group as metabolic activity genes.

[0068] Furthermore, the scoring module specifically includes:

[0069] Using the deconvolution algorithm, calculate the proportion of highly metabolically active cells in all RNA sequence data corresponding to patients according to the gene expression profiles of highly metabolically active cells and low metabolically active cells, and define this proportion as the scoring system.

[0070] The beneficial effects of the present invention are as follows:

[0071] 1. By combining single-cell sequencing technology and large cohort RNA-Seq technology, a scoring system for HNSCC was constructed based on metabolic activity at both the cellular and clinical prognosis levels.

[0072] 2. It provides a reference for further analyzing the metabolic characteristics of HNSCC and how metabolic reprogramming affects the tumor malignancy and drug resistance of HNSCC patients. Description of the Drawings

[0073] Figure 1 It is a flowchart of the method according to the first embodiment of the present invention.

[0074] Figure 2 It is a flowchart for constructing the METArisk scoring system according to the first embodiment of the present invention.

[0075] Figure 3 It is an effect diagram of screening metabolic biomarkers according to the first embodiment of the present invention. Detailed Embodiments

[0076] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further describes the specific embodiments of the present application in detail with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are only for explaining the present application and are not intended to limit the present application. Additionally, it should be noted that for the sake of description, only parts related to the present application are shown in the drawings rather than all the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. When the operations are completed, the process can be terminated, but there may also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0077] The following will clearly describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.

[0078] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and do not limit the number of objects. For example, the first object can be one or more. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0079] The following will combine the accompanying drawings and, through specific embodiments and their application scenarios, provide a detailed description of the embodiments of this application.

[0080] Embodiment 1

[0081] As Figure 1 shown, this embodiment provides a scoring method based on single-cell analysis for evaluating the prognosis of head and neck squamous cell carcinoma, specifically including:

[0082] 1. Obtain public data:

[0083] (1) Obtain single-cell data: Download HNSCC single-cell data with sequence numbers SRP332116 and 228811 from the SRA database of NCBI, including 25 cases of primary cancer tissues and 8 cases of metastatic cancer, and use cellranger (version 4.0.0) to obtain the fastq files in the original data, and obtain the gene barcode matrix through Seurat (version 4.0.4) and pipeline (version 4.0.5) in R language.

[0084] (2) Obtain cohort RNA data: Download cohort RNA-Seq data and clinical information of 486 HNSCC patients from the TCGA and GEO databases

[0085] 2. Data processing: Perform quality control on the single-cell data: (1) Remove cells with ≤200 expressed genes and cells with ≥2500 gene expressions; (2) Remove doublets; (3) Remove genes expressed in fewer than 3 cells; (4) Remove mitochondrial genes; at the same time, use the 'ScaleData' function to convert the expression matrix into a format easy to analyze, and use the 'FindIntegrationAchors' and 'IntegrateData' functions to merge different data sets. Perform quality control on the cohort RNA data: (1) Retain samples not soaked in formalin; (2) Retain samples with complete clinical data; (3) Retain samples with clear clinical outcomes and survival times

[0086] 3. Cell clustering and annotation: All cells were subjected to dimensionality reduction clustering using the FindClusters function in the Seurat package in R language. Cells with ribosomal gene expression exceeding 35% in each subpopulation were filtered out, and the clustered cell populations were visualized using UMAP plots. As Figure 2 shown, FindAllMarker was used to identify highly expressed genes in each subpopulation, and each cell subpopulation was annotated based on the identified highly expressed genes, specifically expressed genes, and previously reported cell marker genes ( Figure 2 A). The annotated cells were divided into subpopulations such as epithelial cells, T cells, B cells, plasma cells, monocytes, mast cells, fibroblasts, and malignant cells.

[0087] 4. Definition of malignant cells: Malignant cells in the samples were taken, subjected to dimensionality reduction clustering, and classified according to the different origins of malignant cells in each group: The cell group with a higher proportion of cells from metastatic cancer samples was defined as the metastatic group, the cell group with a higher proportion of cells from primary cancer samples was defined as the primary group, and the cell group with an equal proportion of cells from metastatic and primary cancer samples was defined as the mixed group ( Figure 2 B).

[0088] 5. Definition of metabolic activity genes: Differentially expressed genes between metastatic group cells and primary group cells were screened using the single-cell-specific method Psedobulk. KEGG-GO enrichment analysis was performed using the differentially expressed genes, and it was found that the differentially expressed genes were enriched in metabolic pathways ( Figure 2 C). These differentially expressed genes were defined as metabolism-related genes (MRG), and GSEA analysis was performed using the clusterProfiler R package. It was found that most metabolism-related genes were enriched in the metastatic cell group ( Figure 2 D). Therefore, the genes enriched in the metastatic cell group were taken and defined as META_ACTIVE genes.

[0089] 6. Cell grouping according to metabolic activity: The AUCell algorithm was used to analyze the activation of META_ACTIVE genes in each cell to represent the metabolic activity of the cells. The cut-off value of metabolic activity was calculated using the "AUCell_exploreThresholds" function. Malignant cells with a metabolic activity higher than the cut-off value were defined as METAactive cells with high metabolic activity, and cells with a metabolic activity lower than the cut-off value were defined as METAsilence cells with low metabolic activity ( Figure 2 E). Cells other than malignant cells in the dimensionality reduction clustering were taken and integrated using the method reported in the 2019 literature (1) and defined as non-malignant cells

[0090] 7. Determining the biological significance of cell metabolic subtypes: By plotting the PCA graphs of METAactive cells and METAsilence cells and comparing them with the PCA graphs of the sample sources, we found that METAactive cells mostly originated from metastatic cancer samples, while METAsilence cells mostly originated from primary cancer samples. Therefore, METAactive cells with high metabolic activity are closely related to the malignant behavior of tumors ( Figure 2 F).

[0091] 8. Establishing the METArisk score based on the relationship between specific cell populations and prognosis: Using the deconvolution algorithm Cibersortx28, based on the single-cell gene expression profiles of METAactive and METAsilence cell populations, we inferred the proportion of METAactive cells in all cells for each patient in the large cohort RNA-Seq data and defined this proportion as the METArisk score ( Figure 2 G).

[0092] 9. Validating the prognostic evaluation ability of the METArisk score for HNSCC patients: According to the METArisk score, HNSCC patients were divided into two subtypes: high / low METArisk. Kaplan-Meier survival analysis was performed on the patients of the two subtypes. The results showed that patients with low METArisk had a better prognosis, while patients with high METArisk had a worse prognosis ( Figure 2 H), indicating that our metabolic score METArisk can evaluate the prognostic risk of patients.

[0093] 10. As Figure 3 shown, screening metabolic biomarkers using METArisk: Through differential analysis, we selected the differentially expressed genes between the high and low METArisk groups, and intersected them with the metabolism-related genes and the differentially expressed genes between the tumor and normal groups to obtain 56 candidate genes ( Figure 3 A). To screen out the metabolic biomarkers that can independently predict the prognosis of head and neck squamous cell carcinoma, we explored the combination methods of different combinations of survival analysis models. By calculating the C-index ( Figure 3 B) of each combination, we finally selected the combination of SVM-RFE ( Figure 3 C) and multivariate Cox PH ( Figure 3 D) analysis, and screened out four genes, EPHX3, FDCSP, FAM3B and PYGL, from the 56 candidate genes as the candidate metabolic biomarkers for research.

[0094] Example 2

[0095] This embodiment provides a scoring system for evaluating the prognosis of head and neck squamous cell carcinoma established based on single-cell analysis, specifically including:

[0096] A sample acquisition module for acquiring sample data, where the sample data includes single-cell data of head and neck squamous cell carcinoma and large cohort RNA data, and the large cohort RNA data includes RNA sequence data and clinical information of head and neck squamous cell carcinoma patients;

[0097] A quality control module for performing quality control on the single-cell data and the large cohort RNA data;

[0098] A clustering and annotation module for clustering and annotating the cells in the single-cell data;

[0099] A definition module for defining malignant cells and metabolic activity genes;

[0100] A cell population module for dividing cells into populations according to metabolic activity;

[0101] A scoring module for establishing the relationship between the proportion of different cell populations and the prognosis of head and neck squamous cell carcinoma as a scoring system based on the correlation between the metabolic activity of the cells after population division and head and neck squamous cell carcinoma;

[0102] A screening module for screening metabolic biomarkers according to the scoring system;

[0103] Performing quality control on the single-cell data specifically includes:

[0104] Removing cells with ≤200 expressed genes and cells with ≥2500 gene expression numbers;

[0105] Removing doublets;

[0106] Removing genes expressed in fewer than 3 cells;

[0107] Removing mitochondrial genes;

[0108] Merging the data sets;

[0109] Clustering and annotating the cells in the single-cell data specifically includes:

[0110] Performing dimensionality reduction clustering on all cells and filtering out cells with ribosomal gene expression exceeding 35% in each subpopulation;

[0111] Searching for highly expressed genes in each subpopulation and annotating each cell subpopulation according to the found highly expressed genes, specifically expressed genes, and reported cell marker genes;

[0112] The annotated cells are divided into subpopulations such as epithelial cells, T cells, B cells, plasma cells, monocytes, mast cells, fibroblasts, and malignant cells.

[0113] Obtaining single-cell data: Download single-cell data from the SRA database of NCBI, and use cellranger (version 4.0.0) to obtain the fastq files in the original data. Obtain the gene barcode matrix through Seurat (version 4.0.4) and pipeline (version 4.0.5) in R language.

[0114] Obtaining cohort RNA data: Download the cohort RNA-Seq data and clinical information of HNSCC patients from the TCGA and GEO databases.

[0115] Quality control of single-cell data: (1) Remove cells with ≤200 expressed genes and cells with ≥2500 gene expressions; (2) Remove doublet cells; (3) Remove genes expressed in fewer than 3 cells; (4) Remove mitochondrial genes; at the same time, use the 'ScaleData' function to convert the expression matrix into an easy-to-analyze format, and use the 'FindIntegrationAchors' and 'IntegrateData' functions to merge different datasets. Quality control of cohort RNA data: (1) Retain samples not soaked in formalin; (2) Retain samples with complete clinical data; (3) Retain samples with clear clinical outcomes and survival times.

[0116] Cell clustering and annotation: Perform dimensionality reduction clustering on all cells through the FindClusters function in the Seurat package in R language, filter out cells with ribosomal gene expression exceeding 35% in each subpopulation, and display the clustered subpopulations with UMAP plots. Use FindAllMarker to find highly expressed genes in each subpopulation, and annotate each cell subpopulation according to the found highly expressed genes, specifically expressed genes, and reported cell marker genes. The annotated cells are divided into subpopulations such as epithelial cells, T cells, B cells, plasma cells, monocytes, mast cells, fibroblasts, and malignant cells.

[0117] Preferably, the definition module includes a malignant cell definition module and a metabolic activity gene definition module;

[0118] The malignant cell definition module is used for:

[0119] Perform dimensionality reduction clustering on the malignant cells in the sample data, and classify according to the origin of the malignant cells in each group:

[0120] The cell group with the most metastatic cancer sample sources is defined as the metastatic group, the cell group with the most primary cancer sample sources is defined as the primary group, and the remaining malignant cell groups with comparable metastatic cancer and primary cancer sample sources are defined as the mixed group. The relevant code is as follows:

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130] The metabolic activity gene definition module is used for:

[0131] Screening for differential genes between metastatic population cells and primary population cells;

[0132] Performing enrichment analysis on the differential genes;

[0133] Defining the differential genes enriched in the metabolic pathway as metabolism-related genes;

[0134] Performing GSEA analysis on the metabolism-related genes, defining the metabolism-related genes enriched in the metastatic cell population as metabolic activity genes. Screen for differential genes between metastatic population cells and primary population cells by the unique method of single cells, Psedobulk. Perform KEGG-GO enrichment analysis with the differential genes and find that the differential genes are enriched in the metabolic pathway. Take these differential genes and define them as metabolism-related genes (MRG), and use the clusterProfiler R package to perform GSEA analysis. It is found that most of the metabolism-related genes are enriched in the metastatic cell population. Then take this part of the genes enriched in the metastatic cell population and define them as META_ACTIVE genes. The relevant code is as follows:

[0135]

[0136]

[0137]

[0138]

[0139]

[0140] Cell population based on metabolic activity: The AUCell algorithm was used to analyze the activation of META_ACTIVE genes in each cell to represent the metabolic activity of the cell. The cut-off value of metabolic activity was calculated by the "AUCell_exploreThresholds" function. Malignant cells with metabolic activity higher than the cut-off value were defined as cells with high metabolic activity, which were defined as METAactive cells, and cells with metabolic activity lower than the cut-off value were defined as cells with low metabolic activity, which were defined as METAsilence cells. Cells other than malignant cells in the dimensionality reduction clustering were taken and integrated by a novel algorithm and defined as non-malignant cells. The relevant code is as follows:

[0141]

[0142]

[0143]

[0144] Preferably, the scoring module specifically includes:

[0145] Using the deconvolution algorithm, according to the gene expression profiles of cells with high and low metabolic activities, the proportion of cells with high metabolic activity corresponding to all RNA sequence data of the patient was calculated, and this proportion was defined as the scoring system. Establish the METArisk score based on the relationship between specific cell populations and prognosis: Using the deconvolution algorithm, according to the single-cell gene expression profiles of METAactive and METAsilence cell populations, the proportion of METAactive cells in all cells of each patient in the large cohort RNA-Seq data was inferred, and this proportion was defined as the METArisk score. The relevant code is as follows:

[0146]

[0147]

[0148] Verify the prognostic evaluation ability of the METArisk score for HNSCC patients: According to the METArisk score, HNSCC patients were divided into two subtypes of high / low METArisk, and Kaplan-meier survival analysis was performed on the patients of the two subtypes. The results showed that patients with low METArisk had a better prognosis, and patients with high METArisk had a worse prognosis, indicating that our metabolic score METArisk can evaluate the prognostic risk of patients. The relevant code is as follows:

[0149]

[0150]

[0151] Using METArisk to screen metabolic biomarkers: Through differential analysis, we selected the differentially expressed genes between the high and low METArisk groups, and intersected them with metabolism-related genes and the differentially expressed genes between tumor and normal groups to obtain 56 candidate genes. To screen out metabolic biomarkers that can independently predict the prognosis of head and neck squamous cell carcinoma, we explored different combinations of survival analysis models. By calculating the C-index of each combination, we finally selected the combination of SVM-RFE and multivariate Cox PH analysis, and screened out four genes, EPHX3, FDCSP, FAM3B, and PYGL, from the 56 candidate genes. The relevant code is as follows:

[0152] BiocManager::install("mixOmics")

[0153] BiocManager::install("CoxBoost")

[0154] library(devtools)

[0155] install_github("binderh / CoxBoost")

[0156] library(survival)

[0157] library(randomForestSRC)

[0158] library(glmnet)

[0159] library(plsRcox)

[0160] library(superpc)

[0161] library(gbm)

[0162] library(CoxBoost)

[0163] library(survivalsvm)

[0164] library(dplyr)

[0165] library(tibble)

[0166] library(BART)

[0167] library(limma)

[0168] library(tidyverse)

[0169] library(dplyr)

[0170] library(miscTools)

[0171] library(compareC)

[0172] library(ggplot2)

[0173] library(ggsci)

[0174] library(tidyr)

[0175] library(ggbreak)

[0176] Sys.setenv(LANGUAGE = "en") # Display English error messages

[0177] options(stringsAsFactors = FALSE) # Prevent chr from being converted to factor

[0178] # Load the dataset

[0179] #tcga<-read.table("TCGA.txt",header = T,sep = "\t",quote = "",check.names = F)

[0180] #GSE57303<-read.table("GSE57303.txt",header = T,sep = "\t",quote = "",check.names = F)

[0181] #GSE62254<-read.table("GSE62254.txt",header = T,sep = "\t",quote = "",check.names = F)

[0182] # Generate a list containing three datasets

[0183] mm<-list(TCGA = tcga,

[0184] GSE65858 = GSE65858,GSE41613 = GSE41613)

[0185] # Data standardization

[0186] mm<-lapply(mm,function(x){ It should be noted that there are some incomplete variable references in the original text (such as `tcga`, `GSE65858`, `GSE41613` which are not defined properly). This translation is based on the text as it is.

[0187] x[,-c(1:3)] <- scale(x[,-c(1:3)])

[0188] return(x)})

[0189] result <- data.frame()

[0190] # TCGA as the training set

[0191] est_data <- mm$TCGA

[0192] # GEO as the validation set

[0193] val_data_list <- mm

[0194] pre_var <- colnames(est_data)[-c(1:3)]

[0195] est_dd <- est_data[,c('OS.time','OS',pre_var)]

[0196] val_dd_list <- lapply(val_data_list, function(x){x[,c('OS.time','OS',pre_var)]})

[0197] # Set the seed number and the number of nodes, where the number of nodes can be adjusted

[0198]

[0199]

[0200]

[0201]

[0202]

[0203]

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212]

[0213]

[0214]

[0215]

[0216]

[0217]

[0218]

[0219]

[0220]

[0221]

[0222]

[0223]

[0224]

[0225]

[0226]

[0227]

[0228]

[0229]

[0230]

[0231]

[0232]

[0233]

[0234]

[0235]

[0236] # Assign the obtained result to the result2 variable for operations

[0237] result2 <- result

[0238] Convert the long data of the result to wide data

[0239] dd2 <- pivot_wider(result2, names_from = 'ID', values_from = 'Cindex') %>% as.data.frame()

[0240] # Define the C-index as numeric

[0241] dd2[,-1] <- apply(dd2[,-1], 2, as.numeric)

[0242] # Calculate the mean of the C-index of each model in the three datasets

[0243] dd2$All <- apply(dd2[,2:4], 1, mean)

[0244] # Calculate the mean of the C-index of each model in the GEO validation set

[0245] dd2$GEO <- apply(dd2[,3:4], 1, mean)

[0246] View the C-index of each model

[0247] head(dd2)

[0248] # Output the C-index result

[0249] # write.table(dd2, "output_C_index.txt", col.names = T, row.names = F, sep = "\t", quote = F)

[0250] dd2 <- read.table('output_C_index.txt', header = T, sep = "\t", check.names = F)

[0251] library(ComplexHeatmap)

[0252] library(circlize)

[0253] library(RColorBrewer)

[0254] # Sort by C-index

[0255] dd2<-dd2[order(dd2$GEO,decreasing=T),]

[0256] # Only draw the C-index heatmap for the GEO validation set

[0257] dt<-dd2[,3:4]

[0258] rownames(dt)<-dd2$Model

[0259] ## Heatmap drawing

[0260]

[0261]

[0262] It should be noted that in this article, the storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.

[0263] The term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such a process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0264] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0265] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

[0266] The above is only the preferred embodiment of the present application and the technical principles applied. The present application is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions that can be made by those skilled in the art will not depart from the protection scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, it may also include more other equivalent embodiments, and the scope of the present application is determined by the scope of the claims.

Claims

1. A scoring method based on single-cell analysis for evaluating the prognosis of head and neck squamous cell carcinoma, characterized in that, it includes: Obtaining sample data, which includes single-cell data of head and neck squamous cell carcinoma and large cohort RNA data. The large cohort RNA data includes RNA sequence data and clinical information of head and neck squamous cell carcinoma patients; Performing quality control on the single-cell data and the large cohort RNA data; Clustering and annotating the cells in the single-cell data; Defining malignant cells and metabolic activity genes; Grouping cells according to metabolic activity; Based on the correlation between the metabolic activity of the grouped cells and head and neck squamous cell carcinoma, establishing the relationship between the proportion of different cell groups and the prognosis of head and neck squamous cell carcinoma; Screening metabolic biomarkers according to the scoring system; The specific quality control of the single-cell data includes: removing cells with ≤200 expressed genes and cells with ≥2500 gene expression numbers; removing doublet cells; removing genes expressed in less than 3 cells; removing mitochondrial genes; merging data sets; The specific clustering and annotation of the cells in the single-cell data include: performing dimensionality reduction clustering on all cells, and filtering out cells with ribosomal gene expression exceeding 35% in each subpopulation; finding highly expressed genes in each subpopulation, and annotating each cell subpopulation according to the found highly expressed genes, specifically expressed genes, and reported cell marker genes; the annotated cells are divided into subpopulations such as epithelial cells, T cells, B cells, plasma cells, monocytes, mast cells, fibroblasts, and malignant cells; The definition of malignant cells includes: performing dimensionality reduction clustering on the malignant cells in the sample data; classifying according to the origin of malignant cells in each group: the cell group with the most metastatic cancer sample sources is defined as the metastatic group, the cell group with the most primary cancer sample sources is defined as the primary group, and the remaining malignant cell groups with equivalent metastatic and primary cancer sample sources are defined as the mixed group; The grouping of cells according to metabolic activity includes: calculating the cut-off value of metabolic activity; defining malignant cells with a metabolic activity higher than the cut-off value as high-metabolic-activity cells, and malignant cells with a metabolic activity lower than the cut-off value as low-metabolic-activity cells; defining cells other than malignant cells in the dimensionality reduction clustering as non-malignant cells; The definition of metabolic activity genes includes: screening differential genes between metastatic group cells and primary group cells; performing enrichment analysis on the differential genes; defining the differential genes enriched in the metabolic pathway as metabolism-related genes; performing GSEA analysis on the metabolism-related genes, and defining the metabolism-related genes enriched in the metastatic cell group as metabolic activity genes; The specific scoring system includes: using the deconvolution algorithm to calculate the proportion of high-metabolic-activity cells corresponding to all RNA sequence data of patients according to the gene expression profiles of high-metabolic-activity cells and low-metabolic-activity cells, and defining this proportion as the scoring system.

2. A scoring system based on single-cell analysis for evaluating the prognosis of head and neck squamous cell carcinoma, characterized in that, it includes: A sample acquisition module for obtaining sample data, which includes single-cell data of head and neck squamous cell carcinoma and large cohort RNA data. The large cohort RNA data includes RNA sequence data and clinical information of head and neck squamous cell carcinoma patients; A quality control module for performing quality control on the single-cell data and bulk RNA data; A clustering and annotation module for clustering and annotating the cells in the single-cell data; A definition module for defining malignant cells and metabolic activity genes; A cell population module for clustering cells according to metabolic activity; A scoring module for establishing the relationship between the proportion of different cell populations and the prognosis of head and neck squamous cell carcinoma based on the correlation between the metabolic activity of the clustered cells and head and neck squamous cell carcinoma as a scoring system; A screening module for screening metabolic biomarkers according to the scoring system; Performing quality control on the single-cell data specifically includes: Removing cells with ≤200 expressed genes and cells with ≥2500 gene expression counts; Removing doublet cells; Removing genes expressed in fewer than 3 cells; Removing mitochondrial genes; Merging the datasets; Clustering and annotating the cells in the single-cell data specifically includes: Performing dimensionality reduction clustering on all cells and filtering out cells with ribosomal gene expression exceeding 35% in each subpopulation; Finding highly expressed genes in each subpopulation and annotating each cell subpopulation according to the found highly expressed genes, specifically expressed genes, and reported cell marker genes; The annotated cells are divided into subpopulations such as epithelial cells, T cells, B cells, plasma cells, monocytes, mast cells, fibroblasts, and malignant cells; The definition module includes a malignant cell definition module and a metabolic activity gene definition module; The malignant cell definition module is used for: performing dimensionality reduction clustering on the malignant cells in the sample data and classifying them according to the origin of the malignant cells in each group: the cell group with the largest number of metastatic cancer sample sources is defined as the metastatic group, the cell group with the largest number of primary cancer sample sources is defined as the primary group, and the remaining malignant cell groups with comparable metastatic and primary cancer sample sources are defined as the mixed group; The metabolic activity gene definition module is used for: screening differential genes between metastatic group cells and primary group cells; performing enrichment analysis on the differential genes; defining the differential genes enriched in the metabolic pathway as metabolism-related genes; performing GSEA analysis on the metabolism-related genes and defining the metabolism-related genes enriched in the metastatic cell group as metabolic activity genes; The scoring module specifically includes: using the deconvolution algorithm to calculate the proportion of highly metabolically active cells corresponding to all RNA sequence data of patients according to the gene expression profiles of highly metabolically active cells and low metabolically active cells, and defining this proportion as the scoring system.

Citation Information

Patent Citations

  • Analysis method based on 10X unicell transcriptome sequencing data

    CN109979538A

  • Method for analyzing energy metabolism of cell population

    CN113677992A