Cancer cell mitochondrial biomarker and screening method based on machine algorithm

By using machine learning algorithms to screen for mitochondrial biomarkers OXCT1, CLPB, SLC25A12, MRPL51, SFXN1, GATM, and TRMT10C genes in colorectal cancer cells, and combining LASSO regression and SVM-RFE regression algorithms, the accuracy and specificity issues in cancer diagnosis and treatment in existing technologies have been resolved, achieving highly efficient cancer diagnosis and treatment results.

CN121171344APending Publication Date: 2025-12-19HUNAN UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511315758.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-09-14
Filing Date
2025-09-15
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing technologies struggle to provide highly accurate, specific, and sensitive mitochondrial biomarkers for cancer cells, and the limitations of single-analytical results and their inherent errors restrict progress in cancer diagnosis and treatment.

Method used

Using machine learning algorithms, we identified mitochondrial biomarkers OXCT1, CLPB, SLC25A12, MRPL51, SFXN1, GATM, and TRMT10C genes in colorectal cancer cells. By combining LASSO regression and SVM-RFE regression algorithms, we screened out characteristic genes, constructed a survival prediction model, and performed multi-dimensional validation.

Benefits of technology

It has achieved highly accurate, specific, and sensitive identification of mitochondrial biomarkers in cancer cells, providing new evidence for cancer diagnosis and treatment, and improving the diagnosis and treatment outcomes of colorectal cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121171344A_ABST
    Figure CN121171344A_ABST
Patent Text Reader

Abstract

The invention discloses a cancer cell mitochondrial biomarker and a screening method based on a machine algorithm, and when the cancer cell is colorectal cancer, the mitochondrial biomarker comprises OXCT1, CLPB, SLC25A12, MRPL51, SFXN1, GATM and TRMT10C genes. The screening method comprises the following steps: (1) standardizing an expression data set of cancer cells to be screened, and then obtaining the expression quantity of mitochondrial related genes in the expression data of the cancer cells; (2) difference analysis: carrying out correlation analysis on cancer cell mitochondrial genes or proteins subjected to differential expression, and selecting gene modules which simultaneously meet differential genes and belong to mitochondrial related genes; (3) functional enrichment analysis; and (4) respectively screening by selecting at least two machine algorithms to respectively obtain a characteristic gene or protein list, and determining a core intersection target. The biomarker disclosed by the invention is high in accuracy, specificity and sensitivity. The method is efficient and comprehensive.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a kind of mitochondrial biomarkers and screening methods, in particular to a kind of cancer cell mitochondrial biomarkers and the screening method based on machine algorithm. BACKGROUND

[0002] When the cell in the body mutates, it will continuously divide without being controlled by the body, and eventually form cancer, which is the collective term for more than 100 related diseases. As a cancer with high mortality worldwide, the incidence of colorectal cancer (CRC) ranks second among various malignant tumors in China, and the mortality rate ranks fifth among various cancers, with a serious disease burden and a threat to human health. CRC is usually diagnosed in the middle and late stages, and the prognosis is poor, so it is urgent to clarify the molecular mechanism of the occurrence and development of colorectal cancer.

[0003] Mitochondria, as the organelle of cellular oxidative phosphorylation energy supply, are particularly important for maintaining the normal function and number of mitochondria and cell stability. Mitochondria play a crucial role in tumorigenesis by supporting the growth and survival of tumor cells in harsh environments, such as nutrient depletion and hypoxia: including mitochondrial damage disrupting the process of oxidative glycolysis balance that cancer cells rely on severely, leading to reduced ATP (adenosine triphosphate) production and gradual loss of tumor cell survival ability; mitochondria are key targets for regulating cancer cell proliferation, while DNA fragmentation, mitochondrial depolarization, and oxidative stress can activate apoptosis. The dynamic changes of reactive oxygen species (ROS) also play a crucial role in tumor cells, and excessive ROS increases apoptosis, while appropriate ROS levels promote cancer cell proliferation.

[0004] In summary, it is urgent to find a cancer cell mitochondrial biomarker and a machine algorithm screening method that is efficient, comprehensive, and eliminates the specificity and error of single analysis results, with high accuracy, specificity and sensitivity of mitochondrial biomarkers. SUMMARY

[0005] The technical problem to be solved by the present application is to overcome the above-mentioned defects existing in the prior art, and to provide a cancer cell mitochondrial biomarker with high accuracy, specificity and sensitivity.

[0006] The technical problem to be solved by the present application is to overcome the above-mentioned defects existing in the prior art, and to provide a cancer cell mitochondrial biomarker with high accuracy, specificity and sensitivity.

[0007] The application solves its technical problems by adopting the technical solution as follows: the mitochondrial biomarker of cancer cells, when the cancer cells are colorectal cancer cells, the mitochondrial biomarker includes OXCT1, CLPB, SLC25A12, MRPL51, SFXN1, GATM and TRMT10C genes. The application discloses specific mitochondrial biomarkers in colorectal cancer cells, and these genes show significant expression differences in colorectal cancer cells, providing new biomarker options for the diagnosis, prognosis and treatment of cancer.

[0008] The application further solves its technical problems by adopting the technical solution as follows: a method for screening mitochondrial biomarkers of cancer cells based on machine algorithms, comprising the following steps: (1) first standardize the expression data set of the cancer cells to be screened, obtain the expression data of the cancer cells, and then obtain the expression amount of the mitochondrial related genes in the expression data of the cancer cells, and obtain the mitochondrial related genes expressed in the cancer cells; (2) perform differential analysis on the mitochondrial related genes expressed in the cancer cells obtained in step (1), and then perform correlation analysis on the differentially expressed mitochondrial genes or proteins of the cancer cells, select the gene module that meets the differential genes and belongs to the mitochondrial related genes at the same time, and obtain the mitochondrial gene module related to the cancer cells; (3) perform functional enrichment analysis on the mitochondrial gene module related to the cancer cells obtained in step (2), and obtain a list of genes or proteins classified according to different functions; (4) select at least two machine algorithms to screen the list of genes or proteins classified according to different functions obtained in step (3), respectively, to obtain a list of characteristic genes or proteins, determine the core intersection target, and obtain the mitochondrial biomarker of the cancer cells.

[0009] The inventive idea of the method of the application is that traditional cancer markers mainly focus on the cell nucleus genome, and less research is conducted on mitochondrial genes, which limits the progress of cancer diagnosis and treatment. The method of the application focuses on the discovery and verification of mitochondrial biomarkers in cancer cells, and is realized by a machine algorithm screening method. The method of the application provides a method including data standardization processing, differential analysis and correlation analysis, functional enrichment analysis and machine algorithm screening, which can efficiently identify mitochondrial biomarkers such as colorectal cancer cells, ensure high accuracy, reliability and comprehensiveness, and provide new scientific basis and tools for precise diagnosis and treatment of colorectal cancer and other cancers.

[0010] Preferably, in step (1), the expression data set of all cancer tissues and para-cancer tissues of the cancer cells to be screened is downloaded from the GEO database.

[0011] Preferably, in step (1), the expression dataset is normalized by the normalizeBetweenArrays method through the perl and R language limma package. The differences between different samples and different genes can be eliminated by normalization.

[0012] Preferably, in step (1), the mitochondrial related genes are downloaded from the MitoCarta3.0 database.

[0013] Preferably, in step (1), the expression amount of the mitochondrial related genes in the expression data is obtained by the R language limma package. The cancer cell biomarkers including only mitochondrial related differential genes are obtained.

[0014] Preferably, in step (2), the differential analysis is performed by the R language limma package and pheatmap package. The up-regulated and down-regulated genes distinguished by the differential analysis are used to predict their high or low expression in tumors. High expression indicates that the gene is an oncogene, and low expression indicates that the gene is a tumor suppressor gene, thereby clarifying the expression trend for subsequent experimental verification.

[0015] Preferably, in step (2), the correlation analysis is performed by the R language corrplot package. After correlation analysis, red represents positive correlation, which indicates that both are oncogenes or tumor suppressor genes; blue represents negative correlation, which indicates that when one gene is a cancer (tumor suppressor) gene, the other correlation gene is a tumor suppressor (cancer) gene.

[0016] Preferably, in step (3), the functional enrichment analysis is performed by the R language enrichplot package through gene ontology.

[0017] Preferably, in step (3), the functional enrichment analysis is performed by the R language ggplot2 package and clusterProfiler package through the Kyoto Encyclopedia of Genes and Genomes. After functional enrichment analysis, the gene ontology enrichment related to molecular function, biological process and cellular component, and the enrichment degree in specific biological pathways are obtained, and the results obtained from the two are further combined for screening.

[0018] Preferably, in step (4), the machine algorithm comprises LASSO regression and / or SVM-RFE regression. The method of the present application can identify the characteristic genes of the relevant diseases through bioinformatics methods combined with multiple machine learning algorithms, and can discover new markers. LASSO regression refers to determining variables by finding the minimum λ of classification error, which is mainly used for screening characteristic variables and constructing the best classification model; SVM-REF is a machine learning method based on support vector machine, which finds the best variable by deleting the feature vector generated by SVM; the combination of the two methods can obtain more accurate disease markers. Therefore, exploring cancer cell mitochondrial biomarkers through machine algorithms has important value for clarifying the molecular mechanism of cancer cells and improving the accuracy of diagnosis and treatment of cancer cells.

[0019] Preferably, the specific screening method of the LASSO regression is: constructing a model through the R language glmnet package, drawing a cvfit graph and a LASSO regression graph, and further drawing a cross-validation graph on the cvfit graph to find the minimum value of the ordinate, that is, the minimum value of the cross-validation error, and determining the characteristic genes screened by the LASSO regression through the R language glmnet package.

[0020] Preferably, the specific screening method of the SVM-RFE regression is: setting ten-fold cross-validation through the R language e1071 package, sorting the importance of the characteristic genes, constructing a model to draw an accuracy graph to find the highest point of accuracy, drawing a cross-validation error graph to find the lowest point of error, and determining the characteristic genes screened by the SVM-RFE regression according to the results of the two through the R language e1071 package.

[0021] Preferably, in step (4), the core intersection target points of the characteristic genes screened by different machine algorithms are determined through the R language VennDiagram package.

[0022] Preferably, the cancer cell mitochondrial biomarkers obtained in step (4) are verified for accuracy, specificity and sensitivity.

[0023] Preferably, the verification method one is: performing differential expression analysis of the cancer cell mitochondrial biomarkers obtained in step (4) in cancer tissue and paracancer tissue.

[0024] Preferably, the verification method two is: obtaining the immunohistochemical data expression of the cancer cell mitochondrial biomarkers obtained in step (4) in human normal tissue and human cancer tissue from the HPA database.

[0025] Preferably, the verification method three is: performing differential expression analysis of mRNA expression level in human normal cells and human cancer cells.

[0026] Preferably, the verification method four is differential analysis of protein expression levels at the level of human normal cells and human cancer cells (Western Blot).

[0027] Preferably, the verification method five is to construct a survival prediction model based on the mitochondrial biomarkers of cancer cells obtained in step (4), and to evaluate the discrimination ability of the model for the survival outcome of patients.

[0028] Preferably, the verification method six is to evaluate the diagnostic efficiency of each mitochondrial biomarker of cancer cells obtained in step (4) by a receiver operating characteristic curve (ROC curve), and to quantitatively reflect the diagnostic ability by the area under the curve (AUC).

[0029] Preferably, the verification method seven is to select different physiological systems of human body as control tumor types, and to perform cross-tumor type specific differential expression analysis and verification on the mitochondrial biomarkers of cancer cells obtained in step (4). For example, in order to verify the specificity of the specific mitochondrial biomarkers in colorectal cancer cells, adrenal cortex cancer (ACC) of the endocrine system, cervical squamous cell carcinoma and adenocarcinoma (CESC) of the gynecological system, acute myeloid leukemia (LAML) of the hematological system, sarcoma (SARC) of special mesenchymal tissue origin and pancreatic adenocarcinoma (PAAD) of the digestive system are selected as control tumor types.

[0030] Preferably, the differential expression analysis in cancer disease tissue and pericancer tissue refers to downloading all sequencing data of the cancer disease tissue and pericancer tissue from the TCGA database, and obtaining the differential expression of the mitochondrial biomarkers of cancer cells in step (4) in the cancer disease tissue and pericancer tissue by R language limma package, R language ggplot2 package and R language ggpubr package.

[0031] Preferably, the immunohistochemical data expression refers to the difference in protein expression level.

[0032] Preferably, the differential expression analysis of mRNA expression level refers to verifying the differential expression of mRNA at the in vitro level of the mitochondrial biomarkers of cancer cells obtained in step (4) by qRT-PCR level at the level of human normal cells and human cancer cells.

[0033] Preferably, the differential analysis of protein expression level refers to verifying the protein expression difference of the mitochondrial biomarkers of cancer cells obtained in step (4) by Western Blot at the level of human normal cells and human cancer cells.

[0034] Preferably, the evaluation of the survival prediction model refers to evaluating the model's ability to distinguish the survival outcome of the patient by calculating the prediction consistency index of the model, the score interval, the total score range, the linear prediction value and the 1-year survival probability index of the patient under different risk states.

[0035] Preferably, the evaluation of the diagnostic efficiency of each biomarker refers to drawing a receiver operating characteristic curve with "1-specificity" as the abscissa representing the false positive rate and "sensitivity" as the ordinate representing the true positive rate, and the area under the curve is closer to 1, the diagnostic efficiency is stronger.

[0036] Preferably, the specific differential expression analysis refers to downloading the sequencing data of cancer tissues and normal tissues of the control tumor types from the TCGA database, processing the data through the R language limma package, the R language ggplot2 package and the R language ggpubr package, and analyzing the expression difference of the mitochondrial biomarkers of the cancer cells in the control tumor types to verify that they are significantly differentially expressed in the cancer cells relative to the control tumor types. In the specificity verification, the parameter settings, tool package versions and screening method steps (1) to (4) of the data processing and differential analysis are consistent to ensure the comparability of the analysis results; through comparison of multiple independent samples, the possibility of false positive differential expression of the mitochondrial biomarkers of the cancer cells in other systemic tumors is excluded to ensure the specificity of the cancer cell such as colorectal cancer diagnosis.

[0037] The beneficial effects of the present application are as follows: (1) The present application verifies the accuracy, specificity and high sensitivity of the mitochondrial biomarkers of the cancer cells for the diagnosis of colorectal cancer from multiple dimensions through the TCGA data set, human cell mRNA and protein level experiments, survival prediction models and ROC curve analysis; (2) Compared with the traditional single machine learning analysis method, the present application combines LASSO and SVM-RFE two machine language analyses as mitochondrial related markers, which not only retains the efficiency and comprehensiveness of machine language, but also eliminates the specificity and error of single analysis results; (3) The present application verifies the protein expression difference of the mitochondrial biomarkers of the cancer cells in human normal intestinal epithelial HIEC cells and human colon adenocarcinoma SW480 cells through Western Blot, and the survival prediction model constructed based on these genes has a certain distinguishing ability for the survival outcome of colorectal cancer patients. According to the ROC curve evaluation, the seven mitochondrial biomarkers of the cancer cells as a whole have diagnostic efficiency and diagnostic value, and can be used as potential diagnostic markers for colorectal cancer; (4) The application is verified by comparison with ACC, CESC, LAML, SARC and PAAD and other tumor types, and it is confirmed that the mitochondrial biomarker of cancer cells in the application only presents significant differential expression in colorectal cancer, and has no significant differential expression in various control tumor types, further strengthening the diagnostic specificity, and providing a basis for accurate diagnosis and treatment of mitochondria-related cancer cells such as colorectal cancer.

[0038] The abbreviations in the application are explained: CRC: colorectal cancer; GEO: Gene Expression Omnibus; HPA: Human Protein Atlas; perl: Practical Extraction and Report Language; MitoCarta: Mammalian Mitochondrial List; GO: Gene Ontology; KEGG: Kyoto Encyclopedia of Genes and Genomes; LASSO: Least Absolute Shrinkage and Selection Operator; SVM-RFE: Support Vector Machine Recursive Feature Elimination for Feature Selection; TCGA: Cancer Genome Atlas; BP: Biological Process; CC: Cellular Component; MF: Molecular Function; TCA: Tricarboxylic Acid Cycle; ROS: Reactive Oxygen Species; qRT-PCR: Real-time Fluorescent Quantitative Reverse Transcription PCR; OXCT1: 3-Oxidized Coenzyme A Transferase 1; CLPB: Clp Protease Complex; SLC25A12: Solute Carrier Family 25 Member 12; MRPL51: Mitochondrial Ribosomal Protein L51; SFXN1: Ferrodoxin 1; GATM: Glycine Amidinotransferase; TRMT10C: TRNA Methyltransferase 10C; HIEC: Human Intestinal Epithelial Cell, which is a human-derived healthy intestinal cell isolated from intestinal epithelium and cultured in the laboratory; SW480: a human colon cancer cell line derived from a tumor tissue of a colon adenocarcinoma patient and widely used for in vitro culture; AUC: Area Under the Curve; ACC: Adrenocortical Carcinoma; CESC: Cervical Squamous Cell Carcinoma and Adenocarcinoma; LAML: Acute Myeloid Leukemia; SARC: Sarcoma; PAAD: Pancreatic Adenocarcinoma. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 is a difference analysis diagram of the mitochondrial related genes expressed in the CRC of step (2) of the embodiment of the application; Figure 2 is a correlation analysis diagram of the mitochondrial genes in the CRC of step (2) of the embodiment of the application; Figure 3 A is a GO function enrichment analysis diagram of the mitochondrial gene module related to colorectal cancer of step (3) of the embodiment of the application, Figure 3 B is a KEGG pathway enrichment analysis diagram of the mitochondrial gene module related to colorectal cancer of step (3) of the embodiment of the application; Figure 4 A is a cvfit and cross-validation error path diagram of genes classified according to different functions of step (4) of the embodiment of the application,Figure 4 B is the LASSO regression graph of the genes classified according to different functions in step (4) of the embodiment of the present application; Figure 5 A is the SVM-RFE regression accuracy graph of the genes classified according to different functions in step (4) of the embodiment of the present application, Figure 5 B is the SVM-RFE regression cross-validation error graph of the genes classified according to different functions in step (4) of the embodiment of the present application; Figure 6 is the core intersection target point graph of the feature genes screened by LASSO and SVM-RFE in step (4) of the embodiment of the present application; Figure 7 is the differential expression analysis graph of the 7 mitochondrial biomarkers of colorectal cancer obtained in step (4) of the embodiment of the present application, 3 of which are up-regulated in colorectal cancer tissues ( Figure 7 A, 7F, 7G) and 4 of which are down-regulated in colorectal cancer tissues adjacent to normal tissues ( Figure 7 B, 7C, 7D, 7E); Figure 8 is the immunohistochemical graph of the protein expression of the 7 mitochondrial biomarkers of colorectal cancer obtained in step (4) of the embodiment of the present application in human normal colorectal tissues and human colorectal cancer tissues, 3 of which are up-regulated in human colorectal cancer tissues ( Figure 8 A, 8B, 8C) and 4 of which are down-regulated in human normal colorectal tissues ( Figure 8 D, 8E, 8F, 8G); Figure 9 is the differential expression analysis graph of the 7 mitochondrial biomarkers of colorectal cancer obtained in step (4) of the embodiment of the present application in human normal intestinal epithelial HIEC cells and human colon adenocarcinoma SW480 cells; Figure 10 is the Western Blot result graph of the 7 mitochondrial biomarkers of colorectal cancer obtained in step (4) of the embodiment of the present application in human normal intestinal epithelial HIEC cells and human colon adenocarcinoma SW480 cells; Figure 11 is the survival prediction model graph constructed by the 7 mitochondrial biomarkers of colorectal cancer obtained in step (4) of the embodiment of the present application; Figure 12 is the ROC curve analysis graph of the 7 mitochondrial biomarkers of colorectal cancer obtained in step (4) of the embodiment of the present application; Figure 13 is the expression difference comparison analysis graph of the 7 mitochondrial biomarkers of colorectal cancer obtained in step (4) of the embodiment of the present application in each control tumor type. DETAILED DESCRIPTION

[0040] The application will be further described below with reference to the embodiments and drawings.

[0041] The raw materials or chemical reagents used in the embodiments of the application are commercially available unless otherwise specified.

[0042] Colorectal cancer mitochondrial biomarker embodiment: The colorectal cancer mitochondrial biomarker includes OXCT1, CLPB, SLC25A12, MRPL51, SFXN1, GATM and TRMT10C genes.

[0043] Method for screening colorectal cancer mitochondrial biomarker based on machine algorithm embodiment: (1) First, the expression data set of all colorectal cancer tissues and adjacent tissues to be screened (GSE81558 data downloaded from the GEO database, a total of 32 samples, of which 23 were colorectal cancer patients in the experimental group and 9 were normal intestinal epithelial control samples in the control group), through perl and R language limma package, using the normalizeBetweenArrays method, the expression data set was standardized, and the colorectal cancer expression data was obtained, and then through the R language limma package, the expression amount of mitochondrial related genes (a total of 1136 mitochondrial related genes downloaded from the MitoCarta3.0 database) in the colorectal cancer expression data was obtained, and 1015 mitochondrial related genes expressed in colorectal cancer were obtained; (2) The 1015 mitochondrial related genes expressed in colorectal cancer obtained in step (1) were subjected to differential analysis by R language limma package and pheatmap package, and a total of 81 differentially expressed genes were obtained (as shown in Figure 1 , of which 45 were up-regulated and 36 were down-regulated), and the 81 differentially expressed colorectal cancer mitochondrial genes were subjected to correlation analysis by R language corrplot package (as shown in Figure 2 , blue represents positive correlation, red represents negative correlation, P value less than 0.05 is represented by “*”, less than 0.01 is represented by “**”, and less than 0.001 is represented by “***”), and the gene module that meets the differential genes and belongs to the mitochondrial related genes was selected, and the colorectal cancer related mitochondrial gene module was obtained; (3) The colorectal cancer related mitochondrial gene module obtained in step (2) was subjected to functional enrichment analysis by GO through R language enrichplot package, and the colorectal cancer related mitochondrial gene module obtained in step (2) was subjected to functional enrichment analysis by KEGG through R language ggplot2 package and clusterProfiler package, and the genes classified according to different functions were obtained; GO functional enrichment analysis (as shown in Figure 3As shown in Figure A, the enzyme is mainly enriched in BP and is associated with processes such as amino acid, carboxylic acid, organic acid, and small molecule catabolism. CC results indicate that it is associated with the mitochondrial matrix, inner and outer membranes, and protein complexes. MF results indicate that it is associated with various enzyme activities and protein activities. KEGG pathway analysis (e.g.) Figure 3 (As shown in B) indicates that the substance is mainly enriched in the TCA cycle, ROS, and various metabolic pathways. (4) LASSO regression and SVM-RFE regression were used to screen the genes obtained in step (3) according to different functional classifications, and the characteristic genes were obtained respectively. The core intersection target of the characteristic genes obtained by different machine algorithms was determined by the VennDiagram package of R language, and the colorectal cancer mitochondrial biomarkers were obtained. The specific screening method for LASSO regression is as follows: A model is constructed using the glmnet package in R language, and cvfit graphs are plotted (e.g., ...). Figure 4 As shown in Figure A) and LASSO regression graph (as shown in Figure A) Figure 4 (As shown in B), further plot the cross-validation graph on the cvfit graph, find the minimum value of the ordinate (here it is 13, representing 13 feature genes), that is, the minimum cross-validation error, and use the glmnet package in R language to determine the 13 feature genes for LASSO regression screening, namely OXCT1, CLPB, SLC25A12, DHTKD1, MRPL51, SFXN1, GATM, CPS1, TOMM5, DNAJC4, TRMT10C, PLGRKT and ARMCX2; The specific screening method for the SVM-RFE regression is as follows: using the R language e1071 package, setting up 10-fold cross-validation, ranking the importance of feature genes, constructing the model, and plotting the accuracy graph (e.g.) Figure 5 As shown in A), 14 points with the highest accuracy were found, and a cross-validation error graph was plotted (as shown in A). Figure 5 As shown in B), a total of 14 points with the lowest error were found. Based on the results of both, the 14 characteristic genes selected by SVM-RFE regression were determined using the e1071 package in R language. They are TRMT10C, GATM, OXCT1, SFXN1, MRPL51, DMGDH, HAO2, UCP2, CLPB, KMO, COX7A1, GLS2, SLC25A12 and CCDC51. Using the VennDiagram package in R, seven core intersection targets were identified between the 13 characteristic genes screened by LASSO regression and the 14 characteristic genes screened by SVM-RFE regression. These targets were OXCT1, CLPB, SLC25A12, MRPL51, SFXN1, GATM, and TRMT10C (e.g., ...). Figure 6 (As shown).

[0044] Accuracy, specificity and sensitivity verification of the colorectal cancer mitochondrial biomarkers obtained in step (4) were performed: Method one: The colorectal cancer mitochondrial biomarkers obtained in step (4) were subjected to differential expression analysis in colorectal cancer disease tissues and colorectal cancer adjacent tissues, that is, all sequencing data of the colorectal cancer disease tissues and colorectal cancer adjacent tissues were downloaded from the TCGA database (a total of 488 CRC data and 42 control group data), and the differential expression of the 7 colorectal cancer mitochondrial biomarkers obtained in step (4) in the colorectal cancer disease tissues and colorectal cancer adjacent tissues was obtained by R language limma package, R language ggplot2 package and R language ggpubr package, wherein CLPB (as shown in Figure 7 A), MRPL51 (as shown in Figure 7 B) and TRMT10C (as shown in Figure 7 C) were highly expressed in colorectal cancer tissues, GATM (as shown in Figure 7 D), OXCT1 (as shown in Figure 7 E), SFXN1 (as shown in Figure 7 F) and SLC25A12 (as shown in Figure 7 G) were highly expressed in colorectal cancer adjacent normal tissues, all with statistically significant differences.

[0045] Method two: The immunohistochemical data expression of the cancer cell mitochondrial biomarkers obtained in step (4) in human normal tissues and human cancer tissues was obtained from the HPA database, that is, the immunohistochemical data expression results of the 7 colorectal cancer mitochondrial biomarkers obtained in step (4) in human normal colorectal tissues and human colorectal cancer tissues were downloaded from the HPA database, wherein CLPB (as shown in Figure 8 A), MRPL51 (as shown in Figure 8 B) and TRMT10C (as shown in Figure 8 C) had enhanced protein expression levels in human colorectal cancer tissues, GATM (as shown in Figure 8 D), OXCT1 (as shown in Figure 8 E), SFXN1 (as shown in Figure 8 F) and SLC25A12 (as shown in Figure 8 G) had enhanced protein expression levels in human normal colorectal tissues, indicating that the clinical tissue level verification results were the same as the high and low expression trend of the markers screened by the machine algorithm, further confirming the accuracy and reliability of the screening method.

[0046] Method three is to analyze the differential expression of mRNA expression level at the level of human normal cells and human cancer cells, that is, at the level of human normal intestinal epithelial HIEC cells and human colon adenocarcinoma SW480 cells, the differential expression of mRNA at the level of in vitro is verified by qRT-PCR, wherein CLPB (as shown in Figure 9 A), MRPL51 (as shown in Figure 9 F) and TRMT10C (as shown in Figure 9 G) are highly expressed in colon cancer SW480 cells; GATM (as shown in Figure 9 B), OXCT1 (as shown in Figure 9 C), SFXN1 (as shown in Figure 9 D) and SLC25A12 (as shown in Figure 9 E) are highly expressed in normal intestinal epithelial HIEC cells, indicating that the cell level verification result is the same as the high and low expression trend of the markers screened by the machine algorithm, further confirming the accuracy and reliability of the screening method.

[0047] Method four is to analyze the differential expression of protein expression level at the level of human normal cells and human cancer cells, that is, at the level of human normal intestinal epithelial HIEC cells and human colon adenocarcinoma SW480 cells, the protein expression difference of the 7 mitochondrial biomarkers obtained in step (4) is verified by Western Blot, wherein CLPB, MRPL51, TRMT10C are highly expressed in colon cancer SW480 cells; GATM, OXCT1, SFXN1, SLC25A12 are highly expressed in normal intestinal epithelial HIEC cells, the protein level difference is consistent with mRNA (as shown in Figure 10 ); that is, in actual application, when CLPB, MRPL51, TRMT10C are highly expressed in cells, and GATM, OXCT1, SFXN1, SLC25A12 are lowly expressed, it can be diagnosed as highly suspected to have colorectal cancer, otherwise, the probability of having colorectal cancer is very low.

[0048] Method five is to construct a survival prediction model based on the 7 mitochondrial biomarkers obtained in step (4), the prediction consistency index (C-index) of the model is 0.572, the 95% confidence interval is 0.518-0.625, and p=0.008, indicating that the model has a significantly better ability to distinguish the survival outcome of colorectal cancer patients than the random level (as shown in Figure 11The score interval of each gene, the total score range (0-240 or so), the linear prediction value (range -1-1), and the 1-year survival probability of patients under different risk states (corresponding to 0.95, 0.9, 0.8, etc. values, representing the cumulative survival probability) are also presented in the figure. These indicators collectively provide a quantitative basis for the model to evaluate the survival risk of patients, helping to judge the prognosis of patients after suffering from colorectal cancer. For example, high-risk patients may need more aggressive treatment options, while low-risk patients can be relatively conservative in treatment strategies, thereby providing an important reference for the treatment and management of colorectal cancer patients. The survival risk indirectly assists in judging the situation related to colorectal cancer, helping to make clinical decisions related to colorectal cancer, etc.

[0049] Method six is: evaluating the diagnostic efficiency of the seven mitochondrial biomarkers of colorectal cancer obtained in step (4) by ROC curve, and quantifying the diagnostic ability by AUC; taking "1-specificity" as the horizontal coordinate, representing the false positive rate, and taking "sensitivity" as the vertical coordinate, representing the true positive rate, to draw the ROC curve. The closer the AUC is to 1, the stronger the diagnostic efficiency. The results show that the AUC of OXCT1 is 0.835, the AUC of CLPB is 0.716, the AUC of SLC25A12 is 0.728, the AUC of MRPL51 is 0.713, the AUC of SFXN1 and GATM is 0.746, and the AUC of TRMT10C is 0.691, all of which are higher than 0.5, indicating that the overall AUC performance of the seven mitochondrial biomarkers of colorectal cancer obtained in step (4) is better than the random level, and they show certain diagnostic value (such as Figure 12 as shown).

[0050] Method seven is: selecting ACC, CESC, LAML, SARC and PAAD as control tumor types, and performing specific differential expression analysis and verification on the seven colorectal cancer mitochondrial biomarkers obtained in step (4); the specific differential expression analysis is: downloading the sequencing data of cancer tissues and normal tissues of the control tumor types (ACC: 79 cancer tissues and 17 normal tissues; CESC: 306 cancer tissues and 3 normal tissues; LAML: 173 cancer tissues and 7 normal tissues; SARC: 265 cancer tissues and 2 normal tissues; PAAD: 178 cancer tissues and 4 normal tissues) from the TCGA database, processing the data by R language limma package, R language ggplot2 package and R language ggpubr package, and analyzing the expression difference of the seven colorectal cancer mitochondrial biomarkers in the control tumor types, so as to verify that the seven colorectal cancer mitochondrial biomarkers present significant specific differential expression in marking colorectal cancer compared with the control tumor types, wherein the parameter settings, tool package versions and screening methods of data processing and differential analysis are consistent with steps (1)-(4). The results show that the seven colorectal cancer mitochondrial biomarkers have no statistically significant difference (P>0.05) between cancer tissues and normal tissues in ACC, CESC, LAML, SARC and PAAD, and only maintain a significant difference (P<0.05) between colorectal cancer tissues and cancer-adjacent tissues in the TCGA database (as shown in Figure 13 ), which directly proves that the seven colorectal cancer mitochondrial biomarkers have no significant differential expression in ACC, CESC, LAML, SARC and PAAD, and only present exclusive significant specific differential expression for colorectal cancer, thereby avoiding misjudgment of other tumors.

[0051] In summary, the sensitivity, specificity and diagnostic value of the seven colorectal cancer mitochondrial biomarkers (OXCT1, CLPB, SLC25A12, MRPL51, SFXN1, GATM and TRMT10C) have been fully proved by multidimensional experimental design and verification, and the exclusivity thereof is clear: in terms of sensitivity, the differential analysis of 488 colorectal cancer tissues and 42 cancer-adjacent tissues in the TCGA database (as shown in Figure 7 ), the clinical immunohistochemical verification in the HPA database (as shown in Figure 8 ), the mRNA (as shown in Figure 9 ) and protein levels (as shown in Figure 10As shown) verification, all present the stable difference pattern of "CLPB, MRPL51, TRMT10C are highly expressed in colorectal cancer, GATM, OXCT1, SFXN1, SLC25A12 are highly expressed in normal tissues / cells", and all differences meet statistical significance (P<0.05), which can effectively detect colorectal cancer and avoid missed diagnosis; in terms of specificity and exclusivity, by selecting ACC, CESC, LAML, SARC, PAAD as control tumor types from TCGA database, after uniform analysis tool and parameters, it is found that the 7 markers have no significant difference (P>0.05) between cancer tissues and normal tissues of control tumor types, and only present specific differential expression in colorectal cancer (such as Figure 13 As shown), which proves that other types of cancer cells will not appear the same difference, effectively excluding cross-cancer interference; in terms of diagnostic value, ROC curve analysis shows that the AUC of single marker is higher than 0.5 (OXCT1 is as high as 0.835), and the whole has good diagnostic performance (such as Figure 12 As shown), and the survival prediction model (C-index=0.572, P=0.008) based on the marker can assist clinical prognosis evaluation, further embodying practical value. Since the marker has significant specificity for colorectal cancer, when facing unknown cancer types, it can be used as one of the core diagnostic basis for colorectal cancer, combined with patient symptoms, imaging examination and other comprehensive judgment for clinical diagnosis, which can greatly improve the diagnostic accuracy.

Claims

1. A cancer cell mitochondrial biomarker, characterized by: When the cancer cell is colorectal cancer, the mitochondrial biomarker thereof includes OXCT1, CLPB, SLC25A12, MRPL51, SFXN1, GATM and TRMT10C genes.

2. A method of screening for cancer cell mitochondrial biomarkers based on machine algorithm as claimed in claim 1, wherein, The method comprises the following steps: (1) standardizing the expression data set of the cancer cell to be screened to obtain cancer cell expression data, and then obtaining the expression amount of the mitochondrial related gene in the cancer cell expression data to obtain the mitochondrial related gene expressed in the cancer cell; (2) performing differential analysis on the mitochondrial related gene expressed in the cancer cell obtained in step (1), and then performing correlation analysis on the obtained differentially expressed cancer cell mitochondrial gene or protein, selecting a gene module that meets the differential gene and belongs to the mitochondrial related gene at the same time to obtain a cancer related mitochondrial gene module; (3) performing functional enrichment analysis on the cancer related mitochondrial gene module obtained in step (2) to obtain a gene or protein list classified according to different functions; (4) selecting at least two machine algorithms to screen the gene or protein list classified according to different functions obtained in step (3) respectively to obtain a feature gene or protein list, determining a core intersection target, and obtaining a cancer cell mitochondrial biomarker.

3. The method for screening mitochondrial biomarkers of cancer cells based on machine algorithms according to claim 2, characterized in that: In step (1), the expression data set of all cancer tissues and para-cancer tissues of the cancer cell to be screened is downloaded from the GEO database; the expression data set is standardized by using the normalizeBetweenArrays method of the perl and R language limma package; and the mitochondrial related gene is downloaded from the MitoCarta3.0 database; The expression amount of the mitochondrial related gene in the expression data is obtained by using the R language limma package.

4. The method of claim 2 or 3, wherein the method is based on a machine learning algorithm. In step (2), the differential analysis is performed by using the R language limma package and pheatmap package; and the correlation analysis is performed by using the R language corrplot package.

5. The method of screening cancer cell mitochondrial biomarkers based on machine algorithm according to any one of claims 2-4, wherein: In step (3), the functional enrichment analysis is performed by using the gene ontology through the R language enrichplot package; and the functional enrichment analysis is performed by using the Kyoto Encyclopedia of Genes and Genomes through the R language ggplot2 package and clusterProfiler package.

6. The method of screening cancer cell mitochondrial biomarkers based on machine algorithm according to any one of claims 2-5, wherein: In step (4), the machine algorithm comprises LASSO regression and / or SVM-RFE regression; the specific screening method of the LASSO regression is: constructing a model by using the R language glmnet package, drawing a cvfit graph and a LASSO regression graph, further drawing a cross-validation graph on the cvfit graph, finding the minimum value of the ordinate, i.e., the minimum value of the cross-validation error, and determining the feature genes screened by the LASSO regression by using the R language glmnet package; the specific screening method of the SVM-RFE regression is: setting ten-fold cross-validation by using the R language e1071 package, sorting the importance of the feature genes, constructing a model to draw an accuracy graph, finding the highest point of the accuracy, drawing a cross-validation error graph, finding the lowest point of the error, and determining the feature genes screened by the SVM-RFE regression according to the results of the two by using the R language e1071 package; and the core intersection target points of the feature genes screened by different machine algorithms are determined by using the R language VennDiagram package.

7. The method of screening for cancer cell mitochondrial biomarkers based on machine algorithm according to any one of claims 2 to 6, wherein: The cancer cell mitochondrial biomarkers obtained in step (4) are verified for accuracy, specificity and sensitivity; method one is to perform differential expression analysis on the cancer cell mitochondrial biomarkers obtained in step (4) in cancer tissues and paracancer tissues; method two is to obtain the immunohistochemical data expression of the cancer cell mitochondrial biomarkers obtained in step (4) in human normal tissues and human cancer tissues from the HPA database; method three is to perform differential expression analysis of mRNA expression levels in human normal cells and human cancer cells; method four is to perform differential analysis of protein expression levels in human normal cells and human cancer cells; method five is to construct a survival prediction model based on the cancer cell mitochondrial biomarkers obtained in step (4) and evaluate the model's ability to distinguish the survival outcomes of patients; method six is to evaluate the diagnostic performance of each cancer cell mitochondrial biomarker obtained in step (4) by using a receiver operating characteristic curve, and the diagnostic ability is quantitatively reflected by the area under the curve; and method seven is to select different physiological system cancers of the human body as control tumor types, and perform cross-tumor type specific differential expression analysis and verification on the cancer cell mitochondrial biomarkers obtained in step (4).

8. The method for screening mitochondrial biomarkers of cancer cells based on machine algorithms according to claim 7, characterized in that: The differential expression analysis in cancer disease tissue and paracancer tissue refers to downloading all sequencing data of the cancer disease tissue and paracancer tissue from the TCGA database, obtaining the differential expression of the mitochondrial biomarkers of cancer cells in the cancer disease tissue and paracancer tissue by R language limma package, R language ggplot2 package and R language ggpubr package; the immunohistochemical data expression refers to the difference in protein expression level; the differential expression analysis of mRNA expression level refers to verifying the differential expression of the mitochondrial biomarkers of cancer cells in the mRNA level in vitro by qRT-PCR level at the level of human normal cells and human cancer cells; the differential analysis of protein expression level refers to verifying the protein expression difference of the mitochondrial biomarkers of cancer cells by Western Blot at the level of human normal cells and human cancer cells; the evaluation of the survival prediction model refers to evaluating the distinguishing ability of the model for the survival outcome of patients by calculating the prediction consistency index, score interval, total score range, linear prediction value and 1-year survival probability index of patients in different risk states; the evaluation of the diagnostic efficiency of each biomarker refers to drawing a receiver operating characteristic curve with "1-specificity" as the abscissa representing the false positive rate and "sensitivity" as the ordinate representing the true positive rate, and the area under the curve is closer to 1, the stronger the diagnostic efficiency; the specific differential expression analysis refers to downloading the sequencing data of the cancer tissue and normal tissue of the control tumor type from the TCGA database, processing the data by R language limma package, R language ggplot2 package and R language ggpubr package, and analyzing the expression difference of the mitochondrial biomarkers of cancer cells in the control tumor type to verify that they present significant specific differential expression in the cancer cells relative to the control tumor type.

Citation Information

Patent Citations

  • Evaluation model for predicting prognosis and adjuvant chemotherapy benefit of colon cancer and application

    CN115011689A

  • Biomarker combination, reagent containing biomarker combination and application of biomarker combination

    CN115678994A

  • Colorectal cancer treatment target and serum marker and application thereof

    CN117169505A

  • Personalized treatment decision-making method and system for intestinal cancer and storage medium comprising personalized treatment decision-making method and system

    CN117219158A

  • Colorectal cancer specific methylation marker and application thereof

    CN117385028A