A cluster molecule-based prediction model for colorectal cancer and application thereof

By constructing a colorectal cancer prediction model based on cluster molecules and using machine learning methods to screen for colorectal cancer-specific protein molecules, the problems of complexity in colorectal cancer screening and low specificity of hematological tests in existing technologies have been solved, achieving efficient and economical early diagnosis and improving the detection rate of colorectal cancer.

CN116246710BActive Publication Date: 2026-02-10广州医科大学附属清远医院(清远市人民医院)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211743182.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-02-10
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

Current colorectal cancer screening methods are complex, expensive, and have low coverage. Furthermore, hematological tests have low molecular specificity, resulting in low early diagnosis rates and missed opportunities for optimal treatment.

Method used

A cluster-molecule-based colorectal cancer prediction model was constructed. By collecting colorectal cancer transcript sequencing data, colorectal cancer-specific expression genes encoding blood proteins were screened. A logistic regression model was constructed using machine learning methods and applied to hematological screening and risk assessment to improve the specificity and accuracy of detection.

Benefits of technology

It significantly improves the objectivity, accuracy, and sensitivity of colorectal cancer diagnosis, increases the accuracy of detection, is suitable for early screening of colorectal cancer, simplifies the examination process, reduces costs, and is easy to promote.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246710B_ABST
    Figure CN116246710B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of biomedicine, and discloses a colorectal cancer prediction model based on cluster molecules and application. The application uses an open shared human protein database and a TCGA public database, carries out intersection analysis by using a machine learning method, screens out a specific molecular set of a group of coded proteins that can be detected in blood and indicate the risk of colorectal cancer, and constructs a logistic regression model by using a GEO colorectal cancer data set. The AUC value of the model for predicting the risk of colorectal cancer is 0.962, and the model can accurately distinguish high-risk and low-risk populations in the test set. The application uses multivariate set effects for modeling, plays a role of multivariate common screening, greatly improves the prediction accuracy and sensitivity of the model, and is simple in the hematological screening method and suitable for popularization in clinical application.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biomedical technology, in particular to a colorectal cancer prediction model based on cluster molecules and application. BACKGROUND

[0002] Colorectal cancer (CRC) is one of the most common malignant tumors in the digestive system, with the third highest incidence and the second highest mortality rate. Due to the low rate of early diagnosis, most patients are in the middle and advanced stages at the time of initial diagnosis, and the prognosis is poor. Currently, the screening of colorectal cancer relies on colonoscopy, colon CT, serological or fecal occult blood test, etc. Colonoscopy or colon CT has high diagnostic accuracy, but the examination process is relatively complex, the cost is high, the popularization rate is low, and the patient needs to prepare in advance for intestinal emptying, which has low patient compliance and relatively high cost. It is difficult to popularize in routine physical examination, and patients often seek medical treatment when they have typical symptoms such as hematochezia, which delays the disease and misses the best treatment window. Fecal occult blood test is convenient to sample and intuitive in symptoms, and easy to attract the attention of patients, but CRC patients usually have advanced disease when the fecal occult blood test is positive. Hematology tests are widely used in clinical practice, and commonly used detection indicators include CEA, CA199, CA242, CA50, etc. These molecules have been proven to be expressed in a variety of tumors and are often used as general cancer warning signals. However, due to the limitations of existing single serological detection indicators, low detection molecule specificity, and patient individual differences, the positive detection rate still needs to be improved. Therefore, developing new technologies and improving the early detection rate of colorectal cancer are current problems to be solved. SUMMARY

[0003] The purpose of the present application is to overcome the deficiencies of the prior art and provide a colorectal cancer prediction model based on cluster molecules and application.

[0004] To achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows:

[0005] In a first aspect, the present application provides a method for constructing a colorectal cancer prediction model based on cluster molecules, comprising the following steps:

[0006] (1) Collecting colorectal cancer transcript sequencing data and colorectal cancer data set; extracting probe values and probe annotations from the colorectal cancer data set, removing batch effects, and obtaining a merged data set;

[0007] (2) Collecting genes that can encode blood proteins;

[0008] (3) Screening differentially expressed genes in the colorectal cancer transcript sequencing data, and verifying the differentially expressed genes using the obtained merged data set;

[0009] (4) using the weight gene co-expression network analysis method to screen the colorectal cancer specific expression genes that can encode blood proteins from the genes in step (2);

[0010] (5) screening the colorectal cancer specific protein encoding genes from the colorectal cancer specific expression genes based on the machine learning method to obtain the colorectal cancer cluster molecules;

[0011] (6) verifying the reliability of the obtained colorectal cancer cluster molecules;

[0012] (7) training the cohort based on the merged data set by using the regression method, and obtaining the joint diagnosis score by multiplying the colorectal cancer cluster molecule expression value by the regression coefficient, that is,.

[0013] According to the open sharing human protein database and the TCGA public database resources, the intersection analysis is carried out by using the machine learning method, the specific molecule set that can encode the blood related proteins and predict the risk of colorectal cancer is screened out, the logical regression model (prediction model) is constructed by using the colorectal cancer data set in the GEO database, and is applied to the hematological screening and risk assessment of colorectal cancer. The abnormal expression genes of tumor cells can be translated into corresponding proteins and released into blood through different pathways. The hematological detection method is selected to improve the compliance of the tested person. The application selects the protein molecule group highly related to CRC by means of the large sample database of CRC, and guarantees the specificity of the detected molecules. In addition, the absolute value of the CRC high expression protein cluster molecule is brought into the regression equation, and the obtained value is used as a comprehensive evaluation index, which significantly improves the objectivity, accuracy and sensitivity of CRC diagnosis, and is more representative.

[0014] As a preferred embodiment of the construction method of the application, in step (1), the transcript sequencing data comes from The Cancer Genome Atlas database; the colorectal cancer data set comes from the independent colorectal cancer data sets GSE9348 and / or GSE41258 in the GENE EXPRESSION OMNIBUS database; the probe value and probe annotation are extracted from the colorectal cancer data sets GSE9348 and / or GSE41258, and the batch effect is removed by using the Sva software package to obtain the merged data set.

[0015] Preferably, in step (2), the data of the genes come from the human protein database HPA and / or the human body fluid protein database HBFP. Preferably, the blood includes whole blood, serum and plasma.

[0016] As a preferred embodiment of the construction method of the application, in step (3), the differential expression genes are screened by using the limma software package, and the screening standard is log2 FC>1.5 and FDR<0.05.

[0017] As a preferred embodiment of the construction method of the present application, in step (4), the weight gene co-expression network analysis method is used to cluster the highly synergistically changed genes to generate corresponding modules, the internal connectivity of the genes in the associated modules and the correlation of the associated modules with the clinicopathological features of colorectal cancer are analyzed, and the core genes in the module with the highest correlation are found out; the core genes are considered to be MEblue and / or MEturquoise.

[0018] As a preferred embodiment of the construction method of the present application, in step (5), the screening is performed using the Lasso regression model, the random forest algorithm and the SVM-RFE algorithm. The colorectal cancer-specific protein-coding genes are highly expressed in colorectal cancer cells, and the proteins encoded by the genes enter the blood or body fluid through different ways; the colorectal cancer cluster molecules can be used as a colorectal cancer screening and risk warning signal. Preferably, the colorectal cancer cluster molecules are obtained by gene level screening, and the application involves protein expression of the cluster molecules

[0019] The protein detection method can be any method known in the art, such as immunological, biological, chemical detection methods and other related protein detection methods.

[0020] As a preferred embodiment of the construction method of the present application, in step (6), the expression of the colorectal cancer cluster molecules in the CRC samples is verified by the limma software package; the area under the ROC curve (AUC) and the 95% confidence interval of the colorectal cancer cluster molecules are calculated by the pROC software.

[0021] As a preferred embodiment of the construction method of the present application, in step (7), the combined data set is randomly divided into a CRC training queue and a verification queue at a ratio of 1:1, a prediction model is constructed in the training queue using the logistics regression method, the stability of the model is verified using the 10-fold cross-validation method, the combined diagnostic score is obtained by multiplying the expression value of the colorectal cancer cluster molecules by the regression coefficient, the AUC of the prediction model is calculated using the training queue, and the accuracy of the prediction model is identified using the verification queue.

[0022] Preferably, the numerical expression of the combined diagnostic score is: Cd score = ∑ (cluster molecule expression value x regression coefficient) + B, wherein the Cd is the combined diagnosis; the cluster molecule expression value is the protein expression value; and the B is the logistic regression constant term, which is automatically generated in the regression analysis.

[0023] Preferably, the prediction model is evaluated by actual joint diagnostic score, specifically, the actual ROC curve area (AUC value) and the prediction accuracy of the subject. The ROC curve is used to judge the accuracy and application value of each cluster molecule in predicting the risk of colorectal cancer: AUC <0.5 indicates that the variable index has no prediction value, 0.5 < AUC <0.7 indicates that the variable index has low prediction accuracy, 0.5 < AUC <0.7 indicates that the variable index has medium prediction accuracy, AUC >0.9 indicates that the evaluation index has high accuracy, and the ideal index is AUC =1.

[0024] In a second aspect, the present application provides a cluster molecule-based colorectal cancer prediction model constructed by the above method.

[0025] In a third aspect, the present application applies the colorectal cancer cluster molecules, including QSOX2, TGFBI, CD44, INHBA, S100A11, VEGFA and MET, in the preparation of colorectal cancer screening and / or prediction reagents. Preferably, the expression of the colorectal cancer cluster molecules is applied in the preparation of colorectal cancer screening and / or prediction reagents, which can be used as an early warning signal and early molecular screening method for colorectal cancer. Among them, the secreted protein encoded by QSOX2 is related to tumor proliferation; TGFBI is a tumor-related secreted protein induced by transforming growth factor beta; VEGFA is related to angiogenesis; CD44 is related to tumor stemness; MET is related to tumor mutant wnt signaling pathway; S100A11 is a member of the S100 protein family and is highly expressed in various tumors; INHBA is a member of the transforming growth factor-beta (TGF-beta) superfamily and is related to tumor angiogenesis, etc.

[0026] In a fourth aspect, the present application applies the cluster molecule-based colorectal cancer prediction model in the preparation of colorectal cancer screening and / or prediction reagents.

[0027] Compared with the prior art, the present application has the following advantages:

[0028] This invention utilizes a CRC prediction model based on the Cd score of colorectal cancer-specific expression cluster molecules for hematological screening. The detection indicators are selected from multiple CRC sequencing databases, resulting in a large sample size and good representativeness. Validated in clinically diagnosed colorectal cancer samples, the model calculates the Cd-score by detecting the expression of protein cluster molecules in the blood of clinical colorectal cancer patients. The AUC is 0.962, with an accuracy of 91.9%. Compared to the clinically used CEA hematological test results (AUC 0.71, accuracy 79.7%), the accuracy of this protein cluster molecule prediction model is improved by 12.2%. Compared to colorectal endoscopy and CT scans, this method is simple, economical, and easy to promote, making it suitable for early CRC screening. Attached Figure Description

[0029] Figure 1 A schematic diagram of the process for screening and predicting blood-specific cluster molecules in CRC patients.

[0030] Figure 2 This is a cluster analysis of differentially expressed genes based on the TCGA-CRC dataset; A is a volcano plot of differentially expressed genes; B is a heatmap of differentially expressed genes.

[0031] Figure 3 A represents weighted gene co-expression network analysis (WGCNA); B represents the WGCNA soft threshold; C represents the gene dendrogram and module categories; and D represents the module gene weight analysis.

[0032] Figure 4 For enrichment analysis of differentially expressed molecules.

[0033] Figure 5 This study investigates the intersection of three algorithms and screens for differentially expressed molecules.

[0034] Figure 6 A is for the molecular validation of differentially expressed genes based on the merged datasets GSE9438 and GSE41258; B is for the validation of differentially expressed gene levels in colorectal cancer; C is for the ROC curves and AUC values ​​of each gene.

[0035] Figure 7 The prediction performance of the cluster molecular model and the CEA model based on the GEO merged dataset is shown in Figure A; A is the ROC curve of the cluster molecular and CEA control models; B is the confusion matrix of the two models.

[0036] Figure 8 A represents the expression of cluster molecules in clinical samples; B represents the detection of cluster molecule mRNA in clinical samples of colorectal cancer; C represents the detection of cluster molecule protein levels in clinical samples of colorectal cancer.

[0037] Figure 9The application of cluster molecular model and CEA in predicting clinical serum samples of CRC patients; A is the ROC curve of the cluster molecular prediction model and the CEA control model; B is the confusion matrix of the cluster molecular prediction model and the CEA control model. Detailed Implementation

[0038] To better illustrate the objectives, technical solutions, and advantages of this invention, the invention will be further described below with reference to specific embodiments. Those skilled in the art should understand that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0039] Unless otherwise specified, the experimental methods used in the examples are conventional methods; the materials and reagents used are commercially available unless otherwise specified.

[0040] Example: Establishing a hematological prediction model based on the molecular expression of colorectal cancer-specific protein clusters

[0041] Blood contains proteins synthesized and secreted by cells in various tissues and organs, or released into the bloodstream through other means. These mainly include cytokines, growth factors, complement, antibodies, peptide hormones, and immunoglobulins—proteins with important physiological functions. The rapid development of omics and bioinformatics technologies has advanced the progress of oncology research. This invention utilizes globally shared public databases, acquiring various malignant tumor sequencing databases for data mining and analysis to screen for specific tumor-specific markers in the blood, aiming to improve clinical diagnostic rates.

[0042] This invention utilizes colorectal cancer transcript sequencing data from the TCGA and GEO databases to screen differentially expressed genes in colorectal cancer. It compares various protein-coding genes in the blood and performs intersection analysis using machine learning methods such as Lasso regression, random forest, and SVM-RFE to identify seven overlapping CRC characteristic protein molecules (QSOX2, TGFBI, CD44, INHBA, CD44, VEGFA, and MET). These are validated using the CRC dataset in the GEO database. A logistic regression method is used to construct an equation, i.e., a predictive model, based on hematological detection values ​​of the "colorectal cancer characteristic protein cluster," which can be used for hematological screening of colorectal cancer. The technical process is as follows: Figure 1 As shown.

[0043] Specifically as follows:

[0044] 1. Patient dataset screening

[0045] Data from three different datasets were used in this study: colorectal cancer (CRC) transcript sequencing data from The Cancer Genome Atlas (TCGA); and two independent colorectal cancer datasets, GSE9348 and GSE41258, from the GENE EXPRESSION OMNIBUS (GEO) database. These three independent datasets covered Asian and European populations. The CRC dataset from TCGA contained a total of 699 samples, including 55 healthy individuals and 644 tumor samples. The GSE9348 dataset, after merging with the GSE41258 dataset, contained a total of 472 samples, with 460 usable samples. Detailed sample information is shown in Table 1.

[0046] Table 1: Summary of Clinical Information from Three Independent Datasets

[0047]

[0048] 2. Identify protein molecules that may exist in the blood.

[0049] A search of the Human Protein Atlas (HPA) database (https: / / www.proteinatlas.org / ) and the Human Body Fluid Protein Database (HBFP) (https: / / bmbl.bmi.osumc.edu / HBFP) identified 1524 genes that encode proteins in the blood.

[0050] 3. Gene expression data processing

[0051] Gene expression analysis was performed on FPKM values ​​of cancer and adjacent normal tissues in the TCGA-CRC dataset using the limma software package. Differentially expressed genes were screened based on log2 FC > 1.5 and FDR < 0.05 (see details). Figure 2 The study identified 325 upregulated genes and 358 downregulated genes in CRC. Two independent CRC datasets, GSE9348 and GSE41258, were selected from the GEO database. Probe values ​​and annotations were extracted from the original databases, batch effects were removed using the Sva software package, and the two datasets were merged to generate a new dataset. The newly generated GEO dataset was used to validate the differentially expressed genes in CRC selected from the TCGA database.

[0052] 4. Screening for colorectal cancer-specific expression genes encoding related proteins in the blood using WGCNA analysis.

[0053] Weighted Gene Co-Expression Network Analysis (WGCNA) was used to analyze gene expression patterns in multiple samples. Genes exhibiting highly co-expressive changes were clustered to generate corresponding modules (gene sets). The connectivity of genes within associated modules and the correlation between associated modules and clinicopathological features were analyzed. Core genes in the modules with the highest correlation (i.e., modules with high weight scores) were identified. This prediction model selected the two modules with the highest weight scores (MEblue and MEturquoise) as candidate gene sets, obtaining 125 upregulated genes highly correlated with CRC clinicopathological features. Further gene enrichment analysis of the obtained upregulated genes demonstrated that most of the upregulated genes were enriched in colorectal tumor-related pathways (see details). Figure 3 , Figure 4 ).

[0054] 5. Further screening of colorectal cancer-specific protein-coding genes based on machine learning methods.

[0055] Using the Lasso regression model, random forest algorithm, and SVM-RFE algorithm, seven CRC synchronously highly expressed molecules with overlapping intersections were screened out: QSOX2, TGFBI, CD44, INHBA, S100A11, VEGFA, and MET, collectively referred to as CRC cluster molecules (see details). Figure 5 ).

[0056] 6. CRC Cluster Molecular Credibility Verification Based on Merged GSE9438 and GSE41258 Datasets

[0057] The expression of CRC cluster molecules in CRC samples was verified using the limma software package, and the results showed that all molecules were upregulated. The area under the ROC curve (AUC) and 95% confidence interval of the CRC cluster molecules were calculated using the pROC software (see details). Figure 6 ).

[0058] 7. Construct a prediction model based on CRC cluster molecule and GEO merged datasets

[0059] The combined datasets GSE9438 and GSE41258 were randomly divided into a CRC training cohort and a validation cohort at a 1:1 ratio. A predictive model was constructed in the training cohort using logistic regression, and its stability was verified using 10-fold cross-validation. The combined diagnostic score was obtained by multiplying the cluster molecular expression value by the regression coefficient, expressed as: Cd score = ∑(cluster molecular expression value * regression coefficient) + B, where the cluster molecular expression value is the protein expression value; Cd: combined diagnosis; and B is the logistic regression constant, automatically generated during regression analysis. The AUC of the predictive model was calculated using the training cohort, and the accuracy was evaluated using the validation cohort. The results showed that the AUC of the training cohort was 0.97, the model prediction AUC of the validation cohort was 0.93, and the model prediction AUC of the combined dataset was 0.95, achieving an accuracy of 90.2%. In the control group, CEA serological testing was used, and the model prediction AUC was 0.76, with an accuracy of 71.5%. These results indicate that the predictive model can more accurately distinguish between cancer and non-cancer populations (see details). Figure 7 ).

[0060] 8. Clinical CRC Sample Predictive Analysis (Model Validation)

[0061] We collected paired cancer and adjacent cancer samples from 80 CRC patients (Qingyuan People's Hospital, Guangdong Province, in accordance with the regulations of the Clinical Trial Review Committee and with informed consent from the patients) (see Table 2 for details), serum samples from 15 healthy individuals, and serum samples from 60 confirmed CRC cases (see Table 3 for details).

[0062] Table 2: Clinicopathological information of patients with colon cancer

[0063]

[0064] Table 3: Clinical and pathological information of serum samples

[0065]

[0066] The expression of seven cluster molecules in CRC tissues was verified at the mRNA and protein levels using quantitative PCR and Western blotting. The results showed that the expression levels of all molecules were upregulated in cancer tissues compared to adjacent normal tissues, demonstrating the heterogeneity and representativeness of this cluster molecule expression at both the gene and protein levels (see details). Figure 8 The expression levels of characteristic genes in serum were detected using an enzyme-linked immunosorbent assay (ELISA) kit to obtain the corresponding expression dataset (see details). Figure 9Logistic regression analysis using this prediction model shows that the cluster molecular prediction model has an AUC of 0.962 and an accuracy of 91.9%, which is significantly higher than the CEA single-factor prediction effect (AUC = 0.71, accuracy 79.7%). This indicates that the cluster molecular prediction model based on this invention can more accurately distinguish between high-risk and low-risk CRC populations.

[0067] The results of this invention show that a CRC prediction model based on the Cd score of colorectal cancer-specific expression cluster molecules can accurately predict the risk of colorectal cancer in hematological screening, providing a new method to improve the detection rate of colorectal cancer and has broad application prospects.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A method for constructing a colorectal cancer prediction model based on clustered molecules, characterized in that, Includes the following steps: (1) Collect colorectal cancer transcript sequencing data and colorectal cancer dataset; The probe values ​​and probe annotations were extracted from the colorectal cancer dataset, and batch effects were removed to obtain the merged dataset. (2) Collect genes that can encode proteins in the blood; (3) Screen the differentially expressed genes in the colorectal cancer transcript sequencing data, and verify the differentially expressed genes using the merged dataset; (4) Use the weighted gene co-expression network analysis method to screen for colorectal cancer-specific expression genes that can encode blood proteins from the genes described in step (2); (5) Based on machine learning methods, colorectal cancer-specific protein-coding genes are screened from the colorectal cancer-specific expressed genes to obtain colorectal cancer cluster molecules; the colorectal cancer cluster molecules include QSOX2, TGFBI, CD44, INHBA, S100A11, VEGFA and MET; (6) Verify the credibility of the obtained colorectal cancer cluster molecules; (7) The merged dataset is randomly divided into a CRC training queue and a validation queue at a ratio of 1:

1. A prediction model is constructed in the training queue using the logistic regression method. At the same time, the stability of the model is verified by the 10-fold cross-validation method. The joint diagnostic score is obtained by multiplying the molecular expression value of the colorectal cancer cluster by the regression coefficient. The AUC of the prediction model is calculated using the training queue, and the accuracy of the prediction model is identified using the validation queue. Thus, the colorectal cancer prediction model is obtained.

2. The construction method according to claim 1, characterized in that, In step (1), the transcript sequencing data comes from The Cancer Genome Atlas database; the colorectal cancer dataset comes from the independent colorectal cancer datasets GSE9348 and / or GSE41258 in the GENE EXPRESSION OMNIBUS database; probe values ​​and probe annotations are extracted from the colorectal cancer datasets GSE9348 and / or GSE41258, and batch effects are removed using the Sva software package to obtain the merged dataset.

3. The construction method according to claim 1, characterized in that, In step (3), the differentially expressed genes are screened using the limma software package, with the screening criteria being log2 FC>1.5 and FDR<0.

05.

4. The construction method according to claim 1, characterized in that, In step (4), the weighted gene co-expression network analysis method is used to cluster highly co-variant genes to generate corresponding modules, analyze the connectivity of genes within the associated modules and the correlation between the associated modules and the clinicopathological features of colorectal cancer, and identify the core gene in the module with the highest correlation; the core gene is MEblue and / or MEturquoise.

5. The construction method according to claim 1, characterized in that, In step (5), the screening is performed using the Lasso regression model, the random forest algorithm, and the SVM-RFE algorithm.

6. The construction method according to claim 1, characterized in that, In step (6), the verification is performed by using the limma software package to verify the expression of the colorectal cancer cluster molecule in the CRC sample; the area under the ROC curve (AUC) and 95% confidence interval of the colorectal cancer cluster molecule are calculated using the pROC software.

7. A colorectal cancer prediction model based on cluster molecules constructed by the construction method according to any one of claims 1 to 6.

8. The application of the cluster molecule-based colorectal cancer prediction model of claim 7 in the preparation of reagents for colorectal cancer screening and / or prediction.