Biomarker for predicting whether individual suffers from colorectal early canceration or not and application of biomarker
High-throughput proteomics technology screens out protein markers related to early colorectal cancer and constructs a multi-marker joint detection model, which solves the problem of lack of high-sensitivity biomarkers for early diagnosis of colorectal cancer in the prior art, and achieves accurate prediction of early colorectal cancer.
Patent Information
- Application Number
- CN202510135653.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2025-06-20
Smart Images

Figure CN120177784A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent with the application number: 202411153096.5 and the invention title "Biomarkers for Early Colorectal Cancer Detection and Their Applications". Technical Field
[0002] The present invention relates to the field of early diagnosis of colorectal cancer. Specifically, it relates to a biomarker for predicting whether an individual has early colorectal cancer and its applications. Background Art
[0003] Colorectal cancer is one of the most common malignant tumors clinically. According to the statistical data of the National Cancer Center, the incidence and mortality rates of colorectal cancer are both among the top 5 of malignant tumors. Affected by various factors such as population aging and dietary structure changes, the incidence of colorectal cancer is also increasing year by year.
[0004] Detecting and identifying colorectal malignant lesions at an early stage is an important measure to reduce the incidence and mortality rates of colorectal cancer and improve the survival of colorectal cancer. For patients diagnosed at the adenoma and stage I colorectal cancer stages, the 5-year survival rate after radical surgery is over 90%; however, the 5-year survival rate of patients with distant metastases at the time of detection is only about 5%. Currently, colonoscopy is the gold standard for detecting colorectal tumors. However, colonoscopy requires advanced instrument equipment and specialized operators, with high technical requirements and high costs. Moreover, the subjects need to perform bowel preparation and have poor compliance, and it is not suitable for repeated examinations and population screening. In addition, there are some other methods for screening colorectal cancer, such as fecal occult blood test (FOBT), but this method has limitations such as being easily interfered by diet, having a high false positive rate, and low sensitivity. The currently commonly used blood markers for colorectal cancer (CEA, CA199) have unsatisfactory diagnostic efficacy for early colorectal cancer and precancerous lesions (advanced adenomas), with insufficient sensitivity and high false positive rates. Although there are also some molecular diagnostic technologies based on blood and feces (such as DNA methylation detection) that have improved the detection sensitivity of early colorectal cancer to a certain extent, the diagnostic sensitivity for advanced adenomas is only about 60%. There is a lack of biomarkers for early diagnosis of colorectal cancer clinically, especially the discovery of high-sensitivity biomarkers for the diagnosis of advanced adenomas is of great significance.
[0005] Proteomics is the science that studies the protein composition, localization, changes and their interaction rules in cells, tissues or organisms, including the study of protein expression patterns and proteome functional patterns. With the development of proteomics technology, high performance liquid chromatography-high resolution tandem mass spectrometry has gradually become the mainstream technology in proteomics, and more and more new tumor markers have been discovered. Although there have been many articles and patents reporting on the discovery of new tumor markers in recent years, they have only remained at the laboratory research stage and are rarely used in clinical applications and market promotion. Moreover, in most cases, a single indicator is far from enough for the in vitro diagnosis of tumors. Only by adopting a combined joint detection form and combining various dimensions of detection can the accuracy of prediction be enhanced. Therefore, finding new biomarkers related to the early diagnosis of colorectal cancer and its precancerous lesions and constructing an early prediction model by combining multiple biomarkers have important clinical value.
[0006] Therefore, the present invention screens relevant protein biomarkers for the diagnosis of early colorectal cancer through high-throughput proteomics technology, and constructs a prediction model for early colorectal cancer by combining multiple biomarkers, which will be of great significance for the early diagnosis and treatment of colorectal cancer. Summary of the Invention
[0007] Aiming at the problems existing in the prior art, the present invention provides a biomarker for predicting whether an individual has early colorectal cancer and its application. By using the method of proteomics, by analyzing the proteins with significantly different abundance levels in the blood of patients with advanced adenomas, early colorectal cancer patients and healthy control groups during the tumor progression of colorectal cancer, biomarkers for predicting the risk of whether an individual has early colorectal cancer (including advanced adenomas and early colorectal cancer) are screened out, and a multi-biomarker combined detection model is further constructed, which can accurately, non-invasively and efficiently predict the risk of an individual having early colorectal cancer and meet the clinical needs.
[0008] On the one hand, the present invention provides the use of a biomarker for preparing a reagent for predicting whether an individual has early colorectal cancer, and the biomarker includes any one or more of trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, growth differentiation factor 15, and the early colorectal cancer includes advanced adenomas and early colorectal cancer.
[0009] In the present invention, the definition of early colorectal cancer is defined as including colorectal cancer + advanced adenomas. That is to say, the biomarkers provided by the present invention can distinguish between people with advanced adenomas and early colorectal cancer and healthy people at the same time.
[0010] Although there are reports on existing markers for differentiating early colorectal cancer, they are generally only used to distinguish between healthy individuals and those with early colorectal cancer, and cannot distinguish patients with advanced adenomas at an earlier stage of lesion progression, resulting in insufficient detection sensitivity.
[0011] The marker provided by the present invention can simultaneously diagnose individuals with advanced adenomas and early colorectal cancer, and distinguish them from healthy individuals, thereby enabling earlier identification of colorectal malignant lesions, which is of great significance for improving the survival rate of colorectal cancer.
[0012] Further, the marker includes trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, and growth differentiation factor 15.
[0013] In this example, using proteomic methods, plasma samples of patients at different stages in the colorectal cancer tumor progression (inflammatory disease - benign polyp - advanced adenoma - early colorectal cancer) and healthy control groups were collected. Different samples were analyzed by high performance liquid chromatography - tandem mass spectrometry (HPLC-MS / MS). Based on orthogonal partial least squares discriminant analysis and significance analysis methods, proteins with significant differences between early colorectal cancer patients and healthy control groups were first screened. Finally, 12 differentially expressed proteins with obvious relevance to colorectal cancer were screened. However, when using these 12 proteins to construct a random forest model for differentiating early colorectal carcinogenesis (including early colorectal cancer and advanced adenoma) and healthy individuals, the diagnostic efficiency was low, and it could not be used to distinguish advanced adenomas and healthy individuals either.
[0014] Therefore, to improve the diagnostic efficiency for early colorectal carcinogenesis (including early colorectal cancer and advanced adenoma) or advanced adenoma, differential proteins were screened again for the advanced adenoma and healthy groups. Finally, the top 10 differentially expressed proteins in terms of importance were screened, and further 7 protein markers were screened using the Boruta algorithm to construct a model, which has good risk prediction ability in the groups of early colorectal cancer, advanced adenoma, and early colorectal carcinogenesis (including early colorectal cancer and advanced adenoma).
[0015] Further, the reagent is used to detect the content of biomarkers in body fluid samples; the body fluid samples include any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.
[0016] In some embodiments, the reagent for predicting colorectal early carcinogenesis is a detection reagent prepared with this biomarker as the detection target, such as sample pretreatment reagents, antigens or antibodies and other biological reagents and kits suitable for the detection of the biomarker; it can also be developed into a standardized reagent or kit suitable for the liquid chromatography-ultraviolet (LC-UV) or liquid chromatography-mass spectrometry (LC-MS) detection of the biomarker, etc.
[0017] In some embodiments, by using the ELISA method, a large-scale verification was carried out on 18 colorectal cancer candidate protein biomarkers, including 12 candidate protein biomarkers screened from early colorectal cancer and 10 candidate protein biomarkers screened from advanced adenomas. Finally, it was found that in the two groups of early colorectal cancer vs. healthy control and advanced adenoma vs. healthy control, 7 novel biomarkers with significant differences were found simultaneously (TFF1, TFF3, IGFBP1, IGFBP4, SERPINA1, OPN, GDF-15), which can be used as candidate biomarkers for the differential diagnosis of colorectal early carcinogenesis (including early colorectal cancer and advanced adenomas) and advanced adenomas; at the same time, the model constructed by the 7 protein biomarkers has good risk prediction ability in the groups of early colorectal cancer, advanced adenoma, and early colorectal cancer + advanced adenoma (colorectal early carcinogenesis). Among them, in the group of colorectal early carcinogenesis vs. healthy control, the AUC value of the model reached 0.896; in the group of early colorectal cancer vs. healthy control, the AUC value of the model reached 0.983; in the group of advanced adenoma vs. healthy control, the AUC value of the model reached 0.807, and the AUC values all reached above 0.8, having a high diagnostic value; and the 7 finally screened biomarkers that contribute significantly to the model are all differential biomarkers included in advanced adenomas, and when screening only from the protein biomarkers of early colorectal cancer, the best 7 biomarkers that contribute significantly to the model of the present invention cannot be screened.
[0018] Therefore, 7 biomarkers are preferably selected in the present invention: trefoil factor 1 (TFF1), trefoil factor 3 (TFF3), insulin-like growth factor binding protein 1 (IGFBP1), insulin-like growth factor binding protein 4 (IGFBP4), serine protease inhibitor A1 (SERPINA1), osteopontin (OPN), growth differentiation factor 15 (GDF-15) to construct a diagnostic model, which has good clinical diagnostic value for colorectal early carcinogenesis, can significantly improve the diagnostic discrimination ability and diagnostic efficacy for colorectal early carcinogenesis (including early colorectal cancer + advanced adenoma) or advanced adenoma, and realize the risk prediction for colorectal early carcinogenesis (including early colorectal cancer + advanced adenoma) or advanced adenoma.
[0019] Further, the reagent is used to detect the presence, relative abundance or concentration of biomarkers in a body fluid sample.
[0020] The present invention screens biomarkers for early colorectal cancer or advanced adenoma from blood. These biomarkers have significant differences in the blood of individuals with early colorectal cancer (including early colorectal cancer + advanced adenoma) and those without colorectal cancer (including but not limited to benign polyps, patients with gastrointestinal inflammatory diseases, and healthy controls). By collecting a blood sample, it is possible to predict or assist in diagnosing the likelihood of an individual having early colorectal cancer or advanced adenoma by detecting these biomarkers in the individual's blood. Alternatively, these biomarkers in the blood of a group can be detected, and then the group can be divided into high-risk or low-risk groups for colorectal cancer.
[0021] Further, the detection method includes radioassay, immunoassay, fluorescence assay, flow cytometry fluorescence assay, latex turbidimetry, biochemical assay, enzymatic assay, hybridization assay, gas chromatography-mass spectrometry, liquid chromatography-mass spectrometry, chromatography, chemiluminescence assay, magnetoelectric assay, or photoelectric conversion assay.
[0022] The presence or absence or the level of the biomarker here is a relative concept. For example, when comparing the diseased group with the non-diseased group, the level of these specific biomarkers is compared relative to the diseased group or the non-diseased group as a reference. There may be some biomarkers whose level in the diseased group is relatively higher than that in the non-diseased group, and this increase is statistically significant, such as a significant or highly significant increase. Therefore, when judging these biomarkers, if it is a single biomarker, if the probability of a certain risk occurrence increases and the level of the biomarker changes, this change may be a relative increase or a relative decrease. The difference in this relative increase or relative decrease is significantly different, and of course, it can also be highly significantly different. Therefore, no matter what means are used for detection, a pre-specified value (cut-off value) can be used as a standard. If the value is higher than this value, it is considered that the level has changed, and such a result can be used for prediction or diagnosis.
[0023] Therefore, in some aspects, the biomarker described in the present invention can be obtained by detecting the content of the biomarker in a sample by any known method, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strip, immunoreaction chip, capillary electrophoresis, infrared spectroscopy, etc. As long as it can be used to detect the content of the protein biomarker in the sample, it can be used for the diagnosis of early colorectal cancer and advanced adenomas. As long as the content of the protein biomarker in the sample can be detected, it can be used to predict or diagnose the probability of the occurrence of a certain disease. It can be understood that the detection here is for the sample of an individual, and then compared with a pre-set standard, and the result of the comparison is used to judge or predict the occurrence status of the disease. For example, it can be used to predict the probability of early colorectal cancer or advanced adenoma. Such prediction or diagnosis is whether it occurs within a certain time. Of course, such detection can be continuous detection, and the progress of the disease can be inferred from the change of the content of certain substances.
[0024] In some ways, the relative abundance is the peak area of the biomarker in the detection spectrum obtained by high performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a certain biomarker measured in a control sample is 500, and the average peak area measured in a sample of a patient with early colorectal cancer or advanced adenoma is 3000, then the abundance of the biomarker in the sample is considered to be 6 times that of the control sample.
[0025] On the other hand, the present invention provides a combination of biomarkers for predicting whether an individual is in the early stage of colorectal cancer, and the combination includes trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, and growth differentiation factor 15.
[0026] On another aspect, the present invention provides a kit for predicting whether an individual has early colorectal cancer, including a detection reagent for the biomarker used in any of the above technical solutions.
[0027] In some ways, the detection reagent is an antibody against the biomarker, and the antibody is a monoclonal antibody.
[0028] In another aspect, the present invention provides a system for predicting whether an individual has early colorectal cancer. The system includes a data analysis module, which is used to analyze the detection values of biomarkers. The biomarkers include any one or more of trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, and growth differentiation factor 15. The early colorectal cancer includes advanced adenoma and early colorectal cancer.
[0029] Further, the biomarkers include trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, and growth differentiation factor 15.
[0030] In some embodiments, first, a random forest is used to construct a model. It is preliminarily confirmed that the change in the concentration of any one of the selected 7 novel biomarkers alone can be used to distinguish patients with early colorectal cancer (early colorectal cancer + advanced adenoma) from healthy people, patients with advanced adenoma from healthy people, and patients with early colorectal cancer from healthy people, indicating that these 7 biomarkers have extremely high diagnostic value.
[0031] In summary, the combined diagnostic model constructed in this embodiment, which includes 7 biomarkers (trefoil factor 1 (TFF1), trefoil factor 3 (TFF3), insulin-like growth factor binding protein 1 (IGFBP1), insulin-like growth factor binding protein 4 (IGFBP4), serine protease inhibitor A1 (SERPINA1), osteopontin (OPN), and growth differentiation factor 15 (GDF-15)), has good clinical diagnostic value for early colorectal cancer, can significantly improve the diagnostic discrimination ability and diagnostic efficiency for early colorectal cancer, early colorectal cancer or advanced adenoma, and achieve accurate diagnosis of early colorectal cancer, early colorectal cancer or advanced adenoma.
[0032] Further, the data analysis module uses the detection values of the biomarkers of known samples as a training set. According to whether they have early colorectal cancer, they are divided into an early colorectal cancer group and a healthy group, and the relationship between the detection values of the early colorectal cancer group and the healthy group is analyzed to construct a model.
[0033] Further, the system also includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of the biomarkers; the data input interface is used to input the detection values of the biomarkers, and the data output interface is used to output the prediction results.
[0034] Further, the detection value is the presence or absence, relative abundance, or concentration value of 7 biomarkers.
[0035] In another aspect, the present invention provides the use of a biomarker for preparing a reagent for predicting whether an individual has advanced adenoma, and the biomarker includes any one or more of trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, growth differentiation factor 15, prion protein, guanylate cyclase activator 2A, and regenerating family member protein 1α.
[0036] Further, the biomarker includes trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, and growth differentiation factor 15.
[0037] Further, the reagent is used to detect the content of the biomarker in a body fluid sample; the body fluid sample includes any one or more of saliva, blood, urine, plasma, serum, and spinal fluid.
[0038] Further, the reagent is used to detect the presence, relative abundance, or concentration of the biomarker in a body fluid sample.
[0039] Further, the detection method includes radioassay, immunoassay, fluorescence assay, flow cytometry fluorescence assay, latex turbidimetry, biochemical assay, enzymatic assay, hybridization assay, gas chromatography-mass spectrometry, liquid chromatography-mass spectrometry, chromatography, chemiluminescence method, magnetoelectric method, or photoelectric conversion method.
[0040] In another aspect, the present invention provides a biomarker combination for predicting whether an individual has advanced adenoma, and the combination includes trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, and growth differentiation factor 15.
[0041] In another aspect, the present invention provides a kit for predicting whether an individual has advanced adenoma, including a detection reagent for the biomarker as described in any one of the above technical solutions.
[0042] In another aspect, the present invention provides a system for predicting whether an individual has advanced adenoma, and the system includes a data analysis module for analyzing the detection value of a biomarker, and the biomarker includes any one or more of trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, growth differentiation factor 15, prion protein, guanylate cyclase activator 2A, and regenerating family member protein 1α.
[0043] Further, the biomarkers include trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, and growth differentiation factor 15.
[0044] Further, the data analysis module uses the detection values of the biomarkers of known samples as a training set, divides them into an advanced adenoma group and a healthy group according to whether the subject has advanced adenoma, analyzes the relationship between the detection values of the advanced adenoma group and the healthy group, and constructs a model.
[0045] Further, the system further includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of the biomarkers; the data input interface is used to input the detection values of the biomarkers, and the data output interface is used to output the prediction results.
[0046] Further, the detection value is the presence or absence, relative abundance, or concentration value of 7 biomarkers.
[0047] Further, the trefoil factor 1 is a protein or amino acid sequence with the UniProt database number P04155; the trefoil factor 3 is a protein or amino acid sequence with the UniProt database number Q07654; the insulin-like growth factor binding protein 1 is a protein or amino acid sequence with the UniProt database number P08833; the insulin-like growth factor binding protein 4 is a protein or amino acid sequence with the UniProt database number P22692; the serine protease inhibitor A1 is a protein or amino acid sequence with the UniProt database number P01009; the osteopontin is a protein or amino acid sequence with the UniProt database number P10451; the growth differentiation factor 15 is a protein or amino acid sequence with the UniProt database number Q99988.
[0048] In another aspect, the present invention provides the use of a biomarker combination for preparing a reagent for distinguishing an individual as early colorectal cancer, advanced adenoma, benign polyp, inflammatory bowel disease, healthy population, or other cancers, and the biomarkers include trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, and growth differentiation factor 15.
[0049] The models constructed above are all binary classification models, such as for distinguishing between "healthy population and early colorectal cancer patients", or "healthy population and advanced adenoma patients", or "healthy population and early colorectal cancer patients" for two classifications, etc.
[0050] The present invention also attempts to construct a six-classification model using these 7 protein markers. The six classifications include early colorectal cancer, advanced adenoma, benign polyp, inflammatory bowel disease, healthy population, and other cancers. It is found that the model constructed by the present invention also has very good diagnostic efficacy in the differentiation of the six classifications. At the same time, multiple different algorithms for constructing the six-classification model are compared, and it is found that the six-classification detection model constructed using the gradient boosting algorithm has the highest diagnostic efficacy and can be used for the efficient differential diagnosis of early colorectal cancer, advanced adenoma, benign polyp, inflammatory bowel disease, healthy population, and other cancers.
[0051] The other cancers include other digestive tract cancers, such as esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, cholangiocarcinoma, etc. The six-classification detection model of the present invention can not only accurately distinguish early colorectal carcinogenesis from early colorectal cancer, advanced adenoma, benign polyp, inflammatory bowel disease, and healthy population, but also demonstrates significant advantages in differentiating early colorectal carcinogenesis from other digestive tract cancers (such as esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, cholangiocarcinoma, etc.). Through this model, early colorectal carcinogenesis can be clearly distinguished from other cancers, fully reflecting its high specificity. This means that the model can accurately identify the unique characteristics of early colorectal carcinogenesis, avoid confusing it with other similar diseases or other cancer types, and thus greatly reduce the possibility of misdiagnosis.
[0052] At the same time, in order to further improve the performance and accuracy of the six-classification model, a preliminary comparison and screening of the model supervised classification algorithm model are carried out. The final results show that the optimal model constructed by gradient boosting (gbm) is selected as the final prediction model for the diagnosis of advanced adenoma or early colorectal carcinogenesis. The performance evaluation score of this algorithm is the best, and the comprehensive diagnostic accuracy for the prediction of each disease type is 0.768, and the consistency is 0.713. And through the 10-fold cross-validation method for training, the optimal hyperparameters of the model are determined as follows: the learning rate is 0.1, the number of decision trees is 150, the maximum tree depth is 3, and the minimum number of samples at the terminal node is 10.
[0053] In order to maximize the benefits while ensuring the diagnostic efficacy of the model, further screening of different combinations of protein markers is carried out. Finally, 7 markers are preferably selected: trefoil factor 1 (TFF1), trefoil factor 3 (TFF3), insulin-like growth factor-binding protein 1 (IGFBP1), insulin-like growth factor-binding protein 4 (IGFBP4), serine protease inhibitor A1 (SERPINA1), osteopontin (OPN), and growth differentiation factor 15 (GDF-15) to construct the diagnostic model. At this time, the highest diagnostic efficacy can be achieved by constructing the model with the fewest markers.
[0054] Perform a more in-depth predictive analysis on the multi-classification model based on the selected gradient boosting algorithm, calculate the predicted probability values, and re-give the diagnostic indicators for each disease to more accurately determine the diagnostic performance and thresholds of the model in different disease classifications. Use the multi-classification model of the gradient boosting (gbm) algorithm in the model group for predictive analysis, and calculate the predicted probability values for 6 classifications in the prediction results (healthy control, inflammatory bowel disease, benign polyp, advanced adenoma, early colorectal cancer, other cancers). The classification with the largest predicted probability value is the final prediction result of the system. The final result is that the accuracy of the model in the model group is 0.761 and the consistency is 0.705. The diagnostic sensitivity for early colorectal cancer is 76.9%, the specificity is 95.9%, the positive predictive value is 78.4%, and the negative predictive value is 95.5%; the diagnostic sensitivity for advanced adenoma is 69.8%, the specificity is 94.5%, the positive predictive value is 71.3%, and the negative predictive value is 94.1%; the diagnostic sensitivity for benign polyp is 74.3%, the specificity is 95.4%, the positive predictive value is 67.7%, and the negative predictive value is 96.6%; the diagnostic sensitivity for inflammatory bowel disease is 67.2%, the specificity is 94.2%, the positive predictive value is 67.7%, and the negative predictive value is 94.1%; the diagnostic sensitivity for other cancers is 77.7%, the specificity is 95.9%, the positive predictive value is 67.3%, and the negative predictive value is 97.5%; the diagnostic sensitivity for healthy control is 83.6%, the specificity is 95.4%, the positive predictive value is 89.0%, and the negative predictive value is 92.9%. Its effect is significantly better than that of the existing biomarker combined prediction model.
[0055] At the same time, to test the performance of the above gradient boosting algorithm in the new sample validation set data and thus more comprehensively evaluate the generalization ability and practical application effect of the model, apply the algorithm constructed based on the model group to the new validation group for verification of the prediction performance. In the new validation group samples, the final results show that: the accuracy is 0.78 and the consistency is 0.729. The diagnostic sensitivity for early colorectal cancer is 79.4%, the specificity is 94.4%, the positive predictive value is 73.5%, and the negative predictive value is 95.9%; the diagnostic sensitivity for advanced adenoma is 73.4%, the specificity is 94.7%, the positive predictive value is 73.4%, and the negative predictive value is 94.7%; the diagnostic sensitivity for benign polyp is 72.7%, the specificity is 95.9%, the positive predictive value is 69.6%, and the negative predictive value is 96.5%; the diagnostic sensitivity for inflammatory bowel disease is 68.3%, the specificity is 96.3%, the positive predictive value is 77.4%, and the negative predictive value is 94.3%; the diagnostic sensitivity for other cancers is 83.3%, the specificity is 96.0%, the positive predictive value is 68.2%, and the negative predictive value is 98.3%; the diagnostic sensitivity for healthy control is 85.0%, the specificity is 96.3%, the positive predictive value is 91.1%, and the negative predictive value is 93.5%.
[0056] In another aspect, the present invention provides a biomarker combination for predicting whether an individual has early colorectal cancer, advanced adenoma, benign polyp, inflammatory bowel disease, healthy population or other cancers, and the combination includes trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin and growth differentiation factor 15.
[0057] In another aspect, the present invention provides a kit for predicting whether an individual has early colorectal cancer, advanced adenoma, benign polyp, inflammatory bowel disease, healthy population or other cancers, and the kit includes a detection reagent for the biomarker used in any one of the above technical solutions.
[0058] In another aspect, the present invention provides a system for predicting whether an individual has early colorectal cancer, advanced adenoma, benign polyp, inflammatory bowel disease, healthy population or other cancers, and the system includes a data analysis module for analyzing the detection values of the biomarkers, and the biomarkers include any one or more of trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, and growth differentiation factor 15.
[0059] Further, the biomarkers include trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin and growth differentiation factor 15.
[0060] Further, the data analysis module uses the detection values of the biomarkers of known samples as a training set, and according to different disease classifications, is divided into an early colorectal cancer group, an advanced adenoma group, a benign polyp group, an inflammatory bowel disease group, a healthy group and an other cancer group, analyzes the relationship between the detection values of the early colorectal cancer group, the advanced adenoma group, the benign polyp group, the inflammatory bowel disease group, the healthy group and the other cancer group, and constructs a model.
[0061] Further, the system further includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values of the biomarkers; the data input interface is used to input the detection values of the biomarkers, and the data output interface is used to output the prediction result.
[0062] The beneficial effects of the present invention are:
[0063] 1. Seven brand-new biomarkers that can predict the risk of early colorectal cancer (early colorectal cancer + advanced adenoma) and advanced adenoma were screened out in the present invention: trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, osteopontin, and growth differentiation factor 15. A new combination of protein markers was developed, which can effectively evaluate and diagnose patients with advanced adenoma of the colorectum and colorectal cancer, effectively distinguish patients with early colorectal cancer from healthy people, patients with early colorectal cancer from healthy people, and patients with advanced adenoma from healthy people, can more accurately identify patients with early colorectal cancer and patients with advanced adenoma, and reduces the risk of misdiagnosis and missed diagnosis compared with traditional detection methods, providing strong support for the early detection and intervention of the disease.
[0064] 2. In the present invention, a binary classification algorithm was adopted, and a model was constructed by the random forest algorithm, which showed significant diagnostic value in the comparison between early colorectal cancer and healthy controls, early colorectal cancer and healthy controls, and advanced adenoma and healthy controls; in the group of early colorectal cancer vs healthy controls, the AUC value of the model reached 0.896; in the group of early colorectal cancer vs healthy controls, the AUC value of the model reached 0.983; in the group of advanced adenoma vs healthy controls, the AUC value of the model reached 0.807, all exceeding 0.8. Compared with the performance of the model constructed by the traditional early colorectal cancer marker combination (ROC was 0.699 in the group of advanced adenoma vs healthy controls), it was improved by 10%, indicating that the model has high diagnostic efficiency, which makes the screening of the disease more efficient and accurate, can timely detect potential patients, and provides a valuable opportunity for early treatment.
[0065] 3. In the present invention, a six-classification algorithm is simultaneously adopted, and a combined differential diagnosis model of 7 biomarkers is constructed through the gradient boosting algorithm. It is found that the diagnostic efficacy of the colorectal cancer diagnosis model constructed with 7 biomarkers including trefoil factor 1, trefoil factor 3, insulin-like growth factor-binding protein 1, insulin-like growth factor-binding protein 4, serine protease inhibitor A1, osteopontin, and growth differentiation factor 15 is optimal, and it can be used to more efficiently predict early colorectal carcinogenesis. The accuracy of the model is 0.78, and the consistency is 0.729. The diagnostic sensitivity for early colorectal cancer is 79.4%, the specificity is 94.4%, the positive predictive value is 73.5%, and the negative predictive value is 95.9%; the diagnostic sensitivity for advanced adenomas is 73.4%, the specificity is 94.7%, the positive predictive value is 73.4%, and the negative predictive value is 94.7%. Its effect is significantly better than the existing diagnostic models, significantly improving the diagnostic accuracy, sensitivity, and specificity, and can be used for efficient differential diagnosis of early colorectal cancer, advanced adenomas, benign polyps, inflammatory diseases, healthy individuals, and other cancer populations, so as to intervene in patients in a timely manner; the multi-classification model can provide more comprehensive and accurate diagnostic basis for clinicians, help formulate personalized treatment plans, and improve the scientificity and effectiveness of medical decisions.
[0066] 4. The combined differential diagnosis model of 7 biomarkers constructed by the present invention is convenient, fast, and the test results are highly consistent with the clinical gold standard test results. At the same time, it significantly reduces the cost of diagnosing early colorectal carcinogenesis and has good application prospects.
[0067] 5. The combined differential diagnosis model of 7 biomarkers constructed by the present invention can accurately identify patients with early colorectal carcinogenesis, which is conducive to early detection and early intervention, promoting the early detection and early treatment of early colorectal carcinogenesis, and meeting the urgent needs of the clinic.
[0068] Detailed description
[0069] (1) Diagnosis or detection
[0070] The diagnosis or detection here refers to the detection or assay of biomarkers in a sample, or the content of the target biomarker, such as the absolute content or relative content, and then it is determined whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease, based on the presence or quantity of the target biomarker. The meanings of diagnosis and detection here can be interchanged. The result of this detection or the result of the diagnosis cannot be directly used as the direct result of having a disease, but is an intermediate result. If a direct result is to be obtained, other auxiliary means such as pathology or anatomy are still required to confirm the presence of a certain disease. For example, the present invention provides a variety of new biomarkers that are associated with early colorectal carcinogenesis, and the changes in the content of these biomarkers are directly related to whether colorectal cancer is present.
[0071] (2) The association of markers or biomarkers or differential proteins with early colorectal carcinogenesis or advanced adenomas
[0072] Markers, biomarkers, and differential proteins have the same meaning in the present invention. The association here refers to the direct relevance of the presence or content change of a certain biomarker in a sample to a specific disease. For example, a relative increase or decrease in content indicates that the possibility of having this disease is relatively higher compared to the healthy population.
[0073] If multiple different markers in a sample appear simultaneously or there are relative changes in their contents, it also indicates that the possibility of having this disease is relatively higher compared to the healthy population. That is to say, among the types of markers, some markers have a strong association with the disease, some markers have a weak association with the disease, or some are even not associated with a specific disease. One or more of those markers with a strong association can be used as markers for diagnosing the disease, and those markers with a weak association can be combined with the strong markers to diagnose a certain disease, increasing the accuracy of the test result.
[0074] Regarding the numerous biomarkers in serum discovered in the present invention, these biomarkers can all be used to distinguish colorectal cancer patients from healthy individuals or those with benign diseases. These markers can be directly detected or diagnosed as individual markers alone. Selecting such a marker indicates that the relative change in the content of this marker has a strong association with early colorectal carcinogenesis. Of course, it can be understood that one or more markers with a strong association with early colorectal carcinogenesis can be selected for simultaneous detection. Normally understood, in some ways, selecting biomarkers with a strong association for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90%, or 95% accuracy. This indicates that these markers can obtain an intermediate value for diagnosing a certain disease, but it does not mean that it can directly confirm the presence of a certain disease.
[0075] Of course, it is also possible to select differential proteins with larger ROC values as diagnostic markers. The so-called strength is generally calculated and confirmed through some algorithms, such as the contribution rate or weight analysis of the marker to colorectal cancer. Such calculation methods can include significance analysis (p-value or FDR value) and fold change. Multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA), and orthogonal partial least squares discriminant analysis (OPLS-DA). Of course, other methods are also included, such as ROC analysis, etc. Of course, other model prediction methods are also feasible. When specifically selecting biomarkers, differential proteins disclosed in the present invention can be selected, or other existing well-known biomarker combinations can be selected or combined for prediction through model methods.
[0076] (3) Definition of disease terms
[0077] Colorectal cancer: Also known as large intestine cancer, it refers to cancer originating from the large intestine epithelium, including colon cancer and rectal cancer. The most common pathological type is adenocarcinoma, and squamous cell carcinoma is extremely rare. In China, rectal cancer is the most common, followed by colon cancer (sigmoid colon, cecum, ascending colon, descending colon, and transverse colon). The treatment of colorectal cancer should follow the principle of individualized treatment. According to the patient's age, physical condition, pathological type of the tumor, and invasion range (stage), appropriate treatment methods should be selected, including radical surgical treatment, chemotherapy, targeted therapy, radiotherapy, etc. The formation of colorectal cancer generally goes through the development process of normal mucosal hyperplasia, advanced adenoma (malignant), and adenocarcinoma (malignant), which generally takes 5 - 10 years. Therefore, early screening, diagnosis, and treatment are the most effective means to reduce the mortality rate of colorectal cancer. In particular, if intervention can be carried out at the polyp adenoma (malignant) stage, the occurrence of colorectal cancer can be effectively prevented. Advanced adenoma (AA) is a precancerous lesion, which refers to a size ≥ 1 cm or having any size of villous component ≥ 25% or high-grade dysplasia. Over time, it is very likely to develop into colorectal cancer. If tumor biomarkers with a certain warning effect can be found at the early stage of the occurrence of colorectal cancer for the diagnosis of colorectal cancer and advanced adenoma (malignant tumor), it is of great significance for improving the treatment effect of patients and improving the prognosis of patients.
[0078] Early colorectal carcinogenesis: In the present invention, early colorectal carcinogenesis includes early colorectal cancer and advanced adenoma.
[0079] Early colorectal cancer: Early colorectal cancer refers to any-sized colorectal epithelial tumor with the invasion depth limited to the mucosa and submucosa, regardless of whether there is lymph node metastasis.
[0080] Adenoma: It is a benign tumor originating from the glandular epithelium of the colorectal mucosa, including colon adenoma and rectal adenoma, and is a common benign intestinal tumor. According to the structural characteristics of adenoma, it can be divided into tubular adenoma, villous adenoma, and tubulovillous adenoma. It can be removed by methods such as endoscopic high-frequency electrocoagulation, laser, microwave coagulation, etc., or surgical resection can be selected, and regular follow-up is required. Those with malignant transformation are treated with other treatments according to the situation (such as radiotherapy, chemotherapy, surgery, etc.).
[0081] Advanced adenoma: It refers to tubular villous adenoma, villous adenoma with a diameter > 1 cm and / or adenoma with high-grade dysplasia. Because it is closely related to the occurrence of colorectal cancer, it is considered a precancerous lesion. Advanced adenoma should be removed in time to prevent it from evolving into colorectal cancer.
[0082] Polyp: A mass protruding from the surface of the colorectum, which can be an adenoma or hyperplasia and hypertrophy of the intestinal mucosa. It is collectively referred to as a polyp before the pathological nature is determined. Benign polyp refers to an adenoma or mass diagnosed as benign by pathological diagnosis. Whether a polyp needs to be surgically removed mainly depends on the size, shape, and nature of the polyp. 1. If the diameter of the polyp is less than 2 cm, generally no surgical resection is required, and regular colonoscopy reexamination is sufficient; 2. If the diameter of the intestinal polyp is greater than 2 cm, surgical resection is recommended; 3. If the intestinal polyp has a tendency to malignant transformation, regardless of the diameter of the polyp, surgical resection is recommended.
[0083] Inflammatory bowel disease: It is a non-specific chronic intestinal inflammatory disease, mainly including ulcerative colitis and Crohn's disease. The former mainly damages the colon and rectum, and the latter can damage any part of the gastrointestinal tract from the mouth to the anus, with the terminal small intestine and colon being more common. The treatment mainly includes drug treatment, endoscopic treatment, and surgical treatment. The treatment goal is to promote the healing of the intestinal mucosa, relieve clinical symptoms, prevent related complications, improve the quality of life of patients, and prevent the recurrence of the disease.
[0084] Therefore, the prediction model provided by the present invention can quickly distinguish patients with early colorectal cancer, advanced adenoma, benign polyp, inflammatory bowel disease, and healthy people from the changes in the concentration of biomarkers in body fluid samples.
[0085] (4) Gold standard for the diagnosis of early colorectal cancer: Clinically, most early colorectal cancer patients have no symptoms and signs, and the diagnosis depends on the standardized colonoscopy examination by qualified physicians, and the histopathology of biopsy tissue is the basis for diagnosis. Description of the Drawings
[0086] Figure 1 It is a volcano plot of differential proteins between early colorectal cancer and healthy controls;
[0087] Figure 2 It is a result chart of the performance analysis of candidate markers for early colorectal cancer;
[0088] Figure 3 Random forest model analysis chart between early colorectal cancer + advanced adenoma and healthy controls constructed for 12 markers;
[0089] Figure 4 Random forest model analysis chart between early colorectal cancer and healthy controls constructed for 12 markers;
[0090] Figure 5 Random forest model analysis chart between advanced adenoma and healthy controls constructed for 12 markers;
[0091] Figure 6 Volcano plot of differential proteins between advanced adenoma and healthy controls;
[0092] Figure 7 Chart of the results of the performance analysis of candidate markers for advanced adenoma;
[0093] Figure 8 Chart of the results of the importance analysis of 18 candidate markers;
[0094] Figure 9 Random forest model performance analysis chart between early colorectal cancer + advanced adenoma and healthy controls constructed for 7 markers;
[0095] Figure 10 Random forest model performance analysis chart between early colorectal cancer and healthy controls constructed for 7 markers;
[0096] Figure 11 Random forest model performance analysis chart between advanced adenoma and healthy controls constructed for 7 markers;
[0097] Figure 12 Random forest model performance analysis chart between early colorectal cancer + advanced adenoma and healthy controls constructed for 7 markers in the validation group;
[0098] Figure 13 Random forest model performance analysis chart between early colorectal cancer and healthy controls constructed for 7 markers in the validation group;
[0099] Figure 14 Random forest model performance analysis chart between advanced adenoma and healthy controls constructed for 7 markers in the validation group;
[0100] Figure 15 Chart of the results of the optimal model performance evaluation (accuracy and consistency) constructed for different marker combinations;
[0101] Figure 16Graph showing the performance evaluation results of the optimal combined diagnostic model constructed in Example 2 in the test group;
[0102] Figure 17 Graph showing the performance evaluation results of the optimal combined diagnostic model constructed in Example 2 in the validation group;
[0103] Figure 18 Graph showing the performance evaluation (accuracy and consistency) results of the optimal models constructed based on 6 different algorithms in Example 3. Detailed implementation manners
[0104] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not impose any limitations on it. The reagents used in this embodiment are all known products and are obtained by purchasing commercially available products.
[0105] Example 1 Screening of biomarkers for early colorectal cancer and advanced adenomas and construction of a model using proteomics
[0106] In this example, using the method of proteomics, plasma samples of patients at different stages in the colorectal cancer tumor progression (inflammatory diseases - benign polyps - advanced adenomas - colorectal cancer) and healthy control populations were collected. Different samples were analyzed by high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS). Based on orthogonal partial least squares discriminant analysis and significance analysis methods, proteins with significant differences between early colorectal cancer and healthy controls were first screened. Finally, 12 differential proteins with obvious relevance to early colorectal cancer were obtained. However, in the random forest model constructed with these 12 proteins, the diagnostic efficacy for early colorectal cancer (early colorectal cancer and advanced adenomas) was low, and the diagnostic efficacy for advanced adenomas was significantly reduced. Therefore, to improve the diagnostic efficacy for early colorectal cancer and advanced adenomas, differential proteins were screened again for advanced adenomas and the healthy group. Finally, 10 differential proteins with the top importance rankings were obtained, and a model was constructed using 7 protein biomarkers further screened by the Boruta algorithm. It has good risk prediction ability in the groups of early colorectal cancer, advanced adenomas, and early colorectal cancer + advanced adenomas (early colorectal cancer). And a multi-marker combined detection model was further constructed according to the gradient boosting algorithm, and the diagnostic efficacy of different models was evaluated by ROC analysis. Finally, it was found that the diagnostic efficacy of the 7 biomarkers of the present invention was the highest and could be used for highly efficient differential diagnosis of early colorectal cancer, advanced adenomas, benign polyps, inflammatory diseases, healthy individuals, and other cancers.
[0107] The specific steps are as follows:
[0108] I. Collection of samples
[0109] From January 2018 to December 2020, our research group collected 150 cases of early colorectal cancer, 50 cases of advanced adenomas, 50 cases of inflammatory bowel disease, 50 cases of benign polyps and 50 healthy controls. All the enrolled patients signed the informed consent form. Among them, the patients with early colorectal cancer, advanced adenomas and benign polyps were diagnosed by colonoscopy and pathological histology. The patients with inflammatory bowel disease were diagnosed by colonoscopy, laboratory tests and combined with clinical diagnosis. The healthy controls were normal people with normal routine physical examinations. Inclusion criteria for patients with early colorectal cancer and advanced adenomas:
[0110] (a) No history of other malignancies;
[0111] (b) Have not received radiotherapy, chemotherapy or anti-tumor treatment;
[0112] (c) Patients without other malignancies or autoimmune diseases.
[0113] The healthy subjects in the control group were selected from the physical examination center; there were no intestinal lesions in colonoscopy screening, no abnormalities in tumor markers and biochemical indexes in laboratory tests, and no history of malignant tumors. After informed consent, all the collected plasma samples were stored in a plasma bank at -80 °C.
[0114] II. Sample processing and enzymatic hydrolysis
[0115] First, the plasma sample was centrifuged on a centrifuge for 15 minutes (15000 xg), and the supernatant was taken, filtered, and then 14 high-abundance proteins were removed by immunoaffinity chromatography. Then, it was concentrated on a centrifuge (4000 xg, 1 hour) using a concentrator tube with a cut-off molecular weight of 3 kDa. The concentrated solution was recovered, and buffer exchange was performed on a centrifuge (1000 xg, 2 minutes) using a desalting column with a cut-off molecular weight of 7 kDa, and the replacement solution was AEX-A (20 mM Tris, 4 M Urea, 3% isopropanol, pH 8.0). Using AEX-A as a blank, the protein concentration in the sample was measured using the BCA method (protein concentration detection method). According to the sample grouping in Table 1, TCEP (Thermo Scientific, CAT#77720) was added to the sample, and the sample was incubated at 37 °C for 30 minutes for protein reduction. Then, the corresponding 6-plex TMT reagent (Thermo Scientific, CAT#90309) was added, and the sample was incubated in the dark at room temperature for 1 hour for the TMT labeling reaction. Then, the sample was subjected to buffer exchange using a Zeba column (Thermo Scientific, CAT#89890), and the replacement solution was AEX-A. After mixing the 6-plex TMT-labeled samples, 2 mL of AEX-A was added to the mixed sample, and the final volume was 5.5 mL. The sample was filtered using a 0.22 μm filter and the 6-plex TMT-labeled sample was separated using a 2D-HPLC system. The collected fractions were freeze-dried, and finally, a mixture of Trypsin-Lysin C enzymes (Thermo Scientific, CAT#A41007) was added, and the sample was incubated at 37 °C for 5 hours for enzymatic digestion of the sample, and 5 μL of 10% TFA (trifluoroacetic acid) was added to terminate the enzymatic reaction. A total of 60 enzymatically digested 2D-HPLC fractions were used for nano-LC-MS / MS analysis.
[0116] Table 1. Sample grouping for proteomics research (40 batches, taking batch 1 as an example)
[0117] Sample Number Sample Grouping Experiment Batch TMT-6plex Control Control Batch1 126 Case 1 Case Batch1 127 Case 2 Case Batch1 128 Case 3 Case Batch1 129 Case 4 Case Batch1 130 Case 5 Case Batch1 131
[0118] III. LC-MS / MS Data Acquisition and Database Search Analysis
[0119] The LC-MS / MS system consisted of a combination of Easy-nLC 1200 (Thermo Scientific) and Q Exactive HFX (Thermo Scientific). Mobile phase A was an aqueous solution containing 0.1% formic acid and 2% acetonitrile; mobile phase B was an aqueous solution containing 0.1% formic acid and 80% acetonitrile. The self-made analytical column had a length of 20 cm and was packed with ReproSil-Pur C18, 1.9 μm particles from Dr. Maisch GmbH. 1 μg of peptide was dissolved in mobile phase A and separated using the EASY-nLC 1200 ultra-high performance liquid chromatography system. The liquid phase gradient was set as follows: 0 - 26 min, 7% - 22% B; 26 - 34 min, 22% - 32% B; 34 - 37 min, 32% - 80% B; 37 - 40 min, 80% B, and the liquid phase flow rate was maintained at 450 nL / min.
[0120] The peptides separated by the high performance liquid chromatography system were injected into the NanoFlex ion source for atomization and then into the Q Exactive HF-X for mass spectrometry analysis. The ion source voltage was set at 2.1 kV, the first-order mass spectrometry scanning range was set at 400 - 1200, and the resolution was 60,000 (MS Resolution); the starting point of the second-order mass spectrometry scanning range was 100 m / z, and the resolution was set at 15,000 (MS2 Resolution). In the data-dependent scanning (DDA) mode, the TOP 20 precursor ions were sequentially introduced into the HCD collision cell for fragmentation and then subjected to second-order mass spectrometry analysis in sequence. The automatic gain control (AGC) was set at 5E4, the signal threshold was set at 1E4, and the maximum injection time was set at 22 ms. To avoid repeated scanning of high-abundance peptides, the dynamic exclusion time for tandem mass spectrometry analysis was set at 30 seconds.
[0121] The mass spectrometry data obtained by LC-MS / MS was searched using Maxquant (v1.6.15.0). The data type was TMT proteomics data based on the quantification of secondary reporter ions, and the requirement for the secondary spectra used for quantification was that the proportion of precursor ions in the primary spectra was greater than 75%. The database source was Homo_sapiens_9606_proteome of the Uniprot database (release: 2021-10-14, sequence: 20614), and a common contaminant library was added to the database. Contaminant proteins were removed during data analysis; the digestion method was set to Trypsin / P; the number of missed cleavage sites was set to 2; the precursor ion mass error tolerances for the First search and Main search were set to 20 ppm and 5 ppm respectively, and the mass error tolerance for the secondary fragment ions was 20 ppm. The fixed modification was cysteine alkylation, and the variable modifications were methionine oxidation and protein N-terminal acetylation. The FDRs for protein identification and PSM identification were both set to 1%.
[0122] IV. Screening of the differential proteins with the highest diagnostic efficacy in early colorectal carcinogenesis
[0123] 1. Screening of differential protein markers in early colorectal cancer
[0124] The differential proteins between early colorectal cancer and the healthy group were screened by combining univariate analysis and multivariate statistical analysis. Univariate analysis mainly included the significance analysis (p-value or FDR value) and fold change of characteristic ions in different groups, and multivariate statistical analysis mainly included principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA), and orthogonal partial least squares discriminant analysis (OPLS-DA). Unsupervised principal component analysis could analyze the separation trend of proteins among groups; supervised orthogonal partial least squares discriminant analysis could analyze the degree of protein differences between groups.
[0125] A total of 3051 proteins were identified and 1631 proteins were quantified, including some newly discovered markers related to early colorectal cancer. For the 1631 protein substances found, protein substances with significant content differences were obtained through analysis. All statistical analyses were completed using R, and the specific R-related information is shown in Table 2.
[0126] Table 2. R and its related information used in the present invention
[0127] Name Version R 3.4.1 Rstudio 1.4.1717 MixOmics 6.10.9 Ropls 1.18.1
[0128] Calculate the Variable Importance for the Projection (VIP) to measure the influence intensity and interpretability of the expression patterns of each protein on the classification and discrimination of each group of samples. Further, perform the Wilcoxon rank-sum test to obtain the corrected p-value (FDR). The results of the volcano plot of differential proteins between early colorectal cancer and healthy controls are as follows Figure 1 shown: In early colorectal cancer vs healthy controls, 57 proteins were significantly upregulated in the serum of early colorectal cancer patients, and 62 proteins were significantly downregulated. The results of the performance analysis of early colorectal cancer candidate markers are shown in Figure 2 , where the abscissa is the AUC obtained from ROC analysis, the ordinate is the VIP value obtained from OPLS-DA analysis, and the size of the point represents the Pvalue calculated by the Wilcoxon test.
[0129] And perform importance ranking on differential proteins through T-test difference analysis and OPLS-DA analysis. According to the importance ranking of markers, in this embodiment, the top 12 differential proteins in terms of importance in early colorectal cancer and healthy groups are listed respectively. The information of the 12 differential proteins is specifically shown in Table 3; meanwhile, establish the ROC curves of the single diagnostic performance of the 12 differential proteins respectively, and judge the quality of the experimental results by the size of the area under the curve (AUC). An AUC of 0.5 indicates that a single protein has no diagnostic value; an AUC greater than 0.5 indicates that a single protein has diagnostic value; the larger the AUC, the higher the diagnostic value of a single protein; similarly, the possible range of the AUC value - the 95% confidence interval, the closer it is to 1, the higher and more reliable the diagnostic value of the protein; at the same time, the closer the ROC sensitivity and specificity are to 100%, the higher the diagnostic efficiency of the method; the cut-off value represents a specific threshold used to distinguish positive and negative results in a diagnostic test. When the cut-off value is too high, it may lead to an increase in false negatives and missed individuals with true diseases; while when the cut-off value is too low, it may lead to an increase in false positives and misjudgment of healthy individuals as diseased. Therefore, a suitable cut-off value can more accurately distinguish patients and healthy people, thereby improving the accuracy of diagnosis; the median importance (medianImp) reflects the intermediate level of the relative importance of differential proteins in differentiating different groups or states in the screening of differential markers. The higher the median importance, the greater the contribution of the protein to the differentiation as a whole:
[0130] Table 3. Twelve differential proteins with the highest importance in early colorectal cancer vs. healthy control group
[0131]
[0132]
[0133] The degree of association between the concentration changes of 12 biomarkers and the presence of early colorectal cancer can be distinguished by the AUC value, 95% confidence interval, sensitivity, specificity, etc. in Table 3, among which the AUC value is the most intuitive and obvious. The higher the AUC value, the more accurately the biomarker can distinguish between early colorectal cancer patients and non-colorectal cancer patients.
[0134] As can be seen from Table 3, there is an obvious association between the concentration changes of 12 biomarkers and the presence of early colorectal cancer. Using any one of the 12 biomarkers alone, the concentration change can be used to distinguish early colorectal cancer from healthy controls.
[0135] Meanwhile, the ELISA (enzyme-linked immunosorbent assay) method was used to further verify the 12 candidate protein biomarkers for colorectal cancer screening, including blood samples from 64 patients with early colorectal cancer, 63 patients with advanced adenomas, and 121 healthy control patients. A random forest algorithm was used to construct models composed of 12 biomarkers respectively. The final performance of the models is as Figures 3 - 5 shown, among which, Figure 3 is the analysis diagram of the random forest model between early colorectal cancer + advanced adenoma vs healthy control constructed by 12 biomarkers, Figure 4 is the analysis diagram of the random forest model between early colorectal cancer vs healthy control constructed by 12 biomarkers, Figure 5 is the analysis diagram of the random forest model between advanced adenoma vs healthy control constructed by 12 biomarkers. It can be seen from Figures 3 - 5 that: the models constructed by the 12 protein biomarkers screened from the early colorectal cancer and healthy control groups have good risk prediction ability in the early colorectal cancer group (AUC = 0.947), the risk prediction ability decreases in the early colorectal cancer + advanced adenoma (early colorectal cancer transformation) group (AUC = 0.824), and the prediction performance decreases significantly in the advanced adenoma group (AUC = 0.699). The AUC is lower than 0.8, so it cannot effectively diagnose advanced adenomas, and its diagnostic value for early colorectal cancer transformation (including early colorectal cancer + advanced adenoma) is also relatively low.
[0136] 2. Re-screening of differential protein biomarkers in advanced adenomas
[0137] Therefore, in order to improve the diagnostic efficacy for patients with early colorectal cancer (early colorectal cancer + advanced adenoma) and patients with advanced adenoma, in this embodiment, differential protein screening was performed again for advanced adenoma and the healthy group. The TMT-labeling quantitative technology strategy based on the mass spectrometry platform was used for the discovery research of early protein markers for colorectal cancer. The research cohort included blood samples from 50 healthy controls and 50 patients with advanced adenoma. Through T-test differential analysis and OPLS-DA analysis, candidate markers were screened out. The specific results are as Figures 6 - 7 shown.
[0138] As can be seen Figure 6 from it, in the advanced adenoma vs healthy group, 16 proteins were significantly up-regulated and 27 proteins were significantly down-regulated in the serum of patients with advanced adenoma. The results of ROC and OPLS-DA analysis are shown in Figure 7 . The abscissa is the AUC obtained from the ROC analysis, the ordinate is the VIP value obtained from the OPLS-DA analysis, and the size of the point represents the P value calculated by the Wilcoxon test. At the same time, the differential proteins were ranked according to their importance. According to the ranking of marker importance, in this embodiment, the top 10 differential proteins with the highest importance in the advanced adenoma and healthy groups were listed respectively. The information of the 10 differential proteins is specifically shown in Table 4;
[0139] Table 4. The top 10 differential proteins with the highest importance in the advanced adenoma vs. healthy control group
[0140]
[0141]
[0142] The degree of association between the concentration changes of the 10 differential proteins and the presence or absence of advanced adenoma can be distinguished by the AUC value, 95% confidence interval, sensitivity, specificity, etc. in Table 4, among which the AUC value is the most intuitive and obvious. The higher the AUC value, the more accurately the differential protein can distinguish the advanced adenoma population from the healthy population.
[0143] As can be seen from Table 4, the concentration changes of the 10 differential proteins are significantly associated with the presence or absence of advanced adenoma. Using any one of the 10 differential proteins alone, its concentration change can be used to distinguish advanced adenoma from healthy controls.
[0144] In this embodiment, the ELISA method was used to conduct a large-scale verification on 18 colorectal cancer candidate protein markers, including 12 candidate protein markers screened from early colorectal cancer and 10 candidate protein markers screened from advanced adenomas. The cohort included blood samples from 327 patients with early colorectal cancer, 322 patients with advanced adenomas, and 605 healthy control patients. The Boruta algorithm was used to evaluate the importance of the 18 candidate markers, and finally 7 markers that made significant contributions to the model were screened and used to construct the final model. The specific results are as Figure 8 shown. As can be seen from Figure 8 , in this embodiment, the 7 markers that were finally screened and made significant contributions to the model are: TFF3, IGFBP4, OPN, SERPINA1, GDF-15, IGFBP1, and TFF1.
[0145] The parameters of the random forest model are specifically shown in Table 5.
[0146] Table 5. Random forest model parameters
[0147]
[0148] For the model composed of the 7 markers constructed using the random forest algorithm, the final performance of the model is as Figures 9 - 11 shown. Among them, Figure 9 is the performance analysis chart of the random forest model for early colorectal cancer + advanced adenoma vs healthy control constructed by 7 markers, Figure 10 is the performance analysis chart of the random forest model for early colorectal cancer vs healthy control constructed by 7 markers, Figure 11 is the performance analysis chart of the random forest model for advanced adenoma vs healthy control constructed by 7 markers. As can be seen from Figures 9 - 11 , when using the 7 protein markers screened to construct the model, it has good risk prediction ability in the groups of early colorectal cancer, advanced adenoma, and early colorectal cancer + advanced adenoma (early colorectal cancer transformation). Among them, in the group of early colorectal cancer transformation vs healthy control, the AUC value of the model reaches 0.896; in the group of early colorectal cancer vs healthy control, the AUC value of the model reaches 0.983; in the group of advanced adenoma vs healthy control, the AUC value of the model reaches 0.807, and the AUC values all reach above 0.8. Moreover, the diagnostic performance for early colorectal cancer transformation (including advanced adenoma and early colorectal cancer) is also significantly improved.
[0149] In summary, the 7 markers that were finally screened and made significant contributions to the model are all differential markers included in advanced adenomas. When screening only from the protein markers of early colorectal cancer, the best 7 markers that make significant contributions to the model of the present invention cannot be screened.
[0150] 3. Validation of the performance of the model constructed with 7 protein markers
[0151] In this example, the ELISA method was also used to further conduct independent cohort validation on the 7 candidate protein markers for colorectal cancer screened. The cohort included blood samples from 86 patients with early colorectal cancer, 130 patients with advanced adenomas, and 173 healthy control patients. The random forest model constructed in the training cohort was used for validation.
[0152] The specific results are as Figures 12 - 14 shown, among which, Figure 12 is the performance analysis diagram of the random forest model between early colorectal cancer + advanced adenoma and healthy control constructed with 7 markers in the validation group, Figure 13 is the performance analysis diagram of the random forest model between early colorectal cancer and healthy control constructed with 7 markers in the validation group, Figure 14 is the performance analysis diagram of the random forest model between advanced adenoma and healthy control constructed with 7 markers in the validation group. It can be seen from Figures 12 - 14 that the prediction results in the validation group are highly consistent with the actual clinical diagnosis results, and have good risk prediction ability in the groups of early colorectal cancer, advanced adenoma, and early colorectal cancer + advanced adenoma (early colorectal cancer transformation). Among them, in the group of early colorectal cancer transformation vs healthy control, the AUC value of the model reaches 0.873; in the group of early colorectal cancer vs healthy control, the AUC value of the model reaches 0.984; in the group of advanced adenoma vs healthy control, the AUC value of the model reaches 0.800, and the AUC values all reach 0.8 or above. Moreover, the diagnostic performance for early colorectal cancer transformation (including advanced adenoma and early colorectal cancer) is also significantly improved.
[0153] Therefore, it can be known that the diagnostic model constructed with 7 markers in the present invention has good prediction performance and accuracy, and has the best diagnostic efficacy.
[0154] It has been confirmed that the specific information of the 7 novel biomarkers (TFF1, TFF3, IGFBP1, IGFBP4, SERPINA1, OPN, GDF-15) that meet the standards, have significant differences and high importance is as follows: Trefoil factor 1 (TFF1) is a protein or amino acid sequence with the UniProt database number P04155; Trefoil factor 3 (TFF3) is a protein or amino acid sequence with the UniProt database number Q07654; Insulin like growth factor binding protein 1 (IGFBP1) is a protein or amino acid sequence with the UniProt database number P08833; Insulin like growth factor binding protein 4 (IGFBP4) is a protein or amino acid sequence with the UniProt database number P22692; Serpin family A member 1 (SERPINA1) is a protein or amino acid sequence with the UniProt database number P01009; Osteopontin (OPN) is a protein or amino acid sequence with the UniProt database number P10451; Growth differentiation factor 15 (GDF-15) is a protein or amino acid sequence with the UniProt database number Q99988.
[0155] Example 2 Construction and Validation of a Six-Class Diagnostic Model
[0156] Therefore, in this example, combinations of seven different biomarkers screened in Example 1 were selected to construct a six-class combined diagnostic model. These models are used to distinguish early colorectal cancer, early-stage colorectal cancer, advanced adenomas, benign polyps, inflammatory bowel disease, and healthy individuals, and specifically include the following processes: (1) Construction and screening of the optimal diagnostic model; (2) Verification of the effectiveness of the optimal diagnostic model. The specific screening process and results are as follows (in the present invention, the binary classification model is constructed using a random forest, and the AUC value is used as an evaluation index; when constructing a six-class model, since multiple categories are involved, the AUC value is usually not applicable, so in the present invention, indicators such as accuracy, consistency, sensitivity, and specificity are used to measure the diagnostic efficacy of the model):
[0157] I. Construction and Screening of the Optimal Diagnostic Model
[0158] 1. Obtain data
[0159] Study population:
[0160] From September 2022 to March 2023, 1962 cases of colorectal cancer test cohort and 390 cases of colorectal cancer validation cohort were collected. All enrolled patients signed informed consent forms. Among them, patients with early colorectal cancer, advanced adenoma, benign polyp, and other cancers were all diagnosed by colonoscopy and pathological histology. Patients with inflammatory bowel disease were diagnosed by colonoscopy, laboratory tests, and combined with clinical diagnosis. Healthy controls were those with normal routine physical examinations, negative tumor markers, and fecal occult blood. In the test group, there were 321 cases of early colorectal cancer, 321 cases of advanced adenoma, 226 cases of benign polyp, 299 cases of inflammatory bowel disease, 602 cases of healthy controls, and 193 cases of other cancers), and in the validation group (64 cases of early colorectal cancer, 64 cases of advanced adenoma, 43 cases of benign polyp, 60 cases of inflammatory bowel disease, 120 cases of healthy controls, and 39 cases of other cancers). The data information is shown in Table 6 (in this example, the other cancers include other digestive tract cancers, such as esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, cholangiocarcinoma, etc. The six-classification detection model of the present invention can not only accurately distinguish early colorectal carcinogenesis from early colorectal cancer, advanced adenoma, benign polyp, inflammatory bowel disease, and healthy people, but also shows significant advantages in differentiating early colorectal carcinogenesis from other digestive tract cancers (such as esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, cholangiocarcinoma, etc.)):
[0161] Table 6. Information of modeling samples
[0162] Grouping Test Group Validation Group Early Colorectal Cancer 321 64 Advanced Adenoma 321 64 Benign Polyp 226 43 Inflammatory Bowel Disease 299 60 Healthy Control 602 120 Other Cancers 193 39
[0163] Inclusion criteria for patients with early colorectal cancer: (a) No history of other malignancies, (b) Surgical treatment within one month after blood collection, and confirmed as colorectal cancer by postoperative pathology. Healthy individuals in the control group were selected from the physical examination center; these individuals were confirmed by endoscopic examination to have no signs of gastric diseases and no history of malignancies. After informed consent, all collected serum samples were stored in a serum bank at -80°C.
[0164] In this example, enzyme-linked immunosorbent assay was performed on the collected serum samples to obtain the concentrations of seven protein markers, namely TFF1, TFF3, IGFBP1, IGFBP4, SERPINA1, OPN, and GDF-15 respectively.
[0165] 2. Statistical analysis of experimental data
[0166] The Shapiro-Wilk test was used to evaluate the normal distribution, and the non-parametric Wilcoxon test was used to analyze the differences in blood biomarker concentrations between colorectal cancer patients and healthy controls in the model group and the test group, respectively. In the model group, a combined method of multiple machine learning methods was used to construct a combined diagnostic model for 7 colorectal cancer biomarkers. The predicted probability values were used to estimate the area under the receiver operating characteristic (ROC) curve (AUC) with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate diagnostic model. Using the test group, the Youden index (YI) was calculated to determine the predicted probability cut-off value for distinguishing colorectal cancer patients from normal controls. In addition, the ROCs of individual biomarkers and different subgroups were constructed and compared. Standard descriptive statistics, such as frequencies, means, medians, positive predictive values (PPV), negative predictive values (NPV), and standard deviations (SD), were calculated to describe the experimental results of the study population. Statistical analysis was performed using R 3.6.1, and p-values less than 0.05 were considered statistically significant.
[0167] 3. Construction steps of the six-class combined diagnostic model (7MP) (1) Preliminary comparison and screening of the supervised classification algorithm model for constructing the optimal diagnostic model
[0168] In this embodiment, in order to screen out the optimal supervised classification algorithm for constructing the prediction model, first, the concentration matrix of the best 7 protein biomarkers was used as the original training data set, and models under different supervised classification algorithms were constructed according to the following steps. By comparing the performances of the constructed different models, the optimal supervised classification algorithm was screened out. The specific process is as follows:
[0169] S101, Use the concentration matrix of the seven protein biomarkers of TFF1, TFF3, IGFBP1, IGFBP4, SERPINA1, OPN, and GDF-15 in the samples of the model group as the original training data set.
[0170] S102, Set the supervised classification algorithm for constructing the prediction model and the grid search range during the hyperparameter optimization of the algorithm. The supervised classification algorithms include: gradient boosting, naive Bayes, support vector machine, neural network, generalized linear, and discriminant analysis, 6 algorithms. In this step, the grid search range for hyperparameter optimization of the model for each algorithm is shown in Table 7 below.
[0171] Table 7. Parameter grid search ranges for 6 algorithms
[0172]
[0173] S103, According to the algorithm and the hyperparameter setting range set in step S102, select one of the algorithms and the corresponding hyperparameter combination method as the parameters for constructing the prediction model.
[0174] S104. Split the original data set into K subsets according to the K-fold cross-validation mechanism. To ensure that the ratio of majority-class samples to minority-class samples in each subset is the same as that in the original data set, the Stratified K-Folds cross-validation mechanism needs to be used for data splitting.
[0175] S105. Select one of the K training data subsets obtained by splitting in step S104 as the validation set Ddev.
[0176] S106. Combine the training data subsets not selected in step S105 to form the training data set D.train.
[0177] S107. Based on the training data set D.train obtained in step S106, construct a prediction model based on the selected supervised classification algorithm and hyperparameters.
[0178] S108. According to the prediction model obtained in step S107, evaluate it on the validation set D.dev to obtain the AUC value, and store the prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 is to evaluate the prediction model obtained in step S107 on the validation set determined in the current iteration, and store both the model and the evaluation result in the prediction model pool for future prediction model selection. The evaluation mentioned in this step can be the AUC value or other reasonable metrics for evaluating the model performance.
[0179] S109. Determine whether each subset has been used as the validation set. Step S109 is to determine whether the K subsets obtained in step S104 have all been used as the validation set for model training. If all subsets have been used as the validation set and completed training, execute step S110; if there are subsets that have not been used as the validation set, execute step S105. This step ensures that each sample in the original data set has been used as the validation set, improves the model stability, and prevents the model from overfitting to a certain subset.
[0180] S110. Take the average value of the AUCs of all the models in the obtained prediction model pool Pool as the final performance evaluation value of the model in this combination method. And store the model parameters and the final performance evaluation AUC value in the optimal model pool Pool.best.
[0181] S111. Determine whether prediction models have been constructed for all combinations of each algorithm and its corresponding hyperparameters. Step S111 is to determine whether all combinations of algorithms and their corresponding hyperparameters obtained in step S102 have been used to construct prediction models. If all combinations have completed model construction, execute step S112; if there are combinations that have not completed model construction, execute step S103.
[0182] S112. From the optimal model pool Pool.best obtained after the iteration in step S111, for each algorithm, select the prediction model with the highest AUC value and store it in the candidate prediction model set M.set for colorectal cancer diagnosis.
[0183] S113. Evaluate the model set M.set obtained in step S112 in the test group D.test to obtain the AUC value. The model with the largest AUC value is used as the final prediction model for colorectal cancer diagnosis.
[0184] By performing the above model construction steps, a total of 6 optimal models under different algorithms are finally obtained. During the modeling process, the 10-fold cross-validation method is adopted, and the performance of the models is evaluated in terms of accuracy, consistency, sensitivity, specificity, etc.
[0185] In the present invention, the test group and the validation group adopt two completely different batches of samples. The test group is composed of known samples, and the inventors only screen for markers from the test group; the samples in the validation group are only used to verify the diagnostic efficacy of the marker combination of the present invention. The specific results are shown in Table 8 and Figure 15 as follows: Among them, the performance evaluation scores of the gradient boosting (gbm) algorithm are all the best (the comprehensive diagnostic accuracy for predicting each disease type is 0.768, and the consistency is 0.713).
[0186] Table 8. Performance evaluation table of models constructed by different algorithms for differentiating different disease groups
[0187]
[0188]
[0189] Based on the above analysis results, in this embodiment, the optimal model constructed by gradient boosting (gbm) is selected as the final prediction model for six-class joint diagnosis. The optimal hyperparameters of the model obtained by training through the 10-fold cross-validation method are: learning rate is 0.1, number of trees is 150, maximum tree depth is 3, and minimum number of samples at terminal nodes is 10.
[0190] (2) Verification of the combined performance of the best joint diagnosis model (7MP)
[0191] To further analyze and study the diagnostic value of the six-classification diagnostic model (gradient boosting) constructed based on biomarkers of different protein combinations, in this embodiment, the diagnostic models constructed based on biomarkers of different protein combinations were compared in terms of performance in the test group. The specific combination forms of different models are shown in Table 9 below.
[0192] Table 9. Combination forms of different diagnostic models
[0193] Number of Combined Tests Optimal Combination Form Two - Item Combined Test - 2MP TFF3 + SERPINA1 Three - Item Combined Test - 3MP TFF1 + IGFBP1 + IGFBP4 Four - Item Combined Test - 4MP TFF1 + IGFBP1 + IGFBP4 + SERPINA1 Five - Item Combined Test - 5MP TFF1 + TFF3 + IGFBP4 + SERPINA1 + OPN Six - Item Combined Test - 6MP TFF1 + TFF3 + IGFBP1 + IGFBP4 + SERPINA1 + OPN Seven - Item Combined Test - 7MP TFF1 + TFF3 + IGFBP1 + IGFBP4 + SERPINA1 + OPN + GDF - 15
[0194] The results are specifically as Figure 16 shown in Table 10. Table 10 shows the comparison results of the performance indicators of different diagnostic models constructed by using the 7 biomarkers screened in Example 1 for six-classification. The calculation methods for the minimum value, first quartile, median, mean, third quartile, and maximum value of accuracy and consistency are as follows: (1) Sort the values of accuracy or consistency from smallest to largest; (2) Minimum value: The first value after sorting; (3) First quartile (Q1): Multiply the number of data by 0.25. If the result is an integer, take the average of the values at this position and the next position; if not, round up to get the position, and the value at this position is Q1; (4) Median: If the number of data is odd, the median is the middle value; if it is even, it is the average of the two middle values; (5) Mean: The sum of all values divided by the number of data; (6) Third quartile (Q3): Multiply the number of data by 0.75, and the processing method is the same as Q1; (7) Maximum value: The last value after sorting.
[0195] Among them, the minimum value and the maximum value can reflect the extreme situations of the data and show the worst and best performances that the model may exhibit; the quartiles can help understand the distribution range and dispersion degree of the data; below Q1 represents a lower performance level, and above Q3 represents a higher performance level; the median can reflect the performance at the middle level; the mean comprehensively reflects the overall average performance. By synthesizing the above statistical values, the overall situation, distribution characteristics, and stability of the model performance can be comprehensively understood, thus providing a strong basis for model selection and optimization.
[0196] Table 10. Performance comparison of diagnostic models constructed based on biomarkers of different protein combinations
[0197]
[0198]
[0199] As can be seen from Table 10, for the six-classification diagnosis model, the seven-item joint detection model (7MP) composed of the best 7 markers of the present invention has the best performance. Therefore, the six-classification gradient boosting model constructed by using these seven protein markers is used as the best combined diagnosis model.
[0200] II. Determination and verification of the diagnostic performance of the best combined diagnosis model (7MP)
[0201] 1. Determination of the diagnostic performance of the joint detection model (7MP)
[0202] In order to more accurately determine the diagnostic performance and threshold of the model constructed in this embodiment for different disease classifications, a multi-classification model of the gradient boosting (gbm) algorithm in the model group is used for predictive analysis in the test group, and the predicted probability values for 6 classifications (healthy control, inflammatory bowel disease, benign polyp, advanced adenoma, early colorectal cancer, other cancers) are calculated. The classification with the largest predicted probability value is the final predicted result of the system.
[0203] Among them, the meanings and calculation methods of each index are as follows:
[0204] The calculation results are specifically as Figure 17 shown: The accuracy of the model in the model group is 0.761, and the consistency is 0.705. The diagnostic sensitivity for early colorectal cancer is 76.9%, the specificity is 95.9%, the positive predictive value is 78.4%, and the negative predictive value is 95.5%; the diagnostic sensitivity for advanced adenoma is 69.8%, the specificity is 94.5%, the positive predictive value is 71.3%, and the negative predictive value is 94.1%; the diagnostic sensitivity for benign polyp is 74.3%, the specificity is 95.4%, the positive predictive value is 67.7%, and the negative predictive value is 96.6%; the diagnostic sensitivity for inflammatory bowel disease is 67.2%, the specificity is 94.2%, the positive predictive value is 67.7%, and the negative predictive value is 94.1%; the diagnostic sensitivity for other cancers is 77.7%, the specificity is 95.9%, the positive predictive value is 67.3%, and the negative predictive value is 97.5%; the diagnostic sensitivity for healthy control is 83.6%, the specificity is 95.4%, the positive predictive value is 89.0%, and the negative predictive value is 92.9%.
[0205] II. Verification of the best combined diagnosis model (7MP)
[0206] Based on the algorithm constructed in the model group, the predictive performance is verified in the verification group (Table 6), and the specific results are as Figure 18As shown, the accuracy is 0.78 and the consistency is 0.729. The diagnostic sensitivity for early colorectal cancer is 79.4%, the specificity is 94.4%, the positive predictive value is 73.5%, and the negative predictive value is 95.9%; the diagnostic sensitivity for advanced adenoma is 73.4%, the specificity is 94.7%, the positive predictive value is 73.4%, and the negative predictive value is 94.7%; the diagnostic sensitivity for benign polyps is 72.7%, the specificity is 95.9%, the positive predictive value is 69.6%, and the negative predictive value is 96.5%; the diagnostic sensitivity for inflammatory bowel disease is 68.3%, the specificity is 96.3%, the positive predictive value is 77.4%, and the negative predictive value is 94.3%; the diagnostic sensitivity for other cancers is 83.3%, the specificity is 96.0%, the positive predictive value is 68.2%, and the negative predictive value is 98.3%; the diagnostic sensitivity for healthy controls is 85.0%, the specificity is 96.3%, the positive predictive value is 91.1%, and the negative predictive value is 93.5%.
[0207] In summary, it can be seen that the combined diagnostic model constructed in this embodiment, which contains 7 biomarkers (trefoil factor 1 (TFF1), trefoil factor 3 (TFF3), insulin-like growth factor binding protein 1 (IGFBP1), insulin-like growth factor binding protein 4 (IGFBP4), serine protease inhibitor A1 (SERPINA1), osteopontin (OPN), growth differentiation factor 15 (GDF-15)), has good diagnostic value for six classifications: early colorectal cancer, advanced adenoma, benign polyps, inflammatory diseases, healthy people, and other cancers.
[0208] All patents and publications mentioned in the specification of the present invention represent the publicly available technologies in the art that can be used in the present invention. All patents and publications cited herein are equally listed in the references, just as each publication is specifically individually referenced. The present invention described herein can be implemented in the absence of any one or more elements, one or more limitations, where such limitations are not specifically stated. For example, in each instance herein, the terms "comprising", "consisting essentially of", and "consisting of" can be replaced by either of the remaining two terms. The so-called "a" herein merely means "one", and does not exclude including only one, nor does it exclude including more than two. The terms and expressions adopted herein are for descriptive purposes and are not limiting, and there is no intention to indicate that the terms and interpretations described in this book exclude any equivalent features, but it is understood that any suitable changes or modifications can be made within the scope of the present invention and the claims. It is understood that the embodiments described in the present invention are all preferred embodiments and features, and any person of ordinary skill in the art can make some changes and variations based on the essence described in the present invention, and these changes and variations are also considered to fall within the scope of the present invention and the scope limited by the independent claims and the dependent claims.
Claims
1. Use of a marker for preparing a reagent for predicting whether an individual suffers from early colorectal cancer, characterized in that: The markers include any one or more combinations of trefoil factor 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, trefoil factor 3, prion protein, growth differentiation factor 15, guanylate cyclase activating factor 2A, insulin-like growth factor binding protein 1, regeneration family member protein 1α, and osteopontin; the early colorectal cancer includes advanced adenoma and early colorectal cancer.
2. The use according to claim 1, characterized in that The markers include Trefoil factor 1.
3. The use according to claim 2, characterized in that The markers include a combination of trefoil factor 3 and serine protease inhibitor A1, or a combination of trefoil factor 1, insulin-like growth factor binding protein 1 and insulin-like growth factor binding protein 4, or a combination of insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4 and serine protease inhibitor A1, or a combination of trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 4, serine protease inhibitor A1 and osteopontin, or a combination of trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1 and osteopontin.
4. The use according to claim 3, characterized in that The markers include a combination of trefoil factor 1, trefoil factor 3, insulin-like growth factor binding protein 1, insulin-like growth factor binding protein 4, serpin A1 and osteopontin.
5. A kit for predicting whether an individual suffers from early colorectal cancer, characterized in that: A detection reagent for a biomarker for use as claimed in any one of claims 1 to 4, wherein the early colorectal cancer includes advanced adenoma and early colorectal cancer.
6. A biomarker combination for predicting whether an individual has early colorectal cancer, characterized in that: The biomarker combination includes any one or more combinations of trefoil factor 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, trefoil factor 3, prion protein, growth differentiation factor 15, guanylate cyclase activating factor 2A, insulin-like growth factor binding protein 1, regeneration family member protein 1α, and osteopontin; the early colorectal cancer includes advanced adenoma and early colorectal cancer.
7. Use of a marker for preparing a reagent for simultaneously predicting whether an individual is in advanced adenoma, benign polyp, early colorectal cancer, healthy, inflammatory bowel disease or other cancer, characterized in that, The markers include any one or more combinations of trefoil factor 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, trefoil factor 3, prion protein, growth differentiation factor 15, guanylate cyclase activating factor 2A, insulin-like growth factor binding protein 1, regeneration family member protein 1α, and osteopontin.
8. A kit for simultaneously predicting whether an individual is in advanced adenoma, benign polyp, early colorectal cancer, healthy, inflammatory bowel disease or other cancer, characterized in that: A detection reagent comprising a marker for use as described in claim 7, wherein the early colorectal cancer includes advanced adenoma and early colorectal cancer.
9. A biomarker combination for simultaneously predicting whether an individual is in advanced adenoma, benign polyp, early colorectal cancer, healthy, inflammatory bowel disease or other cancer, characterized in that: The biomarker combination includes any one or more combinations of trefoil factor 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, trefoil factor 3, prion protein, growth differentiation factor 15, guanylate cyclase activating factor 2A, insulin-like growth factor binding protein 1, regeneration family member protein 1α, and osteopontin; the early colorectal cancer includes advanced adenoma and early colorectal cancer.
10. A system for simultaneously predicting whether an individual is in advanced adenoma, benign polyp, early colorectal cancer, healthy, inflammatory bowel disease or other cancer, characterized in that: The system includes a data analysis module, which is used to analyze the detection values of markers, wherein the markers include any one or more combinations of trefoil factor 1, insulin-like growth factor binding protein 4, serine protease inhibitor A1, trefoil factor 3, prion protein, growth differentiation factor 15, guanylate cyclase activating factor 2A, insulin-like growth factor binding protein 1, regeneration family member protein 1α, and osteopontin; the early colorectal cancer includes advanced adenoma and early colorectal cancer.