Biomarker for predicting recurrence risk of colorectal cancer and application of biomarker
By screening and constructing a multi-marker joint detection model for the risk of recurrence of colorectal cancer, and using proteomics technology to analyze blood samples, the problem of insufficient diagnostic efficacy in the existing technology is solved, and accurate prediction and early identification of the risk of recurrence of colorectal cancer is achieved, reducing the risk of recurrence.
Patent Information
- Application Number
- CN202510724976.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The lack of high-sensitivity proteomic biomarkers in the prior art is used to predict the risk of colorectal cancer recurrence, resulting in insufficient diagnostic efficacy and inability to effectively distinguish high-risk patients. Most diagnostic methods rely on genetic testing, the process is cumbersome and the instrument is highly dependent.
Proteins such as CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, TFF3, etc. were screened through proteomics, and a multi-marker joint detection model was constructed. Blood samples were analyzed using high-performance liquid chromatography-tandem mass spectrometry technology, and differential proteins related to the risk of recurrence of colorectal cancer were screened out, and a diagnostic model was constructed to distinguish recurrence risk.
Accurate, non-invasive and efficient prediction of the risk of recurrence of colorectal cancer is achieved, the risks of misdiagnosis and missed diagnosis are reduced, the ability to identify high-risk patients is improved, the opportunity for early intervention is provided, and the risk of recurrence or metastasis is reduced.
Smart Images

Figure CN120254282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of proteomic screening for colorectal cancer diagnostic markers, and more particularly, to a biomarker for predicting the recurrence risk of colorectal cancer and its applications. Background Art
[0002] Colorectal cancer (CRC) is one of the most common malignant tumors in clinical practice.
[0003] Surgical resection is the main means for CRC patients to achieve long-term survival. However, the recurrence or metastasis of colorectal cancer after surgery is the leading cause of patient death. Currently, the main treatment for CRC is radical tumor resection. Especially for stage I / II colorectal cancer patients, radical resection is usually directly performed with the main goal of cure, without the need for extensive clearance, and generally no adjuvant radiotherapy or chemotherapy is carried out. However, stage I / II CRC patients still have the possibility of recurrence after surgical treatment. If patients who do not benefit from surgical treatment or experience progression can have risk prediction and timely adjustment of treatment plans (such as adjuvant radiotherapy and chemotherapy, secondary surgical resection, targeted therapy, or immunotherapy, etc.), the overall survival rate and quality of life of patients can be significantly improved. Although there are some reports on markers for evaluating the prognosis and recurrence risk of colorectal cancer, they are basically based on gene-level evaluation and detection, such as based on ctDNA, DNA methylation detection, etc. Not only is the detection process more cumbersome and highly instrument-dependent, but the diagnostic sensitivity is still not high. There is a lack of biomarkers for the diagnosis of colorectal cancer recurrence risk in clinical practice. In particular, the discovery of highly sensitive proteomic biomarkers for the diagnosis of colorectal cancer recurrence risk is of great significance.
[0004] Proteomics is the science that studies the protein composition, localization, changes, and their interaction rules in cells, tissues, or organisms, including the study of protein expression patterns and proteome functional patterns. With the development of mass spectrometry technology, liquid chromatography-tandem mass spectrometry (LC-MS / MS) has become the most important tool in proteomic research. The development of proteomics is of great significance for finding disease diagnostic markers, screening drug targets, toxicology research, etc., and has thus been widely applied in medical research. Although there have been many articles and patents reporting on the discovery of novel tumor markers in recent years, they only remain at the laboratory research stage and are rarely applied clinically or promoted in the market. Moreover, in most cases, a single indicator is far from sufficient for the in vitro diagnosis of tumor recurrence risk. Only by adopting a combined detection form and combining various dimensions of detection can the prediction accuracy be enhanced. Therefore, finding new markers related to the diagnosis of colorectal cancer recurrence risk and constructing a prediction model by combining multiple markers have important clinical value. Summary of the Invention
[0005] In view of the problems existing in the prior art, the present invention provides a biomarker for predicting the recurrence risk of colorectal cancer and its application. By using proteomics methods, proteins with significantly different abundance levels in the blood of two groups of patients with colorectal cancer, namely those with recurrence and those without recurrence after surgical treatment, are analyzed to screen out biomarkers that can be used to predict the recurrence risk of colorectal cancer. Furthermore, a multi-marker combined detection model is constructed, which can accurately, non-invasively, and efficiently predict the recurrence risk of colorectal cancer and meet the clinical needs.
[0006] On the one hand, the present invention provides the use of a biomarker for preparing a reagent for predicting the recurrence risk of colorectal cancer, and the biomarker includes any one or more of CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, and TFF3.
[0007] Predicting the recurrence risk of colorectal cancer is completely different from predicting whether a person has colorectal cancer. In previous studies, our research group studied a large number of proteomic biomarkers for predicting early colorectal cancer and constructed a model for predicting early colorectal cancer. However, when this model was used to predict the recurrence risk of colorectal cancer, its diagnostic efficacy was poor, and it was unable to directly distinguish patients with high recurrence risk of colorectal cancer, resulting in insufficient detection sensitivity for diagnosis. Therefore, the screening of proteomic biomarkers for the recurrence risk of colorectal cancer was carried out again, hoping to find a more effective method for predicting the recurrence risk of colorectal cancer prognosis.
[0008] The proteomic biomarkers provided by the present invention can accurately predict the recurrence risk of colorectal cancer patients after treatment, thereby enabling earlier identification of high-risk recurrence populations, helping doctors anticipate in advance whether timely adjustment of treatment plans is needed, reducing the risk of recurrence or metastasis of colorectal cancer, prolonging the survival period, and truly benefiting colorectal cancer patients.
[0009] The present invention uses proteomics methods to collect plasma samples from patients with colorectal cancer who have recurrence within a short period (such as 1 to 5 years) and those who do not have recurrence within a short period after colorectal cancer treatment. Different samples are analyzed by high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS). Based on the orthogonal partial least squares discriminant analysis and significance analysis methods, proteins with significant differences between colorectal cancer recurrence patients and non-recurrence patients are first screened. Finally, 10 differential proteins with obvious relevance to the recurrence risk of colorectal cancer are screened out. These 10 proteins can be used to distinguish colorectal cancer recurrence patients and non-recurrence patients and have a certain diagnostic efficacy, thus can be used to predict the recurrence risk of colorectal cancer.
[0010] Among them, the CD36 is a protein or amino acid sequence with the UniProt database number P16671; PLIN1 is a protein or amino acid sequence with the UniProt database number O60240; ITIH3 is a protein or amino acid sequence with the UniProt database number Q06033; CD74 is a protein or amino acid sequence with the UniProt database number P04233; MGAM2 is a protein or amino acid sequence with the UniProt database number Q2M2H8; RNASE1 is a protein or amino acid sequence with the UniProt database number P07998; SELL is a protein or amino acid sequence with the UniProt database number P14151; TFF1 is a protein or amino acid sequence with the UniProt database number P04155; ORM2 is a protein or amino acid sequence with the UniProt database number P19652; TFF3 is a protein or amino acid sequence with the UniProt database number Q07654.
[0011] The present invention also surprisingly discovers that some of the protein markers obtained through proteomic screening are known markers that can be used for other cancers. For example, TFF3 has been reported to be used for the prediction of lung cancer, but this screening discovers that this marker can also be used for the prediction of the recurrence risk after colorectal cancer surgery. It can be seen that for many different cancers, their protein markers are not completely separated or irrelevant. In fact, there are many cross - relationships or influences. Many protein markers can be used for the prediction of early - stage cancer, and also for the prognostic diagnosis of cancer. Even many protein markers can be used for the prediction and diagnosis of different stages of multiple different cancers. Therefore, there are still many brand - new functions in the field of proteomics waiting to be explored, and the market prospect is very broad.
[0012] Furthermore, the markers include CD36, PLIN1, ITIH3, MGAM2, RNASE1, and SELL.
[0013] To improve the diagnostic efficiency for the recurrence risk of colorectal cancer, it is also necessary to combine different differential proteins to construct a diagnostic model. Sort according to the importance obtained from the screening, and respectively select different numbers of differential proteins with higher rankings for combination. Finally, 6 protein markers are screened. Based on these 6 protein markers, a model is constructed, which has good risk prediction ability in the diagnosis of the recurrence risk of colorectal cancer. The data from the detection of clinical colorectal cancer samples shows that just using these 6 biomarkers to predict the recurrence risk of colorectal cancer, the AUC value can reach 0.945, and its effect is significantly better than the effect of the existing combination of multiple biomarkers in predicting the recurrence risk of colorectal cancer.
[0014] Further, the recurrence risk refers to recurrence within three years after colorectal cancer treatment; the recurrence of colorectal cancer includes in-situ or adjacent area recurrence of colorectal cancer, and metastasis of colorectal cancer.
[0015] It can be understood that the "within three years" in "recurrence within three years after colorectal cancer treatment" here is not an absolute and unchanging time node. It is only a time node summarized based on current clinical experience. As time goes by or with the improvement of other treatment methods, this time node or the length of time may change. For example, recurrence after colorectal cancer treatment may occur within one year, within 360 days, within half a year, within 180 days, within two years, within two and a half years, within three and a half years, etc.
[0016] In some ways, the recurrence of colorectal cancer means that after surgical treatment with radical resection of the tumor, the patient has in-situ or adjacent area recurrence of colorectal cancer within three years, or metastasis of colorectal cancer occurs, such as liver metastasis, lung metastasis, peritoneal metastasis, bone metastasis, brain metastasis, etc.
[0017] Further, the reagent is used to detect the content of a biomarker in a body fluid sample; the body fluid sample includes any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.
[0018] In some embodiments, the reagent for predicting the recurrence risk of colorectal cancer is a detection reagent prepared with this biomarker as the detection target, such as sample pretreatment reagents, antigens or antibodies, and other biological reagents and kits suitable for the detection of the biomarker; it can also be developed into a standardized reagent or kit suitable for the biomarker, etc.
[0019] Further, the reagent is used to detect the presence, relative abundance or concentration of a biomarker in a body fluid sample.
[0020] The present invention screens biomarkers for predicting the recurrence risk of colorectal cancer from blood. There are significant differences in the blood of people at high risk and low risk of colorectal cancer recurrence for these biomarkers. By collecting a blood sample, it is possible to predict or assist in diagnosing the possibility of colorectal cancer recurrence in an individual by detecting these biomarkers in the individual's blood, or it is possible to detect these biomarkers in the blood of a certain group, and then divide this group into people at high risk and low risk of colorectal cancer recurrence.
[0021] Further, the detection method includes radiometric methods, immunological methods, fluorescence methods, flow-through fluorescence methods, latex turbidimetry, biochemical methods, enzymatic methods, hybridization methods, gas chromatography-mass spectrometry, liquid chromatography-mass spectrometry, chromatography, chemiluminescence methods, magnetoelectric methods or photoelectric conversion methods.
[0022] The presence, absence, or high or low content of the biomarker here is a relative concept. For example, when comparing the high-risk group of colorectal cancer recurrence and the low-risk group of colorectal cancer recurrence, the content of these specific biomarkers is compared with the high-risk group and the low-risk group of colorectal cancer recurrence as the benchmarks. There may be some biomarkers with a relatively higher content in the high-risk group of colorectal cancer recurrence than in the low-risk group of colorectal cancer recurrence, and this increase is statistically significant, such as a significant or highly significant increase. Therefore, when judging these biomarkers, if it is a single biomarker, if the probability of a certain risk occurrence increases and the content of the biomarker changes, this change may be a relative increase or may also be a relative decrease, and the difference between this relative increase or relative decrease is significantly different, and of course it can also be highly significantly different. Therefore, no matter what means are used for detection, a pre-specified value (cut-off value) can be used as the standard. If it is higher than this value, it is considered that the content has changed, and such results can all be used for prediction or diagnostic value.
[0023] Therefore, in some aspects, the biomarker described in the present invention can be obtained by detecting the content of the biomarker in a sample by any known method, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strip, immunoreaction chip, capillary electrophoresis, infrared spectroscopy, etc. As long as it can be used to detect the content of the protein biomarker in the sample, it can be used for the diagnosis of the high-risk group of colorectal cancer recurrence and the low-risk group of colorectal cancer recurrence. As long as the content of the protein biomarker in the sample can be detected, it can be used to predict or diagnose the probability of the occurrence of a certain disease. It can be understood that this detection is for the individual sample, and then compared with the pre-set standard, and the comparison result is used to judge or predict the occurrence status of the disease. For example, it can be used to predict the probability of colorectal cancer recurrence. Such prediction or diagnosis is whether it will occur within a certain time. Of course, such detection can be continuous detection, and the progress of the disease can be inferred from the change in the content of certain substances.
[0024] In some ways, the relative abundance is the peak area of the biomarker in the detection spectrum obtained by high performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a certain biomarker measured in the control sample is 100 and the average peak area measured in the sample of the high-risk group of colorectal cancer recurrence is 600, then it is considered that the abundance of the biomarker in the sample is 6 times that of the control sample.
[0025] Furthermore, the colorectal cancer patient is a stage I / II CRC patient.
[0026] On the other hand, the present invention provides a kit for predicting the recurrence risk of colorectal cancer, and the kit includes a detection reagent for the biomarker as described above in the use.
[0027] On another aspect, the present invention provides a biomarker combination for predicting the recurrence risk of colorectal cancer, and the biomarker combination includes CD36, PLIN1, ITIH3, MGAM2, RNASE1, and SELL.
[0028] On another aspect, the present invention provides a system for predicting the recurrence risk of colorectal cancer, and the system includes a data analysis module for analyzing the detection values of biomarkers, and the biomarkers include any one or more of CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, and TFF3.
[0029] Further, the biomarkers include CD36, PLIN1, ITIH3, MGAM2, RNASE1, and SELL; the recurrence risk refers to recurrence within three years after surgical treatment of colorectal cancer; the recurrence of colorectal cancer includes in-situ or adjacent region recurrence of colorectal cancer and metastasis of colorectal cancer.
[0030] Further, the data analysis module uses the detection values of biomarkers of known samples as a training set, and according to whether colorectal cancer recurs, it is divided into a colorectal cancer recurrence group and a colorectal cancer non-recurrence group, analyzes the relationship between the detection values of the colorectal cancer recurrence group and the colorectal cancer non-recurrence group, and constructs a model.
[0031] In some embodiments, a combined diagnostic model for predicting the recurrence risk of colorectal cancer is constructed by combining multiple machine learning methods, and it is preliminarily confirmed that for any one of the selected novel biomarkers alone, the change in its concentration can be used to distinguish high-risk and low-risk populations of colorectal cancer recurrence, indicating that these biomarkers have extremely high diagnostic value.
[0032] In some embodiments, the equation of the constructed model is:
[0033]
[0034] Wherein, Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 6), Xi represents the detection value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker, and b is a constant 1.6999441; the coefficients of the 6 biomarkers are:
[0035]
[0036] When Y ≤ 0.5, the risk of colorectal cancer recurrence in the subject to be tested is low; when Y > 0.5, the risk of colorectal cancer recurrence in the subject to be tested is high.
[0037] Furthermore, the system further includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0038] Furthermore, the detection value is the presence or absence, relative abundance, or concentration value of each biomarker.
[0039] On the other hand, the present invention provides the use of a biomarker for preparing a reagent for predicting non-recurrence, in-situ recurrence, or distant metastasis in colorectal cancer patients after surgery, and the biomarker includes any one or more of CD36, PLIN1, ITIH3, MGAM2, RNASE1, SELL, ORM2, TFF1, and TFF3.
[0040] The present invention attempts to use protein biomarkers for predicting recurrence risk to further distinguish between in-situ recurrence patients and distant metastasis patients, and finds that each protein biomarker can be used to distinguish among non-recurrence, in-situ recurrence, and distant metastasis in colorectal cancer patients after surgery. The distant metastasis includes patients with only distant metastasis, as well as patients with both in-situ recurrence and distant metastasis.
[0041] When a six-classification combined diagnosis model is constructed using 6 biomarkers, namely CD36, PLIN1, ITIH3, MGAM2, RNASE1, and SELL, the accuracy of distinguishing between in-situ recurrence patients and distant metastasis patients reaches about 70%.
[0042] When 3 more biomarkers, namely ORM2, TFF1, and TFF3, are further added on the basis of the 6 biomarkers to construct a nine-classification combined diagnosis model containing 9 protein biomarkers, the diagnostic efficacy for distinguishing between in-situ recurrence and distant metastasis in colorectal cancer patients after surgical treatment can be further improved, and the accuracy reaches about 84%. It can be directly used for the early prediction of non-recurrence, in-situ recurrence, or distant metastasis in colorectal cancer patients after surgery.
[0043] Furthermore, the reagent is used to predict non-recurrence, in-situ recurrence, or distant metastasis in colorectal cancer patients within three years after surgery.
[0044] Furthermore, the reagent is used to detect the content of biomarkers in a body fluid sample; the body fluid sample includes any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.
[0045] Further, the reagent is used to detect the presence, relative abundance or concentration of biomarkers in a body fluid sample.
[0046] In another aspect, the present invention provides a kit for predicting non - recurrence, in - situ recurrence or distant metastasis after surgery in colorectal cancer patients, and the kit includes a detection reagent for the biomarker as described above.
[0047] In another aspect, the present invention provides a biomarker combination for predicting non - recurrence, in - situ recurrence or distant metastasis after surgery in colorectal cancer patients, and the combination includes CD36, PLIN1, ITIH3, MGAM2, RNASE1, SELL, ORM2, TFF1 and TFF3.
[0048] In another aspect, the present invention provides a system for predicting non - recurrence, in - situ recurrence or distant metastasis after colorectal cancer surgery. The system includes a data analysis module, and the data analysis module is used to analyze the detection values of biomarkers, and the biomarkers include any one or more of CD36, PLIN1, ITIH3, MGAM2, RNASE1, SELL, ORM2, TFF1, TFF3.
[0049] Further, the distant metastasis also includes simultaneous in - situ recurrence and distant metastasis, and as long as distant metastasis occurs, it is classified into the distant metastasis group.
[0050] Further, the data analysis module uses the detection values of biomarkers of known samples as a training set, and according to the situation after surgery of colorectal cancer patients, it is divided into a non - recurrence group, an in - situ recurrence group and a distant metastasis group, analyzes the relationship between the detection values of the non - recurrence group, the in - situ recurrence group and the distant metastasis group, and constructs a model.
[0051] Further, the model is constructed based on the gradient boosting algorithm.
[0052] The gradient boosting algorithm does not output a model formula and a cut - off value like generalized linear models. All calculations are directly completed by machine learning. The detection values can be directly input into the software system, and the prediction results can be directly obtained.
[0053] Further, the system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0054] The beneficial effects of the present invention are as follows:
[0055] 1. The present invention has screened 10 brand-new biomarkers that can predict the recurrence risk of colorectal cancer, developed a new combination of protein biomarkers, which can effectively evaluate and diagnose patients at high risk of colorectal cancer recurrence, effectively distinguish between high-risk and low-risk populations for colorectal cancer recurrence, and can more accurately identify patients at high risk of colorectal cancer recurrence. Compared with traditional detection methods, it reduces the risk of misdiagnosis and missed diagnosis, providing strong support for the early detection and intervention of the disease.
[0056] 2. The combined differential diagnosis model of 6 biomarkers constructed by the present invention is convenient, fast, and the detection results are highly consistent with the results of the clinical gold standard test. At the same time, it significantly reduces the cost of predicting the recurrence risk of colorectal cancer and has good application prospects.
[0057] 3. Based on the screened biomarkers that predict the recurrence risk of colorectal cancer, a three-classification model that can simultaneously distinguish between the non-recurrence group, in-situ recurrence group, and distant metastasis group is further constructed, providing a more effective and accurate prediction and diagnosis mode.
[0058] Detailed description
[0059] (1) Diagnosis or detection
[0060] Here, the diagnosis or detection refers to the detection or assay of biomarkers in a sample, or the content of the target biomarker, such as the absolute content or relative content, and then it is indicated whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease, based on the presence or quantity of the target biomarker. The meanings of diagnosis and detection here can be interchanged. The result of this detection or the result of the diagnosis cannot be directly used as the direct result of having a disease, but is an intermediate result. If a direct result is to be obtained, other auxiliary means such as pathology or anatomy are required to confirm the presence of a certain disease. For example, the present invention provides a variety of new biomarkers related to the recurrence risk of colorectal cancer, and the change in the content of these biomarkers has a direct correlation with whether the patient belongs to the high-risk population for colorectal cancer recurrence.
[0061] (2) The connection between the biomarker or biological marker or differential protein and the recurrence risk of colorectal cancer
[0062] Biomarker, biological marker, and differential protein have the same meaning in the present invention. Here, the connection means that the appearance or change in the content of a certain biomarker in a sample has a direct correlation with a specific disease. For example, a relative increase or decrease in content indicates that the possibility of having this disease is relatively higher than that of the healthy population.
[0063] If multiple different biomarkers in a sample appear simultaneously or there are relative changes in their contents, it indicates that the likelihood of having this disease is relatively higher compared to the healthy population. That is to say, among the types of biomarkers, some biomarkers have a strong correlation with the disease, some have a weak correlation with the disease, or some may even have no correlation with a specific disease. One or more of those biomarkers with a strong correlation can be used as biomarkers for diagnosing the disease, and those with a weak correlation can be combined with the strong ones to diagnose a certain disease, increasing the accuracy of the test results.
[0064] Regarding the numerous biomarkers found in the serum in the present invention, these biomarkers can all be used to distinguish between high-risk and low-risk populations for colorectal cancer recurrence. These biomarkers can be used alone as individual biomarkers for direct detection or diagnosis. Selecting such a biomarker indicates that the relative change in the content of this biomarker has a strong correlation with the risk of colorectal cancer recurrence. Of course, it can be understood that simultaneous detection of one or more biomarkers with a strong correlation with the risk of colorectal cancer recurrence can be selected. Normally understood, in some ways, selecting biomarkers with a strong correlation for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90% or 95% accuracy. Then it can be stated that these biomarkers can obtain an intermediate value for diagnosing a certain disease, but it does not mean that it can directly confirm the presence of a certain disease.
[0065] Of course, it is also possible to select the differential proteins with larger ROC values as biomarkers for diagnosis. The so-called strength or weakness is generally calculated and confirmed through some algorithms, such as the contribution rate or weight analysis of the biomarker to the assessment of the risk of colorectal cancer recurrence. Such calculation methods can include significance analysis (p-value or FDR value) and fold change, and multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA) and orthogonal partial least squares discriminant analysis (OPLS-DA). Of course, other methods are also included, such as ROC analysis, etc. Of course, other model prediction methods are also possible. When specifically selecting biomarkers, the differential proteins disclosed in the present invention can be selected, or other existing well-known biomarker combinations can be selected or combined for prediction through model methods.
[0066] (3) Definition of disease terms
[0067] Colorectal cancer: Also known as large bowel cancer, it refers to cancers originating from the epithelium of the large intestine, including colon cancer and rectal cancer. The most common pathological type is adenocarcinoma, and squamous cell carcinoma is extremely rare. In China, rectal cancer is the most common, followed by colon cancer (sigmoid colon, cecum, ascending colon, descending colon, and transverse colon). The treatment of colorectal cancer should follow the principle of individualized treatment. According to the patient's age, physical condition, pathological type of the tumor, and invasion range (stage), appropriate treatment methods should be selected, including radical surgical treatment, chemotherapy, targeted therapy, radiotherapy, etc.; The formation of colorectal cancer generally goes through the development process of normal mucosal hyperplasia, advanced adenoma (malignant), and adenocarcinoma (malignant), which generally takes 5 - 10 years.
[0068] Recurrence of colorectal cancer: It refers to the phenomenon that after a patient completes radical treatment (surgical resection), after a period of clinical disease-free state (i.e., a period without signs of tumors), cancer cells grow again or tumor lesions (also known as metastases) appear again at the original tumor site or other parts of the body. Recurrence may stem from residual cancer cells that were not completely removed during treatment, or micrometastatic foci that had spread earlier but were not detected. Recurrence of colorectal cancer can occur from several months to several years after treatment, but 80% of recurrences occur within 2 - 3 years after surgery, and the recurrence risk is significantly reduced after 5 years. Metastasis of colorectal cancer refers to metastasis caused by lymphatic, hematogenous, or direct infiltration, etc., including liver metastasis, lung metastasis, peritoneal metastasis, bone metastasis, brain metastasis, etc., and can be further divided into oligometastasis (1 - 3 lesions), extensive metastasis, and so on.
[0069] Early detection and intervention of colorectal cancer recurrence are the keys to improving the prognosis of colorectal cancer. If biomarkers with certain warning effects can be found at the early stage when colorectal cancer has not yet recurred, and the recurrence risk of colorectal cancer is diagnosed, it is of great significance for improving the treatment effect of patients and improving the prognosis of patients.
[0070] Stage I colorectal cancer means that the tumor is confined within the intestinal wall, the tumor invades the submucosa but does not reach the muscular layer, there is no lymph node metastasis, and no distant metastasis.
[0071] Stage II colorectal cancer means that the tumor has invaded the entire intestinal wall or surrounding tissues, the tumor invades the muscularis propria but does not penetrate the intestinal wall, there is no lymph node metastasis, and no distant metastasis.
[0072] (4) The gold standard for the diagnosis of colorectal cancer recurrence: is pathological histological examination (i.e., pathological confirmation of biopsy or surgical resection specimens), observing the presence of cancer cells under a microscope, and clarifying the nature of the tumor in combination with immunohistochemistry or molecular detection. Description of the Drawings
[0073] Figure 1 It is a volcano plot for the differential analysis of protein markers in the high-risk group and low-risk group of colorectal cancer recurrence in Example 1;
[0074] Figure 2 ROC and OPLS-DA analysis result graphs for the high-risk group and low-risk group of colorectal cancer recurrence in Example 1;
[0075] Figure 3 AUC result graphs for models constructed with different hyperparameters in Example 2;
[0076] Figure 4 ROC curve of the combined diagnostic model in the test group in Example 2;
[0077] Figure 5 ROC curve of the combined diagnostic model in the validation group in Example 2;
[0078] Figure 6 Diagnostic performance evaluation graph of the three-class combined diagnostic model in Example 3. Detailed implementation manner
[0079] The present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not impose any limitations on it. The reagents used in this embodiment are all known products and are obtained by purchasing commercially available products.
[0080] Example 1. Screening biomarkers for colorectal cancer recurrence risk using proteomics
[0081] By collecting plasma samples from colorectal cancer patients after radical surgery, enriching low-abundance proteins based on the method of removing high-abundance proteins by immunoaffinity chromatography, detecting the protein abundance in the samples by a high-performance liquid chromatography tandem mass spectrometry device, analyzing the differences in its abundance between patients without recurrence or metastasis within three years of colorectal cancer and patients with recurrence or metastasis within three years of colorectal cancer, and analyzing its diagnostic performance. The specific steps are as follows:
[0082] 1. Sample collection
[0083] A total of about 2 ml of peripheral blood samples were collected from 180 colorectal cancer patients (stage I / II CRC patients) 2 weeks after surgical resection treatment, placed in a vacuum tube containing EDTA anticoagulant, mixed well, centrifuged at 120 g for 10 minutes at room temperature, and the supernatant was taken, repeated twice; then at 360 g for 20 minutes. Then the platelet samples were collected in centrifuge tubes and stored at -80°C. The 180 patients included 102 males and 78 females, with an average age of 58 years (26 - 84 years), 67 stage I CRC patients, and 113 stage II CRC patients. All enrolled patients signed informed consent forms. Among them, all colorectal cancer patients were those diagnosed by pathological histology. The inclusion criteria were: (a) no history of other malignant tumors; (b) no patients with combined other malignant tumors or autoimmune diseases.
[0084] In the following three years, patients were followed up every six months, and regular CD74, colonoscopy, and imaging examinations were conducted. Once recurrence or metastasis was detected, pathological histology confirmation was required. A total of 31 patients with stage I / II CRC had recurrence within three years after surgery, including 18 cases of in-situ local recurrence and 15 cases of distant metastasis (including 8 cases of liver metastasis, 6 cases of lung metastasis, and 1 case of peritoneal metastasis). Then, the samples stored at -80°C were taken out and divided into two groups: 149 samples from patients without recurrence were classified into the low-risk group of CRC recurrence, and 31 samples with recurrence or metastasis were classified into the high-risk group of CRC recurrence for the screening of proteomic markers.
[0085] 2. Sample processing and enzymatic digestion
[0086] First, plasma samples were centrifuged on a centrifuge for 15 minutes (15,000g), and the supernatant was taken and filtered, followed by immunoaffinity chromatography to remove 14 high-abundance proteins. Then, a concentrator with a cut-off molecular weight of 3 kDa was used to concentrate the low abundance to 350 μL on a centrifuge (4000 g, 1 hour). The concentrated solution was recovered, and a desalting column with a cut-off molecular weight of 7 kDa was used to perform buffer exchange (Buffer Exchange) on a centrifuge (1000g, 2 minutes), and the exchange buffer was AEX-A (20 mM Tris, 4M Urea, 3% isopropanol, pH 8.0). Using AEX-A as a blank, the protein concentration in the sample was measured by the BCA method. According to the sample grouping, 25 μL of TCEP was added to the sample, and the sample was incubated at 37°C for 30 minutes for protein reduction. Then, the corresponding TMT 16-plex reagent was added, and the sample was incubated in the dark at room temperature for 1 hour for the TMT labeling reaction. Then, the Zeba column was used to perform buffer exchange on the sample, and the exchange buffer was AEX-A. After mixing the samples labeled with TMT 16-plex, 2 mL of AEX-A was added to the mixed sample, and the final volume was 5.5 mL. The sample was filtered using a 0.22 m filter and the TMT 16-plex labeled sample was separated using a 2D-HPLC system. The collected fractions were lyophilized, and finally, a mixture of Trypsin-Lysin C enzymes was added, and the sample was incubated at 37°C for 5 hours for enzymatic digestion. 5 μL of 10% TFA was added to terminate the enzymatic digestion reaction. A total of 60 enzymatically digested 2D-HPLC fractions were used for nanoLC-MS / MS analysis.
[0087] 3. LC-MS / MS data acquisition and database search analysis
[0088] DIA analysis was performed using a Vanquish Neo system (Thermo Fisher) with a nanoliter flow rate for chromatographic separation. The samples after nanoliter high-performance liquid chromatography separation were subjected to DIA (data-independent) mass spectrometry analysis using an Astral high-resolution mass spectrometer (Thermo Scientific). Detection mode: positive ion, the precursor ion scan range was 380 - 980 m / z, the resolution of the first-stage mass spectrometry was 240,000 at 200 m / z, the Normalized AGC Target was 500%, and the Maximum IT was 5 ms. MS2 used the DIA data acquisition mode, with 299 scan windows set, the Isolation Window was 2 m / z, the HCD Collision Energy was 25 eV, the Normalized AGC Target was 500%, and the Maximum IT was 3 ms.
[0089] 4. Data preprocessing
[0090] The MS / MS data were searched using Maxquant (v1.6.15.0). The data type was DIA proteomics data based on the quantification of secondary reporter ions. The requirement for the MS / MS spectra used for quantification was that the proportion of precursor ions in the first-stage spectra was greater than 75%. The database source was Homo_sapiens_9606_proteome_gene from the Uniprot database (release: 2021-10-14, sequence: 20,437), and a common contaminant library was added to the database. Contaminant proteins were removed during data analysis; the digestion method was set to Trypsin / P; the number of missed cleavage sites was set to 2; the precursor ion mass error tolerances for the First search and Main search were set to 20 ppm and 5 ppm, respectively, and the mass error tolerance for the MS / MS fragment ions was 20 ppm. The fixed modification was cysteine alkylation, and the variable modifications were methionine oxidation and protein N-terminal acetylation. The FDRs for protein identification and PSM identification were both set to 1%.
[0091] 5. Differential analysis
[0092] A combination of univariate analysis and multivariate statistical analysis was used to screen for differential proteins between the high-risk group and low-risk group of colorectal cancer recurrence. Univariate analysis mainly included the significance analysis (p-value or FDR value) and fold change of characteristic molecules in different groups, and multivariate statistical analysis mainly included receiver operating characteristic curve (ROC) analysis and Boruta feature selection based on the random forest algorithm. All statistical analyses were completed using R, and the specific R-related information is shown in Table 1.
[0093] Table 1. R and its related information used in the present invention
[0094]
[0095] The variable importance for the projection (VIP) was calculated to measure the influence intensity and interpretability of the expression patterns of each protein on the classification and discrimination of each group of samples. Further, the Wilcoxon rank-sum test was performed to obtain the corrected p-value (FDR). According to the conditions of FDR < 0.01 and Fold change > 2, 71 down-regulated proteins and 66 up-regulated proteins were screened out (see details in Figure 1 ).
[0096] To evaluate the role of each protein biomarker in the diagnosis and prediction of the recurrence risk of colorectal cancer, in this example, the ROC and Boruta analysis methods were used to evaluate each protein biomarker, and the results are shown in Figure 2 , where the abscissa is the AUC obtained from the ROC analysis, the ordinate is -log10(FDR) calculated by the Wilcoxon test, and the size of the points represents the VIP value obtained from the Boruta analysis. According to the further screening of VIP > 3 and AUC > 0.6, a total of 10 more significant candidate protein biomarkers were found, see details in Table 2.
[0097] Table 2. Differential biomarkers between the high-risk group and the low-risk group of colorectal cancer recurrence
[0098]
[0099] Among them, the smaller the FDR value and / or the larger the VIP value, to a certain extent, it indicates that the difference between the high-risk group and the low-risk group of colorectal cancer recurrence of the protein is more significant, and at the same time, it also indicates that the protein may have higher diagnostic value.
[0100] Example 2. Construction and validation of the recurrence risk model of colorectal cancer
[0101] 1. Models constructed with different biomarker combinations
[0102] Although a single biomarker can also distinguish the recurrence risk of colorectal cancer patients after surgery, generally speaking, combining multiple biomarkers can achieve higher accuracy in discrimination or prediction. However, for a single biomarker with higher accuracy in predicting the recurrence risk of colorectal cancer patients after surgery, its role in the combination with one or more other biomarkers may not necessarily be greater. At the same time, it is not that the more the number of biomarkers, the higher the prediction accuracy (AUC value) of the combination. Therefore, a large number of verification experiments are still needed.
[0103] In this example, a model was constructed and studied for the 10 protein markers of CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, and TFF3 screened in Example 1. The research cohort was blood samples taken 2 weeks after surgical treatment of 390 colorectal cancer patients (stage I / II CRC patients). All enrolled patients signed informed consent forms. Among them, 58 stage I / II CRC patients had recurrence within three years after surgery, including 21 cases of in-situ local recurrence and 37 cases of distant metastasis (including 20 cases of liver metastasis, 9 cases of lung metastasis, 6 cases of peritoneal metastasis, and 2 cases of bone metastasis). Thus, they were divided into two groups: 332 patient samples without recurrence were classified into the low-risk group of colorectal cancer recurrence, and 58 cases with recurrence or metastasis were classified into the high-risk group of colorectal cancer recurrence. They were randomly divided into a test group and a validation group. The test group and the validation group included 166 cases in the low-risk group of colorectal cancer recurrence and 29 cases in the high-risk group of colorectal cancer recurrence, respectively.
[0104] In the test group, a combined diagnostic model of multiple protein markers was constructed using a method that combines multiple machine learning methods. The predicted probability value was used to estimate the area under the receiver operator characteristic (ROC) curve (AUC) with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate diagnostic model. Using the test group, the Youden index (YI) was calculated to determine the cut-off value for predicting the probability of distinguishing between the high-risk group of colorectal cancer recurrence and the low-risk group of colorectal cancer recurrence. In addition, the ROCs of individual markers and different subgroups were constructed and compared. Standard descriptive statistics were calculated, such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD), to describe the experimental results of the study population. Statistical analysis was performed using R 3.6.1, and a p-value less than 0.05 was considered statistically significant.
[0105] The steps for constructing the combined diagnostic model are as follows:
[0106] S101: Among the 10 protein markers of CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, and TFF3 in the samples of the test group, a concentration matrix of randomly selected 2 to 10 markers was used as the original training data set.
[0107] S102: The generalized linear model (glmnet) algorithm was selected for constructing the prediction model and the grid search range during the hyperparameter optimization process of the algorithm. In this step, the grid search range for hyperparameter optimization of each algorithm was set as shown in Table 3.
[0108] Table 3. Parameter grid of glmnet algorithm
[0109]
[0110] S103. Select one of the hyperparameter combination methods according to the algorithm and hyperparameter setting range set in step S102 as the parameters for constructing the prediction model.
[0111] S104. Split the original data set into K subsets according to the K-fold cross-validation mechanism. To ensure that the ratio of majority-class samples to minority-class samples in each subset is the same as that in the original data set, the stratified K-fold cross-validation mechanism needs to be used for data splitting.
[0112] S105. Select one of the K training data subsets obtained by splitting in step S104 as the validation set Ddev.
[0113] S106. Combine the training data subsets not selected in step S105 to form the training data pool Dtrainl.
[0114] S107. Based on the training data set Dtrain obtained in step S106, construct a prediction model based on the selected supervised classification algorithm and hyperparameters.
[0115] S108. According to the prediction model obtained in step S107, evaluate the AUC value on the validation set Ddev, and store the current prognosis prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 is to evaluate the prediction model obtained in step S107 on the validation set determined in the current iteration, and store both the model and the evaluation results in the prediction model pool for use in subsequent prediction model selection. The evaluation mentioned in this step can be the AUC value or other reasonable metrics for evaluating model performance.
[0116] S109. Determine whether each subset has been used as the validation set. Step S109 is to determine whether the K subsets obtained in step S104 have all been used as the validation set for model training. If all subsets have been used as the validation set and completed training, execute step S110; if there are subsets that have not been used as the validation set, execute step S105. This step ensures that each sample in the original data set has been used as the validation set, improves model stability, and prevents the model from overfitting to a certain subset.
[0117] S110. Take the average value of the AUCs of all models in the obtained prediction model pool Pool as the final performance evaluation value of the model for this combination method. And store the model parameters and the final performance evaluation AUC value in the optimal model pool Poolbest.
[0118] S111, Determine whether prediction models have been constructed for all combinations of hyperparameters. Step S111 is to determine whether prediction models have been constructed for all algorithms and their corresponding hyperparameter combinations obtained in step S102. If the construction of models has been completed for all combinations, then execute step S112; if there are combinations for which model construction has not been completed, then execute step S103.
[0119] S113, From the model set Poolbest obtained in step S112, select the model with the largest AUC value as the final prediction model for colorectal cancer recurrence risk diagnosis.
[0120] S114, Repeat all the above steps until modeling has been completed for all combinations of markers.
[0121] By performing the above model construction steps, we obtained the optimal models constructed for all combinations of markers. To compare the performance of the models under these different marker combinations, we used the ROC method to evaluate the AUC values of these models in the test group, and the results are shown in Table 4.
[0122] Table 4. Comparison of the areas under the ROC curves of models constructed with different marker combinations in the test group
[0123]
[0124] Table 4 is sorted according to the markers with larger VIP values and smaller FDR values in Table 2. Starting from the markers ranked higher, markers are selected for combination in sequence, from 2MP to 5MP. The detected AUC value, accuracy, and sensitivity all increase as the number of markers increases. However, when CD74 is added on the basis of 5MP, the diagnostic performance of the constructed model does not continue to increase, and the AUC value even decreases. It is speculated that the addition of CD74 may have introduced noise. Therefore, CD74 was removed, and other markers such as SELL were further added. Finally, it was found that the model constructed by the combination of 6 markers (6MP-2, CD36 + RNASE1 + PLIN1 + ITIH3 + MGAM2 + SELL) has an AUC higher than the average value of the models with other marker combinations, even higher than 8MP, 9MP, and 10MP, and has fewer markers, making it the most optimal model.
[0125] 2. Optimization of model parameters
[0126] For the optimal biomarker combination form CD36 + RNASE1 + PLIN1 + ITIH3 + MGAM2 + SELL, based on this biomarker combination, in this example, models constructed under 9 different combinations of glmnet algorithm hyperparameters were analyzed, and the performance of the models was evaluated by the AUC value (the 10-fold cross-validation method was used to calculate the AUC during the modeling process). The results are shown in Table 5 and Figure 3 as follows.
[0127] Table 5. AUC of models constructed under different combinations of glmnet algorithm hyperparameters
[0128]
[0129] It can be seen from Table 5 that when the combination of glmnet algorithm hyperparameters is alpha = 1 and lambda = 0.0055, the AUC reaches the maximum value of 0.945.
[0130] The equation of the model constructed based on the optimal hyperparameter combination is:
[0131]
[0132] where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 6), Xi represents the detected value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker (Table 6), and b is the constant 1.6999441.
[0133] Table 6. Coefficients of 6 biomarkers in the model
[0134]
[0135] The complete model equation is:
[0136] Y = 0.124 ITIH3 + 7.153 SELL + 4.292 RNASE1 + 5.613 MGAM2 + 6.551 PLIN1 + 7.173 CD36 + 1.6999441
[0137] Determination of the diagnostic threshold for the colorectal cancer combined diagnosis model:
[0138] ## Setting levels: control = case, case = control
[0139] ## Setting direction: controls < case
[0140] The ROC curve was plotted with the predicted values in the test group, and the optimal diagnostic cut-off value of 0.5054961 was set according to the Youden index value. That is, when the predicted value of the diagnostic model ≤ 0.5054961, it is considered that the risk of colorectal cancer recurrence in the subject to be tested is low; when the model predicted value > 0.5054961, it is considered that the risk of colorectal cancer recurrence in the subject to be tested is high. The results are as Figure 4 shown: The AUC of the model in the test group was 0.945, the sensitivity was 99.2%, and the specificity was 98.6%.
[0141] 3. Validation of the combined diagnostic model for colorectal cancer
[0142] The constructed optimal model was validated in the validation group, and the ROC curve was plotted as Figure 5 shown. The AUC of the model in the validation group was 0.937, the sensitivity was 98.4%, and the specificity was 97.6%, which was very close to the diagnostic effect in the test group. It can be seen that the colorectal cancer recurrence risk prediction model constructed by the present invention using 6 protein markers has good prediction performance and accuracy, and has the best diagnostic efficacy.
[0143] Example 3. Construction and validation of a three-class diagnostic model
[0144] In this example, an attempt was made to construct a three-class combined diagnostic model for distinguishing the non-recurrence group, in-situ recurrence group, and distant metastasis group of colorectal cancer prognosis. The specific process includes the following: (1) Construction and screening of the optimal diagnostic model; (2) Verification of the effect of the optimal diagnostic model. The specific screening process and results are as follows (in the present invention, the AUC value was used as the evaluation index for the binary classification model in Example 2; when constructing a three-class model, since multiple categories are involved, the AUC value is usually not applicable, and in this example, indicators such as sensitivity, specificity, accuracy, and consistency were used to measure the diagnostic efficacy of the model):
[0145] I. Construction and screening of the diagnostic model
[0146] For a test cohort of 580 patients with colorectal cancer (patients with stage I / II CRC), all enrolled patients signed an informed consent form. Among them, 87 patients with stage I / II CRC had recurrence within three years after surgery, including 38 cases of local in-situ recurrence and 49 cases of distant metastasis (including 27 cases of liver metastasis, 13 cases of lung metastasis, 7 cases of peritoneal metastasis, and 2 cases of bone metastasis). Thus, they were divided into two groups: 493 patient samples without recurrence were classified into the low-risk group of colorectal cancer recurrence, and 87 cases with recurrence or metastasis were classified into the high-risk group of colorectal cancer recurrence. They were randomly divided into a test group and a validation group. The test group included 300 patients in the low-risk group of colorectal cancer recurrence and 56 patients in the high-risk group of colorectal cancer recurrence, including 25 cases of local in-situ recurrence and 31 cases of distant metastasis (including 18 cases of liver metastasis, 8 cases of lung metastasis, 4 cases of peritoneal metastasis, and 1 case of bone metastasis); the validation group included 193 patients in the low-risk group of colorectal cancer recurrence and 31 patients in the high-risk group of colorectal cancer recurrence, including 13 cases of local in-situ recurrence and 18 cases of distant metastasis (including 9 cases of liver metastasis, 5 cases of lung metastasis, 3 cases of peritoneal metastasis, and 1 case of bone metastasis with in-situ recurrence). In this example, based on the combination form of markers CD36+RNASE1+PLIN1+ITIH3+MGAM2+SELL screened in Example 2, a three-classification detection model was further constructed to effectively distinguish the non-recurrence group (low-risk group), in-situ recurrence group (in-situ recurrence group), and distal metastasis group (metastasis group) of colorectal cancer prognosis. All enrolled patients signed an informed consent form. Among them, all colorectal cancer patients were diagnosed by pathological histology. Inclusion criteria: (a) No history of other malignant tumors; (b) No patients with combined other malignant tumors or autoimmune diseases.
[0147] In this example, LC-MS / MS data collection and detection were performed on the collected serum samples, and the concentrations of six protein markers, namely CD36, RNASE1, PLIN1, ITIH3, MGAM2, and SELL, were obtained respectively.
[0148] The Shapiro-Wilk test was used to evaluate the normal distribution, and the non-parametric Wilcoxon test was used to analyze the differences in blood biomarker concentrations between the colorectal cancer prognosis non-recurrence group (low-risk group), the colorectal cancer prognosis in-situ recurrence group (in-situ recurrence group), and the colorectal cancer prognosis distant metastasis group (metastasis group). A three-class combined diagnostic model of 6 biomarkers was constructed using a combination of machine learning methods. The predicted probability value was used to estimate the area under the receiver operating characteristic (ROC) curve (AUC) with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate diagnostic model. Using the test group, the Youden index (YI) was calculated to determine the predicted probability cut-off value for distinguishing the low-risk group, in-situ recurrence group, and metastasis group. In addition, the ROCs of individual biomarkers and different subgroups were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R 3.6.1, and a p-value less than 0.05 was considered statistically significant.
[0149] In this embodiment, in order to construct an optimal three-class combined diagnostic model, after comparing the models constructed by 6 algorithms including gradient boosting, naive Bayes, support vector machine, neural network, generalized linear, and discriminant analysis, the gradient boosting method was selected as the best supervised classification algorithm for constructing the prediction model. The grid search range for hyperparameter optimization of the gradient boosting method is shown in Table 7 below.
[0150] Table 7. Parameter grid search range of the gradient boosting method
[0151]
[0152] Through optimization and screening in terms of accuracy, consistency, sensitivity, specificity, etc., the optimal parameter combination mode was determined as: interaction.depth 2, n.trees 150, shrinkage 0.1, n.minobsinnode 10.
[0153] Completely different two batches of samples were used for the test group and the validation group. In this embodiment, only the biomarkers were screened and the model was constructed from the test group; the samples of the validation group were only used to verify the diagnostic efficacy of the model. The specific results are shown in Table 8.
[0154] Table 8. Performance evaluation table for the model constructed by the gradient boosting method to distinguish three classes
[0155]
[0156] As can be seen from Table 8, the gradient boosting model constructed based on six protein markers, namely CD36, RNASE1, PLIN1, ITIH3, MGAM2, and SELL, can be used to predict whether colorectal cancer patients will have no recurrence, in-situ recurrence, or distant metastasis after surgery. It also shows that the protein markers screened in the present invention can be used to distinguish the recurrence risk of colorectal cancer patients after surgical treatment, and can also be used to distinguish whether it is in-situ recurrence or distant metastasis when the recurrence risk is high (when there is both in-situ recurrence and distant metastasis, it is also classified into the distant metastasis group).
[0157] II. Combined Performance of the Three-Class Classification Joint Diagnosis Model
[0158] In order to further improve the diagnostic value of the three-class classification diagnostic model (gradient boosting) constructed by biomarker combinations of different proteins, in this embodiment, based on the 10 protein markers screened in Example 1, the performance of the diagnostic models constructed by biomarker combinations of different proteins was compared in the test group. The specific combination forms of different models are shown in Table 9 below.
[0159] Table 9. Combination Forms of Different Diagnostic Models
[0160]
[0161] The results are specifically as Figure 6 shown in and Table 10. Table 10 shows the comparison results of the performance indicators of different diagnostic models constructed by the 10 biomarkers screened in Example 1 for three-class classification. The calculation methods for the minimum value, first quartile, median, mean, third quartile, and maximum value of accuracy and consistency are as follows: (1) Sort the values of accuracy or consistency from smallest to largest; (2) Minimum value: The first value after sorting; (3) First quartile (Q1): Multiply the number of data by 0.25. If the result is an integer, take the average of the values at this position and the next position; if not, round up to get the position, and the value at this position is Q1; (4) Median: If the number of data is odd, the median is the middle value; if it is even, it is the average of the two middle values; (5) Mean: The sum of all values divided by the number of data; (6) Third quartile (Q3): Multiply the number of data by 0.75, and the processing method is the same as Q1; (7) Maximum value: The last value after sorting. Among them, the minimum value and the maximum value can reflect the extreme situations of the data, showing the worst and best performances that the model may exhibit; the quartiles can help understand the distribution range and dispersion degree of the data; below Q1 represents a lower performance level, and above Q3 represents a higher performance level; the median can reflect the performance at the middle level; the mean comprehensively reflects the overall average performance. Combining the above statistical values, the overall situation, distribution characteristics, and stability of the model performance can be comprehensively understood, thus providing a strong basis for model selection and optimization.
[0162] Table 10. Performance Comparison of Diagnostic Models Constructed Based on Different Protein Combination Biomarkers
[0163]
[0164] As can be seen from Table 10, for the three-class diagnostic model, the nine-marker combined detection model (9MP) composed of nine markers has the best performance. This also clearly shows that on the basis of the six protein biomarkers of CD36, RNASE1, PLIN1, ITIH3, MGAM2, and SELL, the continued addition of three markers, namely ORM2, TFF1, and TFF3, has a very obvious improvement effect on the diagnostic efficiency of distinguishing whether colorectal cancer patients will have no recurrence, in-situ recurrence, or distant metastasis after surgical prognosis. Therefore, the three-class gradient boosting model constructed using these nine protein biomarkers (CD36 + RNASE1 + PLIN1 + ITIH3 + MGAM2 + SELL + ORM2 + TFF1 + TFF3) is used as the best combined diagnostic model.
[0165] III. Determination and Verification of the Diagnostic Performance of the Three-Class Combined Diagnostic Model
[0166] 1. Determination of the Diagnostic Performance of the Three-Class Combined Diagnostic Model
[0167] In order to more accurately determine the diagnostic performance and threshold of the model constructed in this embodiment for different disease classifications, a multi-class model with the gradient boosting (gbm) algorithm is used for predictive analysis in the test group, and the predicted probability values for the three-classification (low-risk group, in-situ recurrence group, and metastasis group) are calculated. The classification with the largest predicted probability value is the final prediction result of the system.
[0168] Among them, the meanings and calculation methods of each index are as follows:
[0169] Calculation results: The accuracy of the three-class combined diagnostic model in the test group is 0.84, and the consistency is 0.83. The diagnostic sensitivity for the low-risk group is 96.5%, and the specificity is 97.9%; the diagnostic sensitivity for the in-situ recurrence group is 82.4%, and the specificity is 81.5%; the diagnostic sensitivity for the metastasis group is 84.7%, and the specificity is 89.4%.
[0170] It should be noted that the three-class combined diagnostic model constructed by gradient boosting is a model constructed by machine learning and cannot fit a specific equation formula like a generalized linear model.
[0171] 2. Verification of the Three-Class Combined Diagnostic Model
[0172] Based on the model constructed from the test group, the predictive performance is verified in the validation group, and the specific results are as follows:
[0173] The accuracy is 0.83 and the consistency is 0.83. The diagnostic sensitivity for the low-risk group is 97.2%, and the specificity is 96.8%; the diagnostic sensitivity for the in-situ recurrence group is 83.9%, and the specificity is 80.4%; the diagnostic sensitivity for the metastasis group is 83.8%, and the specificity is 84.6%.
[0174] In summary, the three-class combined diagnostic model constructed in this embodiment, which contains 9 protein markers, has good diagnostic value for the three classifications of the low-risk group, the in-situ recurrence group, and the metastasis group.
[0175] All patents and publications mentioned in the specification of the present invention indicate that these are publicly known technologies in the art and can be used in the present invention. All patents and publications cited herein are equally listed in the references, as if each publication were specifically and individually referenced. The present invention described herein can be implemented in the absence of any one or more elements, one or more limitations, where such limitations are not specifically stated. For example, in each example herein, the terms "comprising", "consisting essentially of", and "consisting of" can be replaced by any one of the remaining two terms. The so-called "a" herein only means "one", and does not exclude including only one, nor does it exclude including more than two. The terms and expressions used herein are for descriptive purposes and are not limiting, and there is no intention to indicate that the terms and explanations described herein exclude any equivalent features, but it can be understood that any suitable changes or modifications can be made within the scope of the present invention and the claims. It can be understood that the embodiments described in the present invention are all preferred embodiments and features, and any person of ordinary skill in the art can make some changes and variations based on the essence described in the present invention, and these changes and variations are also considered to be within the scope of the present invention and the scope limited by the independent claims and the dependent claims.
Claims
1. Use of a marker for preparing a reagent for predicting the recurrence risk of colorectal cancer, characterized in that, The biomarker includes any one or more of CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, and TFF3.
2. The use according to claim 1, wherein, The biomarker includes CD36, PLIN1, ITIH3, MGAM2, RNASE1, and SELL.
3. The use according to claim 2, characterized in that, The recurrence risk refers to recurrence within three years after colorectal cancer treatment; the recurrence of colorectal cancer includes in-situ or adjacent area recurrence of colorectal cancer, and metastasis of colorectal cancer.
4. The use according to claim 2, characterized in that, The reagent is used to detect the content of biomarkers in a body fluid sample; the body fluid sample includes any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.
5. The use according to claim 2, characterized in that, The reagent is used to detect the presence, relative abundance, or concentration of biomarkers in a body fluid sample.
6. A kit for predicting the recurrence risk of colorectal cancer, characterized in that, A detection reagent for biomarkers including the use according to any one of claims 1 to 5.
7. A biomarker combination for predicting the recurrence risk of colorectal cancer, characterized in that, The combination includes CD36, PLIN1, ITIH3, MGAM2, RNASE1, and SELL.
8. A system for predicting the recurrence risk of colorectal cancer, characterized in that, The system includes a data analysis module, and the data analysis module is used to analyze the detection values of biomarkers, and the biomarkers include any one or more of CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, and TFF3.
9. The system according to claim 8, wherein The biomarker includes CD36, PLIN1, ITIH3, MGAM2, RNASE1, and SELL; the recurrence risk refers to recurrence within three years after colorectal cancer treatment; the recurrence of colorectal cancer includes in-situ or adjacent area recurrence of colorectal cancer, and metastasis of colorectal cancer.
10. The system according to claim 9, wherein The data analysis module uses the detection values of biomarkers of known samples as a training set, and is divided into a colorectal cancer recurrence group and a non-recurrence group of colorectal cancer according to whether colorectal cancer recurs, analyzes the relationship between the detection values of the colorectal cancer recurrence group and the non-recurrence group of colorectal cancer, constructs a model, and the equation of the model is: , where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 6), Xi represents the measured value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant 1.6999441; the coefficients of the biomarkers are as follows: , When Y ≤ 0.5, the recurrence risk of colorectal cancer in the subject to be tested is low; when Y > 0.5, the recurrence risk of colorectal cancer in the subject to be tested is high.
Citation Information
Patent Citations
Colorectal cancer detection model construction method and system and biomarker
CN116519954A
Biomarker combination and application thereof in prediction and / or diagnosis of colorectal cancer
CN117587128A
Colorectal cancer marker and application thereof
CN118048455A
Product for diagnosing colorectal cancer based on proteomics and application
CN119757585A
Methods and biomarkers for analysis of colorectal cancer
US20140302100A1