Biomarkers for predicting the risk of colorectal cancer recurrence and their applications
By screening and constructing a multi-marker joint detection model, using CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, TFF3 and other proteins, the problem of inaccurate diagnosis of colorectal cancer recurrence risk in the prior art was solved, and efficient and accurate prediction of colorectal cancer recurrence risk was achieved, and the ability to identify high-risk patients in the early stage was provided.
Patent Information
- Application Number
- CN202510724976.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The lack of high-sensitivity proteomic markers in the prior art is used to diagnose the risk of recurrence of colorectal cancer, resulting in inaccurate prediction of recurrence risk in patients after surgery, and the inability to effectively distinguish between high-risk and low-risk groups. The detection process is cumbersome and the degree of instrument dependence is high.
Proteins such as CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, TFF3 and other proteins were screened as biomarkers by proteomics. Combined with high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS/MS) and orthogonal partial least squares discriminant analysis, a multimarker joint detection model was constructed to predict the risk of recurrence of colorectal cancer.
It achieves accurate, non-invasive and efficient prediction of the risk of recurrence of colorectal cancer, reduces the risk of misdiagnosis and missed diagnosis, can identify people at high risk of recurrence early, help doctors adjust treatment plans, and prolong survival.
Smart Images

Figure CN120254282B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of proteomics screening of colorectal cancer diagnostic markers, and in particular, to a biomarker for predicting the risk of colorectal cancer recurrence and its application. Background Art
[0002] Colorectal cancer (CRC) is one of the most common malignant tumors in clinical practice.
[0003] Surgical resection is the primary means of achieving long-term survival for CRC patients, but postoperative recurrence or metastasis of colorectal cancer is the leading cause of death. Currently, the mainstay of CRC treatment is radical resection, particularly for stage I / II colorectal cancer. Radical resection is typically performed directly, with the primary goal of cure, without the need for extended clearance and generally without adjuvant chemoradiotherapy. However, recurrence is still possible in stage I / II CRC patients after surgical treatment. Risk prediction for patients who do not benefit from surgery or whose disease progresses, allowing for timely adjustment of treatment options (such as adjuvant chemoradiotherapy, secondary surgical resection, targeted therapy, or immunotherapy), can significantly improve overall survival and quality of life. Although some biomarkers for assessing the risk of recurrence in colorectal cancer prognosis have been reported, these are primarily based on genetic assessments, such as ctDNA and DNA methylation assays. These assessments are not only cumbersome and instrument-dependent, but also lack high diagnostic sensitivity. Clinically, biomarkers for diagnosing colorectal cancer recurrence risk are lacking, and the discovery of highly sensitive proteomic biomarkers for diagnosing colorectal cancer recurrence risk is of great significance.
[0004] Proteomics is the study of protein composition, localization, changes, and interactions within cells, tissues, or organisms, encompassing the study of protein expression patterns and proteome functional patterns. With the advancement of mass spectrometry, liquid chromatography coupled to mass spectrometry (LC-MS / MS) has become the predominant tool in proteomics research. The advancement of proteomics is crucial for identifying disease diagnostic markers, screening drug targets, and conducting toxicology studies, leading to its widespread application in medical research. Despite numerous publications and patents reporting on the discovery of tumor markers in recent years, these findings remain largely laboratory research, with limited clinical application and market penetration. Furthermore, in most cases, a single indicator is insufficient for in vitro diagnosis of tumor recurrence risk. Only a combination of multiple diagnostic tests, integrating multiple dimensions, can enhance predictive accuracy. Therefore, identifying new biomarkers for the diagnosis of colorectal cancer recurrence risk and constructing predictive models combining multiple markers holds significant clinical value. Summary of the Invention
[0005] In response to the problems existing in the prior art, the present invention provides a biomarker for predicting the risk of recurrence of colorectal cancer and its application. By using the proteomics method, by analyzing the proteins with significantly different abundance levels in the blood of two groups of people with colorectal cancer recurrence and non-recurrence after surgical treatment, biomarkers that can be used to predict the risk of recurrence of colorectal cancer are screened out, and a multi-marker combined detection model is further constructed, which can achieve accurate, non-invasive and efficient prediction of the risk of recurrence of colorectal cancer to meet clinical needs.
[0006] On the one hand, the present invention provides a use of a marker for preparing a reagent for predicting the risk of recurrence of colorectal cancer, wherein the marker includes any one or more of CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, and TFF3.
[0007] Predicting the risk of colorectal cancer recurrence is a completely different diagnostic process from predicting whether or not one has colorectal cancer. Our research group previously studied a large number of proteomic markers for predicting early colorectal cancer and constructed a model for predicting early colorectal cancer. However, when this model was applied to the risk of colorectal cancer recurrence, its diagnostic efficacy was poor and it was unable to directly distinguish patients at high risk of colorectal cancer recurrence, resulting in insufficient diagnostic sensitivity. Therefore, we are again screening for proteomic markers of colorectal cancer recurrence risk, hoping to find a more effective method for predicting the prognosis and recurrence risk of colorectal cancer.
[0008] The proteomic markers provided by the present invention can accurately predict the risk of recurrence of colorectal cancer patients after treatment, thereby enabling earlier identification of people at high risk of recurrence. This can help doctors predict in advance whether treatment plans need to be adjusted in a timely manner, reduce the risk of colorectal cancer recurrence or metastasis, prolong survival, and truly benefit colorectal cancer patients.
[0009] The present invention uses proteomics methods to collect plasma samples from patients who experience recurrence of colorectal cancer within a short period of time (for example, within 1 to 5 years) after treatment, as well as patients who do not experience recurrence within a short period of time. The different samples are analyzed using high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS). Based on orthogonal partial least squares discriminant analysis and significance analysis methods, proteins with significant differences between patients with recurrent colorectal cancer and patients without recurrence are first screened. Ultimately, 10 differential proteins with a significant correlation with the risk of colorectal cancer recurrence are screened out. These 10 proteins can be used to distinguish patients with recurrent colorectal cancer from patients without recurrence, have a certain diagnostic efficacy, and can thus be used to predict the risk of colorectal cancer recurrence.
[0010] Among them, the CD36 is a protein or amino acid sequence with a UniProt database number of P16671; PLIN1 is a protein or amino acid sequence with a UniProt database number of O60240; ITIH3 is a protein or amino acid sequence with a UniProt database number of Q06033; CD74 is a protein or amino acid sequence with a UniProt database number of P04233; MGAM2 is a protein or amino acid sequence with a UniProt database number of Q2M2H8; RNASE1 is a protein or amino acid sequence with a UniProt database number of P07998; SELL is a protein or amino acid sequence with a UniProt database number of P14151; TFF1 is a protein or amino acid sequence with a UniProt database number of P04155; ORM2 is a protein or amino acid sequence with a UniProt database number of P19652; and TFF3 is a protein or amino acid sequence with a UniProt database number of Q07654.
[0011] The present inventors also surprisingly discovered that some of the protein markers obtained through proteomic screening are already known markers for other cancers. For example, TFF3, previously reported for lung cancer prediction, was found in this screening to be also useful for predicting recurrence risk after colorectal cancer surgery. This indicates that the protein markers for many different cancers are not completely separate or unrelated; in fact, there are many cross-relationships or influences. Many protein markers can be used for both early cancer prediction and prognostic diagnosis, and even for prediction and diagnosis of multiple different cancers at different stages. Therefore, the field of proteomics has many new capabilities yet to be explored, and the market prospects are very broad.
[0012] Furthermore, the markers include CD36, PLIN1, ITIH3, MGAM2, RNASE1 and SELL.
[0013] To improve the diagnostic efficacy of colorectal cancer recurrence risk, it is necessary to combine different differential proteins to construct a diagnostic model. These differential proteins are ranked according to their importance, and different numbers of the top-ranked differential proteins are selected for combination. Ultimately, six protein markers were screened. The model constructed based on these six protein markers has good risk prediction capabilities in the diagnosis of colorectal cancer recurrence risk. Data from testing clinical colorectal cancer samples showed that the AUC value of predicting colorectal cancer recurrence risk using only these six biomarkers can reach 0.945, which is significantly better than the effect of combining multiple existing biomarkers to predict colorectal cancer recurrence risk.
[0014] Furthermore, the recurrence risk refers to recurrence within three years after colorectal cancer treatment; the colorectal cancer recurrence includes colorectal cancer recurrence in situ or adjacent areas, and colorectal cancer metastasis.
[0015] It can be understood that the "within three years" in "recurrence within three years after colorectal cancer treatment" here is not an absolute and unchanging time node, but a time node currently summarized based on clinical experience. With the passage of time or the improvement of other treatment methods, this time node or the length of time may change. For example, colorectal cancer may recur after treatment within one year, within 360 days, or within half a year, 180 days, or within two years, within two and a half years, within three and a half years, etc.
[0016] In some methods, the colorectal cancer recurrence refers to the recurrence of colorectal cancer in situ or in adjacent areas within three years after radical tumor resection, or the occurrence of colorectal cancer metastasis, such as liver metastasis, lung metastasis, peritoneal metastasis, bone metastasis, brain metastasis, etc.
[0017] Furthermore, the reagent is used to detect the content of biomarkers in a body fluid sample; the body fluid sample includes any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.
[0018] In some embodiments, the reagent for predicting the risk of colorectal cancer recurrence is a detection reagent prepared with the biomarker as the detection target, such as sample pretreatment reagents, antigens or antibodies, and other biological reagents and kits suitable for the detection of the biomarker; it can also be developed into a standardized reagent or kit suitable for the biomarker.
[0019] Furthermore, the reagent is used to detect the presence or relative abundance or concentration of biomarkers in a body fluid sample.
[0020] The present invention uses blood screening to identify biomarkers that predict the risk of colorectal cancer recurrence. These biomarkers show significant differences in the blood of people at high risk of colorectal cancer recurrence and people at low risk of colorectal cancer recurrence. By collecting blood samples, these biomarkers in the individual's blood can be detected to predict or assist in diagnosing the possibility of colorectal cancer recurrence in the individual, or these biomarkers in the blood of a certain group can be detected, and then the group can be divided into people at high risk of colorectal cancer recurrence and people at low risk of colorectal cancer recurrence.
[0021] Furthermore, the detection method includes a radiometric method, an immunological method, a fluorescence method, a flow cytometry method, a latex turbidimetry method, a biochemical method, an enzymatic method, a hybridization method, a gas chromatography-mass spectrometry method, a liquid chromatography-mass spectrometry method, a chromatography method, a chemiluminescence method, a magnetoelectric method or a photoelectric conversion method.
[0022] The presence or absence of a marker, or the level of a marker, is a relative concept. For example, when comparing a high-risk group for colorectal cancer recurrence with a low-risk group, the levels of these specific markers are compared relative to the baseline of the high-risk and low-risk groups. For some markers, the high-risk group may have higher levels than the low-risk group, and this increase may be statistically significant, such as a significant or highly significant increase. Therefore, when assessing the presence of a single marker, if the probability of a particular risk increases, the marker's level may change. This change may be a relative increase or decrease, and such a relative increase or decrease may be considered significant, or even highly significant. Therefore, regardless of the testing method, a predetermined cutoff value can be used as a standard. A value above this cutoff value is considered a change in the level, and such a result can be used for prognostic or diagnostic purposes.
[0023] Therefore, in some aspects, the markers described herein can be obtained by detecting the marker content in a sample using any known method, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strips, immunoreaction chips, capillary electrophoresis, infrared spectroscopy, and the like. As long as the protein marker content in a sample can be detected, it can be used to diagnose high-risk and low-risk groups for colorectal cancer recurrence. As long as the protein marker content in a sample can be detected, it can be used to predict or diagnose the probability of a certain disease. It will be understood that the detection here involves testing an individual sample, then comparing it with a pre-set standard, and using the comparison results to determine or predict the disease status. For example, it can be used to predict the probability of colorectal cancer recurrence. This prediction or diagnosis is based on whether or not colorectal cancer recurrence will occur within a certain period of time. Of course, such detection can be continuous, and the progression of the disease can be inferred as the content of certain substances changes.
[0024] In some embodiments, the relative abundance is the peak area of the biomarker in a detection spectrum obtained by high-performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a biomarker measured in a control sample is 100 and the average peak area measured in a high-risk group for colorectal cancer recurrence is 600, then the abundance of the biomarker in the sample is considered to be 6 times that in the control sample.
[0025] Furthermore, the colorectal cancer patient is a stage I / II CRC patient.
[0026] In another aspect, the present invention provides a kit for predicting the risk of recurrence of colorectal cancer, wherein the kit comprises a detection reagent for the biomarker for the purpose described above.
[0027] In another aspect, the present invention provides a biomarker combination for predicting the risk of recurrence of colorectal cancer, wherein the biomarker combination includes CD36, PLIN1, ITIH3, MGAM2, RNASE1 and SELL.
[0028] On the other hand, the present invention provides a system for predicting the risk of recurrence of colorectal cancer, the system comprising a data analysis module, the data analysis module being used to analyze the detection values of markers, the markers comprising any one or more of CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, and TFF3.
[0029] Furthermore, the markers include CD36, PLIN1, ITIH3, MGAM2, RNASE1 and SELL; the recurrence risk refers to recurrence within three years after surgical treatment of colorectal cancer; the colorectal cancer recurrence includes colorectal cancer recurrence in situ or adjacent areas, and colorectal cancer metastasis.
[0030] Furthermore, the data analysis module uses the detection values of markers of known samples as a training set, and divides them into a colorectal cancer recurrence group and a colorectal cancer non-recurrence group according to whether the colorectal cancer recurs, analyzes the relationship between the detection values of the colorectal cancer recurrence group and the colorectal cancer non-recurrence group, and constructs a model.
[0031] In some embodiments, a combined diagnostic model for predicting the risk of colorectal cancer recurrence is constructed by combining multiple machine learning methods, and it is preliminarily confirmed that the concentration changes of any one of the screened biomarkers alone can be used to distinguish between people at high risk and low risk of colorectal recurrence, indicating that these biomarkers have extremely high diagnostic value.
[0032] In some embodiments, the equation of the constructed model is:
[0033]
[0034] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m=6), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant of 1.6999441. The coefficients of the six biomarkers are:
[0035]
[0036] When Y≤0.5, the risk of colorectal cancer recurrence of the subject is low; when Y>0.5, the risk of colorectal cancer recurrence of the subject is high.
[0037] Furthermore, the system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0038] Furthermore, the detection value is the presence or absence, relative abundance or concentration value of each biomarker.
[0039] On the other hand, the present invention provides a use of a marker for preparing a reagent for predicting whether a colorectal cancer patient will not relapse after surgery, relapse in situ, or have distant metastasis, wherein the marker comprises any one or more of CD36, PLIN1, ITIH3, MGAM2, RNASE1, SELL, ORM2, TFF1, and TFF3.
[0040] The present invention attempts to apply protein markers used to predict recurrence risk to differentiate between patients with primary recurrence and those with distant metastasis. It was found that each protein marker can be used to distinguish between patients with colorectal cancer who have no recurrence after surgery, those with primary recurrence, and those with distant metastasis. The term "distant metastasis" encompasses patients with only distant metastasis as well as those with both primary recurrence and distant metastasis.
[0041] When a three-classification combined diagnostic model was constructed using six markers, CD36, PLIN1, ITIH3, MGAM2, RNASE1, and SELL, the accuracy of distinguishing between patients with in situ recurrence and patients with distant metastasis reached about 70%.
[0042] By adding ORM2, TFF1, and TFF3 to the six markers, a three-category combined diagnostic model containing nine protein markers was constructed. The diagnostic efficacy of this model, which distinguishes between primary recurrence and distant metastasis in colorectal cancer patients after surgery, can be further improved, reaching an accuracy of approximately 84%. This model can be directly used to predict whether colorectal cancer patients will not relapse after surgery, will relapse in situ, or will experience distant metastasis.
[0043] Furthermore, the reagent is used to predict whether colorectal cancer patients will not relapse, relapse in situ or have distant metastasis within three years after surgery.
[0044] Furthermore, the reagent is used to detect the content of biomarkers in a body fluid sample; the body fluid sample includes any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.
[0045] Furthermore, the reagent is used to detect the presence or relative abundance or concentration of biomarkers in a body fluid sample.
[0046] In another aspect, the present invention provides a kit for predicting whether colorectal cancer patients will not relapse after surgery, will relapse in situ, or will have distant metastasis. The kit comprises a detection reagent for the biomarker for the purpose described above.
[0047] In another aspect, the present invention provides a biomarker combination for predicting whether colorectal cancer patients will not relapse after surgery, relapse in situ, or have distant metastasis, wherein the combination includes CD36, PLIN1, ITIH3, MGAM2, RNASE1, SELL, ORM2, TFF1, and TFF3.
[0048] On the other hand, the present invention provides a system for predicting whether colorectal cancer will not recur after surgery, recur in situ, or metastasize to distant sites. The system includes a data analysis module, which is used to analyze the detection values of markers, and the markers include any one or more of CD36, PLIN1, ITIH3, MGAM2, RNASE1, SELL, ORM2, TFF1, and TFF3.
[0049] Furthermore, the distant metastasis also includes simultaneous in situ recurrence and distant metastasis. As long as distant metastasis occurs, it will be classified into the distant metastasis group.
[0050] Furthermore, the data analysis module uses the detection values of markers of known samples as a training set, and divides colorectal cancer patients into a non-recurrence group, an in situ recurrence group, and a distant metastasis group according to their post-operative conditions. The relationship between the detection values of the non-recurrence group, the in situ recurrence group, and the distant metastasis group is analyzed to construct a model.
[0051] Furthermore, the model is constructed based on the gradient boosting algorithm.
[0052] Unlike generalized linear regression, the gradient boosting algorithm cannot output model formulas and cutoff values. All calculations are done directly by machine learning. The test values can be directly input into the software system to obtain the prediction results.
[0053] Furthermore, the system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0054] The beneficial effects of the present invention are:
[0055] 1. The present invention has screened 10 new biomarkers that can predict the risk of colorectal cancer recurrence and developed a new protein marker combination. It can effectively evaluate and diagnose patients at high risk of colorectal cancer recurrence, effectively distinguish between high-risk and low-risk groups for colorectal cancer recurrence, and more accurately identify patients at high risk of colorectal cancer recurrence. Compared with traditional detection methods, it reduces the risk of misdiagnosis and missed diagnosis, and provides strong support for early detection and intervention of the disease.
[0056] 2. The combined differential diagnosis model of six biomarkers constructed in the present invention is convenient and fast, and the test results are highly consistent with the clinical gold standard test results. At the same time, it significantly reduces the cost of predicting the risk of colorectal cancer recurrence and has good application prospects.
[0057] 3. Based on the biomarkers screened to predict the risk of colorectal cancer recurrence, a three-classification model was further constructed that can simultaneously distinguish between the non-recurrence group, the in situ recurrence group and the distant metastasis group, providing a more effective and accurate predictive diagnostic model.
[0058] Detailed description
[0059] (1) Diagnosis or testing
[0060] The diagnosis or detection here refers to the detection or testing of biomarkers in a sample, or the content of a target biomarker, such as the absolute content or relative content, and then the presence or amount of the target marker is used to indicate whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease. The meanings of diagnosis and detection here are interchangeable. The result of such a test or diagnosis cannot be directly used as a direct result of being ill, but is an intermediate result. If a direct result is obtained, other auxiliary means such as pathology or anatomy are required to confirm that the patient has a certain disease. For example, the present invention provides a variety of new biomarkers related to the risk of recurrence of colorectal cancer, and changes in the content of these markers are directly correlated with whether the patient belongs to a group at high risk of colorectal cancer recurrence.
[0061] (2) Association between markers, biomarkers, or differentially expressed proteins and the risk of colorectal cancer recurrence
[0062] The terms "marker," "biomarker," and "differential protein" have the same meaning in this invention. Association here refers to a direct correlation between the presence or change in the level of a biomarker in a sample and a specific disease. For example, a relative increase or decrease in the level indicates a higher likelihood of the individual having the disease compared to a healthy population.
[0063] The simultaneous presence of multiple markers in a sample, or the relative changes in their levels, indicate a higher likelihood of the individual having the disease compared to healthy individuals. This means that among marker types, some are strongly associated with disease, while others are weakly associated, or even unrelated to a particular disease. One or more markers with strong correlations can be used as diagnostic markers, while markers with weaker correlations can be combined with stronger markers to diagnose a disease, increasing the accuracy of test results.
[0064] For the numerous biomarkers in serum discovered by the present invention, these markers can be used to distinguish between people at high risk of colorectal cancer recurrence and those at low risk. The markers here can be used alone as single markers for direct detection or diagnosis. The selection of such markers indicates that the relative change in the content of the marker is strongly correlated with the risk of colorectal cancer recurrence. Of course, it is understandable that one or more markers with a strong correlation with the risk of colorectal cancer recurrence can be selected for simultaneous detection. It is normal to understand that in some ways, selecting biomarkers with strong correlation for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90% or 95% accuracy, which means that these markers can obtain intermediate values for diagnosing a certain disease, but it does not mean that a certain disease can be directly confirmed.
[0065] Of course, differentially expressed proteins with larger ROC values can also be selected as diagnostic markers. The so-called strength is generally calculated and confirmed using algorithms, such as the contribution rate or weight analysis of markers to colorectal cancer recurrence risk assessment. Such calculation methods can include significance analysis (p-value or FDR value) and fold change. Multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA), and orthogonal partial least squares discriminant analysis (OPLS-DA), and of course other methods such as ROC analysis are also included. Of course, other model prediction methods are also possible. When selecting specific biomarkers, the differentially expressed proteins disclosed in this invention can be selected, or other existing well-known marker combinations can be selected or combined to make predictions through model methods.
[0066] (3) Definition of disease terms
[0067] Colorectal cancer: also known as large intestinal cancer, refers to cancers that originate from the large intestinal epithelium, including colon cancer and rectal cancer. Adenocarcinoma is the most common pathological type, with squamous cell carcinoma being a rare occurrence. In my country, rectal cancer is the most common, followed by colon cancer (sigmoid colon, cecum, ascending colon, descending colon, and transverse colon). The treatment of colorectal cancer should be individualized. Appropriate treatment methods, including radical surgery, chemotherapy, targeted therapy, and radiotherapy, are selected based on the patient's age, constitution, pathological type of the tumor, and extent of invasion (staging). The development of colorectal cancer generally progresses through normal mucosal hyperplasia, advanced adenoma (malignant), and adenocarcinoma (malignant), and generally takes 5-10 years.
[0068] Colorectal cancer recurrence: refers to the phenomenon that after the patient completes radical treatment (surgical resection), after a period of clinical disease-free state (i.e., a period of no tumor signs), cancer cells or tumor lesions (also known as metastasis) appear again at the primary tumor site or other parts of the body. Recurrence may be due to residual cancer cells that were not completely eliminated during treatment, or micrometastases that have spread but not been detected. Colorectal cancer recurrence can occur months to years after treatment, but 80% of recurrences occur within 2-3 years after surgery, and the risk of recurrence decreases significantly after 5 years. Colorectal cancer metastasis refers to metastasis caused by lymphatic, hematogenous or direct infiltration, including liver metastasis, lung metastasis, peritoneal metastasis, bone metastasis, brain metastasis, etc., which can be divided into oligometastasis (1 to 3 lesions), extensive metastasis, etc.
[0069] Early detection and intervention of colorectal cancer recurrence are key to improving the prognosis of colorectal cancer. Identifying biomarkers with a certain warning function in the early stages of colorectal cancer, before it recurs, and diagnosing the risk of colorectal cancer recurrence will be of great significance for improving patient treatment outcomes and prognosis.
[0070] Stage I colorectal cancer means that the tumor is confined to the intestinal wall, the tumor invades the submucosa but does not reach the muscularis, there is no lymph node metastasis, and there is no distant metastasis.
[0071] Stage II colorectal cancer means that the tumor has invaded the entire layer of the intestinal wall or surrounding tissues, the tumor has invaded the muscularis propria but has not penetrated the intestinal wall, there is no lymph node metastasis, and there is no distant metastasis.
[0072] (4) The gold standard for diagnosing colorectal cancer recurrence is histopathological examination (i.e., pathological confirmation of biopsy or surgical resection specimens), which observes the presence of cancer cells under a microscope and combines immunohistochemistry or molecular testing to clarify the nature of the tumor. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 This is a volcano plot of the differential analysis of protein markers in the high-risk group and low-risk group for colorectal cancer recurrence in Example 1;
[0074] Figure 2 Graphs showing the ROC and OPLS-DA analysis results for the high-risk and low-risk groups for colorectal cancer recurrence in Example 1;
[0075] Figure 3 AUC results for models constructed with different hyperparameters in Example 2;
[0076] Figure 4 : is the ROC curve of the combined diagnosis model in the test group in Example 2;
[0077] Figure 5 is the ROC curve of the combined diagnostic model in Example 2 in the validation group;
[0078] Figure 6 This is a diagram for evaluating the diagnostic performance of the three-class combined diagnosis model in Example 3. DETAILED DESCRIPTION
[0079] The present invention will be described in further detail below in conjunction with the accompanying drawings and Examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not serve to limit the present invention in any way. The reagents used in this example are all known products and were obtained by purchasing commercially available products.
[0080] Example 1: Screening for biomarkers of colorectal cancer recurrence risk using proteomics
[0081] Plasma samples were collected from patients undergoing radical surgery for colorectal cancer. Low-abundance proteins were enriched using immunoaffinity chromatography to remove high-abundance proteins. Protein abundance in the samples was detected using a tandem high-performance liquid chromatography-mass spectrometry device. The differences in protein abundance between patients with no colorectal cancer recurrence or metastasis within three years and those with colorectal cancer recurrence or metastasis within three years were analyzed, and the diagnostic performance was analyzed. The specific steps are as follows:
[0082] 1. Sample collection
[0083] Peripheral blood samples (approximately 2 ml) were collected from 180 patients with colorectal cancer (stage I / II CRC) 2 weeks after surgical resection. Blood samples were placed in vacuum tubes containing EDTA anticoagulant, mixed, and centrifuged twice at 120g for 10 minutes at room temperature. The supernatant was collected and then centrifuged at 360g for 20 minutes. Platelet samples were then collected in centrifuge tubes and stored at -80°C. The 180 patients included 102 males and 78 females with a mean age of 58 years (range, 26-84 years). Sixty-seven patients had stage I CRC and 113 had stage II CRC. All patients provided informed consent. Patients with colorectal cancer had histologically confirmed disease. Inclusion criteria included: (a) no history of other malignancies; and (b) no concurrent malignancies or autoimmune diseases.
[0084] For the next three years, patients were followed up every six months, with regular CD74, colonoscopy, and imaging examinations. Any recurrence or metastasis was confirmed by pathological histology. A total of 31 patients with stage I / II CRC relapsed within three years after surgery, including 18 with local in situ recurrence and 15 with distant metastasis (including 8 liver metastases, 6 lung metastases, and 1 peritoneal metastasis). Samples were then removed and stored at -80°C and divided into two groups: 149 patients without recurrence were classified as having a low risk of colorectal cancer recurrence, and 31 patients with recurrence or metastasis were classified as having a high risk of colorectal cancer recurrence for proteomic biomarker screening.
[0085] 2. Sample processing and enzymatic hydrolysis
[0086] First, plasma samples were centrifuged for 15 minutes at 15,000 g. The supernatant was filtered and subjected to immunoaffinity chromatography to isolate 14 highly abundant proteins. Low-abundance proteins were then concentrated to 350 μL using a 3 kDa cutoff concentrator at 4000 g for 1 hour. The recovered concentrate was then subjected to buffer exchange (AEX-A) using a 7 kDa cutoff desalting column at 1000 g for 2 minutes. Protein concentrations were determined using the BCA assay using AEX-A as a blank. According to sample grouping, 25 μL of TCEP was added to the samples, and the samples were incubated at 37°C for 30 minutes for protein reduction. TMT labeling was then performed by adding the corresponding TMT 16-plex reagent and incubating at room temperature in the dark for 1 hour. The sample was then buffer exchanged using a Zeba column with AEX-A. The TMT 16-plex labeled samples were mixed, and 2 mL of AEX-A was added to the mixed samples, bringing the final volume to 5.5 mL. The samples were filtered through a 0.22 µm filter and separated using a 2D-HPLC system. The collected fractions were freeze-dried, and finally, Trypsin-Lysin C enzyme cocktail was added. The samples were digested by incubation at 37°C for 5 hours, and the digestion reaction was terminated by the addition of 5 μL of 10% TFA. A total of 60 2D-HPLC fractions were used for nanoLC-MS / MS analysis.
[0087] 3. LC-MS / MS data acquisition and database analysis
[0088] DIA analysis was performed using a nanoflow Vanquish Neo system (Thermo Fisher Scientific). Samples separated by nano-HPLC were analyzed by DIA (data-independent) mass spectrometry on an Astral high-resolution mass spectrometer (Thermo Scientific). Detection mode: positive ionization, precursor ion scan range: 380-980 m / z, primary mass spectrometer resolution: 240,000 at 200 m / z, Normalized AGC Target: 500%, Maximum IT: 5 ms. MS2 acquisition mode: DIA, with 299 scan windows, an Isolation Window of 2 m / z, an HCD Collision Energy of 25 eV, a Normalized AGC Target of 500%, and a Maximum IT: 3 ms.
[0089] 4. Data Preprocessing
[0090] Secondary mass spectrometry data were retrieved using Maxquant (v1.6.15.0). The data type is DIA proteomics data based on secondary reporter ion quantification. The secondary spectrum used for quantification requires that the parent ion accounts for more than 75% in the primary spectrum. The database comes from the Homo_sapiens_9606_proteome_gene (release: 2021-10-14, sequence: 20,437) of the Uniprot database, and a common contamination library is added to the database. Contaminating proteins are deleted during data analysis; the enzyme cleavage method is set to Trypsin / P; the number of missed cleavage sites is set to 2; the parent ion mass error tolerance of the First search and Main search is set to 20 ppm and 5 ppm, respectively, and the mass error tolerance of the secondary fragment ion is 20 ppm. The fixed modification is cysteine alkylation, and the variable modification is methionine oxidation and acetylation of the protein N-terminus. The FDR for protein identification and PSM identification is set to 1%.
[0091] 5. Difference Analysis
[0092] A combination of univariate and multivariate statistical analyses was used to screen for differentially expressed proteins between high- and low-risk groups for colorectal cancer recurrence. Univariate analysis primarily included significance analysis (p-value or FDR value) and fold change analysis of signature molecules across different groups. Multivariate statistical analysis primarily included receiver operating characteristic (ROC) curve analysis and Boruta signature screening based on the random forest algorithm. All statistical analyses were performed using R. Detailed R information is provided in Table 1.
[0093] Table 1. R used in the present invention and related information
[0094]
[0095] The variable importance for the projection (VIP) was calculated to measure the influence and explanatory power of each protein expression pattern on the classification and discrimination of each group of samples. The Wilcoxon rank sum test was further performed to obtain the corrected p value (FDR). According to the conditions of FDR < 0.01 and Fold change > 2, 71 down-regulated proteins and 66 up-regulated proteins were screened (see Figure 1 ).
[0096] In order to evaluate the role of each protein marker in the diagnosis and prediction of colorectal cancer recurrence risk, this example used ROC and Boruta analysis methods to evaluate each protein marker. The results are shown in Figure 2 The horizontal axis represents the AUC obtained from ROC analysis, and the vertical axis represents the -log10 (FDR) calculated by the Wilcoxon test. The size of the dots represents the VIP value obtained from Boruta analysis. Further screening based on VIP > 3 and AUC > 0.6 identified 10 more significant candidate protein biomarkers, as detailed in Table 2.
[0097] Table 2. Differential markers between high-risk and low-risk groups for colorectal cancer recurrence
[0098]
[0099] Among them, the smaller the FDR value and / or the larger the VIP value, to a certain extent, it indicates that the difference in protein between the high-risk group and the low-risk group of colorectal cancer recurrence is more significant, and it also indicates that the protein may have a higher diagnostic value.
[0100] Example 2: Construction and validation of a colorectal cancer recurrence risk model
[0101] 1. Models constructed using combinations of different markers
[0102] While a single biomarker can differentiate the risk of recurrence after surgery for colorectal cancer, combining multiple biomarkers generally offers greater accuracy in differentiation or prediction. However, a single biomarker that more accurately predicts the risk of recurrence after surgery for colorectal cancer may not necessarily play a greater role in the combination when combined with one or more other biomarkers. Furthermore, the greater the number of biomarkers, the higher the predictive accuracy (AUC value) of the combination. Therefore, extensive validation experiments are still needed.
[0103] This example constructed and studied a model for the 10 protein markers screened in Example 1: CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, and TFF3. The study cohort consisted of 390 patients with colorectal cancer (stage I / II CRC) who underwent surgery two weeks after surgery. All enrolled patients provided informed consent. Fifty-eight patients with stage I / II CRC relapsed within three years of surgery, including 21 with local in situ recurrence and 37 with distant metastases (including 20 liver metastases, 9 lung metastases, 6 peritoneal metastases, and 2 bone metastases). The samples were then divided into two groups: 332 patients without relapse were assigned to a low-risk group for colorectal cancer recurrence, and 58 patients with recurrence or metastasis were assigned to a high-risk group for colorectal cancer recurrence. The samples were then randomly divided into a test group and a validation group, consisting of 166 patients in the low-risk group and 29 patients in the high-risk group, respectively.
[0104] In the test group, a combined diagnostic model of multiple protein markers was constructed using a combination of multiple machine learning methods. The predicted probability values were used to estimate the area under the receiver operator characteristic (ROC) curve (AUC) with a 95% confidence interval (CI) to evaluate the discriminatory ability of the multivariate diagnostic model. Using the test group, the Youden index (YI) was calculated to determine the cut-off value for the predicted probability of distinguishing the high-risk group for colorectal cancer recurrence from the low-risk group for colorectal cancer recurrence. In addition, the ROCs of single markers and different subgroups were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV) and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R3.6.1, and a p-value less than 0.05 was considered statistically significant.
[0105] The steps for building the joint diagnosis model are:
[0106] S101: Randomly select 2 to 10 marker concentration matrices from the 10 protein markers CD36, PLIN1, ITIH3, CD74, MGAM2, RNASE1, SELL, TFF1, ORM2, and TFF3 in the samples of the test group as the original training data set.
[0107] S102: Select the generalized linear model (glmnet) algorithm for constructing the prediction model, and the grid search range for optimizing the algorithm's hyperparameters. In this step, the grid search range for model hyperparameter optimization is set for each algorithm as shown in Table 3.
[0108] Table 3. Parameter grid of glmnet algorithm
[0109]
[0110] S103: According to the algorithm and hyperparameter setting range set in step S102, one of the hyperparameter combinations is selected as the parameters for constructing the prediction model.
[0111] S104: Split the original dataset into K subsets using a K-fold cross validation mechanism. To ensure that the ratio of majority class samples to minority class samples in each subset is the same as in the original dataset, a Stratified K-Folds cross validation mechanism is used for data segmentation.
[0112] S105 , according to the K training data subsets obtained by segmentation in step S104 , one of the subsets is selected as a validation set Ddev.
[0113] S106: Merge the training data subsets not selected in step S105 to form a training data pool Dtrain1.
[0114] S107 , building a prediction model based on the selected supervised classification algorithm and hyperparameters according to the training data set Dtrain obtained in step S106 .
[0115] S108: Evaluate the prediction model obtained in step S107 on the validation set Ddev to obtain an AUC value, and store the current prognosis prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 involves evaluating the prediction model obtained in step S107 on the validation set determined in the current iteration, and storing both the model and the evaluation results in the prediction model pool for future selection and use by the base prediction model. The evaluation mentioned in this step can be an AUC value or other reasonable metric for evaluating model performance.
[0116] S109: Determine whether all subsets have been used as validation sets. Step S109 determines whether all K subsets obtained in step S104 have been used as validation sets and trained on the model. If all subsets have been used as validation sets and training has been completed, proceed to step S110; if any subsets have not been used as validation sets, proceed to step S105. This step ensures that every sample in the original dataset has been used as a validation set, improving model stability and preventing overfitting of the model to a particular subset.
[0117] S110: The average AUC value of all models in the prediction model pool Pool is used as the final performance evaluation value of the combined model. The model parameters and the final performance evaluation AUC value are stored in the optimal model pool Poolbest.
[0118] S111: Determine whether all hyperparameter combinations have been used to construct prediction models. Step S111 determines whether prediction models have been constructed for all algorithms and corresponding hyperparameter combinations obtained in step S102. If all combinations have been used to construct models, step S112 is executed. If any combination has not been used to construct models, step S103 is executed.
[0119] S113 , selecting the model with the largest AUC value from the model set Poolbest obtained in step S112 as the final prediction model for colorectal cancer recurrence risk diagnosis.
[0120] S114, repeat all the above steps until all combinations of markers are modeled.
[0121] By executing the above model-building steps, we obtained the optimal models for all combinations of markers. To compare the performance of the models under these different marker combinations, we used the receiver operating characteristic (ROC) method to evaluate the AUC values of these models in the test group. The results are shown in Table 4.
[0122] Table 4. Comparison of the area under the ROC curve of the models constructed by different marker combinations in the test group
[0123]
[0124] Table 4 shows the ranking of the markers in Table 2 according to their highest VIP values and lowest FDR values. Starting with the top-ranked markers, marker combinations were selected sequentially. From 2MP to 5MP, the AUC values, accuracy, and sensitivity of the detection increased with the addition of more markers. However, when CD74 was added to 5MP, the diagnostic performance of the constructed model did not continue to improve, and the AUC value actually decreased. It is speculated that the addition of CD74 may have introduced noise. Therefore, CD74 was removed, and markers such as SELL were further added. Finally, it was found that the model constructed with the six-marker combination (6MP-2, CD36 + RNASE1 + PLIN1 + ITIH3 + MGAM2 + SELL) had a higher AUC than the model means of other marker combinations, even higher than 8MP, 9MP, and 10MP. Furthermore, the model with fewer markers was the most preferred.
[0125] 2. Optimization of model parameters
[0126] For the optimal marker combination CD36+ RNASE1+ PLIN1+ ITIH3+ MGAM2+ SELL, based on this marker combination, this example analyzed the models constructed under 9 different combinations of glmnet algorithm hyperparameters, and evaluated the model performance by AUC value (AUC was calculated using 10-fold cross-validation method during the modeling process). The results are shown in Table 5 and Figure 3 shown.
[0127] Table 5. AUC of the constructed model under different hyperparameter combinations of the glmnet algorithm
[0128]
[0129] According to Table 5, when the glmnet algorithm hyperparameter combination is alpha = 1, lambda = 0.0055, the AUC reaches the maximum value of 0.945.
[0130] The equation for building a model based on the optimal hyperparameter combination is:
[0131]
[0132] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 6), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker (Table 6), and b is a constant of 1.6999441.
[0133] Table 6. Coefficients of the six biomarkers in the model
[0134]
[0135] The complete model equation is:
[0136] Y=0.124 ITIH3+7.153 SELL+4.292 RNASE1+5.613 MGAM2+6.551 PLIN1+7.173CD36+ 1.6999441
[0137] Determination of diagnostic threshold of the combined diagnosis model for colorectal cancer:
[0138] ## Setting levels: control = case, case = control
[0139] ## Setting direction: controls < case
[0140] The ROC curve was drawn using the predicted values in the test group, and the optimal diagnostic cutoff value of 0.5054961 was set based on the Youden index value. That is, when the predicted value of the diagnostic model is ≤0.5054961, the risk of colorectal cancer recurrence in the patient is considered low; when the predicted value of the model is >0.5054961, the risk of colorectal cancer recurrence in the patient is considered high. Figure 4 As shown: The model has an AUC of 0.945, a sensitivity of 99.2%, and a specificity of 98.6% in the test group.
[0141] 3. Validation of the combined diagnostic model for colorectal cancer
[0142] The optimal model constructed was verified in the validation group and the ROC curve was drawn as follows Figure 5 As shown, the model achieved an AUC of 0.937, a sensitivity of 98.4%, and a specificity of 97.6% in the validation group, which is very close to the diagnostic performance in the test group. This indicates that the colorectal cancer recurrence risk prediction model constructed using six protein markers has good predictive performance and accuracy, and has the best diagnostic efficacy.
[0143] Example 3: Construction and verification of a three-category diagnostic model
[0144] This example attempts to construct a three-category combined diagnostic model for distinguishing between a colorectal cancer prognosis non-recurrence group, a colorectal cancer prognosis in situ recurrence group, and a colorectal cancer prognosis distant metastasis group. The specific process includes the following: (1) construction and screening of the optimal diagnostic model; (2) validation of the effectiveness of the optimal diagnostic model. The specific screening process and results are as follows (in the present invention, the two-category model in Example 2 uses the AUC value as the evaluation indicator; when a three-category model is constructed, since multiple categories are involved, the AUC value is generally not applicable. In this example, indicators such as sensitivity, specificity, accuracy, and consistency are used to measure the diagnostic efficacy of the model):
[0145] 1. Construction and screening of diagnostic models
[0146] A cohort of 580 patients with colorectal cancer (stage I / II CRC) was included, and all enrolled patients provided written informed consent. Among these patients, 87 stage I / II CRC patients experienced recurrence within three years after surgery, including 38 local recurrences in situ and 49 distant metastases (including 27 liver metastases, 13 lung metastases, 7 peritoneal metastases, and 2 bone metastases). They were divided into two groups: 493 patient samples without recurrence were classified as low-risk group for colorectal cancer recurrence, and 87 patients with recurrence or metastasis were classified as high-risk group for colorectal cancer recurrence. They were randomly divided into a test group and a validation group. The test group included 300 patients with low-risk group for colorectal cancer recurrence and 56 patients with high-risk group for colorectal cancer recurrence, including 25 patients with local recurrence in situ and 31 patients with distant metastasis (including 18 liver metastases, 8 lung metastases, 4 peritoneal metastases and 1 bone metastasis); the validation group included 193 patients with low-risk group for colorectal cancer recurrence and 31 patients with high-risk group for colorectal cancer recurrence, including 13 patients with local recurrence in situ and 18 patients with distant metastasis (including 9 liver metastases, 5 lung metastases, 3 peritoneal metastases and 1 bone metastasis with in situ recurrence). This example aims to build upon the marker combination CD36+RNASE1+PLIN1+ITIH3+MGAM2+SELL screened in Example 2 to further construct a three-category detection model that can effectively differentiate between the colorectal cancer prognosis group with no recurrence (low-risk group), the colorectal cancer prognosis group with in situ recurrence (in situ recurrence group), and the colorectal cancer prognosis group with distant metastasis (metastasis group). All enrolled patients provided informed consent. Patients with colorectal cancer were diagnosed histopathologically. Inclusion criteria included: (a) no history of other malignancies; and (b) no concurrent malignancies or autoimmune diseases.
[0147] In this example, LC-MS / MS data acquisition and detection were performed on the collected serum samples to obtain the concentrations of six protein markers: CD36, RNASE1, PLIN1, ITIH3, MGAM2, and SELL.
[0148] The Shapiro-Wilk test was used to assess normal distribution, and the nonparametric Wilcoxon test was used to analyze differences in blood marker concentrations between the colorectal cancer prognosis group without recurrence (low-risk group), the colorectal cancer prognosis group with in situ recurrence (in situ recurrence group), and the colorectal cancer prognosis group with distant metastasis (metastasis group). A three-class combined diagnostic model for the six markers was constructed using a combination of machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) with 95% confidence intervals (CI) was estimated using the predicted probability values to assess the discriminatory ability of the multivariate diagnostic model. Using the test set, the Youden index (YI) was calculated to determine the predicted probability cutoff value for distinguishing the low-risk group, in situ recurrence group, and metastasis group. In addition, ROCs for individual markers and different subgroups were constructed and compared. Standard descriptive statistics, such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD), were calculated to describe the experimental results of the study population. Statistical analysis was performed using R3.6.1, and p values less than 0.05 were considered statistically significant.
[0149] In this embodiment, in order to construct the optimal three-class joint diagnosis model, after comparing the six algorithms of gradient boosting, naive Bayes, support vector machine, neural network, generalized linear, and discriminant analysis, the gradient boosting method was selected as the best supervised classification algorithm for constructing the prediction model. The grid search range for hyperparameter optimization of the gradient boosting method model is shown in Table 7 below.
[0150] Table 7. Parameter grid search range of gradient boosting method
[0151]
[0152] Through optimization screening in terms of accuracy, consistency, sensitivity, specificity, etc., the optimal parameter combination mode was determined to be: interaction.depth 2, n.trees 150, shrinkage 0.1, n.minobsinnode 10.
[0153] The test and validation groups used two completely different batches of samples. This example only screened markers and constructed models from the test group; the samples from the validation group were only used to verify the diagnostic efficacy of the model. The specific results are shown in Table 8.
[0154] Table 8. Performance evaluation table of the gradient boosting method to build a model to distinguish three categories
[0155]
[0156] Table 8 shows that a gradient boosting model constructed based on six protein markers, CD36, RNASE1, PLIN1, ITIH3, MGAM2, and SELL, can be used to predict whether colorectal cancer patients will experience no recurrence, in situ recurrence, or distant metastasis after surgery. This also demonstrates that the protein markers screened by this invention can be used to distinguish the risk of recurrence in colorectal cancer patients after surgical treatment and, when the risk of recurrence is high, to distinguish between in situ recurrence and distant metastasis (patients with both in situ recurrence and distant metastasis are also classified as having distant metastasis).
[0157] Combined performance of two- and three-class joint diagnosis models
[0158] To further enhance the diagnostic value of three-category diagnostic models (gradient boosting) constructed using different protein combinations of biomarkers, this example compared the performance of diagnostic models constructed using different protein combinations of biomarkers in a test group, based on the 10 protein markers screened in Example 1. The specific combinations of the different models are shown in Table 9.
[0159] Table 9. Combinations of different diagnostic models
[0160]
[0161] The results are as follows Figure 6 As shown in Table 10, Table 10 is a comparison of the performance indicators of different diagnostic models constructed by the 10 biomarkers screened in Example 1 for three categories. The calculation method of the minimum value, first quartile, median, mean, third quartile and maximum value of accuracy and consistency is as follows: (1) Sort the accuracy or consistency values from small to large; (2) Minimum value: the first value after sorting; (3) First quartile (Q1): multiply the number of data by 0.25. If the result is an integer, take the average of the values at this position and the next position; if it is not an integer, round up to get the position, and the value at this position is Q1; (4) Median: if the number of data is odd, the median is the middle value; if it is even, it is the average of the two middle values; (5) Mean: the sum of all values divided by the number of data; (6) Third quartile (Q3): multiply the number of data by 0.75, and process it in the same way as Q1; (7) Maximum value: the last value after sorting. The minimum and maximum values reflect data extremes, demonstrating the worst and best possible model performance. Quartiles help understand the data's distribution and dispersion. Q1 and below indicate lower performance, while Q3 and above indicate higher performance. The median reflects intermediate performance, and the mean comprehensively reflects the overall average performance. By combining these statistical values, we can gain a comprehensive understanding of the overall performance, distribution characteristics, and stability of the model, providing a strong basis for model selection and optimization.
[0162] Table 10. Performance comparison of diagnostic models based on different protein combination biomarkers
[0163]
[0164] Table 10 shows that for the three-category diagnostic model, the nine-marker combined model (9MP) performed best. This clearly demonstrates that the addition of ORM2, TFF1, and TFF3 to the six protein markers of CD36, RNASE1, PLIN1, ITIH3, MGAM2, and SELL significantly improves the diagnostic efficacy of differentiating colorectal cancer patients from those with postoperative recurrence, in situ recurrence, or distant metastasis. Therefore, the three-category gradient boosting model constructed with these nine protein markers (CD36 + RNASE1 + PLIN1 + ITIH3 + MGAM2 + SELL + ORM2 + TFF1 + TFF3) was selected as the optimal combined diagnostic model.
[0165] 3. Diagnostic Performance Measurement and Validation of the Three-Classification Joint Diagnosis Model
[0166] 1. Diagnostic performance measurement of the three-classification joint diagnosis model
[0167] In order to more accurately determine the diagnostic performance and thresholds of the model constructed in this embodiment for different disease classifications, a multi-classification model of the gradient boosting (GBM) algorithm was used to perform predictive analysis in the test group, and the predicted results were calculated as the predicted probability values of the three categories (low-risk group, in situ recurrence group, and metastasis group). The category with the largest predicted probability value was the final prediction result of the system.
[0168] The meaning and calculation method of each indicator are as follows:
[0169] Results: The three-category combined diagnostic model had an accuracy of 0.84 and a consistency of 0.83 in the test group. The diagnostic sensitivity for the low-risk group was 96.5% and the specificity was 97.9%. The diagnostic sensitivity for the primary recurrence group was 82.4% and the specificity was 81.5%. The diagnostic sensitivity for the metastasis group was 84.7% and the specificity was 89.4%.
[0170] It should be noted that the three-classification joint diagnosis model constructed by gradient boosting is a model constructed by machine learning and cannot fit a specific equation formula like a generalized linear model.
[0171] 2. Validation of the three-classification joint diagnosis model
[0172] The prediction performance of the model built based on the test group was verified in the validation group. The specific results are as follows:
[0173] The accuracy was 0.83, and the consistency was 0.83. The diagnostic sensitivity for the low-risk group was 97.2%, and the specificity was 96.8%. The diagnostic sensitivity for the primary recurrence group was 83.9%, and the specificity was 80.4%. The diagnostic sensitivity for the metastasis group was 83.8%, and the specificity was 84.6%.
[0174] In summary, the three-category combined diagnostic model containing 9 protein markers constructed in this example has good diagnostic value for the three categories of low-risk group, in situ recurrence group and metastasis group.
[0175] All patents and publications cited in this specification are intended to indicate that they are state of the art and that the present invention may be used. All patents and publications cited herein are incorporated by reference in their entirety, as if each publication were specifically incorporated by reference. The invention described herein may be practiced in the absence of any element or elements, limitation or limitations, unless otherwise specified. For example, in each instance, the terms "comprising," "consisting essentially of," and "consisting of" may be replaced with either of the other two terms. The term "a" or "an" herein simply means "one" and does not exclude the inclusion of only one or more. The terms and expressions used herein are intended to be descriptive, not limiting, and are not intended to exclude any equivalent features. However, it is understood that any suitable changes or modifications may be made within the scope of the present invention and the appended claims. It is understood that the embodiments described herein are preferred embodiments and features, and that modifications and variations can be made by persons of ordinary skill in the art based on the spirit of the present invention. Such modifications and variations are considered to be within the scope of the present invention and the scope of the independent and appended claims.
Claims
1. Use of a reagent for detecting a marker for preparing a reagent for predicting whether a colorectal cancer patient will not relapse, relapse in situ, or metastasize within three years after surgery, characterized in that: The markers consist of CD36, PLIN1, ITIH3, MGAM2, RNASE1, SELL, ORM2, TFF1, and TFF3.
2. A kit for predicting whether colorectal cancer patients will not relapse, relapse in situ, or metastasize within three years after surgery, characterized in that: A detection reagent for a biomarker for use as claimed in claim 1.
3. A biomarker combination for predicting whether colorectal cancer patients will not relapse, relapse in situ, or metastasize within three years after surgery, characterized in that: The combination consists of CD36, PLIN1, ITIH3, MGAM2, RNASE1, SELL, ORM2, TFF1, and TFF3.
4. A system for predicting whether colorectal cancer patients will not relapse, relapse in situ, or metastasize within three years after surgery, characterized in that: The system includes a data analysis module, which is used to analyze the detection values of markers, and the markers are composed of CD36, PLIN1, ITIH3, MGAM2, RNASE1, SELL, ORM2, TFF1, and TFF3.
5. The system according to claim 4, wherein: The data analysis module uses the detection values of markers of known samples as a training set, and divides colorectal cancer patients into a non-recurrence group, an in situ recurrence group, and a distant metastasis group according to their post-operative conditions. The relationship between the detection values of the non-recurrence group, the in situ recurrence group, and the distant metastasis group is analyzed to construct a model; the model is constructed based on a gradient boosting algorithm; the system also includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output prediction results.
Citation Information
Patent Citations
Colorectal cancer detection model construction method and system and biomarker
CN116519954A
Product for diagnosing colorectal cancer based on proteomics and application
CN119757585A