A biomarker for predicting the risk of recurrence or metastasis of endometrial cancer and its application
By screening and constructing a prediction model based on biomarkers such as MUC16, CTSG, B3GNT2, CORO1A and SFTPB, the problem of insufficient diagnosis of endometrial cancer recurrence risk is solved, efficient and accurate risk prediction is achieved, and the treatment effect and quality of life of patients are improved.
Patent Information
- Application Number
- CN202510725313.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The lack of high sensitivity biomarkers in the prior art is used to predict the risk of recurrence of endometrial cancer, resulting in insufficient diagnosis of recurrence risk in patients after surgical treatment, affecting the adjustment of treatment plans and patient prognosis.
Biomarkers such as MUC16, CTSG, B3GNT2, CORO1A and SFTPB were screened out by proteomics, and patients' blood samples were analyzed in combination with high-performance liquid chromatography-tandem mass spectrometry technology to construct a predictive model to accurately predict the risk of recurrence or metastasis of endometrial cancer.
Accurate, non-invasive and efficient prediction of the risk of recurrence or metastasis of endometrial cancer is achieved, which improves the accuracy of diagnosis and the possibility of personalized treatment, reduces the risk of misdiagnosis and missed diagnosis, and improves the treatment effect and quality of life of patients.
Smart Images

Figure CN120254284B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedicine technology, and in particular to a biomarker for predicting the risk of recurrence or metastasis of endometrial cancer and its application. Background Art
[0002] Endometrial cancer (EC) is a cancer that develops from the malignant transformation of epithelial cells in the endometrium. It primarily originates from glandular cells in the endometrium. These cells proliferate and differentiate normally under the regulation of estrogen and progesterone. However, under certain pathological conditions, these cells may undergo genetic mutations, abnormal proliferation, and ultimately develop into malignant tumors.
[0003] Endometrial cancer is one of the most common gynecological malignancies of the female reproductive system, second only to cervical cancer. Its incidence is higher in developed countries, particularly among postmenopausal women. The average age of onset is around 60, but there has been a trend toward younger onset in recent years. Clinical manifestations include abnormal vaginal bleeding, specifically postmenopausal bleeding or menstrual irregularities such as increased menstrual flow and prolonged periods. Some patients may also experience vaginal discharge, which may be serous, bloody, or purulent.
[0004] The incidence and mortality of endometrial cancer are increasing annually, and recurrence is one of the main reasons for poor prognosis. Traditionally, prognosis assessment for EC is based primarily on clinicopathological parameters. However, endometrioid carcinoma, the most common pathological type, can show variability in prognosis even among patients with similar stage and grade based on clinicopathological parameters. Therefore, relying solely on clinicopathological parameters for prognostic assessment still has significant limitations.
[0005] Currently, the primary treatment for endometrial cancer is radical resection, especially for patients with stage I / II endometrial cancer. Radical resection is typically performed directly, with the primary goal of cure, without the need for extended clearance and generally without adjuvant chemoradiotherapy. However, approximately 5%-20% of patients with stage I / II endometrial cancer still experience recurrence after surgery. If risk prediction can be performed for patients who do not benefit from surgical treatment or whose disease progresses, and treatment options are adjusted promptly (such as adjuvant chemoradiotherapy, secondary surgical resection, targeted therapy, or immunotherapy), overall survival and quality of life can be significantly improved. Clinically, there is a lack of biomarkers for diagnosing the risk of endometrial cancer recurrence, and the discovery of highly sensitive proteomic biomarkers for diagnosing the risk of endometrial cancer recurrence is of great significance.
[0006] Proteomics is the study of protein composition, localization, changes, and interactions within cells, tissues, or organisms. This includes the study of protein expression patterns and proteome functional patterns. With the advancement of mass spectrometry, liquid chromatography coupled to mass spectrometry (LC-MS / MS) has become the primary tool in proteomics research. The development of proteomics is crucial for identifying disease diagnostic markers, screening drug targets, and conducting toxicology studies, leading to its widespread application in medical research.
[0007] Therefore, it is of great clinical value to find new diagnostic markers related to the risk of recurrence of endometrial cancer and to combine multiple markers to construct a prediction model. Summary of the Invention
[0008] In response to the problems existing in the prior art, the present invention provides a biomarker for predicting the risk of recurrence or metastasis of endometrial cancer and its application. By using proteomics methods, the differences in protein abundance in the blood of patients with recurrent and non-recurrent endometrial cancer after surgical treatment are analyzed to screen out biomarkers that can be used to predict the risk of recurrence of endometrial cancer, and further construct an endometrial cancer recurrence risk prediction model, which can accurately, non-invasively and efficiently predict the risk of recurrence or metastasis of endometrial cancer, provide a new tool for the management of endometrial cancer patients, and improve patients' treatment and prognosis monitoring.
[0009] A first aspect of the present invention provides use of a substance for detecting biomarkers in preparing a product for predicting the risk of recurrence or metastasis of endometrial cancer, wherein the biomarkers include one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB.
[0010] The present invention provides biomarkers obtained by proteomics that can accurately predict the risk of recurrence or metastasis of endometrial cancer, which helps doctors judge the severity of the patient's condition and whether the treatment plan needs to be adjusted, thereby providing patients with more accurate treatment recommendations and truly benefiting endometrial cancer patients.
[0011] The present invention utilizes a proteomics approach to collect plasma samples from patients who experience a short-term recurrence (e.g., within 1 to 5 years) of endometrial cancer treatment, as well as patients who do not experience a short-term recurrence. The different samples are analyzed using high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS). Based on orthogonal partial least squares discriminant analysis and significance analysis, proteins with significant differences between recurrent and non-recurrent endometrial cancer are first screened. Five differentially expressed proteins with a significant correlation with the risk of endometrial cancer recurrence are identified. These five proteins can be used to distinguish whether endometrial cancer patients will relapse after surgical treatment, demonstrating a certain diagnostic efficacy.
[0012] Among them, the MUC16 is the protein or amino acid sequence with the UniProt database number Q8WXI7; CTSG is the protein or amino acid sequence with the UniProt database number P08311; B3GNT2 is the protein or amino acid sequence with the UniProt database number Q9NY97; CORO1A is the protein or amino acid sequence with the UniProt database number P31146; SFTPB is the protein or amino acid sequence with the UniProt database number P07988.
[0013] The present inventors surprisingly discovered that some of the protein markers obtained through proteomic screening are already known markers for other cancers. For example, SFTPB, previously reported for predicting non-small cell lung cancer, was found in this screening to be also useful for predicting the risk of recurrence or metastasis in endometrial cancer. This demonstrates that the protein markers for many different tumors are not completely separate or unrelated; in fact, they exhibit numerous cross-relationships or influences. Many protein markers can be used for both early-stage cancer prediction and prognostic diagnosis, and even for predicting and diagnosing a wide range of cancers at different stages. Therefore, the field of proteomics holds many new capabilities yet to be explored, and the market prospects are vast.
[0014] In some embodiments, the biomarker can be one marker or a combination of several markers, such as a combination of two markers, a combination of three markers, a combination of four markers, or a combination of five markers.
[0015] In some specific embodiments, the biomarkers comprise a combination of two or more markers, such as a combination of three or more markers. When the combination of markers is used to construct a prediction model for the risk of endometrial cancer recurrence or metastasis, the AUC value is 0.713-0.945, the sensitivity is 72.5%-95.9%, and the specificity is 73.3%-96.1%.
[0016] Furthermore, the biomarkers include MUC16, CTSG, B3GNT2 and CORO1A.
[0017] To study the diagnostic efficacy of endometrial cancer recurrence risk, it was necessary to combine different differentially expressed proteins to construct a diagnostic model. The identified differentially expressed proteins were ranked according to their importance, and different numbers of the top-ranked differentially expressed proteins were selected and combined to obtain five protein markers. After verification, it was found that the model based on these five protein markers had good risk prediction capabilities in diagnosing endometrial cancer recurrence risk. Data from endometrial cancer recurrence risk samples showed that using only these four biomarkers to predict endometrial cancer recurrence risk achieved an AUC value of 0.945, demonstrating good diagnostic performance.
[0018] In some embodiments, the product is selected from one or more of a reagent, a kit, a chip, a probe, or a membrane strip. The product is a product that detects the biomarkers described above, and includes, for example, sample pretreatment reagents, antigens, or antibodies, and other biological reagents and kits suitable for detecting the biomarkers. The product can also be developed into standardized reagents or kits, chips, probes, or membrane strips suitable for detecting the biomarkers.
[0019] In some embodiments, the product is used to predict patients' risk of recurrence or metastasis of endometrial cancer.
[0020] Furthermore, the risk of recurrence or metastasis refers to the recurrence of endometrial cancer within three years after surgical treatment; the recurrence or metastasis of endometrial cancer includes recurrence of endometrial cancer in situ or adjacent areas, and metastasis of endometrial cancer.
[0021] It is understandable that the "within three years" in "recurrence or metastasis within three years after treatment of endometrial cancer" is not an absolute and unchanging time node. It is only a time node currently summarized based on clinical experience. With the passage of time or the improvement of other treatment methods, this time node or the length of time may change. For example, colorectal cancer may recur after treatment within one year, within 360 days, or within half a year, 180 days, or within two years, within two and a half years, within three and a half years, etc.
[0022] In some embodiments, the recurrence or metastasis of endometrial cancer refers to the recurrence of endometrial cancer in situ or in adjacent areas within three years after radical tumor resection, or the occurrence of endometrial cancer metastasis, such as liver metastasis, lung metastasis, peritoneal metastasis, bone metastasis, etc.
[0023] In some embodiments, the product is used to detect the amount of a biomarker in a biological sample.
[0024] In some specific embodiments, the biological sample is selected from one or more of saliva, blood, urine, plasma, serum and cerebrospinal fluid.
[0025] In some specific embodiments, the detected amount includes the presence or relative abundance or concentration of a biomarker.
[0026] The present invention uses blood screening to identify biomarkers that predict the risk of endometrial cancer recurrence. These biomarkers show significant differences in the blood of people with and without endometrial cancer recurrence. By collecting blood samples, these biomarkers can be detected in the individual's blood to predict or assist in diagnosing whether the individual has endometrial cancer recurrence or not.
[0027] Furthermore, the detection method generally includes a radiometric method, an immunological method, a fluorescence method, a flow cytometry, a latex turbidimetry method, a biochemical method, an enzymatic method, a hybridization method, a gas chromatography-mass spectrometry method, a liquid chromatography-mass spectrometry method, a chromatography method, a chemiluminescence method, a magnetoelectric method or a photoelectric conversion method.
[0028] The presence or absence of biomarkers, or the level of these biomarkers, is a relative concept. For example, when comparing endometrial cancer recurrence and non-recurrence groups, the levels of these specific biomarkers are compared relative to the baseline of the recurrence or non-recurrence groups. It's possible that certain biomarkers may be higher in recurrence than in the non-recurrence group, and this increase can be statistically significant, such as a significant or highly significant increase. Therefore, when assessing the presence of these biomarkers, if a single biomarker indicates an increased probability of a certain risk, its level may change. This change can be a relative increase or decrease, and this relative increase or decrease can be significant, or even highly significant. Therefore, regardless of the method used for detection, a predetermined value (cut-off value) can be used as a standard. A value above this value is considered a change in the level, and such a result can be used for prognostic or diagnostic purposes.
[0029] Therefore, in some aspects, the biomarkers described in the present invention can be obtained by detecting the marker content in a sample by any known method, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strips, immunoreaction chips, capillary electrophoresis, infrared spectroscopy, etc. As long as it can be used to detect the protein marker content in a sample, it can be used to diagnose the recurrence and non-recurrence of endometrial cancer. As long as the protein marker content in a sample can be detected, it can be used to predict or diagnose the probability of a certain disease. It is understood that the detection here is to detect an individual sample and then compare it with a pre-set standard. The comparison result is used to judge or predict the occurrence status of the disease. For example, it can be used to predict the probability of the risk of endometrial cancer recurrence. This prediction or diagnosis is whether it will occur within a certain period of time. Of course, such detection can be continuous detection, and the progress of the disease can be inferred as the content of certain substances changes.
[0030] In some embodiments, the relative abundance is the peak area of the biomarker in a detection spectrum obtained by high-performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a biomarker measured in a control sample is 300 and the average peak area measured in an endometrial cancer recurrence sample is 1800, then the abundance of the biomarker in the sample is considered to be 6 times that in the control sample.
[0031] A second aspect of the present invention provides the use of proteins as biomarkers in the preparation of products for predicting the risk of endometrial cancer recurrence or metastasis. The proteins include one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB. These MUC16, CTSG, B3GNT2, CORO1A, and SFTPB genes can be used to assist in determining the risk of endometrial cancer recurrence or metastasis, assess drug efficacy, and more. The inventors have discovered that MUC16, CTSG, B3GNT2, CORO1A, and SFTPB are closely associated with the risk of endometrial cancer recurrence or metastasis.
[0032] A third aspect of the present invention provides a product for predicting the risk of recurrence or metastasis of endometrial cancer, comprising a substance for detecting biomarkers, wherein the biomarkers include one or more of MUC16, CTSG, B3GNT2, CORO1A and SFTPB.
[0033] In some embodiments, the biomarkers include MUC16, CTSG, B3GNT2, and CORO1A.
[0034] A fourth aspect of the present invention provides a biomarker combination for predicting the risk of recurrence or metastasis of endometrial cancer, comprising MUC16, CTSG, B3GNT2, and CORO1A.
[0035] A fifth aspect of the present invention provides a method for constructing a risk prediction model for endometrial cancer recurrence or metastasis for purposes other than disease diagnosis, comprising the following steps:
[0036] 1) constructing a sample dataset based on the detection amount of biomarkers in the biological sample, wherein the biomarkers include one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB;
[0037] 2) dividing the data set into a test set and a training set, and constructing and training the endometrial cancer recurrence or metastasis risk prediction model through machine learning methods.
[0038] In some embodiments, the biomarkers include MUC16, CTSG, B3GNT2, and CORO1A.
[0039] In some embodiments, the machine learning method is selected from at least one of a gradient boosting algorithm, a random forest algorithm, a support vector machine algorithm, a decision tree algorithm, a K-nearest neighbor algorithm, a logistic regression algorithm, and a neural network algorithm. For example, the gradient boosting algorithm can be selected. Unlike generalized linear regression, which outputs a model formula and cutoff value, the gradient boosting algorithm performs all calculations directly through machine learning. Test values can be directly input into the software system to directly obtain prediction results.
[0040] In some embodiments, multiple machine learning methods are used to construct a risk prediction model for endometrial cancer recurrence or metastasis, and it is preliminarily confirmed that the concentration changes of any one of the screened new biomarkers can be used to distinguish between people with endometrial cancer recurrence and those without recurrence, indicating that these biomarkers have extremely high diagnostic value.
[0041] The present invention discovered that by detecting the detection amount of biomarkers in biological samples and then inputting the detection amount into the formula of the endometrial cancer recurrence or metastasis risk prediction model, the prediction score Y of the prediction model is obtained, and the prediction score Y is compared with the threshold (cut off) defined by the Youden Index. If the prediction score Y is greater than the threshold (cut off), it is judged that the endometrial cancer has recurred; if the prediction score y is less than or equal to the threshold (cut off), it is judged that the endometrial cancer has not recurred.
[0042] In some embodiments, the equation for the constructed endometrial cancer recurrence risk prediction model is:
[0043]
[0044] Wherein, Y is the prediction score, i represents the i-th biomarker, m represents the number of biomarkers (m=5), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant of 3.8638532; the coefficients of the four biomarkers are:
[0045]
[0046] When Y ≤ the cutoff value, the patient has no recurrence of endometrial cancer. When Y > the cutoff value, the patient has recurrence of endometrial cancer. Specifically, the cutoff value is 0.4773883.
[0047] In some embodiments, the method further includes step 3) testing the endometrial cancer recurrence or metastasis risk prediction model using a test set. After training, the trained prediction model is validated using the test set, and the effectiveness of the prediction model is evaluated using AUC, specificity, and sensitivity as evaluation metrics.
[0048] The prediction model constructed based on the combination of 4MP in the present invention can distinguish between recurrence and non-recurrence of endometrial cancer after surgery. Its AUC in the training group is 0.945, sensitivity is 0.959, and specificity is 0.961; the AUC in the training group is 0.924, sensitivity is 0.872, and specificity is 0.846.
[0049] A sixth aspect of the present invention provides a device for predicting the risk of recurrence or metastasis of endometrial cancer, comprising a data acquisition unit and a calculation unit;
[0050] The data acquisition unit is used to obtain the detection amount data of the biomarkers as described above in the biological sample of the subject, wherein the biomarkers include one or more of MUC16, CTSG, B3GNT2, CORO1A and SFTPB;
[0051] The calculation unit is used to calculate and output the predicted score of the subject's risk of endometrial cancer recurrence or metastasis based on the detection data, and make a judgment based on the threshold value.
[0052] In some embodiments, the biomarkers include MUC16, CTSG, B3GNT2, and CORO1A.
[0053] In some embodiments, the prediction device further includes a data storage unit and a data output unit; the data storage unit is used to store the detection amount (or detection value) of the biomarker; the data input interface is used to input the detection value of the biomarker, and the data output unit is used to output the prediction result.
[0054] Furthermore, the detection value is the presence or absence, relative abundance or concentration value of each biomarker.
[0055] The seventh aspect of the present invention provides a device comprising a processor and a memory, wherein the memory is used to store a computer program, and is characterized in that the processor is used to execute the computer program stored in the memory so that the device performs the prediction method as described above or the construction method as described above.
[0056] An eighth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is processed, the prediction method as described above or the construction method as described above is executed.
[0057] In some embodiments, the computer-readable storage medium includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0058] In some implementations, being processed refers to being executed by one or more processors.
[0059] A ninth aspect of the present invention provides a method for predicting the risk of recurrence or metastasis of endometrial cancer, comprising the following steps:
[0060] S1. Obtaining detection amount data of biomarkers in a biological sample of a subject, wherein the biomarkers include one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB;
[0061] S2. Apply the endometrial cancer recurrence risk prediction model obtained by the construction method described above to process the detection data to output a prediction result of the endometrial cancer recurrence or metastasis risk.
[0062] In some embodiments, the biomarkers include MUC16, CTSG, B3GNT2, and CORO1A.
[0063] In another aspect, the present invention provides a system for predicting the recurrence risk of endometrial cancer, the system comprising a data analysis module for analyzing detection values of markers, wherein the markers include MUC16, CTSG, B3GNT2, and CORO1A.
[0064] Furthermore, the data analysis module uses the detection values of markers of known samples as a training set, and divides them into an endometrial cancer recurrence group and an endometrial cancer non-recurrence group according to whether the endometrial cancer recurs, analyzes the relationship between the detection values of the endometrial cancer recurrence group and the endometrial cancer non-recurrence group, and constructs a model.
[0065] In some embodiments, a combined diagnostic model for predicting the risk of endometrial cancer recurrence is constructed by combining multiple machine learning methods, and it is preliminarily confirmed that the concentration changes of any one of the screened new biomarkers alone can be used to distinguish between people at high risk and low risk of endometrial cancer recurrence, indicating that these biomarkers have extremely high diagnostic value.
[0066] In some embodiments, the equation of the constructed model is:
[0067]
[0068] Wherein, Y is the prediction score, i represents the i-th biomarker, m represents the number of biomarkers (m=5), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant of 3.8638532; the coefficients of the four biomarkers are:
[0069]
[0070] When Y ≤ the cutoff value, the patient has no recurrence of endometrial cancer. When Y > the cutoff value, the patient has recurrence of endometrial cancer. Specifically, the cutoff value is 0.4773883.
[0071] Furthermore, the system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0072] In another aspect, the present invention provides a use of a marker for preparing a reagent for predicting whether a patient with endometrial cancer will not relapse after surgery, have an in situ recurrence, or have distant metastasis, wherein the marker comprises any one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB.
[0073] In another aspect, the present invention provides a kit for predicting whether endometrial cancer patients will not relapse after surgery, will relapse in situ, or will have distant metastasis. The kit comprises a detection reagent for the biomarker for the purpose described above.
[0074] In another aspect, the present invention provides a biomarker combination for predicting whether endometrial cancer patients will not relapse after surgery, relapse in situ, or have distant metastasis, wherein the combination comprises any one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB.
[0075] In another aspect, the present invention provides a system for predicting whether endometrial cancer will not recur after surgery, recur in situ, or metastasize to distant sites. The system includes a data analysis module for analyzing the detection values of markers, including MUC16, CTSG, B3GNT2, CORO1A, and SFTPB.
[0076] Furthermore, the distant metastasis also includes simultaneous in situ recurrence and distant metastasis. As long as distant metastasis occurs, it will be classified into the distant metastasis group.
[0077] Furthermore, the data analysis module uses the detection values of markers of known samples as a training set, and divides liver cancer patients into a non-recurrence group, an in situ recurrence group, and a distant metastasis group according to their post-operative conditions. The relationship between the detection values of the non-recurrence group, the in situ recurrence group, and the distant metastasis group is analyzed to construct a model.
[0078] Furthermore, the model is constructed based on the gradient boosting algorithm.
[0079] Unlike generalized linear regression, the gradient boosting algorithm cannot output model formulas and cutoff values. All calculations are done directly by machine learning. The test values can be directly input into the software system to obtain the prediction results.
[0080] Furthermore, the system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0081] The beneficial effects of the present invention are:
[0082] 1. The present invention has screened 5 new biomarkers that can predict the risk of recurrence or metastasis of endometrial cancer, and developed a new combination of protein markers that can effectively assess and diagnose the risk of recurrence or metastasis of endometrial cancer, effectively distinguish between recurrence and non-recurrence of endometrial cancer, and effectively distinguish between in situ recurrence and distant metastasis of endometrial cancer. This not only improves the diagnostic accuracy of endometrial cancer recurrence and metastasis, but also provides an important biomarker basis for personalized treatment and prognosis monitoring of endometrial cancer patients. In addition, the discovery and application of these markers are expected to improve the treatment effect and quality of life of endometrial cancer patients, and provide a scientific basis for the early diagnosis and treatment strategy selection of endometrial cancer. Compared with traditional detection methods, it reduces the risk of misdiagnosis and missed diagnosis, and provides strong support for early detection and intervention of the disease.
[0083] 2. The four-biomarker combined differential diagnosis model constructed by this invention is convenient and rapid, and its test results are highly consistent with the clinical gold standard test results. It also significantly reduces the cost of predicting the risk of endometrial cancer recurrence or metastasis, and has excellent application prospects. The combination of markers of this invention is superior to the diagnosis of broad-spectrum tumor markers (CEA, CA125) or a broad-spectrum tumor marker combination model.
[0084] 3. Based on the biomarkers screened for predicting the risk of recurrence or metastasis of endometrial cancer, a three-classification model was further constructed that can simultaneously distinguish between the non-recurrence group, the in situ recurrence group, and the distant metastasis group, providing a more effective and accurate predictive diagnostic model.
[0085] Detailed description
[0086] (1) Diagnosis or testing
[0087] The diagnosis or detection here refers to the detection or analysis of biomarkers in a sample, or the content of a target biomarker, such as the absolute content or relative content, and then the presence or amount of the target marker is used to indicate whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease. The meanings of diagnosis and detection here are interchangeable. The result of such a test or diagnosis cannot be directly used as a direct result of the disease, but is an intermediate result. If a direct result is obtained, other auxiliary means such as pathology or anatomy are required to confirm that the patient has a certain disease. For example, the present invention provides a variety of new biomarkers related to the risk of recurrence or metastasis of endometrial cancer. Changes in the content of these markers are directly correlated with whether the patient belongs to the risk group for recurrence or metastasis of endometrial cancer.
[0088] (2) Association between markers, biomarkers, or differentially expressed proteins and the risk of recurrence or metastasis of endometrial cancer
[0089] The terms "marker," "biomarker," and "differential protein" have the same meaning in this invention. Association here refers to a direct correlation between the presence or change in the level of a biomarker in a sample and a specific disease. For example, a relative increase or decrease in the level indicates a higher likelihood of the individual having the disease compared to a healthy population.
[0090] The simultaneous presence of multiple markers in a sample, or the relative changes in their levels, indicate a higher likelihood of the individual having the disease compared to healthy individuals. This means that among marker types, some are strongly associated with disease, while others are weakly associated, or even unrelated to a particular disease. One or more markers with strong correlations can be used as diagnostic markers, while markers with weaker correlations can be combined with stronger markers to diagnose a disease, increasing the accuracy of test results.
[0091] Regarding the numerous biomarkers found in serum by the present invention, these markers can be used to distinguish between recurrent and non-recurrent endometrial cancer, as well as between non-recurrent, in situ recurrence, and distant metastasis. The markers herein can be used alone as individual markers for direct detection or diagnosis. The selection of such markers indicates that the relative change in the marker's level is strongly correlated with the risk of recurrence or metastasis of endometrial cancer. Of course, it is understood that one or more markers with a strong correlation with the risk of recurrence or metastasis of endometrial cancer can be selected for simultaneous detection. It is generally understood that, in some approaches, selecting biomarkers with strong correlations for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90%, or 95%, indicating that these markers can achieve intermediate values for diagnosing a certain disease, but does not directly confirm the presence of a certain disease.
[0092] Of course, the differential protein with the larger ROC value can also be selected as a diagnostic marker. The so-called strength is generally calculated and confirmed by some algorithms, such as the contribution rate or weight analysis of the marker to the risk prediction of recurrence or metastasis of endometrial cancer. Such calculation methods can be significance analysis (p value or FDR value) and fold change (Fold change). Multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA) and orthogonal partial least squares discriminant analysis (OPLS-DA), and of course other methods, such as ROC analysis, etc. Of course, other model prediction methods are also possible. When specifically selecting biomarkers, the differential proteins disclosed in the present invention can be selected, or other existing well-known marker combinations can be selected or combined to make predictions through model methods.
[0093] (3) Definition of disease terms
[0094] Endometrial cancer: Endometrial cancer is a malignant tumor that originates in the endometrium and is one of the most common types of gynecological malignancies. It can affect women of any age, but most cases occur in menopausal or postmenopausal women. The main pathological types of endometrial cancer include endometrioid adenocarcinoma, serous adenocarcinoma, and clear cell carcinoma. Typical symptoms include abnormal vaginal bleeding, vaginal discharge, and lower abdominal pain. Because early symptoms are often subtle, endometrial cancer is often discovered in its late stages, making treatment more difficult and the prognosis poorer. The diagnosis of endometrial cancer typically requires a combination of imaging studies (such as ultrasound and MRI) and pathological examination (obtaining tissue samples through fractional dilation and curettage or hysteroscopic biopsy). Stage I / II endometrial cancer is an early-stage lesion confined to the uterus. The tumor has spread from the uterine body downward to the cervical stroma, but has not extended beyond the uterus and has not involved the paracervical tissue or vagina. Surgery is curative, but subsequent treatment must be individualized based on the pathological findings.
[0095] Endometrial cancer recurrence occurs when cancer cells grow again after initial treatment. This recurrence can occur at the original site of the endometrial cancer (in situ recurrence) or elsewhere in the body (distant metastasis). Most endometrial cancer recurrences (approximately 70%-80%) occur within 2-3 years after surgery. The risk of recurrence decreases with time after surgery, reaching a significantly lower risk after 5 years.
[0096] Endometrial cancer metastasis refers to the process by which malignant tumor cells spread from their original site to other parts of the body. These cells travel through the blood or lymphatic system and form new tumors in new locations. Metastasis is categorized as both in situ and distant metastasis.
[0097] Localized metastasis of endometrial cancer: Localized metastasis refers to the recurrence of cancer cells in or near the original site of cancer. This may include the uterus, ovaries, fallopian tubes, or other structures in the pelvis. Localized metastasis usually means that the cancer has not spread widely to other parts of the body, but it still requires aggressive treatment to prevent further spread.
[0098] Distant metastasis of endometrial cancer refers to the spread of tumor cells to more distant sites in the body, such as the peritoneum, lymph nodes, lungs, liver, or other organs. Common distant metastatic sites of endometrial cancer include the peritoneum (causing symptoms such as ascites), lymph nodes, lungs, and liver. Distant metastasis usually indicates that the cancer has progressed to an advanced stage, making treatment more difficult and the prognosis poorer.
[0099] Therefore, early detection and intervention of endometrial cancer recurrence or metastasis are crucial for improving treatment efficacy and increasing survival rates. The biomarkers of the present invention have accurate and specific predictive efficacy for the recurrence or metastasis of endometrial cancer, which is beneficial to improving the treatment effect of patients and has important significance for improving the prognosis of patients.
[0100] (4) The gold standard for the diagnosis of endometrial cancer is pathological diagnosis (i.e., pathological confirmation of a puncture biopsy or surgical resection specimen), which involves observing the presence of cancer cells under a microscope and combining immunohistochemistry or molecular testing to clarify the nature of the tumor. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] Figure 1 This is a volcano plot of the differential analysis of protein markers in the high-risk and low-risk groups for endometrial cancer recurrence in Example 1.
[0102] Figure 2 Graphs showing the ROC and OPLS-DA analysis results for the high-risk and low-risk groups for endometrial cancer recurrence in Example 1.
[0103] Figure 3This is a graph showing the hyperparameter combination optimization results of the glmnet algorithm in Example 2.
[0104] Figure 4 2 is the ROC curve of the endometrial cancer recurrence risk prediction model in the training group in Example 2.
[0105] Figure 5 2 is the ROC curve of the endometrial cancer recurrence risk prediction model in the training group in Example 2. DETAILED DESCRIPTION
[0106] The present invention will be described in further detail below in conjunction with the accompanying drawings and Examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not serve to limit the present invention in any way. The reagents used in this example are all known products and were obtained by purchasing commercially available products.
[0107] Example 1: Screening for biomarkers of endometrial cancer recurrence risk using proteomics
[0108] Plasma samples were collected from patients with endometrial cancer after radical surgery. Low-abundance proteins were enriched by immunoaffinity chromatography to remove high-abundance proteins. Protein abundance in the samples was detected by HPLC-MS / MS. The difference in protein abundance between patients with and without recurrent endometrial cancer was analyzed to screen for differentially expressed proteins. The specific steps are as follows:
[0109] 1.1 Sample Collection
[0110] Blood samples were collected from 150 patients with endometrial cancer (stage I / II) 2 weeks after surgical resection. All endometrial cancer samples were biopsied and pathologically confirmed. Approximately 2 ml of peripheral blood was collected from the subjects, placed in a vacuum tube containing EDTA anticoagulant, mixed, and centrifuged twice at 120 g for 10 minutes at room temperature. The supernatant was removed and the mixture was centrifuged at 360 g for 20 minutes. Platelet samples were then collected in centrifuge tubes and stored at -80°C until further use. All 150 patients were female, with an average age of 55 years (range, 30-79 years). Eighty-four patients had stage I disease, and 66 had stage II disease. All enrolled patients provided written informed consent. All endometrial cancer patients had histopathological confirmation. Inclusion criteria included: (a) no history of other malignancies; and (b) no concurrent malignancies or autoimmune diseases.
[0111] Over the next three years, patients were followed up every six months, with regular CA-125 and imaging examinations. Once recurrence or metastasis was detected, pathological histology was required to confirm the diagnosis. A total of 18 stage I / II patients relapsed within three years after surgery, of which 10 were local recurrences in situ and 8 were distant metastases (including 5 peritoneal metastases, 2 lymph node metastases, and 1 liver metastasis). Samples stored at -80°C were then taken out and divided into two groups: 132 patient samples without recurrence were classified as a low-risk group for colorectal endometrial cancer recurrence, and 18 patient samples with recurrence or metastasis were classified as a high-risk group for colorectal endometrial cancer recurrence for proteomic biomarker screening.
[0112] 1.2 Sample processing and enzymatic hydrolysis
[0113] 1) First, plasma samples were centrifuged for 15 minutes (15,000 g), and the supernatant was filtered and then subjected to immunoaffinity chromatography to remove 14 highly abundant proteins.
[0114] 2) The low-abundance fraction was concentrated to 350 μL using a 3 kDa cut-off concentration tube in a centrifuge (4000 g, 1 hour).
[0115] 3) The concentrate was recovered and buffer exchanged using a 7 kDa cutoff desalting column in a centrifuge (1000 g, 2 minutes) using AEX-A (20 mM Tris, 4 M urea, 3% isopropanol, pH 8.0).
[0116] 4) Determine the protein concentration in the sample using the biuret assay (BCA) with AEX-A as a blank. According to the sample grouping, add 25 μL of TCEP to the sample and incubate at 37°C for 30 minutes to reduce the protein. Then, add the corresponding TMT 16-plex reagent and incubate at room temperature in the dark for 1 hour to perform the TMT labeling reaction. Then, perform buffer exchange of the sample using Zeba columns with AEX-A. After mixing the TMT 16-plex-labeled sample, add 2 mL of AEX-A to the mixed sample to a final volume of 5.5 mL.
[0117] 5) Filter the sample through a 0.22 µm filter and separate the TMT 16-plex-labeled sample using a 2D-HPLC system. Freeze-dry the collected fractions and add Trypsin-Lysin C enzyme mix. Incubate the sample at 37°C for 5 hours to digest the sample. Terminate the digestion reaction by adding 5 µL of 10% TFA.
[0118] 6) A total of 60 2D-HPLC fractions after enzymatic digestion were used for nanoLC-MS / MS analysis.
[0119] 1.3 LC-MS / MS Data Acquisition
[0120] Each sample obtained in step 1.2 was separated using the nanoliter flow rate liquid chromatography system Easy nLC-1200 and detected online using a high-resolution mass spectrometer Q Exactive HF-X. The details are as follows:
[0121] 1) Separation: Mobile phase A consisted of 0.1% formic acid in water, and mobile phase B consisted of 0.1% formic acid in acetonitrile (80% acetonitrile, 20% water). The chromatographic column consisted of an enrichment column and an analytical column equilibrated with 100% mobile phase A. The sample was loaded via an autosampler onto the enrichment column (100 μm ID × 4 cmL, C18, 3 μm, 100 Å) and separated on the analytical column (75 μm ID × 25 cmL, C18, 3 μm, 100 Å) at a flow rate of 300 nL / min.
[0122] 2) After chromatographic separation, samples were analyzed by mass spectrometry using a Q Exactive HF-X mass spectrometer. Positive ion detection was used, with a parent ion scan range of 350–1800 m / z, a primary MS resolution of 120,000 at 200 m / z, an AGC (Automatic Gain Control) target of 3e6, a maximum IT of 50 ms, and a dynamic exclusion time of 40 s. The mass-to-charge ratios of peptides and peptide fragments were acquired using data-dependent acquisition (DDA): 20 secondary MS / MS spectra (MS2 scans) were collected after each full scan (MS2 scan). The MS2 Activation Type is HCD, the Isolation window is 0.7 m / z, the secondary mass spectrometry resolution is 30,000 at 200 m / z, the AGC target is 1e5, the Maximum IT is 65 ms, the Fixed first mass is 110.0 m / z, the Normalized Collision Energy is 32 eV, the Minimum AGC target is 2.00e4, the Charge exclusion is 1, 6-8, >8, the Multiple charge states is one charge state only, the Peptidematch is preferred, and the Exclude isotopes is on.
[0123] 1.4 Data Preprocessing
[0124] Secondary mass spectrometry data were retrieved using Maxquant (v1.6.15.0). The data type is DIA proteomics data based on secondary reporter ion quantification. The secondary spectrum used for quantification requires that the parent ion accounts for more than 75% in the primary spectrum. The database comes from the Homo_sapiens_9606_proteome_gene (release: 2021-10-14, sequence: 20,437) of the Uniprot database, and a common contamination library is added to the database. Contaminating proteins are deleted during data analysis; the enzyme cleavage method is set to Trypsin / P; the number of missed cleavage sites is set to 2; the parent ion mass error tolerance of the First search and Main search is set to 20 ppm and 5 ppm, respectively, and the mass error tolerance of the secondary fragment ion is 20 ppm. The fixed modification is cysteine alkylation, and the variable modification is methionine oxidation and acetylation of the protein N-terminus. The FDR for protein identification and PSM identification is set to 1%.
[0125] 1.5 Gap Analysis
[0126] We screened for differentially expressed proteins and transcripts using a combination of univariate and multivariate statistical analyses. Univariate analysis primarily included significance analysis (p-value or FDR value) and fold change analysis of signature molecules across different groups. Multivariate statistical analysis primarily included receiver operating characteristic (ROC) curve analysis and Boruta signature screening based on the random forest algorithm. All statistical analyses were performed using R. Specific R information is provided in Table 1.
[0127] Table 1. R and related information
[0128]
[0129] The variable importance for the projection (VIP) was calculated to measure the influence and explanatory power of each protein expression pattern on the classification and discrimination of each group of samples. The Wilcoxon rank sum test was further performed to obtain the corrected p value (FDR). According to the conditions of FDR < 0.01 and Fold change > 2, 63 down-regulated proteins and 57 up-regulated proteins were screened (see Figure 1 ).
[0130] To evaluate the role of each protein marker in the diagnosis and prediction of endometrial cancer recurrence risk, ROC and Boruta analysis methods were used to evaluate each protein marker. The results are shown in Figure 2The horizontal axis represents the AUC obtained from ROC analysis, and the vertical axis represents the -log10 (FDR) calculated by the Wilcoxon test. The size of the dots represents the VIP value obtained from Boruta analysis. Further screening based on VIP > 3 and AUC > 0.6 identified five more significant candidate protein biomarkers, as detailed in Table 2.
[0131] Table 2 Differential markers of recurrence risk in endometrial cancer
[0132] Among them, the smaller the FDR value and / or the larger the VIP value, to a certain extent, it indicates that the difference between the high risk of endometrial cancer recurrence and the low risk of endometrial cancer recurrence is more significant, and it also indicates that the protein may have a higher diagnostic value.
[0133] Example 2: Construction and validation of a model for predicting the risk of recurrence of endometrial cancer
[0134]
[0135] In this example, the five protein markers screened and obtained in Example 1 were combined to construct a recurrence risk prediction model for research.
[0136] While a single biomarker can differentiate the risk of recurrence in patients with endometrial cancer after surgery, combining multiple biomarkers generally offers greater accuracy in differentiation or prediction. However, a single biomarker with a higher accuracy in predicting endometrial cancer recurrence risk may not necessarily be more effective when combined with one or more other biomarkers. Furthermore, a greater number of biomarkers does not necessarily equate to a higher predictive accuracy (AUC) for the combination. Therefore, extensive validation experiments are still needed.
[0137] 2.1 Get Data
[0138] A total of 252 blood samples of patients with endometrial cancer (stage I / II) were collected 2 weeks after surgical treatment. All enrolled patients signed informed consent forms, of which 32 patients relapsed within three years after surgery (14 were local recurrences in situ and 18 were metastatic). They were divided into two groups: 220 patient samples without recurrence were classified as the low-risk group for endometrial cancer recurrence, and 32 patients with recurrence or metastasis were classified as the high-risk group for endometrial cancer recurrence. They were randomly divided into a training group and a test group. The training group included 110 patients in the low-risk group for endometrial cancer recurrence and 16 patients in the high-risk group for endometrial cancer recurrence, and the test group included 110 patients in the low-risk group for endometrial cancer recurrence and 16 patients in the high-risk group for endometrial cancer recurrence.
[0139] The same steps as steps 1.2 and 1.3 in Example 1 were used to obtain the concentrations of five proteins in serum from 250 samples.
[0140] 2.2 Data Statistical Analysis
[0141] In the training group, a combined diagnostic model for multiple endometrial cancer recurrence risk markers was constructed using a combination of multiple machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using predicted probability values with 95% confidence intervals (CI) to assess the discriminative ability of the multivariate diagnostic model.
[0142] Using the training set, the Youden index (YI) was calculated to determine the cut-off value for distinguishing the predicted probability of endometrial cancer recurrence in a high-risk group from that in a low-risk group. In addition, the ROC of the models formed by the combination of different markers was constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV) and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R3.6.1, and a p value of less than 0.05 was considered statistically significant.
[0143] 2.3 Construction of prediction model
[0144] The specific steps for constructing a model for predicting the recurrence risk of endometrial cancer are as follows:
[0145] S101. Randomly select concentration matrices of 2 to 5 protein markers from the 5 protein markers of the samples in the training group as the original training data set.
[0146] S102: Select the generalized linear model (glmnet) algorithm for building the prediction model and the grid search range during the algorithm's hyperparameter optimization process. In this step, the grid search range for the model's hyperparameter optimization is set for each algorithm as shown in Table 3.
[0147] Table 3. Parameter grid of glmnet algorithm
[0148]
[0149] S103. According to the algorithm and hyperparameter setting range set in step S102, one of the hyperparameter combinations is selected as a parameter for constructing a prediction model.
[0150] S104: Split the original dataset into K subsets using a K-fold cross validation mechanism. To ensure that the ratio of majority class samples to minority class samples in each subset is the same as in the original dataset, a Stratified K-Folds cross validation mechanism is used for data segmentation.
[0151] S105 , selecting one of the K training data subsets obtained by segmentation in step S104 as a validation set Ddev.
[0152] S106: Combine the training data subsets not selected in step S105 to form a training data pool Dtrain1.
[0153] S107 . According to the training data set Dtrain obtained in step S106 , a prediction model is constructed based on the selected supervised classification algorithm and hyperparameters.
[0154] S108: Evaluate the prediction model obtained in step S107 on the validation set Ddev to obtain an AUC value, and store the current prognosis prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 involves evaluating the prediction model obtained in step S107 on the validation set determined in the current iteration, and storing both the model and the evaluation results in the prediction model pool for future selection and use by the base prediction model. The evaluation mentioned in this step can be an AUC value or other reasonable metric for evaluating model performance.
[0155] S109: Determine whether all subsets have been used as validation sets. Step S109 determines whether all K subsets obtained in step S104 have been used as validation sets and trained on the model. If all subsets have been used as validation sets and training has been completed, proceed to step S110; if any subsets have not been used as validation sets, proceed to step S105. This step ensures that every sample in the original dataset has been used as a validation set, improving model stability and preventing the model from overfitting to a particular subset.
[0156] S110: The average AUC value of all models in the prediction model pool Pool is used as the final performance evaluation value of the combined model. The model parameters and the final performance evaluation AUC value are stored in the optimal model pool Poolbest.
[0157] S111: Determine whether all hyperparameter combinations have been used to construct prediction models. Step S111 determines whether all algorithms and corresponding hyperparameter combinations obtained in step S102 have been used to construct prediction models. If all combinations have been used to construct models, step S112 is executed; if any combination has not been used to construct models, step S103 is executed.
[0158] S112. After completing step S111, it is necessary to check whether all hyperparameter combinations have been used to build prediction models. If it is found that there are still hyperparameter combinations for which models have not been built, it is necessary to return to step S103 and continue to build the remaining models. If all hyperparameter combinations have been used to build models, it is possible to continue to step S113 and select the best model from the model pool.
[0159] S113. From the model set Poolbest obtained in step S112, select the model with the largest AUC value as the final prediction model for the risk of recurrence of endometrial cancer.
[0160] S114. Repeat all the above steps until all combinations of markers are modeled.
[0161] By executing the above prediction model construction steps, the optimal models for all combinations of markers were obtained. To compare the performance of the models under these different marker combinations, the receiver operating characteristic (ROC) method was used to evaluate the AUC values of these prediction models in the training set. The results are shown in Table 4.
[0162] Table 4 Comparison of the area under the ROC curve of the models constructed with different marker combinations in the training group
[0163]
[0164] Table 4 shows the AUC, sensitivity, and specificity of the models formed by different marker combinations. It was found that the combination (4MP-2) MUC16+CTSG+B3GNT2+CORO1A had the highest AUC, sensitivity, and specificity. Further addition could not further improve the diagnostic efficacy. In addition, 4MP-2 has fewer markers and is the most preferred model.
[0165] 2.4 Optimization of model parameters
[0166] For the optimal marker combination of MUC16+CTSG+B3GNT2+CORO1A in step 2.3, we analyzed the models constructed under 9 different combinations of glmnet algorithm hyperparameters based on this marker combination, and evaluated the model performance by AUC value (AUC was calculated using 10-fold cross-validation during the modeling process). The results are shown in Tables 5 and Figure 3 shown.
[0167] Table 5 AUC of the constructed model under different combinations of glmnet algorithm hyperparameters
[0168]
[0169] As can be seen from Table 5, when the hyperparameter combination of the glmnet algorithm is alpha=1, lambda=0.0547, the AUC reaches the maximum value of 0.945.
[0170] The equation for building a model based on the optimal hyperparameter combination is:
[0171]
[0172] Where Y is the prediction score, i represents the i-th biomarker, m represents the number of biomarkers (m = 5), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker (see Table 6), and b is a constant of 3.8638532.
[0173] Table 6 Coefficients of the five biomarkers in the model
[0174]
[0175] The complete model equation is: y = 4.967 MUC16 + 6.255 CTSG + 2.163 B3GNT2 + 7.522 CORO1A + 3.8638532
[0176] Determination of diagnostic threshold for the endometrial cancer recurrence risk prediction model:
[0177] The ROC curve was drawn with the prediction scores in the training group, and the optimal prediction cutoff value was set to 0.4773883 according to the Youden index value. The results are shown in Figure 4 When the prediction score of the prediction model is ≤0.4773883, the patient is considered to have a low risk of endometrial cancer recurrence; when the prediction score of the prediction model is >0.4773883, the patient is considered to have a high risk of endometrial cancer recurrence. Figure 4 It can be seen that the AUC of the prediction model in the training group was 0.945, the sensitivity was 0.959, and the specificity was 0.961.
[0178] 2.6 Validation of the Endometrial Cancer Recurrence Risk Prediction Model
[0179] The optimal model constructed was verified in the training group and the ROC curve was drawn as follows Figure 5 shown.
[0180] from Figure 5 It can be seen that the AUC of the model in the training group is 0.921, the sensitivity is 0.872, and the specificity is 0.846.
[0181] In summary, the endometrial cancer recurrence risk prediction model constructed using four protein markers has good predictive performance and accuracy and the best diagnostic efficacy.
[0182] Example 3: Construction and verification of a three-category diagnostic model
[0183] This example attempts to construct a three-category combined diagnostic model for distinguishing between the endometrial cancer prognosis non-recurrence group, the endometrial cancer prognosis in situ recurrence group, and the endometrial cancer prognosis distant metastasis group. The specific process includes the following: (1) construction and screening of the optimal diagnostic model; (2) validation of the effectiveness of the optimal diagnostic model. The specific screening process and results are as follows (in the present invention, the two-classification model in Example 2 uses the AUC value as the evaluation indicator; when a three-classification model is constructed, since multiple categories are involved, the AUC value is generally not applicable. In this example, the diagnostic efficacy of the model is measured using indicators such as sensitivity, specificity, accuracy, and consistency):
[0184] 3.1 Construction and screening of the three-category diagnostic model
[0185] For the test cohort of 382 endometrial cancer patients, all enrolled patients signed informed consent. Among them, 34 patients relapsed within two years after surgery (16 of them were local recurrences in situ and 18 of them had metastases). They were divided into two groups: 348 patient samples without recurrence were classified as the low-risk group for endometrial cancer recurrence, and 34 patients with recurrence or metastasis were classified as the high-risk group for endometrial cancer recurrence. They were randomly divided into training group and test group. The training group included 174 patients in the low-risk group for endometrial cancer recurrence and 17 patients in the high-risk group for endometrial cancer recurrence, including 8 patients with in situ recurrence and 9 patients with metastasis; the training group included 174 patients in the low-risk group for endometrial cancer recurrence and 17 patients in the high-risk group for endometrial cancer recurrence, including 8 patients with in situ recurrence and 9 patients with metastasis.
[0186] In this example, based on the marker combination MUC16 + CTSG + B3GNT2 + CORO1A screened in Example 2, a three-category detection model was constructed that can effectively distinguish between the endometrial cancer prognosis non-recurrence group (low-risk group), the endometrial cancer prognosis in situ recurrence group (in situ recurrence group), and the endometrial cancer prognosis distant metastasis group (metastasis group). All enrolled patients signed informed consent. Patients with endometrial cancer were pathologically confirmed, and inclusion criteria included: (a) no history of other malignancies; (b) no concurrent malignancies or autoimmune diseases.
[0187] The collected serum samples were subjected to LC-MS / MS data acquisition and detection to obtain the concentrations of four protein markers, MUC16, CTSG, B3GNT and CORO1A.
[0188] The Shapiro-Wilk test was used to assess normal distribution, and the nonparametric Wilcoxon test was used to analyze differences in blood marker concentrations between the endometrial cancer prognosis non-recurrence group (low-risk group), the endometrial cancer prognosis in situ recurrence group (in situ recurrence group), and the endometrial cancer prognosis distant metastasis group (metastasis group). A three-class combined diagnostic model for the four markers was constructed using a combination of machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using the predicted probability values with 95% confidence intervals (CI) to assess the discriminatory ability of the multivariate diagnostic model. Using the training set, the Youden index (YI) was calculated to determine the predicted probability cut-off value for distinguishing the endometrial cancer prognosis non-recurrence group (low-risk group), the endometrial cancer prognosis in situ recurrence group (in situ recurrence group), and the endometrial cancer prognosis distant metastasis group (metastasis group). In addition, the ROC values of the models constructed with different marker combinations were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV) and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R3.6.1, and a p value of less than 0.05 was considered statistically significant.
[0189] In this embodiment, in order to construct the optimal three-class joint diagnosis model, after comparing the six algorithms of gradient boosting, naive Bayes, support vector machine, neural network, generalized linear, and discriminant analysis, the gradient boosting method was selected as the best supervised classification algorithm for constructing the prediction model. The grid search range for hyperparameter optimization of the gradient boosting method model is shown in Table 7 below.
[0190] Table 7. Parameter grid search range of gradient boosting method
[0191]
[0192] Through optimization screening in terms of accuracy, consistency, sensitivity, specificity, etc., the optimal parameter combination mode was determined to be: interaction.depth 2, n.trees 150, shrinkage 0.1, n.minobsinnode 10.
[0193] The training group and the training group used two completely different batches of samples. In this example, only the marker screening and model construction were performed from the training group; the samples from the training group were only used to verify the diagnostic efficacy of the model. The specific results are shown in Table 8.
[0194] Table 8. Performance evaluation table of the gradient boosting method to build a model to distinguish three categories
[0195]
[0196] As can be seen from Table 8, the gradient boosting model constructed based on the four protein markers MUC16, CTSG, B3GNT2, and CORO1A can be used to predict whether endometrial cancer patients will have no recurrence, in situ recurrence, or distant metastasis after surgery. This also demonstrates that the protein markers screened by the present invention can be used to distinguish the risk of recurrence in endometrial cancer patients after surgical treatment, and can also be used to distinguish in situ recurrence from distant metastasis when the risk of recurrence is high (patients with both in situ recurrence and metastasis are also included in the metastasis group). However, the diagnostic efficacy for distinguishing between the in situ recurrence group and the metastasis group is still not ideal.
[0197] 3.2 Combined performance of the three-class joint diagnosis model
[0198] To further enhance the diagnostic value of three-category diagnostic models (gradient boosting) constructed using different protein combinations of biomarkers, this example compared the performance of diagnostic models constructed using different protein combinations of biomarkers in a training set based on the five protein markers screened in Example 1. The specific combinations of the different models are shown in Table 9.
[0199] Table 9. Combinations of different diagnostic models
[0200]
[0201] For the three-category classification, the performance index comparison results of different diagnostic models constructed by the five biomarkers screened in Example 1 are shown. The calculation method of the minimum value, first quartile, median, mean, third quartile, and maximum value of accuracy and consistency are as follows: (1) Sort the accuracy or consistency values from small to large; (2) Minimum value: the first value after sorting; (3) First quartile (Q1): multiply the number of data by 0.25. If the result is an integer, take the average of the values at this position and the next position; if it is not an integer, round up to get the position, and the value at this position is Q1; (4) Median: if the number of data is odd, the median is the middle value; if it is even, it is the average of the two middle values; (5) Mean: the sum of all values divided by the number of data; (6) Third quartile (Q3): multiply the number of data by 0.75 and process it in the same way as Q1; (7) Maximum value: the last value after sorting. The minimum and maximum values reflect data extremes, demonstrating the worst and best possible model performance. Quartiles help understand the data's distribution and dispersion. Q1 and below indicate lower performance, while Q3 and above indicate higher performance. The median reflects intermediate performance, and the mean comprehensively reflects the overall average performance. By combining these statistical values, we can gain a comprehensive understanding of the overall performance, distribution characteristics, and stability of the model, providing a strong basis for model selection and optimization.
[0202] Table 10. Performance comparison of diagnostic models based on different protein combination biomarkers
[0203]
[0204] As shown in Table 10, the nine-item joint test model (5MP), consisting of five markers, performed best for the three-category diagnostic model. This also clearly shows that adding SFTPB to the four-item combination of MUC16, CTSG, B3GNT2, and CORO1A has a certain improvement in the diagnostic efficacy of differentiating endometrial cancer surgical prognosis from non-recurrence, in situ recurrence, and distant metastasis. Therefore, the three-category gradient boosting model 5MP constructed using these five protein markers (MUC16, CTSG, B3GNT2, CORO1A, and SFTPB) was selected as the optimal joint diagnostic model.
[0205] 3.2 Diagnostic Performance Measurement and Verification of the Three-Classification Joint Diagnosis Model
[0206] 1) Diagnostic performance measurement of the three-classification combined diagnosis model
[0207] In order to more accurately determine the diagnostic performance and threshold of the model constructed in this embodiment for different disease classifications, a multi-classification model of the gradient boosting (GBM) algorithm was used to perform predictive analysis in the training group, and the predicted results were calculated as the predicted probability values of three categories (the prognosis of endometrial cancer surgery will be no recurrence, in situ recurrence, or distant metastasis), among which the category with the largest predicted probability value is the final prediction result of the system.
[0208] The meaning and calculation method of each indicator are as follows:
[0209] Results: The three-class combined diagnosis model achieved an accuracy of 0.82 and a consistency of 0.81 in the training group. The diagnostic sensitivity for endometrial cancer patients without recurrence after surgery was 93.6% and the specificity was 94.5%. The diagnostic sensitivity for patients with in situ recurrence after surgery was 83.6% and the specificity was 82.3%. The diagnostic sensitivity for patients with distant metastasis after surgery was 83.2% and the specificity was 83.9%.
[0210] It should be noted that the three-classification joint diagnosis model constructed by gradient boosting is a model constructed by machine learning and cannot fit a specific equation formula like a generalized linear model.
[0211] 2) Validation of the three-classification combined diagnosis model
[0212] The prediction performance of the model built based on the training group was verified in the test group. The specific results are as follows:
[0213] The accuracy was 0.81, and the consistency was 0.81. The sensitivity for diagnosing non-recurrence of endometrial cancer after surgery was 93.8%, and the specificity was 92.8%. The sensitivity for diagnosing in situ recurrence of endometrial cancer after surgery was 82.5%, and the specificity was 81.9%. The sensitivity for diagnosing distant metastasis of endometrial cancer after surgery was 82.4%, and the specificity was 82.5%.
[0214] In summary, the three-category combined diagnostic model containing five protein markers constructed in this example has good diagnostic value for the three categories of endometrial cancer non-recurrence group after surgery, endometrial cancer in situ recurrence group after surgery, and endometrial cancer distant metastasis group after surgery.
[0215] It can be understood that the embodiments described in the present invention are some preferred embodiments and features. Any person skilled in the art can make some changes and variations based on the essence of the description of the present invention. These changes and variations are also considered to fall within the scope of the present invention and the scope limited by the independent claims and the appended claims.
Claims
1. Use of a reagent for detecting a marker for preparing a reagent for predicting whether endometrial cancer patients will not relapse within two years after surgery, whether they will relapse in situ or have distant metastasis, characterized in that: The markers consist of MUC16, CTSG, B3GNT2, CORO1A and SFTPB.
2. A kit for predicting whether endometrial cancer patients will not relapse, relapse in situ, or metastasize within two years after surgery, characterized in that: The kit comprises a detection reagent for the biomarker for use according to claim 1.
3. A biomarker combination for predicting whether endometrial cancer patients will not relapse, relapse in situ, or metastasize within two years after surgery, characterized in that: The panel consists of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB.
4. A system for predicting whether endometrial cancer will not recur, recur in situ, or metastasize within two years after surgery, characterized in that: The system includes a data analysis module for analyzing the detection values of markers, wherein the markers consist of MUC16, CTSG, B3GNT2, CORO1A and SFTPB.
5. The system according to claim 4, wherein: The data analysis module uses the detection values of markers of known samples as a training set, and divides endometrial cancer patients into a non-recurrence group, an in situ recurrence group, and a distant metastasis group according to their post-operative conditions. The relationship between the detection values of the non-recurrence group, the in situ recurrence group, and the distant metastasis group is analyzed to construct a model; the model is constructed based on a gradient boosting algorithm; the system also includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output prediction results.
Citation Information
Patent Citations
Use of HE4 and other biochemical markers for assessment of endometrial and uterine cancers
CN101473041A
Compositions and methods for treating and diagnosing cancer
CN1852974A