Biomarker for predicting recurrence or metastasis risk of endometrial cancer and application of biomarker
By screening and constructing biomarker prediction models based on MUC16, CTSG, B3GNT2, CORO1A and SFTPB, the problem of inaccurate risk assessment of endometrial cancer recurrence in the prior art was solved, efficient and accurate prediction and personalized treatment recommendations were achieved, and patients' treatment effect and quality of life were improved.
Patent Information
- Application Number
- CN202510725313.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The lack of high sensitivity proteomic biomarkers in the prior art are used to predict the risk of recurrence of endometrial cancer, resulting in inaccurate assessment of recurrence risk in patients after surgical treatment, affecting the adjustment of treatment plans and patient prognosis.
The difference in protein abundance in blood in patients with endometrial cancer after surgery was analyzed by proteomics, and biomarkers such as MUC16, CTSG, B3GNT2, CORO1A and SFTPB were screened out to construct a prediction model, and high-performance liquid chromatography-tandem mass spectrometry technology was used for detection, and a prediction model was constructed in combination with machine learning algorithms.
It achieves accurate, non-invasive and efficient prediction of the risk of recurrence of endometrial cancer, improves the sensitivity and specificity of diagnosis, reduces the risks of misdiagnosis and missed diagnosis, provides personalized treatment suggestions, and improves the treatment effect and quality of life of patients.
Smart Images

Figure CN120254284A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biological medicine technology, and particularly relates to a biomarker for predicting the recurrence or metastasis risk of endometrial cancer and its application. Background Art
[0002] Endometrial cancer (EC) refers to cancer formed by the malignant transformation of epithelial cells in the endometrium. It mainly originates from glandular cells in the endometrium, which normally proliferate and differentiate under the regulation of estrogen and progesterone. However, under certain pathological conditions, the cells may undergo gene mutations and abnormal hyperplasia, and ultimately develop into malignant tumors.
[0003] Endometrial cancer is one of the most common gynecological malignancies in the female reproductive system, second only to cervical cancer. Its incidence is relatively high in developed countries, especially more common in postmenopausal women, with an average onset age of about 60 years old, but there is a trend of younger age in recent years. The clinical manifestations are: abnormal vaginal bleeding, specifically manifested as postmenopausal vaginal bleeding or menstrual disorders, such as increased menstrual volume, prolonged menstrual period, etc.; some patients may have vaginal discharge, and the discharge can be serous, bloody or purulent, etc.
[0004] The incidence and mortality of endometrial cancer have been increasing year by year. Recurrence is one of the main reasons leading to poor prognosis of patients. Traditionally, the prognosis assessment of EC is mainly based on clinicopathological parameters. However, for endometrioid carcinoma, which is the most common pathological type, patients with similar stages and grades based on clinicopathological parameters may also have different prognoses. There are still great limitations in relying solely on clinicopathological parameters for prognosis assessment.
[0005] Currently, the main treatment method for endometrial cancer is radical tumor resection. Especially for patients with stage I / II endometrial cancer, radical resection is usually directly performed, mainly for cure, without the need for extensive clearance, and generally no adjuvant radiotherapy or chemotherapy is carried out. However, about 5%-20% of patients with stage I / II endometrial cancer still relapse after surgery. If patients who do not benefit from surgery or have progression can be predicted for risk and the treatment plan can be adjusted in a timely manner (such as adjuvant radiotherapy or chemotherapy, secondary surgical resection, targeted therapy or immunotherapy, etc.), the overall survival rate and quality of life of patients can be significantly improved. There is a lack of biomarkers for diagnosing the recurrence risk of endometrial cancer clinically, especially the discovery of highly sensitive proteomic biomarkers for diagnosing the recurrence risk of endometrial cancer is of great significance.
[0006] Proteomics is the science that studies the protein composition, localization, changes, and their interaction laws in cells, tissues, or organisms, including the study of protein expression patterns and proteome functional patterns. With the development of mass spectrometry technology, liquid chromatography-tandem mass spectrometry (LC-MS / MS) has become the most important tool in proteomics research. The development of proteomics is of great significance for finding disease diagnostic markers, screening drug targets, toxicology research, etc., and has thus been widely applied in medical research.
[0007] Therefore, finding new markers related to the diagnosis of endometrial cancer recurrence risk and constructing a prediction model by combining multiple markers have important clinical value. Summary of the Invention
[0008] Aiming at the problems existing in the prior art, the present invention provides a biomarker for predicting the recurrence or metastasis risk of endometrial cancer and its application. By using the method of proteomics, through analyzing the abundance differences of proteins in the blood of endometrial cancer patients with recurrence and non-recurrence after surgical treatment, biomarkers capable of predicting the recurrence risk of endometrial cancer are screened out, and a prediction model for endometrial cancer recurrence risk is further constructed. It can accurately, non-invasively, and efficiently predict the recurrence or metastasis risk of endometrial cancer, provide a new tool for the management of endometrial cancer patients, and improve the treatment and prognosis monitoring of patients.
[0009] The first aspect of the present invention provides the use of a substance for detecting a biomarker in the preparation of a product for predicting the recurrence or metastasis risk of endometrial cancer, and the biomarker includes one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB.
[0010] The biomarker obtained by the present invention through proteomics can accurately predict the recurrence or metastasis risk of endometrial cancer, which is beneficial for doctors to judge the severity of the patient's condition and whether treatment plans need to be adjusted, so as to provide more accurate treatment suggestions for patients and truly benefit endometrial cancer patients.
[0011] The present invention uses the method of proteomics to collect plasma samples of patients with endometrial cancer who experience recurrence within a short period (such as 1 to 5 years) and those who do not experience recurrence within a short period after treatment. By analyzing different samples through high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS), based on the orthogonal partial least squares discriminant analysis and significance analysis methods, proteins with significant differences between endometrial cancer recurrence and non-recurrence are first screened, and 5 differential proteins with obvious relevance to the recurrence risk of endometrial cancer are obtained. These 5 proteins can be used to distinguish whether endometrial cancer patients will experience recurrence after surgical treatment and have certain diagnostic efficacy.
[0012] Among them, the MUC16 is a protein or amino acid sequence with the UniProt database number Q8WXI7; CTSG is a protein or amino acid sequence with the UniProt database number P08311; B3GNT2 is a protein or amino acid sequence with the UniProt database number Q9NY97; CORO1A is a protein or amino acid sequence with the UniProt database number P31146; SFTPB is a protein or amino acid sequence with the UniProt database number P07988.
[0013] The present invention also surprisingly discovers that some of the protein markers obtained through proteomic screening are known markers that can be used for other cancers. For example, SFTPB has been reported to be used for the prediction of non-small cell lung cancer, but this screening finds that this marker can also be used for the prediction of the recurrence or metastasis risk of endometrial cancer. It can be seen that for many different tumors, their protein markers are not completely separated or irrelevant. In fact, there are many cross relationships or influences. Many protein markers can be used for the prediction of early cancer, and also for the prognostic diagnosis of cancer. Even many protein markers can be used for the prediction and diagnosis of different stages of many different cancers. Therefore, there are still many brand-new functions in the field of proteomics waiting to be explored, and the market prospect is very broad.
[0014] In some embodiments, the biomarker can be 1 marker or a combination of several markers, such as a combination of 2 markers, a combination of 3 markers, a combination of 4 markers, or a combination of 5 markers.
[0015] In some specific embodiments, the biomarker comprises a combination of more than 2 markers, such as a combination of more than 3 markers. When a combination of markers constructs a prediction model for the recurrence or metastasis risk of endometrial cancer, the AUC value is 0.713 - 0.945, the sensitivity is 72.5% - 95.9%, and the specificity is 73.3% - 96.1%.
[0016] Furthermore, the biomarker includes MUC16, CTSG, B3GNT2, and CORO1A.
[0017] To study the diagnostic efficacy for the recurrence risk of endometrial cancer, it is also necessary to combine different differential proteins to construct a diagnostic model, rank them according to the importance obtained from the screening, and separately select different numbers of differential proteins with higher rankings for combination to obtain 5 protein markers. After verification, it is found that a model constructed based on these 5 protein markers has good risk prediction ability in the diagnosis of the recurrence risk of endometrial cancer. The data of samples for detecting the recurrence risk of endometrial cancer show that just using these 4 biomarkers to predict the recurrence risk of endometrial cancer, the AUC value can reach 0.945, and the diagnostic performance is good.
[0018] In some embodiments, the product is selected from one or more of a reagent, a kit, a chip, a probe, or a membrane strip. The product is a product targeted at detecting the above-mentioned biomarker, and includes, for example, biological reagents and kits suitable for detecting the biomarker, such as sample pretreatment reagents, antigens, or antibodies; it can also be developed into a standardized reagent or kit, a chip, a probe, or a membrane strip suitable for the biomarker.
[0019] In some embodiments, the product is used for predicting patients at risk of endometrial cancer recurrence or metastasis.
[0020] Furthermore, the risk of recurrence or metastasis refers to recurrence of endometrial cancer within three years after surgical treatment; the recurrence or metastasis of endometrial cancer includes in-situ or adjacent region recurrence of endometrial cancer, and metastasis of endometrial cancer.
[0021] It can be understood that the "within three years" in "recurrence or metastasis within three years after endometrial cancer treatment" here is not an absolute and unchanging time node, but only a time node summarized based on current clinical experience. With the passage of time or the improvement of other treatment methods, this time node or the length of time may change. For example, it may be within one year, within 360 days, within six months, within 180 days, within two years, within two and a half years, within three and a half years, etc. It is possible to have recurrence of colorectal cancer after treatment.
[0022] In some ways, the recurrence or metastasis of endometrial cancer means that after surgical treatment with radical tumor resection, the patient has in-situ or adjacent region recurrence of endometrial cancer within three years, or metastasis of endometrial cancer occurs, such as liver metastasis, lung metastasis, peritoneal metastasis, bone metastasis, etc.
[0023] In some embodiments, the product is used to detect the detected amount of a biomarker in a biological sample.
[0024] In some specific embodiments, the biological sample is selected from one or more of saliva, blood, urine, plasma, serum, and spinal fluid.
[0025] In some specific embodiments, the detected amount includes the presence or absence, relative abundance, or concentration of the biomarker.
[0026] The present invention has screened biomarkers for predicting the risk of endometrial cancer recurrence from blood. These biomarkers have significant differences in the blood of the population with and without endometrial cancer recurrence. By collecting a blood sample, it is possible to predict or assist in diagnosing whether an individual has endometrial cancer recurrence or not by detecting these biomarkers in the individual's blood.
[0027] Further, the detection methods generally include radiological methods, immunological methods, fluorescence methods, flow fluorescence methods, latex turbidimetry, biochemical methods, enzymatic methods, hybridization methods, gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), chromatography, chemiluminescence methods, magnetoelectric methods or photoelectric conversion methods.
[0028] The presence, absence or level of these biomarkers here is a relative concept. For example, when comparing the recurrence group and non-recurrence group of endometrial cancer, the levels of these specific biomarkers are compared with the recurrence group or non-recurrence group of endometrial cancer as a reference. There may be some biomarkers with relatively higher levels in endometrial cancer recurrence compared to the non-recurrence group of endometrial cancer, and this increase is statistically significant, such as a significant or highly significant increase. Therefore, when judging these biomarkers, if it is a single biomarker, if the probability of a certain risk occurrence increases and the level of this biomarker changes, this change may be a relative increase or may also be a relative decrease, and the difference in this relative increase or relative decrease is significantly different, and of course it can also be highly significantly different. Therefore, no matter what means are used for detection, a pre-specified value (cut-off value) can be used as a standard. If it is higher than this value, it is considered that the level has changed, and such results can all be used for prediction or diagnosis.
[0029] Therefore, in some aspects, the biomarkers of the present invention can be obtained by detecting the biomarker levels in a sample by any known method, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strips, immunoreaction chips, capillary electrophoresis, infrared spectroscopy, etc. As long as it can be used to detect the protein biomarker levels in a sample, it can be used for the diagnosis of endometrial cancer recurrence and non-recurrence. As long as the protein biomarker levels in the sample can be detected, it can be used to predict or diagnose the probability of occurrence of a certain disease. It can be understood that this detection is for individual samples, and then compared with a pre-set standard, and the comparison results are used to judge or predict the occurrence status of the disease. For example, it can be used to predict the probability of endometrial cancer recurrence risk. Such prediction or diagnosis is whether it will occur within a certain time. Of course, such detection can be continuous detection, and the progress of the disease can be inferred from the changes in the levels of certain substances.
[0030] In some ways, the relative abundance is the peak area of the biomarker in the detection map obtained by high performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a certain biomarker measured in a control sample is 300, and the average peak area measured in a recurrent endometrial cancer sample is 1800, then the abundance of this biomarker in the sample is considered to be 6 times that in the control sample.
[0031] The second aspect of the present invention provides the use of a protein as a biomarker in the preparation of a product for predicting the risk of recurrence or metastasis of endometrial cancer, and the protein includes one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB. The genes of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB can be used for the auxiliary judgment of the risk of recurrence or metastasis of endometrial cancer, the evaluation of drug efficacy, etc. The inventors of the present invention have found that MUC16, CTSG, B3GNT2, CORO1A, and SFTPB are closely related to the risk of recurrence or metastasis of endometrial cancer.
[0032] The third aspect of the present invention provides a product for predicting the risk of recurrence or metastasis of endometrial cancer, including a substance for detecting a biomarker, and the biomarker includes one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB.
[0033] In some ways, the biomarker includes MUC16, CTSG, B3GNT2, and CORO1A.
[0034] The fourth aspect of the present invention provides a biomarker combination for predicting the risk of recurrence or metastasis of endometrial cancer, including MUC16, CTSG, B3GNT2, and CORO1A.
[0035] The fifth aspect of the present invention provides a method for constructing a prediction model for the risk of recurrence or metastasis of endometrial cancer for non-disease diagnosis purposes, including the following:
[0036] 1) Construct a sample data set based on the detection amount of biomarkers in biological samples, and the biomarkers include one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB;
[0037] 2) Divide the data set into a test set and a training set, and construct and train the prediction model for the risk of recurrence or metastasis of endometrial cancer through machine learning methods.
[0038] In some embodiments, the biomarker includes MUC16, CTSG, B3GNT2, and CORO1A.
[0039] In some embodiments, the machine learning method is selected from at least one of gradient boosting algorithm, random forest algorithm, support vector machine algorithm, decision tree algorithm, K-nearest neighbor algorithm, logistic regression algorithm, and neural network algorithm. For example, it can be selected from the gradient boosting algorithm. Unlike generalized linear models that can output model formulas and cut-off values, all calculations in the gradient boosting algorithm are directly completed by machine learning. By directly inputting the detection values into the software system, the prediction results can be directly obtained.
[0040] In some embodiments, multiple machine learning methods are used to construct a prediction model for the recurrence or metastasis risk of endometrial cancer. It is preliminarily confirmed that for any one of the newly identified biomarkers selected alone, the change in its concentration can be used to distinguish between the population with recurrent endometrial cancer and the non-recurrent population, indicating that these biomarkers have extremely high diagnostic value.
[0041] The present invention discovers that by detecting the detected amount of biomarkers in a biological sample, then inputting the detected amount into the formula of the prediction model for the recurrence or metastasis risk of endometrial cancer to obtain the prediction score Y of the prediction model, and comparing this prediction score Y with the threshold (cut-off) defined by the Youden index. If the prediction score Y > the threshold (cut-off), it is judged as recurrent endometrial cancer; if the prediction score Y ≤ the threshold (cut-off), it is judged as non-recurrent endometrial cancer.
[0042] In some embodiments, the equation of the prediction model for the recurrence risk of endometrial cancer is:
[0043]
[0044] Where Y is the prediction score, i represents the i-th biomarker, m represents the number of biomarkers (m = 5), Xi represents the detected value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker, and b is a constant 3.8638532; the coefficients of the 4 biomarkers are:
[0045]
[0046] When Y ≤ the cut-off value, the subject to be tested is non-recurrent endometrial cancer; when Y > the cut-off value, the subject to be tested is recurrent endometrial cancer. Specifically, the cut-off value is 0.4773883.
[0047] In some embodiments, it further includes step 3) testing the prediction model for the recurrence or metastasis risk of endometrial cancer using a test set. After training is completed, the trained prediction model is verified using a test set. At the same time, the AUC value, specificity, and sensitivity are used as evaluation indicators to evaluate the effect of the prediction model.
[0048] In the present invention, the prediction model constructed based on the combination of 4MP can distinguish between recurrence and non-recurrence after endometrial cancer surgery, with an AUC of 0.945, a sensitivity of 0.959, and a specificity of 0.961 in the training group; and an AUC of 0.924, a sensitivity of 0.872, and a specificity of 0.846 in the training group.
[0049] The sixth aspect of the present invention provides a device for predicting the risk of recurrence or metastasis of endometrial cancer, including a data acquisition unit and a calculation unit;
[0050] The data acquisition unit is used to obtain the detection amount data of the biomarkers in the biological sample of the subject as described above, and the biomarkers include one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB;
[0051] The calculation unit is used to calculate and output a prediction score for the risk of recurrence or metastasis of endometrial cancer for the subject based on the detection amount data, and make a judgment according to the threshold
[0052] In some embodiments, the biomarkers include MUC16, CTSG, B3GNT2, and CORO1A.
[0053] In some embodiments, the prediction device further includes a data storage unit and a data output unit; the data storage unit is used to store the detection amount (or detection value) of the biomarker; the data input interface is used to input the detection value of the biomarker, and the data output unit is used to output the prediction result.
[0054] Furthermore, the detection value is the presence or absence, relative abundance, or concentration value of each biomarker.
[0055] The seventh aspect of the present invention provides a device, including a processor and a memory, where the memory is used to store a computer program, and it is characterized in that the processor is used to execute the computer program stored in the memory so that the device executes the prediction method or the construction method as described above.
[0056] The eighth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program, when processed, executes the prediction method or the construction method as described above.
[0057] In some embodiments, the computer-readable storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0058] In some embodiments, the being processed refers to being executed by one or more processors.
[0059] The ninth aspect of the present invention provides a method for predicting the risk of recurrence or metastasis of endometrial cancer, comprising the following steps:
[0060] S1. Obtain the detection amount data of biomarkers in the biological sample of the subject, wherein the biomarkers include one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB;
[0061] S2. Process the detection amount data by using the endometrial cancer recurrence risk prediction model obtained by the construction method as described above to output the prediction result of the risk of recurrence or metastasis of endometrial cancer.
[0062] In some embodiments, the biomarkers include MUC16, CTSG, B3GNT2, and CORO1A.
[0063] In another aspect, the present invention provides a system for predicting the recurrence risk of endometrial cancer, the system comprising a data analysis module, and the data analysis module is used for analyzing the detection values of biomarkers, wherein the biomarkers include MUC16, CTSG, B3GNT2, and CORO1A.
[0064] Further, the data analysis module uses the detection values of biomarkers of known samples as a training set, and divides them into an endometrial cancer recurrence group and an endometrial cancer non-recurrence group according to whether endometrial cancer recurs, analyzes the relationship between the detection values of the endometrial cancer recurrence group and the endometrial cancer non-recurrence group, and constructs a model.
[0065] In some embodiments, a combined diagnostic model for predicting the recurrence risk of endometrial cancer is constructed by combining multiple machine learning methods, and it is preliminarily confirmed that the change in the concentration of any one of the selected novel biomarkers alone can be used to distinguish high-risk and low-risk populations of endometrial cancer recurrence, indicating that these biomarkers have extremely high diagnostic value.
[0066] In some embodiments, the equation of the constructed model is:
[0067]
[0068] Wherein, Y is the predicted score, i represents the i-th biomarker, m represents the number of biomarkers (m = 5), Xi represents the detection value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker, and b is a constant 3.8638532; the coefficients of the 4 biomarkers are:
[0069]
[0070] When Y ≤ the cut-off value, the subject has no recurrence of endometrial cancer; when Y > the cut-off value, the subject has recurrence of endometrial cancer. Specifically, the cut-off value is 0.4773883.
[0071] Furthermore, the system further includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0072] On the other hand, the present invention provides a use of a biomarker for preparing a reagent for predicting non-recurrence, in-situ recurrence or distant metastasis of endometrial cancer patients after surgery, and the biomarker includes any one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB.
[0073] On the other hand, the present invention provides a kit for predicting non-recurrence, in-situ recurrence or distant metastasis of endometrial cancer patients after surgery, and the kit includes a detection reagent for the biomarker as described above.
[0074] On the other hand, the present invention provides a biomarker combination for predicting non-recurrence, in-situ recurrence or distant metastasis of endometrial cancer patients after surgery, and the combination includes any one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB.
[0075] On the other hand, the present invention provides a system for predicting non-recurrence, in-situ recurrence or distant metastasis of endometrial cancer after surgery, and the system includes a data analysis module, and the data analysis module is used to analyze the detection values of biomarkers, and the biomarkers include MUC16, CTSG, B3GNT2, CORO1A, and SFTPB.
[0076] Furthermore, the distant metastasis also includes in-situ recurrence and distant metastasis at the same time, and as long as distant metastasis occurs, it is classified into the distant metastasis group.
[0077] Furthermore, the data analysis module uses the detection values of biomarkers of known samples as a training set, and according to the situation of liver cancer patients after surgery, it is divided into a non-recurrence group, an in-situ recurrence group, and a distant metastasis group, analyzes the relationship between the detection values of the non-recurrence group, the in-situ recurrence group, and the distant metastasis group, and constructs a model.
[0078] Furthermore, the model is constructed based on the gradient boosting algorithm.
[0079] Unlike generalized linear models, the gradient boosting algorithm cannot output the model formula and cut-off value. All calculations are directly performed by machine learning. By directly inputting the detection values into the software system, the prediction results can be directly obtained.
[0080] Furthermore, the system further includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0081] The beneficial effects of the present invention are as follows:
[0082] 1. The present invention has screened 5 new biomarkers that can predict the risk of recurrence or metastasis of endometrial cancer, and developed a new combination of protein biomarkers, which can effectively evaluate and diagnose the risk of recurrence or metastasis of endometrial cancer, effectively distinguish between recurrent and non-recurrent endometrial cancer, and effectively distinguish between in-situ recurrence and distant metastasis of endometrial cancer. This not only improves the diagnostic accuracy of recurrence and metastasis of endometrial cancer, but also provides an important biomarker basis for personalized treatment and prognostic monitoring of endometrial cancer patients. In addition, the discovery and application of these biomarkers are expected to improve the treatment effect and quality of life of endometrial cancer patients, and provide a scientific basis for the early diagnosis and treatment strategy selection of endometrial cancer. Compared with traditional detection methods, the risk of misdiagnosis and missed diagnosis is reduced, providing strong support for the early detection and intervention of diseases.
[0083] 2. The combined differential diagnosis model of 4 biomarkers constructed by the present invention is convenient, fast, and the detection results are highly consistent with the clinical gold standard detection results. At the same time, the cost of predicting the risk of recurrence or metastasis of endometrial cancer is significantly reduced, and it has good application prospects. The combination of biomarkers of the present invention is superior to the diagnosis of broad-spectrum tumor markers (CEA, CA125) or the combined model of broad-spectrum tumor markers.
[0084] 3. Based on the screened biomarkers for predicting the risk of recurrence or metastasis of endometrial cancer, a three-classification model that can simultaneously distinguish between the non-recurrence group, the in-situ recurrence group, and the distant metastasis group is further constructed, providing a more effective and accurate prediction and diagnosis mode.
[0085] Detailed description
[0086] (1) Diagnosis or detection
[0087] The diagnosis or detection here refers to the detection or assay of biomarkers in a sample, or the content of the target biomarker, such as the absolute content or relative content, and then it is determined whether the individual providing the sample may have or suffer from a certain disease, or the likelihood of having a certain disease, based on the presence or quantity of the target biomarker. The meanings of diagnosis and detection here can be interchanged. The result of this detection or the result of the diagnosis cannot be directly used as the direct result of having a disease, but is an intermediate result. To obtain the direct result, other auxiliary means such as pathology or anatomy are required to confirm the presence of a certain disease. For example, the present invention provides a variety of new biomarkers related to the risk of recurrence or metastasis of endometrial cancer, and the change in the content of these biomarkers has a direct correlation with whether the patient belongs to the population at risk of recurrence or metastasis of endometrial cancer.
[0088] (2) The association between the biomarker or biomarker or differential protein and the risk of recurrence or metastasis of endometrial cancer
[0089] The biomarker, biomarker, and differential protein have the same meaning in the present invention. The association here refers to the direct correlation between the appearance or change in the content of a certain biomarker in a sample and a specific disease. For example, a relative increase or decrease in the content indicates a relatively higher likelihood of having this disease compared to the healthy population.
[0090] If multiple different biomarkers in the sample appear simultaneously or there is a relative change in their content, it also indicates a relatively higher likelihood of having this disease compared to the healthy population. That is to say, among the types of biomarkers, some biomarkers have a strong association with the disease, some have a weak association with the disease, or some even have no association with a specific disease. One or more of those biomarkers with a strong association can be used as biomarkers for diagnosing the disease, and those with a weak association can be combined with the strong biomarkers to diagnose a certain disease, increasing the accuracy of the detection result.
[0091] For the numerous biomarkers found in the serum of the present invention, these biomarkers can all be used to distinguish between recurrent and non-recurrent endometrial cancer, as well as non-recurrent, in-situ recurrence, and distant metastasis. These biomarkers can be directly detected or diagnosed individually as single biomarkers. Selecting such a biomarker indicates that the relative change in the content of this biomarker has a strong correlation with the risk of recurrence or metastasis of endometrial cancer. Of course, it is understandable that simultaneous detection of one or more biomarkers with a strong correlation with the risk of recurrence or metastasis of endometrial cancer can be selected. Normally understood, in some ways, selecting biomarkers with a strong correlation for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90%, or 95% accuracy. This indicates that these biomarkers can obtain an intermediate value for diagnosing a certain disease, but it does not mean that it can directly confirm the presence of a certain disease.
[0092] Of course, differential proteins with a larger ROC value can also be selected as diagnostic biomarkers. The so-called strength or weakness is generally calculated and confirmed through some algorithms, such as the contribution rate or weight analysis of the biomarker to the prediction of the risk of recurrence or metastasis of endometrial cancer. Such calculation methods can include significance analysis (p-value or FDR value) and fold change. Multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA), and orthogonal partial least squares discriminant analysis (OPLS-DA). Of course, other methods are also included, such as ROC analysis, etc. Of course, other model prediction methods are also possible. When specifically selecting biomarkers, the differential proteins disclosed in the present invention can be selected, or other existing well-known biomarker combinations can be selected or combined for prediction through model methods.
[0093] (3) Definition of disease terms
[0094] Endometrial cancer: Endometrial cancer is a malignant tumor that originates from the endometrium of women and is one of the most common types of gynecological malignancies. Endometrial cancer can affect women of any age, but most cases occur in women during menopause or after menopause. The main pathological types of endometrial cancer include endometrioid adenocarcinoma, serous adenocarcinoma, clear cell carcinoma, etc. Its typical symptoms include abnormal vaginal bleeding, vaginal discharge, and lower abdominal pain, etc. Due to the lack of obvious early symptoms, endometrial cancer is often discovered at an advanced stage, which makes treatment more difficult and the prognosis is poor. The diagnosis of endometrial cancer usually requires a combination of imaging examinations (such as ultrasound, MRI) and pathological examinations (obtaining tissue samples through fractional curettage or hysteroscopic biopsy). Stage I / II endometrial cancer is an early lesion confined to the uterus. The tumor spreads downward from the uterine body to the cervical stroma but does not extend beyond the uterus and does not involve the parametrium or vagina. Surgery can be curative, but subsequent treatment needs to be individualized according to the pathological results.
[0095] Recurrence of endometrial cancer: It refers to the situation where cancer cells grow again after initial treatment. This recurrence can occur at the primary site of endometrial cancer (in-situ recurrence) or in other parts of the body (distant metastasis). Most recurrences of endometrial cancer (about 70%-80%) occur within 2-3 years after surgery. As the time after surgery extends, the risk of recurrence gradually decreases, and the risk of recurrence is significantly reduced after 5 years.
[0096] Metastasis of endometrial cancer: It refers to the process by which malignant tumor cells spread from the primary site to other parts of the body. These cells move through the bloodstream or lymphatic system and form new tumors in new locations. It is divided into in-situ metastasis and distant metastasis.
[0097] In-situ metastasis of endometrial cancer: Local metastasis refers to the recurrence of tumor cells at the primary site or in the adjacent area, which may include the uterus, ovaries, fallopian tubes, or other structures within the pelvis. In-situ metastasis usually means that the cancer has not spread widely to other parts of the body, but active treatment is still required to prevent further spread.
[0098] Distant metastasis of endometrial cancer: It refers to the situation where tumor cells have spread to more distant parts of the body, such as the peritoneum, lymph nodes, lungs, liver, or other organs. Common distant metastasis sites of endometrial cancer include the peritoneum (resulting in ascites symptoms), lymph nodes, lungs, and liver. Distant metastasis usually indicates that the cancer has progressed to an advanced stage, with increased treatment difficulty and a poor prognosis.
[0099] Therefore, early detection and intervention for the recurrence or metastasis of endometrial cancer are crucial for improving the treatment effect and increasing the survival rate. The biomarker of the present invention has an accurate and specific predictive effect on the recurrence or metastasis of endometrial cancer, which is beneficial to improving the treatment effect of patients and has important significance for improving the prognosis of patients.
[0100] (4) The gold standard for the diagnosis of endometrial cancer: It is pathological diagnosis (i.e., pathological confirmation by puncture biopsy or surgical resection specimens), observing the presence of cancer cells under a microscope, and clarifying the nature of the tumor in combination with immunohistochemistry or molecular detection. Description of the Drawings
[0101] Figure 1 It is a volcano plot for the differential analysis of protein markers in the high-risk group and low-risk group of endometrial cancer recurrence in Example 1.
[0102] Figure 2 It is a graph showing the ROC and OPLS-DA analysis results in the high-risk group and low-risk group of endometrial cancer recurrence in Example 1.
[0103] Figure 3It is the optimization result graph of the hyperparameter combination of the glmnet algorithm in Example 2.
[0104] Figure 4 It is the ROC curve graph of the endometrial cancer recurrence risk prediction model in the training group in Example 2.
[0105] Figure 5 It is the ROC curve graph of the endometrial cancer recurrence risk prediction model in the training group in Example 2. Specific implementation manners
[0106] The present invention will be further described in detail below with reference to the drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not limit it in any way. The reagents used in this embodiment are all known products and are obtained by purchasing commercially available products.
[0107] Example 1. Screening biomarkers for endometrial cancer recurrence risk using proteomics
[0108] By collecting plasma samples of endometrial cancer patients after radical surgery, enriching low-abundance proteins based on the method of removing high-abundance proteins by immunoaffinity chromatography, detecting the protein abundance in the samples by a high-performance liquid chromatography-mass spectrometry tandem device, and analyzing the differences in their abundances between endometrial cancer recurrence and non-recurrence patients, differential proteins are screened out. The specific steps are as follows:
[0109] 1.1 Sample collection
[0110] A total of 150 blood samples of endometrial cancer patients (stage I / II patients) were collected 2 weeks after surgical resection treatment. All endometrial cancer patients were confirmed by pathology of living tissues. Approximately 2 ml of peripheral blood samples of the subjects were collected, placed in a vacuum tube containing EDTA anticoagulant, mixed well, centrifuged at 120 g for 10 minutes at room temperature, and the supernatant was taken, repeated twice. Then, centrifuged at 360 g for 20 minutes. Then, platelet samples were collected in centrifuge tubes and stored at -80 °C for standby. All 150 patients were female, with an average age of 55 years (30 - 79 years), 84 stage I patients, and 66 stage II patients. All enrolled patients signed informed consent forms. Among them, all endometrial cancers were patients diagnosed by pathological histology. Inclusion criteria: (a) No history of other malignant tumors; (b) No patients with combined other malignant tumors or autoimmune diseases.
[0111] In the following three years, patients were followed up every six months, with regular CA-125 and imaging examinations. Once recurrence or metastasis was detected, pathological histological confirmation was required. A total of 18 stage I / II patients had recurrence within three years after surgery, including 10 cases of local in-situ recurrence and 8 cases of distant metastasis (including 5 cases of peritoneal metastasis, 2 cases of lymph node metastasis, and 1 case of liver metastasis). Then, the samples stored at -80°C were taken out and divided into two groups: the samples of 132 patients without recurrence were classified into the low-risk group of colorectal endometrial cancer recurrence, and the samples of 18 patients with recurrence or metastasis were classified into the high-risk group of colorectal endometrial cancer recurrence, which were used for the screening of proteomic biomarkers.
[0112] 1.2 Sample processing and enzymatic digestion
[0113] 1) First, the plasma samples were centrifuged on a centrifuge for 15 minutes (15,000 g), and the supernatant was taken and filtered, followed by immunoaffinity chromatography to remove 14 high-abundance proteins.
[0114] 2) Then, a concentrator with a cut-off molecular weight of 3 kDa was used to concentrate the low abundance to 350 μL on a centrifuge (4000 g, 1 hour).
[0115] 3) The concentrated solution was recovered, and a desalting column with a cut-off molecular weight of 7 kDa was used to perform buffer exchange on a centrifuge (1000 g, 2 minutes). The replacement solution was AEX-A (20 mM Tris, 4 M urea, 3% isopropanol, pH 8.0).
[0116] 4) Using AEX-A as a blank, the protein concentration in the samples was measured using the BCA method. According to the sample grouping, 25 μL of TCEP was added to the samples, and the samples were incubated at 37°C for 30 minutes for protein reduction. Then, the corresponding TMT16-plex reagent was added, and the samples were incubated in the dark at room temperature for 1 hour for the TMT labeling reaction. Then, the samples were subjected to buffer exchange using a Zeba column, and the replacement solution was AEX-A. After mixing the samples labeled with TMT 16-plex, 2 mL of AEX-A was added to the mixed samples, and the final volume was 5.5 mL.
[0117] 5) The samples were filtered using a 0.22 m filter and the samples labeled with TMT 16-plex were separated using a 2D-HPLC system. The collected fractions were freeze-dried, and finally, a mixture of Trypsin-Lysin C enzymes was added, and the samples were incubated at 37°C for 5 hours for enzymatic digestion. 5 μL of 10% TFA was added to terminate the enzymatic digestion reaction.
[0118] 6) A total of 60 enzymatically digested 2D-HPLC fractions were used for nanoLC-MS / MS analysis.
[0119] 1.3 LC-MS / MS Data Acquisition
[0120] Each sample obtained in Step 1.2 was separated using a nanoflow liquid chromatography system, Easy nLC-1200, and detected by on-line coupling with a high-resolution mass spectrometer, Q Exactive HF-X. The specific steps are as follows:
[0121] 1) Separation: Mobile phase - A was 0.1% formic acid aqueous solution, and mobile phase - B was 0.1% formic acid acetonitrile aqueous solution (80% acetonitrile and 20% water). The chromatographic column consisted of a trapping column and an analytical column and was equilibrated with 100% mobile phase - A. The sample was loaded onto the trapping column (100 μm ID × 4 cm L, C18, 3 μm, 100 Å) by an autosampler and separated by the analytical column (75 μm ID × 25 cm L, C18, 3 μm, 100 Å) at a flow rate of 300 nL / min.
[0122] 2) The sample after chromatographic separation was subjected to mass spectrometry analysis using a Q Exactive HF-X mass spectrometer. The detection mode was positive ion mode. The parent ion scan range was 350 - 1800 m / z. The resolution of the first-stage mass spectrometry was 120,000 at 200 m / z. The AGC (Automatic gain control) target was 3e6, and the Maximum IT was 50 ms. The dynamic exclusion time was 40 s. The mass-to-charge ratios of polypeptides and polypeptide fragments were collected according to the data-dependent acquisition (DDA) method: 20 secondary spectra (MS / MS, MS2 scan) were collected after each full scan (first-stage mass spectrometry). The MS2 ActivationType was HCD, the Isolation window was 0.7 m / z, the resolution of the second-stage mass spectrometry was 30,000 at 200 m / z, the AGC target was 1e5, the Maximum IT was 65 ms, the Fixed first mass was 110.0 m / z, the Normalized Collision Energy was 32 ev, the Minimum AGC target was 2.00e4, Charge exclusion was 1, 6 - 8, >8, Multiple charge states was one charge state only, Peptide match was preferred, and Exclude isotopes was on.
[0123] 1.4 Data Preprocessing
[0124] The MS / MS data was searched using Maxquant (v1.6.15.0). The data type was DIA proteomics data based on MS / MS reporter ion quantification. For the MS / MS spectra used for quantification, the proportion of precursor ions in the MS1 spectra was required to be greater than 75%. The database source was Homo_sapiens_9606_proteome_gene from the Uniprot database (release: 2021-10-14, sequence: 20,437), and a common contaminant library was added to the database. Contaminant proteins were removed during data analysis. The digestion method was set to Trypsin / P; the maximum number of missed cleavage sites was set to 2; the mass error tolerances for precursor ions in the first search and main search were set to 20 ppm and 5 ppm, respectively, and the mass error tolerance for MS / MS fragment ions was 20 ppm. The fixed modification was cysteine alkylation, and the variable modifications were methionine oxidation and protein N-terminal acetylation. The false discovery rates (FDRs) for protein identification and PSM identification were both set to 1%.
[0125] 1.5 Differential analysis
[0126] A combination of univariate analysis and multivariate statistical analysis was used to screen for differential proteins and transcripts. Univariate analysis mainly included the significance analysis (p-value or FDR value) and fold change of characteristic molecules in different groups, and multivariate statistical analysis mainly included receiver operating characteristic curve (ROC) analysis and Boruta feature selection based on the random forest algorithm. All statistical analyses were performed using R. The specific R-related information is shown in Table 1.
[0127] Table 1. R and its related information
[0128]
[0129] The variable importance for the projection (VIP) was calculated to measure the influence intensity and explanatory ability of the expression patterns of each protein on the classification and discrimination of each group of samples. Further, the Wilcoxon rank sum test was performed to obtain the corrected p-value (FDR). Sixty-three downregulated proteins and 57 upregulated proteins were screened according to the conditions of FDR < 0.01 and Fold change > 2 (see details in Figure 1 ).
[0130] To evaluate the role of each protein biomarker in the diagnosis and prediction of the recurrence risk of endometrial cancer, the ROC and Boruta analysis methods were used to evaluate each protein biomarker. The results are shown in Figure 2The abscissa is the AUC obtained from the ROC analysis, the ordinate is the -log10(FDR) calculated by the Wilcoxon test, and the size of the points represents the VIP value obtained from the Boruta analysis. Further screening was performed according to VIP > 3 and AUC > 0.6, and a total of 5 more significant candidate protein markers were found, as shown in Table 2 for details.
[0131] Table 2 Differential markers for the recurrence risk of endometrial cancer
[0132] Among them, the smaller the FDR value and / or the larger the VIP value, to a certain extent, it indicates that the difference in the protein between the high-risk and low-risk groups of endometrial cancer recurrence is more significant, and at the same time, it also indicates that the protein may have higher diagnostic value.
[0133] Example 2. Construction and verification of a prediction model for the recurrence risk of endometrial cancer
[0134]
[0135] In this example, a recurrence risk prediction model was constructed based on the combination of the 5 protein markers screened in Example 1 for research.
[0136] Although a single biomarker can also distinguish the recurrence risk of endometrial cancer patients after surgery, generally speaking, combining multiple biomarkers can achieve higher accuracy in discrimination or prediction. However, for a single biomarker with higher accuracy in predicting the recurrence risk of endometrial cancer, its role in the combination with one or more other biomarkers may not necessarily be greater. At the same time, it is not the case that the more biomarkers there are, the higher the prediction accuracy (AUC value) of the combination. Therefore, a large number of verification experiments are still needed.
[0137] 2.1 Data acquisition
[0138] A total of 252 blood samples were collected from endometrial cancer (stage I / II patients) 2 weeks after surgical treatment. All enrolled patients signed informed consent forms. Among them, 32 patients had recurrence within three years after surgery (14 had local in-situ recurrence and 18 had metastasis). Thus, they were divided into two groups: 220 samples from non-recurrent patients were classified into the low-risk group of endometrial cancer recurrence, and 32 samples from patients with recurrence or metastasis were classified into the high-risk group of endometrial cancer recurrence. They were randomly divided into a training group and a test group. The training group included 110 samples from the low-risk group of endometrial cancer recurrence and 16 samples from the high-risk group of endometrial cancer recurrence. The test group included 110 samples from the low-risk group of endometrial cancer recurrence and 16 samples from the high-risk group of endometrial cancer recurrence.
[0139] The concentrations of 5 proteins in the serum of 250 samples were obtained using the same steps as in steps 1.2 and 1.3 of Example 1.
[0140] 2.2 Data statistical analysis
[0141] In the training group, a combined diagnostic model of multiple recurrence risk markers for endometrial cancer was constructed using a combination of multiple machine learning methods. The predicted probability values were used to estimate the area under the receiver operator characteristic (ROC) curve with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate diagnostic model.
[0142] Using the training group, the Youden index (YI) was calculated to determine the cut-off value for predicting the probability of differentiating between the high-risk group and the low-risk group of endometrial cancer recurrence. In addition, the ROCs of the models formed by combining different markers were constructed and compared. Standard descriptive statistical data such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R 3.6.1, and a p-value less than 0.05 was considered statistically significant.
[0143] 2.3 Construction of the prediction model
[0144] The steps for constructing the endometrial cancer recurrence risk prediction model are as follows:
[0145] S101. From the 5 protein markers of the samples in the training group, randomly select the concentration matrices of 2 to 5 markers as the original training data set.
[0146] S102. Select the generalized linear model (glmnet) algorithm for constructing the prediction model and the grid search range during the hyperparameter optimization process of the algorithm. In this step, the grid search range for hyperparameter optimization of each algorithm is set as shown in Table 3.
[0147] Table 3. Parameter grid of the glmnet algorithm
[0148]
[0149] S103. According to the algorithm and the hyperparameter setting range set in step S102, select one of the hyperparameter combination methods as the parameters for constructing the prediction model.
[0150] S104. Split the original data set into K subsets according to the K-fold cross-validation mechanism. To ensure that the proportion of majority-class samples and minority-class samples in each subset is the same as that in the original data set, the stratified K-fold cross-validation mechanism needs to be used for data splitting.
[0151] S105. Select one of the K training data subsets obtained by splitting according to step S104 as the validation set Ddev.
[0152] S106. Merge the training data subsets not selected in step S105 to form the training data pool Dtrainl.
[0153] S107. Based on the training dataset Dtrain obtained in step S106, construct a prediction model based on the selected supervised classification algorithm and hyperparameters.
[0154] S108. According to the prediction model obtained in step S107, evaluate it on the validation set Ddev to obtain the AUC value, and store the current prognosis prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 is to evaluate according to the prediction model obtained in step S107 on the validation set determined in the current iteration, and store both the model and the evaluation result in the prediction model pool for future use in selecting the base prediction model. The evaluation mentioned in this step can be the AUC value or other reasonable metrics for evaluating the model performance.
[0155] S109. Determine whether each subset has been used as the validation set. Step S109 is to determine whether the K subsets obtained in step S104 have all been used as the validation set for model training. If all subsets have been used as the validation set and the training is completed, execute step S110; if there is a subset that has not been used as the validation set, execute step S105. This step ensures that each sample in the original dataset has been used as the validation set, improves the model stability, and prevents the model from overfitting to a certain subset.
[0156] S110. Take the average value of the AUCs of all models in the obtained prediction model pool Pool as the final performance evaluation value of the model in this combination method. And store the model parameters and the final performance evaluation AUC value in the optimal model pool Poolbest.
[0157] S111. Determine whether prediction models have been constructed for each hyperparameter combination method. Step S111 is to determine whether prediction models have been constructed for all the algorithms and the corresponding hyperparameter combination methods obtained in step S102. If all combination methods have completed the model construction, execute step S112; if there is a combination method that has not completed the model construction, execute step S103.
[0158] S112. After completing step S111, it is necessary to check whether all hyperparameter combination methods have been used to construct the prediction model. If it is found that there are still hyperparameter combinations for which the model has not been constructed, then it is necessary to return to step S103 to continue constructing the remaining models. If models have been constructed for all hyperparameter combinations, then step S113 can be continued to select the best model from the model pool.
[0159] S113. From the model set Poolbest obtained in step S112, select the model with the largest AUC value as the final prediction model for the recurrence risk of endometrial cancer.
[0160] S114. Repeat all the above steps until modeling is completed for all combination forms of the markers.
[0161] By performing the above prediction model construction steps, the optimal models constructed for all combination forms of the markers are obtained. To compare the performance of the models under these different marker combination forms, the ROC method was used to evaluate the AUC values of these prediction models in the training group, and the results are shown in Table 4.
[0162] Table 4 Comparison of the areas under the ROC curves of models constructed with different marker combinations in the training group
[0163]
[0164] Table 4 shows the AUC, sensitivity, and specificity of the models formed by different marker combinations. It is found that the combination (4MP-2) MUC16+CTSG+B3GNT2+CORO1A has the highest AUC, sensitivity, and specificity. Continuing to increase the markers cannot further improve the diagnostic efficacy, and the number of 4MP-2 markers is less, so it is the most preferred model.
[0165] 2.4 Optimization of model parameters
[0166] For the optimal marker combination form MUC16+CTSG+B3GNT2+CORO1A in step 2.3, based on this marker combination, models constructed under 9 different combinations of glmnet algorithm hyperparameters were analyzed, and the performance of the models was evaluated by the AUC value (the AUC was calculated using the 10-fold cross-validation method during the modeling process), and the results are shown in Table 5 and Figure 3 as shown.
[0167] Table 5 AUC of models constructed under different combinations of glmnet algorithm hyperparameters
[0168]
[0169] As can be seen from Table 5, when the hyperparameter combination of the glmnet algorithm is alpha = 1 and lambda = 0.0547, the AUC reaches the maximum value of 0.945.
[0170] The equation for constructing the model based on the optimal hyperparameter combination is:
[0171]
[0172] Among them, Y is the predicted score, i represents the i-th biomarker, m represents the number of biomarkers (m = 5), Xi represents the measured value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker (see Table 6), and b is the constant 3.8638532.
[0173] Table 6 Coefficients of 5 biomarkers in the model
[0174]
[0175] The complete model equation is: y = 4.967 MUC16 + 6.255 CTSG + 2.163 B3GNT2 + 7.522 CORO1A + 3.8638532
[0176] Determination of the diagnostic threshold for the endometrial cancer recurrence risk prediction model:
[0177] The ROC curve was plotted using the predicted scores in the training group, and the optimal prediction cut-off value of 0.4773883 was set according to the Youden index value. The results are shown in Figure 4 . When the predicted score of the prediction model ≤ 0.4773883, it is considered that the recurrence risk of endometrial cancer in the subject to be tested is low; when the predicted score of the prediction model > 0.4773883, it is considered that the subject to be tested has a high recurrence risk of endometrial cancer. As can be seen from Figure 4 , the AUC of the prediction model in the training group is 0.945, the sensitivity is 0.959, and the specificity is 0.961.
[0178] 2.6 Validation of the endometrial cancer recurrence risk prediction model
[0179] The constructed optimal model was validated in the training group, and the ROC curve was plotted as shown in Figure 5 .
[0180] As can be seen from Figure 5 , the AUC of the model in the training group is 0.921, the sensitivity is 0.872, and the specificity is 0.846.
[0181] In summary, the endometrial cancer recurrence risk prediction model constructed using 4 protein biomarkers has good prediction performance and accuracy, and has the best diagnostic efficacy.
[0182] Example 3: Construction and Validation of a Three - Classification Diagnostic Model
[0183] In this example, an attempt was made to construct a three - classification combined diagnostic model for differentiating the non - recurrence group, in - situ recurrence group, and distant metastasis group of endometrial cancer prognosis. The specific process includes the following: (1) Construction and screening of the optimal diagnostic model; (2) Validation of the effect of the optimal diagnostic model. The specific screening process and results are as follows (in the present invention, the AUC value was used as the evaluation index for the binary classification model in Example 2; when constructing a three - classification model, since it involves multiple categories, the AUC value is usually not applicable, and in this example, indicators such as sensitivity, specificity, accuracy, and consistency were used to measure the diagnostic efficacy of the model):
[0184] 3.1 Construction and Screening of the Three - Classification Diagnostic Model
[0185] For the test cohort of 382 endometrial cancer patients, all enrolled patients signed informed consent forms. Among them, 34 patients had a recurrence within two years after surgery (16 had in - situ local recurrence and 18 had metastases). Thus, they were divided into two groups: 348 patient samples without recurrence were classified as the low - risk group of endometrial cancer recurrence, and 34 patients with recurrence or metastasis were classified as the high - risk group of endometrial cancer recurrence. They were randomly divided into a training group and a test group. The training group included 174 patients in the low - risk group of endometrial cancer recurrence and 17 patients in the high - risk group of endometrial cancer recurrence, including 8 cases of in - situ recurrence and 9 cases of metastasis; the training group included 174 patients in the low - risk group of endometrial cancer recurrence and 17 patients in the high - risk group of endometrial cancer recurrence, including 8 cases of in - situ recurrence and 9 cases of metastasis.
[0186] Based on the combination form of markers MUC16 + CTSG + B3GNT2 + CORO1A screened in Example 2, this example further constructed a three - classification detection model that can effectively distinguish the non - recurrence group (low - risk group), in - situ recurrence group (in - situ recurrence group), and distant metastasis group (metastasis group) of endometrial cancer prognosis. All enrolled patients signed informed consent forms. Among them, all endometrial cancer patients were confirmed by pathological histology. Inclusion criteria: (a) No history of other malignant tumors; (b) No patients with combined other malignant tumors or autoimmune diseases.
[0187] LC - MS / MS data collection and detection were performed on the collected serum samples to obtain the concentrations of four protein markers, namely MUC16, CTSG, B3GNT, and CORO1A.
[0188] The Shapiro-Wilk test was used to evaluate the normal distribution, and the non-parametric Wilcoxon test was used to analyze the differences in blood biomarker concentrations between the non-recurrence group (low-risk group), in-situ recurrence group (in-situ recurrence group), and distant metastasis group (metastasis group) of endometrial cancer prognosis. A three-class combined diagnostic model of four biomarkers was constructed using a method combining machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using the predicted probability value with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate diagnostic model. Using the training group, the Youden index (YI) was calculated to determine the predicted probability cut-off value for distinguishing the non-recurrence group (low-risk group), in-situ recurrence group (in-situ recurrence group), and distant metastasis group (metastasis group) of endometrial cancer prognosis. In addition, the ROCs of models constructed with different biomarker combinations were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R 3.6.1, and a p-value less than 0.05 was considered statistically significant.
[0189] In this embodiment, in order to construct an optimal three-class combined diagnostic model, after comparing the models constructed by six algorithms including gradient boosting, naive Bayes, support vector machine, neural network, generalized linear, and discriminant analysis, the gradient boosting method was selected as the best supervised classification algorithm for constructing the prediction model. The grid search range for hyperparameter optimization of the gradient boosting method is shown in Table 7 below.
[0190] Table 7. Parameter grid search range of the gradient boosting method
[0191]
[0192] Through optimization and screening in terms of accuracy, consistency, sensitivity, specificity, etc., the optimal parameter combination pattern was determined as: interaction.depth 2, n.trees 150, shrinkage 0.1, n.minobsinnode 10.
[0193] The training group and the training group used two completely different batches of samples. In this embodiment, only the screening of biomarkers and the construction of the model were carried out from the training group; the samples of the training group were only used to verify the diagnostic efficacy of the model. The specific results are shown in Table 8.
[0194] Table 8. Performance evaluation table for distinguishing three classes by the model constructed by the gradient boosting method
[0195]
[0196] As can be seen from Table 8, the gradient boosting model constructed based on the four protein markers of MUC16, CTSG, B3GNT2, and CORO1A can be used to predict whether endometrial cancer patients will have no recurrence, in-situ recurrence, or distant metastasis after surgery. It also shows that the protein markers screened in the present invention can be used to distinguish the recurrence risk of endometrial cancer patients after surgical treatment, and can also be used to distinguish whether it is in-situ recurrence or distant metastasis when the recurrence risk is high (when there is both in-situ recurrence and metastasis, it is also classified into the metastasis group). However, the diagnostic efficacy for distinguishing between the in-situ recurrence group and the metastasis group is not yet ideal.
[0197] 3.2 Combined Performance of the Three-Class Classification Diagnostic Model
[0198] In order to further improve the diagnostic value of the three-class classification diagnostic model (gradient boosting) constructed by biomarker combinations of different proteins, in this example, based on the 5 protein markers screened in Example 1, the performance of the diagnostic models constructed by biomarker combinations of different proteins was compared in the training group. The specific combination forms of different models are shown in Table 9.
[0199] Table 9. Combination Forms of Different Diagnostic Models
[0200]
[0201] Comparison results of performance indicators of different diagnostic models constructed by the 5 biomarkers screened in Example 1 for three-class classification. The calculation methods for the minimum value, first quartile, median, mean, third quartile, and maximum value of accuracy and consistency are as follows: (1) Sort the values of accuracy or consistency from smallest to largest; (2) Minimum value: The first value after sorting; (3) First quartile (Q1): Multiply the number of data by 0.25. If the result is an integer, take the average of the values at this position and the next position; if not, round up to get the position, and the value at this position is Q1; (4) Median: If the number of data is odd, the median is the middle value; if it is even, it is the average of the two middle values; (5) Mean: The sum of all values divided by the number of data; (6) Third quartile (Q3): Multiply the number of data by 0.75, and the processing method is the same as Q1; (7) Maximum value: The last value after sorting. Among them, the minimum and maximum values can reflect the extreme situations of the data and show the worst and best performances that the model may exhibit; the quartiles can help understand the distribution range and dispersion degree of the data; below Q1 represents a lower performance level, and above Q3 represents a higher performance level; the median can reflect the performance at the middle level; the mean comprehensively reflects the overall average performance. Combining the above statistical values, the overall situation, distribution characteristics, and stability of the model performance can be comprehensively understood, providing a strong basis for model selection and optimization.
[0202] Table 10. Performance comparison of diagnostic models constructed based on different protein combination biomarkers
[0203]
[0204] As can be seen from Table 10, for the three-class diagnostic model, the nine-item combined detection model (5MP) composed of 5 biomarkers has the best performance. This also clearly shows that on the basis of these four biomarkers of MUC16 + CTSG + B3GNT2 + CORO1A, adding the biomarker SFTPB can improve the diagnostic efficiency for distinguishing non-recurrence, in-situ recurrence, or distant metastasis of the surgical prognosis of endometrial cancer to a certain extent. Therefore, the three-class gradient boosting model 5MP constructed by these 5 protein biomarkers (MUC16 + CTSG + B3GNT2 + CORO1A + SFTPB) is used as the best combined diagnostic model.
[0205] 3.2 Determination and verification of the diagnostic performance of the three-class combined diagnostic model
[0206] 1) Determination of the diagnostic performance of the three-class combined diagnostic model
[0207] To more accurately determine the diagnostic performance and thresholds of the model constructed in this embodiment for different disease classifications, a multi-classification model based on the gradient boosting (gbm) algorithm was used to perform predictive analysis in the training group. The predicted probability values for the three-class classification (no recurrence, in-situ recurrence, or distant metastasis after endometrial cancer surgery) were calculated, and the classification with the highest predicted probability value was the final prediction result of the system.
[0208] Among them, the meanings and calculation methods of each index are as follows:
[0209] Calculation results: The accuracy of the three-class combined diagnostic model in the training group was 0.82, and the consistency was 0.81. The diagnostic sensitivity for the group without recurrence after endometrial cancer surgery was 93.6%, and the specificity was 94.5%; the diagnostic sensitivity for the in-situ recurrence group after endometrial cancer surgery was 83.6%, and the specificity was 82.3%; the diagnostic sensitivity for the distant metastasis group after endometrial cancer surgery was 83.2%, and the specificity was 83.9%.
[0210] It should be noted that the three-class combined diagnostic model constructed by gradient boosting is a model constructed by machine learning and cannot fit a specific equation formula like a generalized linear model.
[0211] 2) Verification of the three-class combined diagnostic model
[0212] Based on the model constructed from the training group, the predictive performance was verified in the test group, and the specific results are as follows:
[0213] The accuracy was 0.81, and the consistency was 0.81. The diagnostic sensitivity for the group without recurrence after endometrial cancer surgery was 93.8%, and the specificity was 92.8%; the diagnostic sensitivity for the in-situ recurrence group after endometrial cancer surgery was 82.5%, and the specificity was 81.9%; the diagnostic sensitivity for the distant metastasis group after endometrial cancer surgery was 82.4%, and the specificity was 82.5%.
[0214] In summary, the three-class combined diagnostic model constructed in this embodiment, which includes 5 protein markers, has good diagnostic value for the three classifications of the group without recurrence after endometrial cancer surgery, the in-situ recurrence group after endometrial cancer surgery, and the distant metastasis group after endometrial cancer surgery.
[0215] It can be understood that the described embodiments of the present invention are all preferred embodiments and features. Any person of ordinary skill in the art can make some changes and variations according to the essence described in the present invention, and these changes and variations are also considered to be within the scope of the present invention and the scope limited by the independent claims and the dependent claims.
Claims
1. Use of a biomarker in the preparation of a reagent for predicting the risk of recurrence or metastasis of endometrial cancer, wherein the biomarker comprises one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB.
2. The use according to claim 1, characterized in that, The biomarker comprises MUC16, CTSG, B3GNT2, and CORO1A.
3. The use according to claim 2, wherein The biological sample is selected from one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid; and / or, the detected amount comprises the presence or absence, relative abundance, or concentration of the biomarker.
4. The use according to claim 2, wherein The risk of recurrence or metastasis refers to recurrence or metastasis within three years after the treatment of endometrial cancer.
5. A biomarker combination for predicting the risk of recurrence or metastasis of endometrial cancer, characterized in that, Comprises MUC16, CTSG, B3GNT2, and CORO1A.
6. A method for constructing a risk prediction model for recurrence or metastasis of endometrial cancer for non-diagnostic purposes, characterized in that, Comprises the following: 1) Construct a data set based on the detected amount of the biomarker in the biological sample, wherein the biomarker comprises one or more of MUC16, CTSG, B3GNT2, CORO1A, and SFTPB; 2) Divide the data set into a test set and a training set, and construct and train the endometrial cancer recurrence or metastasis risk prediction model by machine learning methods.
7. A device for predicting the risk of recurrence or metastasis of endometrial cancer, characterized in that, Comprises a data acquisition unit and a calculation unit; The data acquisition unit is used to obtain the detected amount data of the biomarker in the biological sample of the subject, wherein the biomarker comprises MUC16, CTSG, B3GNT2, and CORO1A; The calculation unit is used to calculate and output the predicted score of the risk of recurrence or metastasis of endometrial cancer for the subject based on the detected amount data, and make a judgment according to the cut-off value.
8. An apparatus, comprising a processor and a memory, the memory being configured to store a computer program, characterized in that, The processor is used to execute the computer program stored in the memory, so that the device executes the construction method as described in claim 6.
9. A system for predicting the recurrence risk of endometrial cancer, characterized in that, The system comprises a data analysis module, and the data analysis module is used to analyze the detection value of the biomarker, wherein the biomarker comprises MUC16, CTSG, B3GNT2, and CORO1A.
10. The system according to claim 9, characterized in that, The data analysis module uses the detection values of the biomarkers of the known samples as the training set, and divides them into an endometrial cancer recurrence group and an endometrial cancer non-recurrence group according to whether endometrial cancer recurs, analyzes the relationship between the detection values of the endometrial cancer recurrence group and the endometrial cancer non-recurrence group, and constructs a model. The equation of the model is: , where Y is the predicted score, i represents the i-th biomarker, m represents the number of biomarkers (m = 5), Xi represents the measured value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant 3.8638532; the coefficients of the 4 biomarkers are as follows: , When Y ≤ cut-off value, the subject to be tested has no recurrence of endometrial cancer; when Y > cut-off value, the subject to be tested has recurrence of endometrial cancer. Specifically, the cut-off value is 0.4773883.
Citation Information
Patent Citations
Use of HE4 and other biochemical markers for assessment of endometrial and uterine cancers
CN101473041A
Kit used for assistant determination of clinical early stage endometrial carcinoma lymphonodus metastasis
CN108761079A
Marker molecules at different time points in endometrial secretion period and screening method thereof
CN113295815A
Compositions and methods for treating and diagnosing cancer
CN1852974A
Method for genome editing
US20120192298A1