Biomarker for predicting recurrence or metastasis risk of breast cancer and application of biomarker

By screening and constructing a biomarker-based risk prediction model for breast cancer recurrence or metastasis, the problem of insufficient accuracy in the prior art is solved, efficient and accurate prediction of breast cancer recurrence or metastasis risk is achieved, and the rate of misdiagnosis and missed diagnosis and treatment costs are reduced.

CN120254283AActive Publication Date: 2025-07-04HANGZHOU GUANGKE ANDE BIOTECHNOLOGY CO LTD

Patent Information

Application Number
CN202510725139.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the risk of breast cancer recurrence or metastasis, resulting in over-treatment or insufficient treatment, increasing patient burden and waste of medical resources.

Method used

Biomarkers such as FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1 and KLHL22 were screened through proteomics to construct a risk prediction model for breast cancer recurrence or metastasis, and use high-performance liquid chromatography-tandem mass spectrometry technology to analyze it and combine machine learning algorithms for prediction.

Benefits of technology

Accurate, non-invasive and efficient prediction of the risk of recurrence or metastasis of breast cancer is achieved, and can distinguish between in situ recurrence and distant metastasis, reduce misdiagnosis and misdiagnosis, reduce costs, and improve diagnostic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120254283A_ABST
    Figure CN120254283A_ABST
Patent Text Reader

Abstract

The invention discloses a biomarker for predicting recurrence and metastasis risks of breast cancer and application of the biomarker. The invention discloses application of a substance for detecting a biomarker in preparation of a product for predicting the recurrence or metastasis risk of breast cancer. The biomarker comprises one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1 and KLHL22. The invention further constructs a breast cancer recurrence or metastasis risk prediction model. According to the biomarker and the breast cancer recurrence or metastasis risk prediction model, the breast cancer recurrence or metastasis risk can be accurately, noninvasively and efficiently predicted, and clinical requirements can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of proteomic screening for breast cancer diagnostic markers. Specifically, it relates to a biomarker for predicting the risk of breast cancer recurrence or metastasis and its applications. Background Art

[0002] Breast Cancer is one of the most common malignant tumors in women, with its incidence accounting for 23% of all tumors. There are approximately 1.6 million new cases globally each year, making it the second leading cause of cancer death among women worldwide. The key to the diagnosis and treatment of breast cancer lies in early detection, early diagnosis, and early treatment. Early detection and timely treatment of breast cancer can significantly improve the cure rate of patients, increase the chances of breast conservation and axillary preservation, and thus improve the quality of life of patients. However, despite the continuous development and certain achievements of current screening programs and adjuvant therapies, a considerable proportion of breast cancer patients still die due to cancer metastasis.

[0003] Surgical resection remains the main treatment method for breast cancer patients, especially for early and locally advanced breast cancer. However, approximately 30% of breast cancer patients will experience recurrence within 2 to 5 years after surgery. Breast cancer recurrence is divided into two categories: local recurrence and distant metastasis. Local recurrence refers to the recurrence of cancer cells in the breast area or regional lymph nodes, while distant metastasis is the spread of cancer cells through blood vessels to internal organs such as important organs like the lungs, liver, or brain. If patients who do not benefit from surgical treatment or experience progression can have their risk predicted and the treatment plan adjusted in a timely manner (such as adjuvant radiotherapy and chemotherapy, secondary surgical resection, targeted therapy, or immunotherapy, etc.), the overall survival rate and quality of life of breast cancer patients can be significantly improved.

[0004] At present, the assessment of the possibility of breast cancer recurrence or distant metastasis mainly relies on regular follow-up, which often leads to over-treatment or under-treatment. Adopting the same intensity of treatment plan for all patients will cause some patients to bear unnecessary treatment side effects, while some other patients cannot obtain the due treatment effect. This not only burdens the patients and their families but also causes a waste of medical resources. For postoperative patients, the uncertainty of recurrence is an even greater torture and suffering.

[0005] Proteomics is a science that studies the protein composition, localization, changes, and their interaction rules in cells, tissues, or organisms, covering the research of protein expression patterns and proteome function patterns. With the continuous progress of mass spectrometry technology, liquid chromatography-mass spectrometry (LC-MS / MS) has become the core tool for proteomics research. The development of proteomics is of great significance for finding disease diagnostic markers, screening drug targets, and conducting toxicological research, etc., so it is widely used in the field of medical research. Although there have been many articles and patents reporting the discovery of novel tumor markers in recent years, most of them remain at the laboratory research stage, with less clinical application and market promotion. Moreover, in most cases, it is far from enough to rely solely on a single indicator for in vitro diagnosis of the risk of tumor recurrence. Only by adopting a combined joint detection form and combining detection items from different dimensions can the prediction accuracy be effectively improved.

[0006] In summary, developing tumor recurrence risk markers with greater clinical significance to effectively analyze tumor recurrence or metastasis is of extremely important significance for the diagnosis and treatment of breast cancer. Summary of the Invention

[0007] Aiming at the problems existing in the prior art, the present invention provides a biomarker for predicting the risk of breast cancer recurrence or metastasis and its application. By using the method of proteomics, by analyzing the proteins with significantly different abundance levels in the blood of two groups of breast cancer patients with and without recurrence after surgical treatment, biomarkers for predicting the risk of breast cancer recurrence are screened out, and further a prediction model and a prediction device for breast cancer recurrence risk or metastasis are constructed based on the biomarkers, which can accurately, non-invasively, and efficiently predict the risk of breast cancer recurrence; in addition, the biomarker, prediction model, and prediction device of the present invention can also effectively distinguish in-situ recurrence from distant metastasis, so they can meet the clinical needs.

[0008] On the one hand, the present invention provides the use of substances for detecting biomarkers in the preparation of products for the risk of breast cancer or metastasis recurrence, and the biomarkers include one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.

[0009] The proteomic biomarkers provided by the present invention can accurately predict the recurrence or metastasis risk of breast cancer patients after surgical treatment, so as to more early identify high-risk recurrence or high-risk metastasis populations, and can help doctors pre-judge in advance whether it is necessary to adjust the treatment plan in time, reduce the risk of breast cancer recurrence or metastasis, and truly benefit breast cancer patients.

[0010] The present invention uses proteomics methods to collect plasma samples from patients who relapse within a short period (such as 1 to 5 years) and those who do not relapse within a short period after breast cancer surgery. Different samples are analyzed by high performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS). Based on orthogonal partial least squares discriminant analysis and significance analysis methods, proteins with significant differences between breast cancer relapse patients and non-relapse patients are first screened. Finally, 10 differential proteins with obvious correlation with the recurrence risk of breast cancer are screened out. These 10 proteins can be used to distinguish breast cancer relapse patients from non-relapse patients, in-situ relapse patients from distant metastasis patients, and have a certain diagnostic efficacy, thus can be used to predict the recurrence or metastasis risk of breast cancer.

[0011] Among them, the FBLN5 is the protein or amino acid sequence with the UniProt database number Q9UBX5; MUC16 is the protein or amino acid sequence with the UniProt database number Q8WXI7; ORM1 is the protein or amino acid sequence with the UniProt database number P02763; ADH1B is the protein or amino acid sequence with the UniProt database number P00325; CD36 is the protein or amino acid sequence with the UniProt database number P16671; KRT19 is the protein or amino acid sequence with the UniProt database number P08727; DCD is the protein or amino acid sequence with the UniProt database number P81605; DEFA3 is the protein or amino acid sequence with the UniProt database number P59666; RNASE1 is the protein or amino acid sequence with the UniProt database number P07998; KLHL22 is the protein or amino acid sequence with the UniProt database number Q53GT1.

[0012] The present invention also surprisingly finds that some of the protein markers obtained by proteomics screening are known markers that can be used for other cancers. For example, ADH1B has been reported to be used for predicting laryngeal cancer, pharyngeal cancer, nasal cancer, pancreatic cancer and esophageal cancer, but this screening finds that this marker can also be used for predicting the recurrence or metastasis risk after breast cancer surgery. It can be seen that for many different cancers, their protein markers are not completely separated or irrelevant. In fact, there are many cross relationships or influences. Many protein markers can be used for predicting cancer at an early stage, and also for prognostic diagnosis of cancer. Even many protein markers can be used for predicting and diagnosing different stages of many different cancers. Therefore, there are still many new functions to be explored in the field of proteomics, and the market prospect is very broad.

[0013] Furthermore, the markers include FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1.

[0014] To improve the diagnostic efficacy for the risk of breast cancer recurrence or metastasis, it is also necessary to combine different differential proteins to construct a prediction model, rank them according to the importance obtained by screening, respectively select different numbers of differential proteins with higher rankings for combination, and finally screen out 9 protein markers. Based on these 9 protein markers, a model is constructed, which has good risk prediction ability in the diagnosis of the risk of breast cancer recurrence or metastasis. The data from the detection of clinical breast cancer samples show that just using these 9 biomarkers to predict the risk of breast cancer recurrence, the AUC value can reach 0.954, indicating good diagnostic performance.

[0015] Furthermore, the recurrence or metastasis risk refers to recurrence or metastasis within three years after breast cancer treatment; the recurrence or metastasis of breast cancer includes in-situ or adjacent area recurrence of breast cancer and metastasis of breast cancer.

[0016] It can be understood that the short term in "recurrence or metastasis within three years after breast cancer treatment" here is not an absolute and unchanging time node. It is just a time node summarized based on current clinical experience. As time goes by or with the improvement of other treatment methods, this time node or the length of time may change. For example, it may be within one year, within 360 days, or within half a year, within 180 days, or within two years, within two and a half years, within three and a half years, etc. It is possible to have recurrence or metastasis of breast cancer after treatment.

[0017] In some embodiments, the recurrence or metastasis of breast cancer means that after surgical treatment by radical tumor resection, the patient has in-situ or adjacent area recurrence of breast cancer within three years, or has metastasis of breast cancer, such as lymph node metastasis, bone metastasis, etc.

[0018] In some embodiments, the product is used for predicting patients at risk of breast cancer recurrence or metastasis.

[0019] In some embodiments, the product is a detection reagent prepared with these biomarkers as the detection target, such as sample pretreatment reagents, antigens or antibodies, and other biological reagents and kits suitable for the detection of these biomarkers; it can also be developed into standardized reagents or kits suitable for these biomarkers. Specifically, the product can be one or more of reagents, kits, chips, probes, or membrane strips.

[0020] In some embodiments, the product is used for detecting the content of biomarkers in body fluid samples; the body fluid samples include any one or more of saliva, blood, urine, plasma, serum, and spinal fluid.

[0021] Furthermore, the product is used to detect the presence, relative abundance or concentration of biomarkers in a body fluid sample.

[0022] The present invention screens for biomarkers for predicting the risk of breast cancer recurrence or metastasis from blood. There are significant differences in the blood of high-risk and low-risk breast cancer recurrence populations for these biomarkers. By collecting a blood sample, it is possible to predict or assist in diagnosing the likelihood of breast cancer recurrence or metastasis in an individual by detecting these biomarkers in the individual's blood, or it is possible to detect these biomarkers in the blood of a certain population, and then divide the population into high-risk and low-risk breast cancer recurrence populations, in-situ recurrence risk and distant metastasis risk.

[0023] Furthermore, the detection method includes a radio method, an immuno method, a fluorescence method, a flow fluorescence method, a latex turbidimetry method, a biochemical method, an enzymatic method, a hybridization method, a gas chromatography-mass spectrometry method, a liquid chromatography-mass spectrometry method, a chromatography method, a chemiluminescence method, a magnetoelectric method or a photoelectric conversion method.

[0024] The presence or absence or the level of the biomarker here is a relative concept. For example, when comparing the high-risk breast cancer recurrence group and the low-risk breast cancer recurrence group, the levels of these specific biomarkers are compared with the high-risk breast cancer recurrence group and the low-risk breast cancer recurrence group as a reference. There may be some biomarkers with relatively higher levels in the high-risk breast cancer recurrence group than in the low-risk breast cancer recurrence group, and this increase is statistically different, such as a significant or highly significant increase. Therefore, when judging these biomarkers, if it is a single biomarker, if the probability of a certain risk occurrence increases and the level of the biomarker changes, this change may be a relative increase or may also be a relative decrease, and the difference in this relative increase or relative decrease is significantly different, and of course it can also be highly significantly different. Therefore, no matter what means are used for detection, a pre-specified value (cut-off value) can be used as a standard. Those higher than this value are considered to have a change in level, and having such a result can all be used for prediction or diagnostic value.

[0025] Therefore, in some aspects, the biomarker described in the present invention can be obtained by detecting the content of the biomarker in a sample through any known method, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strip, immunoreaction chip, capillary electrophoresis, infrared spectroscopy, etc. As long as it can be used to detect the content of the protein biomarker in a sample, it can be used for the diagnosis of the high-risk group of breast cancer recurrence and the low-risk group of breast cancer recurrence. As long as the content of the protein biomarker in the sample can be detected, it can be used to predict or diagnose the probability of the occurrence of a certain disease. It can be understood that the detection here is for the individual sample, and then compared with a pre-set standard, and the result of the comparison is used to judge or predict the occurrence status of the disease. For example, it can be used to predict the probability of breast cancer recurrence. Such prediction or diagnosis is whether it will occur within a certain time. Of course, such detection can be continuous detection, and the progress of the disease can be inferred from the change in the content of certain substances.

[0026] In some embodiments, the relative abundance is the peak area of the biomarker in the detection spectrum obtained by high performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a certain biomarker measured in a control sample is 100, and the average peak area measured in the sample of patients in the high-risk group of breast cancer recurrence is 600, then it is considered that the abundance of the biomarker in the sample is 6 times that in the control sample.

[0027] In some embodiments, the biomarker can be 1 biomarker or a combination of several biomarkers, such as a combination of 2 biomarkers, a combination of 3 biomarkers, a combination of 4 biomarkers, a combination of 5 biomarkers, a combination of 6 biomarkers, a combination of 7 biomarkers, a combination of 8 biomarkers, a combination of 9 biomarkers, a combination of 10 biomarkers. When the combination of biomarkers is from 2 to 10, the AUC value of the constructed prediction model is 0.797 - 0.954, and both the sensitivity and specificity are relatively high.

[0028] The second aspect of the present invention provides the use of a protein as a biomarker in the preparation of a product for predicting the risk of breast cancer recurrence or metastasis, wherein the protein is selected from one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22. The genes FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22 can be used for the auxiliary judgment of the risk of breast cancer recurrence or metastasis, the evaluation of drug efficacy, etc. The inventors of the present invention have found that FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22 are closely related to the risk of breast cancer recurrence or metastasis.

[0029] The third aspect of the present invention provides a product for predicting the risk of breast cancer recurrence or metastasis, including a substance for detecting a biomarker, wherein the biomarker includes one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.

[0030] The fourth aspect of the present invention provides a biomarker combination for predicting the risk of breast cancer recurrence or metastasis, including one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.

[0031] The fifth aspect of the present invention provides a method for constructing a breast cancer recurrence or metastasis risk prediction model for non-disease diagnosis purposes, including the following:

[0032] 1) Construct a sample data set based on the detection amount of biomarkers in biological samples, wherein the biomarkers include one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22;

[0033] 2) Divide the data set into a test set and a training set, and construct and train the breast cancer recurrence or metastasis risk prediction model through a machine learning method.

[0034] In some embodiments, the biomarkers are FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1; the recurrence risk refers to recurrence within three years after breast cancer surgery; the breast cancer recurrence includes in-situ or adjacent area recurrence of breast cancer and breast cancer metastasis.

[0035] Further, the biological samples are biological samples of breast cancer recurrence and biological samples of non-recurrence of breast cancer. By analyzing the relationship between the detection values of biomarkers in biological samples of breast cancer recurrence and biological samples of non-recurrence of breast cancer, a model is constructed.

[0036] In some embodiments, the machine learning method is selected from at least one of gradient boosting algorithm, random forest algorithm, support vector machine algorithm, decision tree algorithm, K-nearest neighbor algorithm, logistic regression algorithm, and neural network algorithm.

[0037] In some embodiments, the prediction model is a generalized linear model (GLM), and its equation is:

[0038]

[0039] In the formula, Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 9), Xi represents the detection value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker, and b is a constant 6.3464595; the coefficients of the 9 biomarkers are:

[0040]

[0041] When Y ≤ 0.513363, the risk of breast cancer recurrence or metastasis of the subject to be tested is low; when Y > 0.513363, the risk of breast cancer recurrence or metastasis of the subject to be tested is high.

[0042] The fifth aspect of the present invention provides a device for predicting the risk of breast cancer recurrence or metastasis, including a data acquisition unit and a calculation unit;

[0043] The data acquisition unit is used to obtain the detection amount data of the biomarkers in the biological samples of the subject, and the biomarkers include one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22;

[0044] The calculation unit is used to calculate and output the prediction score of the subject's risk of breast cancer recurrence based on the detection amount data and make a judgment according to the threshold.

[0045] In some embodiments, the biomarkers include FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1.

[0046] In some embodiments, the prediction device further includes a data storage unit and a data output unit; the data storage unit is used to store the detected amount (or detection value) of the biomarker; the data input interface is used to input the detection value of the biomarker, and the data output unit is used to output the prediction result.

[0047] In some embodiments, the detection value is the presence or absence, relative abundance or concentration value of each biomarker.

[0048] The sixth aspect of the present invention provides a device, including a processor and a memory, the memory is used to store a computer program, and it is characterized in that the processor is used to execute the computer program stored in the memory, so that the device executes the prediction method or the construction method as described above.

[0049] The seventh aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program, when processed, executes the prediction method or the construction method as described above.

[0050] In some embodiments, the computer-readable storage medium includes: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs.

[0051] In some embodiments, the being processed means being executed by one or more processors.

[0052] The eighth aspect of the present invention provides a method for predicting the risk of breast cancer recurrence or metastasis, including the following steps:

[0053] S1. Obtain the detected amount data of biomarkers in the biological sample of the subject, and the biomarkers include one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1 and KLHL22;

[0054] S2. Process the detected amount data by using the breast cancer recurrence or metastasis risk prediction model obtained by the construction method as described above to output the breast cancer recurrence or metastasis risk prediction result.

[0055] In some embodiments, the biomarkers include FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3 and RNASE1.

[0056] In another aspect, the present invention provides a system for predicting the recurrence risk of breast cancer. The system includes a data analysis module, which is used to analyze the detection values of biomarkers, and the biomarkers include FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1.

[0057] Further, the data analysis module uses the detection values of biomarkers of known samples as a training set. According to whether breast cancer recurs, it is divided into a breast cancer recurrence group and a non-recurrence group of breast cancer, and analyzes the relationship between the detection values of the breast cancer recurrence group and the non-recurrence group of breast cancer to construct a model.

[0058] In some embodiments, a combined diagnostic model for predicting the recurrence risk of breast cancer is constructed by combining multiple machine learning methods, and it is preliminarily confirmed that the change in the concentration of any one of the selected novel biomarkers alone can be used to distinguish between high-risk and low-risk populations of breast cancer recurrence, indicating that these biomarkers have extremely high diagnostic value.

[0059] In some embodiments, the equation of the constructed model is:

[0060] Among them, Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 9), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant 6.3464595; the coefficients of the 9 biomarkers are:

[0061]

[0062] When Y ≤ 0.513363, the risk of recurrence or metastasis of breast cancer in the tested person is low; when Y > 0.513363, the risk of recurrence or metastasis of breast cancer in the tested person is high.

[0063] Further, the system further includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.

[0064] In another aspect, the present invention provides the use of a biomarker for preparing a reagent for predicting non-recurrence, in-situ recurrence, or distant metastasis in breast cancer patients after surgery. The biomarker includes any one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.

[0065] The present invention attempts to use protein markers for predicting recurrence risk to further distinguish between patients with in-situ recurrence and those with distant metastasis, and finds that each protein marker can be used to distinguish among three categories of breast cancer patients after surgery: non-recurrence, in-situ recurrence, and distant metastasis. The distant metastasis includes patients with only distant metastasis, as well as patients with both in-situ recurrence and distant metastasis.

[0066] Further, the reagent is used to predict non-recurrence, in-situ recurrence, or distant metastasis in breast cancer patients within three years after surgery.

[0067] Further, the reagent is used to detect the content of biomarkers in a body fluid sample; the body fluid sample includes any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.

[0068] Further, the reagent is used to detect the presence, absence, relative abundance, or concentration of biomarkers in a body fluid sample.

[0069] On the other hand, the present invention provides a kit for predicting non-recurrence, in-situ recurrence, or distant metastasis in breast cancer patients after surgery, and the kit includes a detection reagent for a biomarker as described above.

[0070] On the other hand, the present invention provides a biomarker combination for predicting non-recurrence, in-situ recurrence, or distant metastasis in breast cancer patients after surgery, and the combination includes CD36, PLIN1, ITIH3, MGAM2, RNASE1, SELL, ORM2, TFF1, and TFF3.

[0071] On the other hand, the present invention provides a system for predicting non-recurrence, in-situ recurrence, or distant metastasis after breast cancer surgery, and the system includes a data analysis module for analyzing the detection values of markers, and the markers include any one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.

[0072] Further, the distant metastasis also includes both in-situ recurrence and distant metastasis at the same time, and as long as distant metastasis occurs, it is classified into the distant metastasis group.

[0073] Further, the data analysis module uses the detection values of markers of known samples as a training set, and according to the situation of breast cancer patients after surgery, divides them into a non-recurrence group, an in-situ recurrence group, and a distant metastasis group, analyzes the relationship between the detection values of the non-recurrence group, the in-situ recurrence group, and the distant metastasis group, and constructs a model.

[0074] Further, the model is constructed based on the gradient boosting algorithm.

[0075] Unlike generalized linear models, the gradient boosting algorithm does not output a model formula and a cut-off value. All calculations are directly performed by machine learning. By directly inputting the test values into the software system, the prediction results can be directly obtained.

[0076] Furthermore, the system further includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the test values of biomarkers; the data input interface is used to input the test values of biomarkers, and the data output interface is used to output the prediction results.

[0077] The beneficial effects of the present invention are as follows:

[0078] 1. The present invention has screened 10 biomarkers that can predict the recurrence risk of breast cancer, developed a new combination of protein biomarkers, which can effectively evaluate and diagnose patients at risk of breast cancer recurrence, effectively distinguish between high-risk and low-risk populations of breast cancer recurrence, can more accurately identify breast cancer recurrence, and compared with traditional detection methods, reduces the risk of misdiagnosis and missed diagnosis, providing strong support for the early detection and intervention of diseases.

[0079] 2. The combined discrimination prediction model of 9 biomarkers constructed by the present invention is convenient, fast, and the test results are highly consistent with the clinical gold standard test results. At the same time, it significantly reduces the cost of predicting the recurrence risk of breast cancer and has good application prospects.

[0080] 3. Based on the biomarkers of the present invention, a three-classification model that can simultaneously distinguish between non-recurrence group, in-situ recurrence group, and metastatic group has been further constructed, providing a more effective and accurate prediction and diagnosis mode.

[0081] Detailed description

[0082] (1) Diagnosis or prediction

[0083] Here, diagnosis or prediction refers to detecting or analyzing the biomarkers in a sample, or the content of the target biomarker, such as the absolute content or relative content, and then indicating whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease, based on the presence or quantity of the target biomarker. Here, the meanings of diagnosis and prediction can be interchanged. The result of this prediction or diagnosis cannot be directly used as the direct result of having a disease, but is an intermediate result. If a direct result is to be obtained, other auxiliary means such as pathology or anatomy are required to confirm the presence of a certain disease. For example, the present invention provides a variety of new biomarkers related to the recurrence or metastasis risk of breast cancer, and the changes in the content of these biomarkers are directly related to whether the patient belongs to the population at risk of breast cancer recurrence or metastasis.

[0084] (2) Association between markers, biomarkers or differential proteins and the risk of breast cancer recurrence or metastasis

[0085] Markers, biomarkers and differential proteins have the same meaning in the present invention. The association here refers to the direct relevance between the presence or content change of a certain biomarker in a sample and a specific disease. For example, a relative increase or decrease in content indicates a relatively higher likelihood of having this disease compared to the healthy population.

[0086] If multiple different markers in a sample appear simultaneously or there are relative changes in their contents, it also indicates a relatively higher likelihood of having this disease compared to the healthy population. That is to say, among the types of markers, some markers have a strong association with the disease, some have a weak association, and some may even have no association with a specific disease. One or more of those markers with a strong association can be used as markers for diagnosing the disease, and those with a weak association can be combined with the strong markers to diagnose a certain disease, increasing the accuracy of the test results.

[0087] Regarding the numerous biomarkers found in the serum in the present invention, these biomarkers can all be used to distinguish between high-risk and low-risk populations for breast cancer recurrence, and in-situ metastasis and distant metastasis. These markers can be directly detected or diagnosed as individual markers. Selecting such a marker indicates that the relative change in the content of this marker has a strong association with the risk of breast cancer recurrence or metastasis. Of course, it can be understood that simultaneous detection of one or more markers with a strong association with the risk of breast cancer recurrence or metastasis can be selected. Normally understood, in some embodiments, selecting biomarkers with a strong association for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90% or 95% accuracy. Then it can be shown that these markers can obtain an intermediate value for diagnosing a certain disease, but it does not mean that it can directly confirm having a certain disease.

[0088] Of course, it is also possible to select differentially expressed proteins with larger ROC values as diagnostic markers. The so-called strength is generally calculated and confirmed through some algorithms, such as the contribution rate or weight analysis of the marker to the assessment of the risk of breast cancer recurrence or metastasis. Such calculation methods can include significance analysis (p-value or FDR value) and fold change. Multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA), and orthogonal partial least squares discriminant analysis (OPLS-DA). Of course, other methods are also included, such as ROC analysis. Of course, other model prediction methods are also possible. When specifically selecting biomarkers, differentially expressed proteins disclosed in the present invention can be selected, or other existing well-known biomarker combinations can be selected or combined for prediction through model methods.

[0089] (3) Definition of disease terms

[0090] Breast cancer: Breast cancer is a malignant tumor that occurs in breast tissue and usually originates from breast ducts or lobules. Early symptoms of breast cancer may include breast lumps, nipple discharge, enlarged axillary lymph nodes, etc.

[0091] Surgical resection is the preferred treatment for breast cancer. Tumor resection in the breast or total mastectomy is selected according to the lesion situation of the breast tumor, and central lymph node dissection or cervical lymph node dissection is selected according to the cervical lymph node metastasis situation.

[0092] Breast cancer recurrence: It refers to the phenomenon that after a patient completes radical treatment (surgical resection), after a period of clinical disease-free state (i.e., a period without signs of tumors), cancer cells grow again or tumor lesions (also known as metastases) appear at the original tumor site or other parts of the body. Recurrence may stem from residual cancer cells that were not completely removed during treatment or micrometastases that had spread earlier but were not detected. Breast cancer recurrence can occur from several months to several years after treatment, but 80% of recurrences occur within 2 - 3 years after surgery, and the recurrence risk is significantly reduced after 5 years. Breast cancer recurrence is closely related to tumor stage, treatment standardization, and postoperative management. Although its prognosis is relatively good, some patients may still experience recurrence. Usually, the risk needs to be reduced through regular monitoring, standardized treatment, and long-term follow-up, which is rather cumbersome. The main metastasis route of breast cancer metastasis is lymph node metastasis, and a small number also show distant metastases, such as lung metastasis, bone metastasis, etc. After recurrence or metastasis, treatment plans such as surgery, radiotherapy, chemotherapy, or targeted therapy need to be selected according to the specific situation. If detected early, the recurrence situation can be controlled as soon as possible, significantly improving the prognosis.

[0093] (4) Gold standard for the diagnosis of breast cancer recurrence: It is pathological histological examination (i.e., pathological confirmation of biopsy or surgical resection specimens), observing the presence of cancer cells under a microscope, and clarifying the nature of the tumor in combination with immunohistochemistry or molecular detection. Description of the Drawings

[0094] Figure 1 It is a volcano plot for analyzing the differences in protein markers between the high-risk group and the low-risk group of breast cancer recurrence in Example 1.

[0095] Figure 2 It is a graph showing the ROC and OPLS-DA analysis results for the high-risk group and the low-risk group of breast cancer recurrence in Example 1.

[0096] Figure 3 It is a graph showing the AUC results of models constructed with different hyperparameters in Example 2.

[0097] Figure 4 It is an ROC curve graph of the combined diagnostic model in the training group in Example 2.

[0098] Figure 5 It is an ROC curve graph of the combined diagnostic model in the test group in Example 2. Detailed Description of the Invention

[0099] The present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not impose any limitations on it. The reagents used in this embodiment are all known products and are obtained by purchasing commercially available products.

[0100] Example 1 Screening Biomarkers for Predicting the Risk of Breast Cancer Recurrence Using Proteomics

[0101] By collecting plasma samples from breast cancer patients after radical surgery, enriching low-abundance proteins based on the method of removing high-abundance proteins by immunoaffinity chromatography, detecting the protein abundance in the samples by a high-performance liquid chromatography-tandem mass spectrometry device, analyzing the differences in their abundances between patients without breast cancer recurrence within three years and those with recurrence or breast cancer metastasis, and analyzing their diagnostic performance. The specific steps are as follows:

[0102] 1.1. Sample Collection

[0103] A total of about 2 ml of peripheral blood samples were collected from 100 breast cancer (stages I-III) patients 2 weeks after surgical treatment, placed in a vacuum tube containing EDTA anticoagulant, mixed well, centrifuged at 120 g for 10 minutes at room temperature, and the supernatant was taken, repeated twice; then at 360 g for 20 minutes. Then the platelet samples were collected in centrifuge tubes and stored at -80°C. The average age of the 100 patients was 43 years old (25 - 69 years old). All enrolled patients signed informed consent forms. Among them, all breast cancer patients were pathologically and histologically diagnosed patients, and the inclusion criteria were: (a) no history of other malignant tumors; (b) no patients with combined other malignant tumors or autoimmune diseases.

[0104] In the following three years, the patients were followed up every 3 to 6 months, and regular breast ultrasound, breast X-ray and other imaging examinations were performed. Once recurrence or metastasis was found, it was confirmed by pathological histology. A total of 24 breast cancer patients relapsed within three years after surgery, of which 14 were local recurrences in situ and 10 were accompanied by lymph node metastasis. Then the samples stored at -80℃ were taken out and divided into two groups: 76 patient samples without recurrence were classified as low-risk group for breast cancer recurrence, and 24 patients with recurrence or metastasis were classified as high-risk group for breast cancer recurrence for the screening of proteomic markers.

[0105] 1.2. Sample processing and enzymatic hydrolysis

[0106] First, the plasma samples were centrifuged for 15 minutes (15,000 g), and the supernatant was filtered and immunoaffinity chromatography was performed to remove 14 high-abundance proteins. Then, the low-abundance proteins were concentrated to 350 μL using a 3 kDa cutoff concentrator tube on a centrifuge (4000 g, 1 hour). The concentrate was recovered and the solution was replaced (Buffer Exchange) using a 7 kDa cutoff desalting column on a centrifuge (1000 g, 2 minutes). The replacement solution was AEX-A (20 mM Tris, 4 M Urea, 3% isopropanol, pH 8.0). Using AEX-A as a blank, the protein concentration in the sample was determined using the BCA method. According to the sample grouping, 25 μL TCEP was added to the sample and incubated at 37 ° C for 30 minutes for protein reduction. Then, the corresponding TMT 16-plex reagent was added and incubated at room temperature in the dark for 1 hour for TMT labeling reaction. Afterwards, the sample was buffer exchanged using Zeba columns with AEX-A as the replacement fluid. After the TMT 16-plex labeled samples were mixed, 2 mL of AEX-A was added to the mixed samples to a final volume of 5.5 mL. The samples were filtered using a 0.22 μm filter and the TMT 16-plex labeled samples were separated using a 2D-HPLC system. The collected fractions were freeze-dried and finally Trypsin-Lysin C enzyme mixture was added. The samples were incubated at 37°C for 5 hours to digest the samples, and 5 μL of 10% TFA was added to terminate the digestion reaction. A total of 60 2D-HPLC fractions after digestion were used for nanoLC-MS / MS analysis.

[0107] 1.3 LC-MS / MS Data Acquisition

[0108] Each sample from Step 1.2 was separated using a nano-flow liquid chromatography system Easy nLC-1200 and online connected to a high-resolution mass spectrometer Q Exactive HF-X. Mobile phase - A was 0.1% formic acid aqueous solution, and mobile phase - B was 0.1% formic acid acetonitrile aqueous solution (acetonitrile was 80%, water was 20%). The chromatographic column consisted of an enrichment column and an analytical column and was equilibrated with 100% mobile phase - A. The sample was loaded onto the enrichment column (100μm ID * 4cm L, C18, 3μm, 100A) by an autosampler, separated by the analytical column (75μm ID * 25cm L, C18, 3μm, 100A), and the flow rate was 300 nL / min. After chromatographic separation, the sample was subjected to mass spectrometry analysis using a Q Exactive HF-X mass spectrometer. The detection mode was positive ion, the precursor ion scan range was 350 - 1800 m / z, the resolution of the first-stage mass spectrometry was 120,000 at 200 m / z, the AGC (Automatic gain control) target was 3e6, the Maximum IT was 50ms, and the dynamic exclusion time was 40s. The mass-to-charge ratios of polypeptides and polypeptide fragments were acquired according to the data-dependent acquisition (DDA) method: 20 MS / MS spectra (MS2 scan) were acquired after each full scan (first-stage mass spectrometry). The MS2 ActivationType was HCD, the Isolation window was 0.7 m / z, the resolution of the second-stage mass spectrometry was 30,000 at 200 m / z, the AGC target was 1e5, the Maximum IT was 65ms, the Fixed first mass was 110.0 m / z, the Normalized Collision Energy was 32ev, the Minimum AGC target was 2.00e4, the Charge exclusion was 1, 6 - 8, >8, Multiple charge states was one charge state only, Peptide match was preferred, and Exclude isotopes was on.

[0109] 1.4. Data preprocessing

[0110] The obtained MS / MS data from Step 1.3 was searched using Maxquant (v1.6.15.0). The data type was DIA proteomics data based on MS / MS reporter ion quantification. For the MS / MS spectra used for quantification, the requirement was that the proportion of precursor ions in the MS spectra was greater than 75%. The database source was Homo_sapiens_9606_proteome_gene in the Uniprot database (release: 2021-10-14, sequence: 20,437), and a common contaminant library was added to the database. Contaminant proteins were removed during data analysis; the digestion method was set to Trypsin / P; the maximum number of missed cleavage sites was set to 2; the mass error tolerances for precursor ions in the First search and Main search were set to 20 ppm and 5 ppm respectively, and the mass error tolerance for MS / MS fragment ions was 20 ppm. The fixed modification was cysteine alkylation, and the variable modifications were methionine oxidation and protein N-terminal acetylation. The false discovery rate (FDR) for protein identification and PSM identification was set to 1%.

[0111] 1.5. Differential analysis

[0112] A combination of univariate analysis and multivariate statistical analysis was used to screen for differential proteins. Univariate analysis mainly included the significance analysis (p-value or FDR value) and fold change of characteristic molecules in different groups, and multivariate statistical analysis mainly included receiver operating characteristic curve (ROC) analysis and Boruta feature selection based on the random forest algorithm. All statistical analyses were performed using R. The specific R-related information is shown in Table 1.

[0113] Table 1. R and its related information used in the present invention

[0114]

[0115] The variable importance for the projection (VIP) was calculated to measure the influence intensity and interpretability of the expression patterns of each protein on the classification and discrimination of each group of samples. Further, the Wilcoxon rank sum test was performed to obtain the corrected p-value (FDR). According to the conditions of FDR < 0.01 and Fold change > 2, 62 downregulated proteins and 64 upregulated proteins were screened (see details in Figure 1 ).

[0116] To evaluate the role of each protein biomarker in the diagnosis and prediction of breast cancer recurrence risk, the ROC and Boruta analysis methods were used to evaluate each protein biomarker. The results are shown in Figure 1, the abscissa is the AUC obtained from the ROC analysis, the ordinate is the -log10(FDR) calculated by the Wilcoxon test, and the size of the points represents the VIP value obtained from the Boruta analysis. Further screening was carried out according to VIP > 3 and AUC > 0.6, and a total of 10 more significant candidate protein markers were found, as shown in Table 2.

[0117] Table 2. Candidate Protein Markers

[0118]

[0119] Among them, the smaller the FDR value and / or the larger the VIP value, to a certain extent, it indicates that the difference in the protein between the two groups of high-risk and low-risk breast cancer recurrence is more significant, and at the same time, it also indicates that the protein may have higher diagnostic value.

[0120] Example 2. Construction and Validation of a Breast Cancer Recurrence Risk Model

[0121] In this example, a model constructed with the 10 protein markers DEFA3, ORM1, FBLN5, MUC16, KRT19, CD36, DCD, ADH1B, RNASE1, and KLHL22 screened in Example 1 was studied.

[0122] Although a single biomarker can also distinguish the recurrence risk of breast cancer patients after surgery, generally speaking, combining multiple biomarkers has higher accuracy in discrimination or prediction. However, a single biomarker with higher accuracy in predicting the recurrence risk of breast cancer patients after surgery may not necessarily play a greater role in the combination with one or more other biomarkers. At the same time, it is not that the more the number of biomarkers, the higher the prediction accuracy (AUC value) of the combination. Therefore, a large number of validation experiments are still needed.

[0123] 2.1. Models Constructed with Different Biomarker Combinations

[0124] The research cohort was blood samples of 220 breast cancer patients 2 weeks after surgical treatment. All enrolled patients signed informed consent forms. Among them, 44 breast cancer patients had recurrence within three years after surgery (24 had local in-situ recurrence and 20 had metastases). Thus, they were divided into two groups: 176 samples of patients without recurrence were classified into the low-risk breast cancer recurrence group, and 44 samples of patients with recurrence or metastasis were classified into the high-risk breast cancer recurrence group. They were randomly divided into a training group and a test group. The training group included 88 samples from the low-risk breast cancer recurrence group and 22 samples from the high-risk breast cancer recurrence group, and the test group included 88 samples from the low-risk breast cancer recurrence group and 22 samples from the high-risk breast cancer recurrence group.

[0125] In the training group, a combined method of multiple machine learning methods was used to construct a combined prediction model for multiple protein markers. The predicted probability values were used to estimate the area under the receiver operator characteristic (ROC) curve with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate prediction model. Using the training group, the Youden index (YI) was calculated to determine the cut-off value for predicting the probability of distinguishing between the high-risk group and the low-risk group of breast cancer recurrence. In addition, the ROCs of individual markers and different combinations were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R 3.6.1, and a p-value less than 0.05 was considered statistically significant.

[0126] The steps for constructing the combined prediction model are as follows:

[0127] S101, Among the 10 protein markers of DEFA3, ORM1, FBLN5, MUC16, KRT19, CD36, DCD, ADH1B, RNASE1, and KLHL22 in the training group samples, a concentration matrix of randomly selected 2 to 10 markers was used as the original training data set.

[0128] S102, The generalized linear model (glmnet) algorithm was selected for constructing the prediction model, as well as the grid search range during the hyperparameter optimization process of the algorithm. In this step, the grid search range for hyperparameter optimization of each algorithm was set as shown in Table 3.

[0129] Table 3, Parameter grid of glmnet algorithm

[0130]

[0131] S103, According to the algorithm and the hyperparameter setting range set in step S102, one of the hyperparameter combination methods was selected as the parameters for constructing the prediction model.

[0132] S104, The original data set was split into K subsets according to the K-fold cross-validation mechanism. To ensure that in each subset, the ratio of majority class samples to minority class samples is the same as that of the original data set, the stratified K-fold cross-validation mechanism was used for data splitting.

[0133] S105, According to the K training data subsets obtained by splitting in step S104, one of the subsets was selected as the validation set Ddev.

[0134] S106. Combine the unselected subsets of the training data in step S105 to form the training data pool Dtrainl.

[0135] S107. Based on the training data set Dtrain obtained in step S106, construct a prediction model based on the selected supervised classification algorithm and hyperparameters.

[0136] S108. According to the prediction model obtained in step S107, evaluate it on the validation set Ddev to obtain the AUC value, and store the current prognosis prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 is to evaluate according to the prediction model obtained in step S107 on the validation set determined in the current iteration, and store both the model and the evaluation results in the prediction model pool for future prediction model selection. The evaluation mentioned in this step can be the AUC value or other reasonable metrics for evaluating the model performance.

[0137] S109. Determine whether each subset has been used as the validation set. Step S109 is to determine whether the K subsets obtained in step S104 have all been used as the validation set for model training. If all subsets have been used as the validation set and completed the training, execute step S110; if there is a subset that has not been used as the validation set, execute step S105. This step ensures that every sample in the original data set has been used as the validation set, improves the model stability, and prevents the model from overfitting to a certain subset.

[0138] S110. Take the average value of the AUCs of all the models in the obtained prediction model pool Pool as the final performance evaluation value of the model for this combination method. And store the model parameters and the final performance evaluation AUC value in the optimal model pool Poolbest.

[0139] S111. Determine whether prediction models have been constructed for all combinations of hyperparameters. Step S111 is to determine whether prediction models have been constructed for all the algorithms and the corresponding combinations of hyperparameters obtained in step S102. If all combinations have completed the model construction, execute step S112; if there is a combination that has not completed the model construction, execute step S103.

[0140] S113. Select the model with the largest AUC value from the model set Poolbest obtained in step S112 as the final prediction model for breast cancer recurrence risk diagnosis.

[0141] S114. Repeat all the above steps until all combinations of the markers have completed the modeling.

[0142] By performing the above model construction steps, the optimal model constructed with all combinations of the markers was obtained. To compare the performance of the models under these different marker combinations, the ROC method was used to evaluate the AUC values of these models in the training group, and the results are shown in Table 4.

[0143] Table 4. Comparison of the areas under the ROC curves of the models constructed with different marker combinations in the training group

[0144]

[0145] Table 4 was sorted according to the markers with larger VIP values and smaller FDR values in Table 2. Starting from the markers ranked higher, the markers were selected for combination in sequence, from 2MP to 9MP. The detected AUC values, accuracy, and sensitivity all increased with the increase in the number of markers. However, when continuing to increase the markers on the basis of 9MP, the diagnostic performance of the constructed model did not continue to increase. The diagnostic efficacy of 9MP was similar to that of 10MP, and 9MP required fewer markers and had lower costs. Therefore, the most preferred model was the model constructed with a combination of 9 markers (FBLN5 + MUC16 + ORM1 + ADH1B + CD36 + KRT19 + DCD + DEFA3 + RNASE1).

[0146] 2.2. Optimization of model parameters

[0147] For the optimal marker combination form FBLN5 + MUC16 + ORM1 + ADH1B + CD36 + KRT19 + DCD + DEFA3 + RNASE1, based on this marker combination, models constructed under 9 different combinations of glmnet algorithm hyperparameters were analyzed, and the performance of the models was evaluated by the AUC value (the AUC was calculated using the 10-fold cross-validation method during the modeling process), and the results are shown in Table 5 and Figure 3 as follows.

[0148] Table 5. AUC of the models constructed under different hyperparameter combinations of the glmnet algorithm

[0149]

[0150] As can be seen from Table 5, when the hyperparameter combination of the glmnet algorithm was alpha = 0.10 and lambda = 0.0005, the AUC reached the maximum value of 0.954.

[0151] The equation of the model constructed based on the optimal hyperparameter combination is:

[0152]

[0153] Among them, Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 9, corresponding to the previous composition), Xi represents the detected value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker, and b is the constant 6.3464595; the coefficients of the 9 biomarkers are as follows:

[0154] Table 6. Coefficients of 9 biomarkers in the model

[0155]

[0156] The complete model equation is:

[0157] Y = 7.180FBLN5 + 6.227MUC16 + 1.460ORM1 + 4.025ADH1B + 1.584CD36 + 8.109KRT19 + 3.016DCD + 6.542DEFA3 + 9.213RNASE1 + 6.3464595

[0158] Determination of the diagnostic threshold of the prediction model:

[0159] ## Setting levels: control = case, case = control

[0160] ## Setting direction: controls < case

[0161] The ROC curve was plotted using the predicted values in the training group, and the optimal diagnostic cut-off value of 0.513363 was set according to the Youden index value. That is, when the predicted score of the prediction model ≤ 0.513363, it is considered that the risk of breast cancer recurrence in the subject to be tested is low; when the model predicted score > 0.513363, it is considered that the risk of breast cancer recurrence in the subject to be tested is high. The results are shown in Figure 4 .

[0162] From Figure 4 it can be seen that the AUC of the prediction model in the training group is 0.954, the sensitivity is 0.953, and the specificity is 0.965.

[0163] 2.3. Verification of the prediction model

[0164] The constructed optimal model was verified in the test group, and the ROC curve was plotted as Figure 5 shown.

[0165] From Figure 5 it can be known that the AUC of the prediction model in the test group is 0.932, the sensitivity is 0.941, and the specificity is 0.968, which is very close to the diagnostic effect in the training group.

[0166] Overall, the breast cancer recurrence risk prediction model constructed with 9 protein markers has good prediction performance and accuracy, and the best diagnostic efficacy.

[0167] Example 3 Construction and verification of a three-class prediction model

[0168] In this example, an attempt was made to construct a three-class combined prediction model for distinguishing the non-recurrence group, in-situ recurrence group, and metastatic group of breast cancer prognosis. The specific process includes the following: (1) construction and screening of the best prediction model; (2) verification of the effect of the best prediction model. The specific screening process and results are as follows (in the present invention, the AUC value was used as the evaluation index for the binary classification model in Example 2; when constructing a three-class model, since multiple categories are involved, the AUC value is usually not applicable, and in this example, indicators such as sensitivity, specificity, accuracy, and consistency were used to measure the diagnostic efficacy of the model):

[0169] 3.1 Construction and screening of the prediction model

[0170] For the test cohort of 300 breast cancer patients, all enrolled patients signed informed consent forms. Among them, 64 breast cancer patients relapsed within three years after surgery (30 had in-situ local recurrence, 34 had metastases, and those with both in-situ recurrence and metastases were also classified as having metastases). Thus, they were divided into two groups: 236 patient samples without recurrence were classified into the low-risk group of breast cancer recurrence, and 64 patients with recurrence or metastases were classified into the high-risk group of breast cancer recurrence. They were randomly divided into a training group and a test group. The training group included 118 patients in the low-risk group of breast cancer recurrence and 32 patients in the high-risk group of breast cancer recurrence, including 15 cases of in-situ local recurrence and 17 cases of metastases; the test group included 118 patients in the low-risk group of breast cancer recurrence and 32 patients in the high-risk group of breast cancer recurrence, including 15 cases of in-situ local recurrence and 17 cases of metastases. In this example, it was hoped to further construct a three-class detection model that can effectively distinguish the non-recurrence group (low-risk group), in-situ recurrence group (in-situ recurrence group), and metastatic group (metastatic group, and those with both in-situ recurrence and metastases were also classified into the metastatic group) of breast cancer prognosis on the basis of the marker combination form FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1 screened in Example 2. All enrolled patients signed informed consent forms. Among them, all breast cancer patients were those diagnosed by pathological histology. Inclusion criteria: (a) no history of other malignant tumors; (b) no patients with combined other malignant tumors or autoimmune diseases.

[0171] In this embodiment, LC-MS / MS data collection and detection were performed on the collected serum samples, and the concentrations of nine protein markers, namely FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1, were obtained respectively. The Shapiro Wilk test was used to evaluate the normal distribution, and the non-parametric Wilcoxon test was used to analyze the differences in blood marker concentrations between the breast cancer prognosis non-recurrence group (low-risk group), breast cancer prognosis in-situ recurrence group (in-situ recurrence group), and breast cancer prognosis distant metastasis group (metastasis group). A three-class combined prediction model of nine markers was constructed using a method combining machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using the predicted probability value with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate prediction model. Using the training group, the Youden index (YI) was calculated to determine the predicted probability cut-off value for distinguishing the low-risk group, in-situ recurrence group, and metastasis group. In addition, the ROCs of individual markers and different subgroups were constructed and compared. Standard descriptive statistics, such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD), were calculated to describe the experimental results of the study population. Statistical analysis was performed using R 3.6.1, and a p-value less than 0.05 was considered statistically significant.

[0172] In this embodiment, in order to construct an optimal three-class combined prediction model, after comparing the models constructed by six algorithms, namely gradient boosting, naive Bayes, support vector machine, neural network, generalized linear, and discriminant analysis, the gradient boosting method was selected as the best supervised classification algorithm for constructing the prediction model. The grid search range for hyperparameter optimization of the gradient boosting method is shown in Table 7 below.

[0173] Table 7. Parameter grid search range of the gradient boosting method

[0174]

[0175] Through optimization and screening in terms of accuracy, consistency, sensitivity, specificity, etc., the optimal parameter combination mode was determined as: interaction.depth 1, n.trees 100, shrinkage 0.1, n.minobsinnode 10.

[0176] Completely different two batches of samples were used for the training group and the test group. Only marker screening and model construction were carried out in the training group; the samples of the test group were only used to verify the diagnostic efficacy of the model. The specific results are shown in Table 8.

[0177] Table 8. Three-class diagnostic efficacy

[0178]

[0179] As can be seen from Table 8, the gradient boosting model constructed based on nine protein markers of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1 can be used to predict whether breast cancer patients will have no recurrence, in-situ recurrence, or metastasis after surgery. It also shows that the protein markers screened in the present invention can be used to distinguish the recurrence risk of breast cancer patients after surgical treatment, and can also be used to distinguish whether it is in-situ recurrence or metastasis when the recurrence risk is high (when there is both in-situ recurrence and metastasis, it is also classified into the metastasis group).

[0180] 3.2, Combined Performance of the Three-Class Joint Prediction Model

[0181] In order to further improve the diagnostic value of the three-class prediction model (gradient boosting) constructed by biomarker combinations of different proteins, in this embodiment, based on the 10 protein markers screened in Example 1, the performance of the prediction models constructed by biomarker combinations of different proteins was compared in the training group. The specific combination forms of different models are shown in Table 9.

[0182] Table 9, Combination Forms of Different Prediction Models

[0183]

[0184] The results are specifically shown in Table 10. Table 10 is the comparison result of the performance indicators of different prediction models constructed by using the 10 biomarkers screened in Example 1 for three-class classification. The calculation methods of the minimum value, the first quartile, the median, the mean, the third quartile, and the maximum value of accuracy and consistency are as follows: (1) Sort the values of accuracy or consistency from small to large; (2) Minimum value: the first value after sorting; (3) First quartile (Q1): Multiply the number of data by 0.25. If the result is an integer, take the average of the values at this position and the next position; if it is not an integer, round up to get the position, and the value at this position is Q1; (4) Median: If the number of data is odd, the median is the middle value; if it is even, it is the average of the two middle values; (5) Mean: The sum of all values divided by the number of data; (6) Third quartile (Q3): Multiply the number of data by 0.75, and the processing method is the same as Q1; (7) Maximum value: the last value after sorting. Among them, the minimum value and the maximum value can reflect the extreme situations of the data and show the worst and best performances that the model may have; the quartiles can help understand the distribution range and dispersion degree of the data; below Q1 represents a lower performance level, and above Q3 represents a higher performance level; the median can reflect the performance of the middle level; the mean comprehensively reflects the overall average performance. Based on the above statistical values, the overall situation, distribution characteristics, and stability of the model performance can be comprehensively understood, thus providing a strong basis for model selection and optimization.

[0185] Table 10. Performance Comparison of Prediction Models Constructed Based on Different Protein Combination Biomarkers

[0186]

[0187] As can be seen from Table 10, for the three-class classification prediction model, the ten-item combined detection model (10MP) composed of ten biomarkers has the best performance. This also clearly shows that the constructed 10MP model has a very obvious improvement effect on the diagnostic efficiency of distinguishing whether breast cancer patients will not relapse, have in-situ recurrence, or develop metastasis after surgery. Therefore, it is preferably to use the three-class gradient boosting model constructed by the ten protein biomarkers (FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, KLHL22) as the best combined prediction model.

[0188] 3.3. Determination and Verification of the Diagnostic Performance of the Three-Class Combined Prediction Model (10MP)

[0189] 3.3.1. Determination of the Diagnostic Performance of the Three-Class Combined Prediction Model (10MP)

[0190] To more accurately determine the diagnostic performance and thresholds of the model constructed in this embodiment for different disease classifications, a multi-classification model of the gradient boosting (gbm) algorithm in the model group was used for predictive analysis in the training group, and the predicted probability values for the three-classification (low-risk group, in-situ recurrence group, and metastasis group) were calculated. The classification with the largest predicted probability value is the final prediction result of the system.

[0191] Among them, the meanings and calculation methods of each index are as follows:

[0192] Calculation results: The accuracy of the three-classification combined prediction model (9MP) in the training group is 0.88, and the consistency is 0.89. The diagnostic sensitivity for the low-risk group is 92.5%, and the specificity is 95.9%; the diagnostic sensitivity for the in-situ recurrence group is 85.8%, and the specificity is 86.3%; the diagnostic sensitivity for the metastasis group is 86.2%, and the specificity is 85.6%.

[0193] It should be noted that the three-classification combined prediction model (10MP) constructed by gradient boosting is a model constructed by machine learning and cannot fit a specific equation formula like a generalized linear model.

[0194] 3.3.2. Verification of the three-classification combined prediction model (9MP)

[0195] Based on the model constructed in the training group, the predictive performance was verified in the test group, and the specific results are as follows:

[0196] The accuracy is 0.88, and the consistency is 0.87. The diagnostic sensitivity for the low-risk group is 92.8%, and the specificity is 94.8%; the diagnostic sensitivity for the in-situ recurrence group is 85.6%, and the specificity is 86.8%; the diagnostic sensitivity for the metastasis group is 85.1%, and the specificity is 84.5%.

[0197] In summary, the three-classification combined prediction model containing 10 protein markers constructed in this embodiment can not only accurately judge the recurrence risk of breast cancer, but also better predict the three classifications of the low-risk group, in-situ recurrence group, and metastasis group.

[0198] All patents and publications mentioned in the specification of the present invention indicate that these are publicly available technologies in the art and can be used in the present invention. All patents and publications cited herein are equally listed in the references, just as each publication is specifically individually referenced. The present invention described herein can be implemented in the absence of any one or more elements, one or more limitations, where such limitations are not specifically stated. For example, in each instance herein, the terms "comprising," "consisting essentially of," and "consisting of" can be replaced by either of the remaining two terms. The so-called "a" herein merely means "one," and does not exclude including only one, nor does it exclude including more than two. The terms and expressions used herein are for the purpose of description and are not limiting, and there is no intention to indicate that the terms and interpretations described herein exclude any equivalent features, but it is understood that any suitable changes or modifications can be made within the scope of the present invention and the claims. It is understood that the embodiments described in the present invention are some preferred embodiments and features, and any person of ordinary skill in the art can make some changes and variations based on the essence described in the present invention, and such changes and variations are also considered to be within the scope of the present invention and the scope limited by the independent claims and the dependent claims.

Claims

1. Use of a biomarker in the preparation of a reagent for predicting the risk of breast cancer recurrence or metastasis, characterized in that, The biomarker(s) include(s) one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.

2. The use according to claim 1, characterized in that, The biomarker(s) include(s) FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1.

3. The use according to claim 2, characterized in that, The risk of recurrence or metastasis refers to recurrence or metastasis within three years after breast cancer treatment.

4. The use according to claim 2, characterized in that, The biological sample(s) is / are selected from one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid; and / or, the detected amount includes the presence or absence, relative abundance, or concentration of the biomarker.

5. A biomarker combination for predicting the risk of breast cancer recurrence or metastasis, characterized in that, including FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1.

6. A method for constructing a risk prediction model for breast cancer recurrence or metastasis for non-diagnostic purposes, characterized in that, including the following: 1) Construct a data set based on the detected amount of the biomarker in the biological sample, where the biomarker(s) include(s) FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1; 2) Divide the data set into a test set and a training set, and construct and train the breast cancer recurrence or metastasis risk prediction model through machine learning methods.

7. A device for predicting the risk of breast cancer recurrence or metastasis, characterized in that, including a data acquisition unit and a calculation unit; The data acquisition unit is used to obtain the detected amount data of the biomarker in the biological sample of the subject, where the biomarker(s) include(s) FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1; The calculation unit is used to calculate and output the predicted score of the subject's risk of breast cancer recurrence or metastasis based on the detected amount data, and make a judgment according to the cut-off value.

8. An apparatus, comprising a processor and a memory, the memory being configured to store a computer program, characterized in that The processor is used to execute the computer program stored in the memory, so that the device executes the construction method as described in claim 6.

9. A system for predicting the risk of breast cancer recurrence, characterized in that, The system includes a data analysis module, which is used to analyze the detection value of the biomarker, where the biomarker(s) include(s) FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1.

10. The system according to claim 9, characterized in that, The data analysis module uses the detection values of the biomarkers of known samples as the training set, divides them into a breast cancer recurrence group and a non-recurrence group of breast cancer according to whether breast cancer recurs, analyzes the relationship between the detection values of the breast cancer recurrence group and the non-recurrence group of breast cancer, and constructs a model. The equation of the model is: where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 9), Xi represents the detection value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker, and b is a constant 6.3464595; the coefficients of the 9 biomarkers are: When Y ≤ 0.513363, the risk of breast cancer recurrence or metastasis of the tested person is low; when Y > 0.513363, the risk of breast cancer recurrence or metastasis of the tested person is high.

Citation Information

Patent Citations

  • Methods and kit for the prognosis of breast cancer

    CN101014720A

  • Marker for breast cancer diagnosis and application thereof

    CN115343476A

  • Biomarker for diagnosing breast cancer of to-be-detected person

    CN116759070A

  • Construction method and risk prediction method of breast cancer far-end metastasis risk prediction gene model

    CN116936086A

  • Method for determining prognosis of breast cancer patient after treatment with surgery

    JP2013245960A

Cited By

  • Method for predicting breast cancer in-situ recurrence by using specific microorganisms of cancer tissue

    CN121555667A