A biomarker for predicting the risk of breast cancer recurrence or metastasis and its application
By screening and constructing a proteomic-based biomarker model, the accurate prediction of breast cancer risk of recurrence or metastasis is solved, and the early identification of high-risk patients is achieved, misdiagnosis and misdiagnosis are reduced, and the targeted and efficient treatment is improved.
Patent Information
- Application Number
- CN202510725139.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The prior art is difficult to accurately predict the risk of breast cancer recurrence or metastasis, resulting in overt treatment or insufficient treatment, and burdening patients and medical resources.
Biomarkers such as FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1 and KLHL22 were screened through proteomics to construct a breast cancer recurrence risk prediction model, and blood samples were analyzed using high-performance liquid chromatography-tandem mass spectrometry technology, combining orthogonal partial least squares discriminant analysis and significance analysis to distinguish patients with recurrence and non-relapsed breast cancer.
It has achieved accurate, non-invasive and efficient prediction of the risk of breast cancer recurrence, and can identify high-risk groups early, help adjust treatment plans, reduce the risk of recurrence or metastasis, and improve diagnosis accuracy.
Smart Images

Figure CN120254283B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of proteomics screening of breast cancer diagnostic markers, and in particular to a biomarker for predicting the risk of breast cancer recurrence or metastasis and its application. Background Art
[0002] Breast cancer is one of the most common malignant tumors in women, with an incidence rate accounting for 23% of all tumors. With approximately 1.6 million new cases worldwide each year, it is the second leading cause of cancer death among women worldwide. The key to the diagnosis and treatment of breast cancer lies in early detection, early diagnosis, and early treatment. Early detection and timely treatment of breast cancer can significantly improve the patient's cure rate, increase the chances of breast and axillary conservation, and thus enhance the patient's quality of life. However, despite the continuous development and certain achievements of current screening programs and adjuvant therapies, a considerable proportion of breast cancer patients still die from cancer metastasis.
[0003] Surgical resection remains the main treatment for breast cancer patients, especially for early-stage and locally advanced breast cancer, but about 30% of breast cancer patients will experience recurrence within 2 to 5 years after surgery. Breast cancer recurrence is divided into two categories: local recurrence and distant metastasis. Local recurrence refers to the recurrence of cancer cells in the breast area or regional lymph nodes, while distant metastasis refers to the spread of cancer cells through blood vessels to important organs such as the lungs, liver or brain. If patients who do not benefit from surgical treatment or who have progressed can have their risks predicted and their treatment plans adjusted in a timely manner (such as adjuvant chemoradiotherapy, secondary surgical resection, targeted therapy or immunotherapy, etc.), the overall survival rate and quality of life of breast cancer patients can be significantly improved.
[0004] Currently, assessing the likelihood of breast cancer recurrence or distant metastasis relies primarily on regular follow-up, which often leads to overtreatment or undertreatment. Treating all patients with the same intensity of treatment can cause some to suffer unnecessary side effects while others fail to achieve the desired therapeutic effect. This not only places a burden on patients and their families but also wastes medical resources. For postoperative patients, the uncertainty of recurrence is a tremendous torment and suffering.
[0005] Proteomics is the study of protein composition, localization, changes, and interactions within cells, tissues, or organisms, encompassing both protein expression patterns and proteome functional patterns. With the continuous advancement of mass spectrometry technology, liquid chromatography coupled to mass spectrometry (LC-MS / MS) has become a core tool in proteomics research. The development of proteomics is crucial for identifying disease diagnostic markers, screening drug targets, and conducting toxicology studies, and is therefore widely used in medical research. Although numerous articles and patents have been published in recent years regarding the discovery of tumor markers, these remain largely at the laboratory research stage, with limited clinical application and market penetration. Furthermore, in most cases, relying solely on a single indicator for in vitro diagnosis of tumor recurrence risk is insufficient. Only by combining multiple diagnostic tests across multiple dimensions can predictive accuracy be effectively improved.
[0006] In summary, developing more clinically significant tumor recurrence risk markers to effectively analyze tumor recurrence or metastasis is of great significance for the diagnosis and treatment of breast cancer. Summary of the Invention
[0007] In response to the problems existing in the prior art, the present invention provides a biomarker for predicting the risk of breast cancer recurrence or metastasis and its application. By using proteomics methods, by analyzing proteins with significantly different abundance levels in the blood of two groups of breast cancer patients after surgical treatment, namely, those with recurrent breast cancer and those without recurrence, biomarkers that can be used to predict the risk of breast cancer recurrence are screened out. Further, based on the markers, a breast cancer recurrence risk or metastasis prediction model and prediction device are constructed, which can accurately, non-invasively and efficiently predict the risk of breast cancer recurrence. In addition, the biomarker, prediction model and prediction device of the present invention can also effectively distinguish between in situ recurrence and distant metastasis, and therefore can meet clinical needs.
[0008] In one aspect, the present invention provides the use of a substance for detecting a biomarker in the preparation of a product for measuring the risk of breast cancer or metastasis recurrence, wherein the biomarker comprises one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.
[0009] The proteomic markers provided by the present invention can accurately predict the risk of recurrence or metastasis of breast cancer patients after surgical treatment, thereby enabling earlier identification of people at high risk of recurrence or metastasis. This can help doctors predict in advance whether treatment plans need to be adjusted in a timely manner, reduce the risk of breast cancer recurrence or metastasis, and truly benefit breast cancer patients.
[0010] The present invention utilizes a proteomics approach to collect plasma samples from patients who experience a short-term recurrence (e.g., within 1 to 5 years) after surgical treatment for breast cancer, as well as patients who do not experience a short-term recurrence. The different samples are analyzed using high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS). Based on orthogonal partial least squares discriminant analysis and significance analysis, proteins with significant differences between patients with breast cancer recurrence and those without recurrence are first screened. Ultimately, 10 differentially expressed proteins with a clear correlation with the risk of breast cancer recurrence are identified. These 10 proteins can be used to distinguish between patients with breast cancer recurrence and those without recurrence, and between patients with in situ recurrence and those with distant metastasis. These proteins have a certain diagnostic efficacy and can therefore be used to predict the risk of breast cancer recurrence or metastasis.
[0011] Among them, the FBLN5 is a protein or amino acid sequence with a UniProt database number of Q9UBX5; MUC16 is a protein or amino acid sequence with a UniProt database number of Q8WXI7; ORM1 is a protein or amino acid sequence with a UniProt database number of P02763; ADH1B is a protein or amino acid sequence with a UniProt database number of P00325; CD36 is a protein or amino acid sequence with a UniProt database number of P16671; KRT19 is a protein or amino acid sequence with a UniProt database number of P08727; DCD is a protein or amino acid sequence with a UniProt database number of P81605; DEFA3 is a protein or amino acid sequence with a UniProt database number of P59666; RNASE1 is a protein or amino acid sequence with a UniProt database number of P07998; KLHL22 is a protein or amino acid sequence with a UniProt database number of Q53GT1.
[0012] The present inventors also surprisingly discovered that some of the protein markers obtained through proteomic screening are already known markers for other cancers. For example, ADH1B has been reported to be used for predicting laryngeal cancer, pharyngeal cancer, nasal cancer, pancreatic cancer, and esophageal cancer. However, this screening revealed that this marker can also be used to predict the risk of recurrence or metastasis after surgical treatment of breast cancer. This shows that the protein markers for many different cancers are not completely isolated or unrelated, and in fact, there are many cross-relationships or influences. Many protein markers can be used for both early cancer prediction and prognostic diagnosis, and even many protein markers can be used for prediction and diagnosis of different cancers at different stages. Therefore, the field of proteomics has many new applications that need to be explored, and the market prospects are very broad.
[0013] Furthermore, the markers include FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3 and RNASE1.
[0014] To improve the diagnostic efficacy of breast cancer recurrence or metastasis risk, it is necessary to combine different differentially expressed proteins to construct a predictive model. These proteins are ranked according to their importance, and different numbers of differentially expressed proteins with the highest rankings are selected for combination. Ultimately, nine protein markers were screened. A model based on these nine protein markers has demonstrated strong predictive capabilities for breast cancer recurrence or metastasis risk. Data from clinical breast cancer samples showed that using only these nine biomarkers to predict breast cancer recurrence risk achieved an AUC value of 0.954, demonstrating excellent diagnostic performance.
[0015] Furthermore, the risk of recurrence or metastasis refers to recurrence or metastasis within three years after breast cancer treatment; the recurrence or metastasis of breast cancer includes recurrence of breast cancer in situ or adjacent areas, and metastasis of breast cancer.
[0016] It is understandable that the short term "recurrence or metastasis within three years after breast cancer treatment" here is not an absolute and unchanging time node. It is only a time node summarized based on current clinical experience. With the passage of time or the improvement of other treatment methods, this time node or the length of time may change. For example, breast cancer recurrence or metastasis after treatment may occur within one year, 360 days, six months, 180 days, two years, two and a half years, three and a half years, etc.
[0017] In some embodiments, the breast cancer recurrence or metastasis refers to the recurrence of breast cancer in situ or in adjacent areas within three years after radical tumor resection, or the occurrence of breast cancer metastasis, such as lymph node metastasis, bone metastasis, etc.
[0018] In some embodiments, the product is used to predict the risk of breast cancer recurrence or metastasis in patients.
[0019] In some embodiments, the product is a detection reagent prepared using the biomarker as the detection target, such as a sample pretreatment reagent, antigen, or antibody, or other biological reagent and kit suitable for detecting the biomarker. Standardized reagents or kits suitable for detecting the biomarker can also be developed. Specifically, the product can be one or more of a reagent, kit, chip, probe, or membrane strip.
[0020] In some embodiments, the product is used to detect the content of biomarkers in a body fluid sample; the body fluid sample includes any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.
[0021] Furthermore, the product is used to detect the presence or relative abundance or concentration of biomarkers in body fluid samples.
[0022] The present invention uses blood screening to identify biomarkers that predict the risk of breast cancer recurrence or metastasis. These biomarkers show significant differences in the blood of people at high risk of breast cancer recurrence and people at low risk of breast cancer recurrence. By collecting blood samples, these biomarkers can be detected in the blood of individuals to predict or assist in diagnosing the possibility of breast cancer recurrence or metastasis in that individual. Alternatively, these biomarkers can be detected in the blood of a certain group of people, and then the group can be divided into people at high risk of breast cancer recurrence and people at low risk of breast cancer recurrence, and the risk of in situ recurrence and the risk of distant metastasis.
[0023] Furthermore, the detection method includes a radiometric method, an immunological method, a fluorescence method, a flow cytometry method, a latex turbidimetry method, a biochemical method, an enzymatic method, a hybridization method, a gas chromatography-mass spectrometry method, a liquid chromatography-mass spectrometry method, a chromatography method, a chemiluminescence method, a magnetoelectric method or a photoelectric conversion method.
[0024] The presence or absence of a marker, or the level of a marker, is a relative concept. For example, when comparing a high-risk group for breast cancer recurrence with a low-risk group, the levels of these specific markers are compared relative to the baseline of the high-risk group and the low-risk group. It's possible that the levels of certain markers in the high-risk group are higher than those in the low-risk group, and this increase is statistically significant, such as a significant or highly significant increase. Therefore, when assessing the presence of a single marker, if the probability of a particular risk increases, the marker's level may change. This change may be a relative increase or decrease, and this relative increase or decrease is considered significant, or even highly significant. Therefore, regardless of the testing method, a predetermined cut-off value can be used as a standard. A value above this cut-off value is considered a change in the level, and such a result can be used for prognostic or diagnostic purposes.
[0025] Therefore, in some aspects, the markers described in the present invention can be obtained by detecting the marker content in a sample using any of the currently known methods, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strips, immunoreaction chips, capillary electrophoresis, infrared spectroscopy, etc. As long as they can be used to detect the protein marker content in a sample, they can be used to diagnose high-risk and low-risk breast cancer recurrence groups. As long as the protein marker content in a sample can be detected, it can be used to predict or diagnose the probability of a certain disease. It is understood that the detection here is to test an individual sample and then compare it with a pre-set standard. The comparison result is used to judge or predict the occurrence status of the disease. For example, it can be used to predict the probability of breast cancer recurrence. This prediction or diagnosis is whether it will occur within a certain period of time. Of course, such detection can be continuous detection, and the progress of the disease can be inferred as the content of certain substances changes.
[0026] In some embodiments, the relative abundance is the peak area of the biomarker in the detection spectrum obtained by high-performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a biomarker measured in a control sample is 100 and the average peak area measured in samples from patients at high risk of breast cancer recurrence is 600, then the abundance of the biomarker in the sample is considered to be 6 times that in the control sample.
[0027] In some embodiments, the biomarker can be a single marker or a combination of several markers, such as a combination of 2 markers, a combination of 3 markers, a combination of 4 markers, a combination of 5 markers, a combination of 6 markers, a combination of 7 markers, a combination of 8 markers, a combination of 9 markers, or a combination of 10 markers. When the combination of biomarkers ranges from 2 to 10, the AUC value of the constructed prediction model is 0.797-0.954, with high sensitivity and specificity.
[0028] A second aspect of the present invention provides the use of proteins as biomarkers in the preparation of products for predicting the risk of breast cancer recurrence or metastasis. The proteins are selected from one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22. The FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22 genes can be used to assist in determining the risk of breast cancer recurrence or metastasis, assess drug efficacy, and other purposes. The inventors have discovered that FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22 are closely associated with the risk of breast cancer recurrence or metastasis.
[0029] A third aspect of the present invention provides a product for predicting the risk of breast cancer recurrence or metastasis, comprising a substance for detecting biomarkers, wherein the biomarkers include one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1 and KLHL22.
[0030] A fourth aspect of the present invention provides a biomarker combination for predicting the risk of breast cancer recurrence or metastasis, comprising one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.
[0031] A fifth aspect of the present invention provides a method for constructing a breast cancer recurrence or metastasis risk prediction model for purposes other than disease diagnosis, comprising the following steps:
[0032] 1) constructing a sample dataset based on the detection amount of biomarkers in the biological sample, wherein the biomarkers include one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22;
[0033] 2) dividing the data set into a test set and a training set, and constructing and training the breast cancer recurrence or metastasis risk prediction model through machine learning methods.
[0034] In some embodiments, the biomarkers are FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1; the risk of recurrence refers to recurrence within three years after surgical treatment of breast cancer; and the breast cancer recurrence includes recurrence of breast cancer in situ or adjacent areas and breast cancer metastasis.
[0035] Furthermore, the biological samples are biological samples of breast cancer recurrence and biological samples of breast cancer non-recurrence, and the model is constructed by analyzing the relationship between the detection values of biomarkers in the biological samples of breast cancer recurrence and the biological samples of breast cancer non-recurrence.
[0036] In some embodiments, the machine learning method is selected from at least one of a gradient boosting algorithm, a random forest algorithm, a support vector machine algorithm, a decision tree algorithm, a K-nearest neighbor algorithm, a logistic regression algorithm, and a neural network algorithm.
[0037] In some embodiments, the prediction model is a generalized linear model (GLM) whose equation is:
[0038]
[0039] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m=9), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant of 6.3464595. The coefficients of the 9 biomarkers are:
[0040]
[0041] When Y≤0.513363, the risk of breast cancer recurrence or metastasis of the subject is low; when Y>0.513363, the risk of breast cancer recurrence or metastasis of the subject is high.
[0042] A fifth aspect of the present invention provides a device for predicting the risk of breast cancer recurrence or metastasis, comprising a data acquisition unit and a calculation unit;
[0043] The data acquisition unit is used to obtain the detection amount data of the biomarkers as described above in the biological sample of the subject, wherein the biomarkers include one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1 and KLHL22;
[0044] The calculation unit is used to calculate and output a predicted score for the subject's breast cancer recurrence risk based on the detection data, and make a judgment based on a threshold.
[0045] In some embodiments, the biomarkers include FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1.
[0046] In some embodiments, the prediction device further includes a data storage unit and a data output unit; the data storage unit is used to store the detection amount (or detection value) of the biomarker; the data input interface is used to input the detection value of the biomarker, and the data output unit is used to output the prediction result.
[0047] In some embodiments, the detection value is the presence or absence or relative abundance or concentration value of each biomarker.
[0048] The sixth aspect of the present invention provides a device comprising a processor and a memory, wherein the memory is used to store a computer program, and is characterized in that the processor is used to execute the computer program stored in the memory so that the device performs the prediction method as described above or the construction method as described above.
[0049] A seventh aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when processed, executes the prediction method as described above or the construction method as described above.
[0050] In some embodiments, the computer-readable storage medium includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0051] In some implementations, being processed refers to being executed by one or more processors.
[0052] An eighth aspect of the present invention provides a method for predicting the risk of breast cancer recurrence or metastasis, comprising the following steps:
[0053] S1. Obtaining detection amount data of biomarkers in a biological sample of a subject, wherein the biomarkers include one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22;
[0054] S2. Applying the breast cancer recurrence or metastasis risk prediction model obtained by the construction method described above to process the detection data to output a breast cancer recurrence or metastasis risk prediction result.
[0055] In some embodiments, the biomarkers include FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1.
[0056] In another aspect, the present invention provides a system for predicting the risk of breast cancer recurrence, the system comprising a data analysis module for analyzing the detection values of markers, wherein the markers include FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3 and RNASE1.
[0057] Furthermore, the data analysis module uses the detection values of markers of known samples as a training set, divides the samples into a breast cancer recurrence group and a breast cancer non-recurrence group according to whether the breast cancer recurs, analyzes the relationship between the detection values of the breast cancer recurrence group and the breast cancer non-recurrence group, and constructs a model.
[0058] In some embodiments, a combined diagnostic model for predicting the risk of breast cancer recurrence is constructed by combining multiple machine learning methods, and it is preliminarily confirmed that the concentration changes of any one of the screened biomarkers alone can be used to distinguish between people at high risk and low risk of breast cancer recurrence, indicating that these biomarkers have extremely high diagnostic value.
[0059] In some embodiments, the equation of the constructed model is:
[0060] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m=9), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant of 6.3464595. The coefficients of the 9 biomarkers are:
[0061]
[0062] When Y≤0.513363, the risk of breast cancer recurrence or metastasis of the subject is low; when Y>0.513363, the risk of breast cancer recurrence or metastasis of the subject is high.
[0063] Furthermore, the system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0064] In another aspect, the present invention provides a use of a marker for preparing a reagent for predicting whether a breast cancer patient will not relapse after surgery, relapse in situ, or have distant metastasis, wherein the marker comprises any one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.
[0065] The present invention attempts to apply protein markers used to predict recurrence risk to differentiate between patients with primary recurrence and those with distant metastasis. It was found that each protein marker can be used to distinguish between patients with breast cancer who have no recurrence after surgery, those with primary recurrence, and those with distant metastasis. The term "distant metastasis" encompasses patients with only distant metastasis as well as those with both primary recurrence and distant metastasis.
[0066] Furthermore, the reagent is used to predict whether a breast cancer patient will not relapse, relapse in situ or have distant metastasis within three years after surgery.
[0067] Furthermore, the reagent is used to detect the content of biomarkers in a body fluid sample; the body fluid sample includes any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.
[0068] Furthermore, the reagent is used to detect the presence or relative abundance or concentration of biomarkers in a body fluid sample.
[0069] In another aspect, the present invention provides a kit for predicting whether a breast cancer patient will not relapse after surgery, will relapse in situ, or will metastasize to distant sites. The kit comprises a detection reagent for the biomarker for the purpose described above.
[0070] In another aspect, the present invention provides a biomarker combination for predicting whether a breast cancer patient will not relapse after surgery, relapse in situ, or have distant metastasis, wherein the combination comprises CD36, PLIN1, ITIH3, MGAM2, RNASE1, SELL, ORM2, TFF1, and TFF3.
[0071] In another aspect, the present invention provides a system for predicting whether breast cancer will not recur after surgery, recur in situ, or metastasize to distant sites. The system includes a data analysis module, which is used to analyze the detection values of markers, wherein the markers include any one or more of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.
[0072] Furthermore, the distant metastasis also includes simultaneous in situ recurrence and distant metastasis. As long as distant metastasis occurs, it will be classified into the distant metastasis group.
[0073] Furthermore, the data analysis module uses the detection values of markers of known samples as a training set, and divides breast cancer patients into a non-recurrence group, an in situ recurrence group, and a distant metastasis group according to their post-operative conditions. The relationship between the detection values of the non-recurrence group, the in situ recurrence group, and the distant metastasis group is analyzed to construct a model.
[0074] Furthermore, the model is constructed based on the gradient boosting algorithm.
[0075] Unlike generalized linear regression, the gradient boosting algorithm cannot output model formulas and cutoff values. All calculations are done directly by machine learning. The test values can be directly input into the software system to obtain the prediction results.
[0076] Furthermore, the system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0077] The beneficial effects of the present invention are:
[0078] 1. The present invention has screened 10 biomarkers that can indicate the risk of breast cancer recurrence and developed a new protein marker combination. This can effectively assess and diagnose patients at risk of breast cancer recurrence, effectively distinguish between high-risk and low-risk groups for breast cancer recurrence, and more accurately identify breast cancer recurrence. Compared with traditional detection methods, it reduces the risk of misdiagnosis and missed diagnosis, providing strong support for early detection and intervention of the disease.
[0079] 2. The combined identification and prediction model of nine biomarkers constructed in the present invention is convenient and fast, and the test results are highly consistent with the clinical gold standard test results. At the same time, it significantly reduces the cost of predicting the risk of breast cancer recurrence and has good application prospects.
[0080] 3. Based on the biomarkers of the present invention, a three-classification model was further constructed that can simultaneously distinguish the non-recurrence group, the in situ recurrence group and the metastasis group, providing a more effective and accurate predictive diagnosis model.
[0081] Detailed description
[0082] (1) Diagnosis or prediction
[0083] The diagnosis or prediction here refers to the detection or analysis of biomarkers in a sample, or the content of a target biomarker, such as the absolute content or relative content, and then the presence or amount of the target marker is used to indicate whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease. The meanings of diagnosis and prediction here are interchangeable. The result of such a prediction or diagnosis cannot be directly used as a direct result of the disease, but is an intermediate result. If a direct result is obtained, other auxiliary means such as pathology or anatomy are required to confirm that the patient has a certain disease. For example, the present invention provides a variety of new biomarkers related to the risk of breast cancer recurrence or metastasis. Changes in the content of these markers are directly correlated with whether the patient belongs to the group at risk of breast cancer recurrence or metastasis.
[0084] (2) Association between markers, biomarkers, or differentially expressed proteins and the risk of breast cancer recurrence or metastasis
[0085] The terms "marker," "biomarker," and "differential protein" have the same meaning in this invention. Association here refers to a direct correlation between the presence or change in the level of a biomarker in a sample and a specific disease. For example, a relative increase or decrease in the level indicates a higher likelihood of the individual having the disease compared to a healthy population.
[0086] The simultaneous presence of multiple markers in a sample, or the relative changes in their levels, indicate a higher likelihood of the individual having the disease compared to healthy individuals. This means that among marker types, some are strongly associated with disease, while others are weakly associated, or even unrelated to a particular disease. One or more markers with strong correlations can be used as diagnostic markers, while markers with weaker correlations can be combined with stronger markers to diagnose a disease, increasing the accuracy of test results.
[0087] The numerous biomarkers in serum discovered by the present invention can be used to distinguish between high-risk and low-risk groups for breast cancer recurrence, and between primary and distant metastases. The markers herein can be used alone as individual markers for direct detection or diagnosis. Selecting such markers indicates that the relative change in the marker's content is strongly correlated with the risk of breast cancer recurrence or metastasis. Of course, it is understood that one or more markers with a strong correlation with the risk of breast cancer recurrence or metastasis can be selected for simultaneous detection. It is generally understood that, in some embodiments, selecting biomarkers with a strong correlation for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90% or 95% accuracy, which indicates that these markers can obtain intermediate values for diagnosing a certain disease, but does not mean that a certain disease can be directly confirmed.
[0088] Of course, the differential protein with the larger ROC value can also be selected as a diagnostic marker. The so-called strength is generally calculated and confirmed by some algorithms, such as the contribution rate or weight analysis of the marker and the risk assessment of breast cancer recurrence or metastasis. Such calculation methods can be significance analysis (p value or FDR value) and fold change (Fold change). Multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA) and orthogonal partial least squares discriminant analysis (OPLS-DA), and of course other methods, such as ROC analysis. Of course, other model prediction methods are also possible. When specifically selecting biomarkers, the differential proteins disclosed in the present invention can be selected, or other existing well-known marker combinations can be selected or combined to make predictions through model methods.
[0089] (3) Definition of disease terms
[0090] Breast cancer: Breast cancer is a malignant tumor that develops in breast tissue, usually originating in the milk ducts or lobules. Early symptoms of breast cancer may include breast lumps, nipple discharge, and swollen axillary lymph nodes.
[0091] Surgical resection is the preferred treatment for breast cancer. Depending on the extent of the breast tumor, either lumpectomy or complete mastectomy may be performed. Depending on the presence of cervical lymph node metastasis, either central or cervical lymph node dissection may be performed.
[0092] Breast cancer recurrence occurs when, after a period of clinically disease-free status (i.e., no evidence of tumor), a patient has completed curative treatment (surgical resection) and, following a period of clinically disease-free status (i.e., a period without evidence of tumor), cancer cells reappear at the primary tumor site or elsewhere in the body. Recurrence may result from residual cancer cells that were not completely eliminated during treatment or from previously undetected micrometastases. Breast cancer recurrence can occur months to years after treatment, but 80% of recurrences occur within 2-3 years after surgery, with the risk of recurrence significantly decreasing after 5 years. Breast cancer recurrence is closely related to tumor stage, treatment protocol, and postoperative management. While the prognosis is generally good, some patients may still experience recurrence, requiring regular monitoring, standardized treatment, and long-term follow-up, a complex and tedious approach. Breast cancer metastasizes primarily through the lymph nodes, although distant metastases, such as lung and bone metastases, can also occur in rare cases. Treatment options for recurrence or metastasis include surgery, radiotherapy, chemotherapy, or targeted therapy, tailored to the individual's circumstances. Early detection can help control recurrence and significantly improve the prognosis.
[0093] (4) The gold standard for diagnosing breast cancer recurrence is histopathological examination (i.e., pathological confirmation of biopsy or surgical resection specimens), which observes the presence of cancer cells under a microscope and combines immunohistochemistry or molecular testing to clarify the nature of the tumor. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] Figure 1 This is a volcano plot of the differential analysis of protein markers in the high-risk and low-risk groups for breast cancer recurrence in Example 1.
[0095] Figure 2 Graphs showing the ROC and OPLS-DA analysis results for the high-risk and low-risk groups for breast cancer recurrence in Example 1.
[0096] Figure 3 AUC results of the models constructed with different hyperparameters in Example 2.
[0097] Figure 4 This is the ROC curve diagram of the combined diagnosis model in Example 2 in the training group.
[0098] Figure 5 This is the ROC curve diagram of the combined diagnosis model in Example 2 in the test group. DETAILED DESCRIPTION
[0099] The present invention will be described in further detail below in conjunction with the accompanying drawings and Examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not serve to limit the present invention in any way. The reagents used in this example are all known products and were obtained by purchasing commercially available products.
[0100] Example 1 Screening biomarkers for predicting breast cancer recurrence risk using proteomics
[0101] Plasma samples were collected from breast cancer patients after radical surgery. Low-abundance proteins were enriched by immunoaffinity chromatography to remove high-abundance proteins. Protein abundance in the samples was detected by HPLC-MS / MS. The differences in protein abundance between patients with no recurrence of breast cancer within three years and those with recurrence or metastasis were analyzed, and the diagnostic performance was analyzed. The specific steps are as follows:
[0102] 1.1 Sample Collection
[0103] Peripheral blood samples (approximately 2 ml) were collected from 100 patients with breast cancer (stages I-III) 2 weeks after surgery. Blood samples were placed in vacuum tubes containing EDTA anticoagulant, mixed thoroughly, and centrifuged twice at 120g for 10 minutes at room temperature. The supernatant was collected and the supernatant was removed. Platelet samples were then collected in centrifuge tubes and stored at -80°C. The average age of the 100 patients was 43 years (range, 25-69 years). All enrolled patients provided written informed consent. Patients with breast cancer had a histopathologically confirmed diagnosis. Inclusion criteria included: (a) no history of other malignancies; and (b) no concurrent malignancies or autoimmune diseases.
[0104] Over the next three years, patients were followed up every three to six months, undergoing regular imaging examinations such as breast ultrasound and mammography. Any recurrence or metastasis was confirmed by pathological histology. A total of 24 breast cancer patients experienced recurrence within three years of surgery, 14 of whom had local recurrence in situ and 10 with lymph node metastasis. Samples stored at -80°C were then taken and divided into two groups: 76 samples from patients without recurrence were classified as a low-risk group for breast cancer recurrence, and 24 samples from patients with recurrence or metastasis were classified as a high-risk group for breast cancer recurrence. These samples were used for proteomic biomarker screening.
[0105] 1.2. Sample processing and enzymatic hydrolysis
[0106] First, plasma samples were centrifuged for 15 minutes at 15,000 g. The supernatant was filtered and then subjected to immunoaffinity chromatography to isolate 14 highly abundant proteins. Low-abundance proteins were then concentrated to 350 μL using a 3 kDa cutoff concentrator at 4000 g for 1 hour. The recovered concentrate was then subjected to buffer exchange using a 7 kDa cutoff desalting column at 1000 g for 2 minutes using AEX-A (20 mM Tris, 4 M Urea, 3% isopropanol, pH 8.0). Protein concentrations were determined using the BCA assay using AEX-A as a blank. According to sample grouping, 25 μL of TCEP was added to the samples, and the samples were incubated at 37°C for 30 minutes for protein reduction. TMT labeling was then performed by adding the corresponding TMT 16-plex reagent and incubating at room temperature in the dark for 1 hour. The sample was then buffer exchanged using a Zeba column with AEX-A. The TMT 16-plex-labeled samples were mixed, and 2 mL of AEX-A was added to the mixed samples, bringing the final volume to 5.5 mL. The samples were filtered through a 0.22 µm filter and separated using a 2D-HPLC system. The collected fractions were freeze-dried, and finally, Trypsin-Lysin C enzyme cocktail was added. The samples were digested by incubation at 37°C for 5 hours, and the digestion reaction was terminated by the addition of 5 μL of 10% TFA. A total of 60 2D-HPLC fractions were used for nanoLC-MS / MS analysis.
[0107] 1.3 LC-MS / MS Data Acquisition
[0108] Each sample from step 1.2 was separated using an Easy nLC-1200 liquid chromatography system at a nanoliter flow rate and coupled online to a Q Exactive HF-X high-resolution mass spectrometer. Mobile phase A consisted of 0.1% formic acid in water, and mobile phase B consisted of 0.1% formic acid in acetonitrile (80% acetonitrile, 20% water). The chromatographic columns consisted of an enrichment column and an analytical column equilibrated with 100% mobile phase A. Samples were loaded onto the enrichment column (100 μm ID x 4 cmL, C18, 3 μm, 100 Å) via an autosampler and separated on the analytical column (75 μm ID x 25 cmL, C18, 3 μm, 100 Å) at a flow rate of 300 nL / min. After chromatographic separation, the samples were analyzed by mass spectrometry on a Q Exactive HF-X mass spectrometer. Positive ion detection was used, with a precursor ion scan range of 350-1800 m / z. The primary mass spectrometer resolution was 120,000 at 200 m / z, with an AGC (Automatic Gain Control) target of 3e6, a Maximum Intensity (IT) of 50 ms, and a dynamic exclusion time of 40 s. The mass-to-charge ratios of peptides and peptide fragments were acquired using data-dependent acquisition (DDA): 20 secondary MS / MS spectra (MS2 scans) were collected after each full scan (MS / MS). The MS2 Activation Type is HCD, the Isolation window is 0.7 m / z, the secondary mass spectrometry resolution is 30,000 at 200 m / z, the AGC target is 1e5, the Maximum IT is 65 ms, the Fixed first mass is 110.0 m / z, the Normalized Collision Energy is 32 eV, the Minimum AGC target is 2.00e4, the Charge exclusion is 1, 6-8, >8, the Multiple charge states is one charge state only, the Peptide match is preferred, and the Exclude isotopes is on.
[0109] 1.4 Data Preprocessing
[0110] The secondary mass spectrometry data obtained in step 1.3 were searched using Maxquant (v1.6.15.0). The data type is DIA proteomics data based on secondary reporter ion quantification. The secondary spectrum used for quantification requires that the parent ion accounts for more than 75% in the primary spectrum. The database comes from the Homo_sapiens_9606_proteome_gene (release: 2021-10-14, sequence: 20,437) of the Uniprot database, and a common contamination library is added to the database. Contaminating proteins are deleted during data analysis; the enzyme digestion method is set to Trypsin / P; the number of missed cut sites is set to 2; the parent ion mass error tolerance of First search and Mainsearch is set to 20 ppm and 5 ppm, respectively, and the mass error tolerance of secondary fragment ions is 20 ppm. The fixed modification is cysteine alkylation, and the variable modification is methionine oxidation and protein N-terminal acetylation. The FDR for protein identification and PSM identification is set to 1%.
[0111] 1.5. Difference Analysis
[0112] Differential protein screening was performed using a combination of univariate and multivariate statistical analyses. Univariate analysis primarily included significance analysis (p-value or FDR value) and fold change analysis of signature molecules across different groups. Multivariate statistical analysis primarily included receiver operating characteristic (ROC) curve analysis and Boruta signature screening based on the random forest algorithm. All statistical analyses were performed using R. Specific R information is provided in Table 1.
[0113] Table 1. R used in the present invention and related information
[0114]
[0115] The variable importance for the projection (VIP) was calculated to measure the influence and explanatory power of each protein expression pattern on the classification and discrimination of each group of samples. The Wilcoxon rank sum test was further performed to obtain the corrected p value (FDR). According to the conditions of FDR < 0.01 and Fold change > 2, 62 down-regulated proteins and 64 up-regulated proteins were screened (see Figure 1 ).
[0116] In order to evaluate the role of each protein marker in the diagnosis and prediction of breast cancer recurrence risk, ROC and Boruta analysis methods were used to evaluate each protein marker. The results are shown in Figure 1The horizontal axis represents the AUC obtained from ROC analysis, and the vertical axis represents the -log10 (FDR) calculated by the Wilcoxon test. The size of the dots represents the VIP value obtained from Boruta analysis. Further screening based on VIP > 3 and AUC > 0.6 identified 10 more significant candidate protein biomarkers, as detailed in Table 2.
[0117] Table 2. Candidate protein markers
[0118]
[0119] Among them, the smaller the FDR value and / or the larger the VIP value, to a certain extent, it indicates that the difference of the protein between the high-risk and low-risk breast cancer recurrence groups is more significant, and it also indicates that the protein may have a higher diagnostic value.
[0120] Example 2: Construction and validation of a breast cancer recurrence risk model
[0121] In this example, the model constructed using the 10 protein markers DEFA3, ORM1, FBLN5, MUC16, KRT19, CD36, DCD, ADH1B, RNASE1, and KLHL22 screened in Example 1 was studied.
[0122] While a single biomarker can differentiate post-operative breast cancer recurrence risk, combining multiple biomarkers generally offers greater accuracy in differentiation or prediction. However, a single biomarker that demonstrates a higher accuracy in predicting post-operative breast cancer recurrence risk may not necessarily be more effective when combined with one or more other biomarkers. Furthermore, a greater number of biomarkers does not necessarily equate to a higher prediction accuracy (AUC value) for the combination. Therefore, extensive validation experiments are still necessary.
[0123] 2.1 Models constructed using different marker combinations
[0124] The research cohort consisted of 220 breast cancer patients whose blood samples were collected two weeks after surgery. All enrolled patients signed informed consent forms, of which 44 breast cancer patients relapsed within three years after surgery (24 were local recurrences in situ and 20 were metastatic). The samples were divided into two groups: 176 patient samples without recurrence were classified as a low-risk group for breast cancer recurrence, and 44 samples with recurrence or metastasis were classified as a high-risk group for breast cancer recurrence. The samples were randomly divided into a training group and a test group. The training group included 88 patients in the low-risk group for breast cancer recurrence and 22 patients in the high-risk group for breast cancer recurrence, and the test group included 88 patients in the low-risk group for breast cancer recurrence and 22 patients in the high-risk group for breast cancer recurrence.
[0125] In the training group, a combination of multiple machine learning methods was used to construct a joint prediction model for multiple protein markers. The predicted probability values were used to estimate the area under the receiver operator characteristic (ROC) curve (AUC) with a 95% confidence interval (CI) to evaluate the discriminatory ability of the multivariate prediction model. Using the training group, the Youden index (YI) was calculated to determine the cut-off value for the predicted probability of distinguishing between the high-risk group for breast cancer recurrence and the low-risk group for breast cancer recurrence. In addition, ROCs for single markers and different combinations were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV) and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R3.6.1, and a p-value less than 0.05 was considered statistically significant.
[0126] The steps for building the joint prediction model are:
[0127] S101, randomly select 2 to 10 marker concentration matrices from the 10 protein markers DEFA3, ORM1, FBLN5, MUC16, KRT19, CD36, DCD, ADH1B, RNASE1, and KLHL22 in the training group samples as the original training data set.
[0128] S102, select the generalized linear model (glmnet) algorithm for building the prediction model, and the grid search range during the algorithm hyperparameter optimization process. In this step, the grid search range for the model hyperparameter optimization is set for each algorithm as shown in Table 3.
[0129] Table 3. Parameter grid of glmnet algorithm
[0130]
[0131] S103: According to the algorithm and hyperparameter setting range set in step S102, one of the hyperparameter combinations is selected as the parameters for constructing the prediction model.
[0132] S104: Split the original dataset into K subsets using a K-fold cross validation mechanism. To ensure that the ratio of majority class samples to minority class samples in each subset is the same as in the original dataset, a Stratified K-Folds cross validation mechanism is used for data segmentation.
[0133] S105 , according to the K training data subsets obtained by segmentation in step S104 , one of the subsets is selected as a validation set Ddev.
[0134] S106: Merge the training data subsets not selected in step S105 to form a training data pool Dtrain1.
[0135] S107 , building a prediction model based on the selected supervised classification algorithm and hyperparameters according to the training data set Dtrain obtained in step S106 .
[0136] S108: Evaluate the prediction model obtained in step S107 on the validation set Ddev to obtain an AUC value. The current prognosis prediction model and the corresponding AUC value are stored in the prediction model pool Pool. Step S108 involves evaluating the prediction model obtained in step S107 on the validation set determined in the current iteration. The model and evaluation results are stored in the prediction model pool for future prediction model selection. The evaluation mentioned in this step can be an AUC value or other reasonable metric for evaluating model performance.
[0137] S109: Determine whether all subsets have been used as validation sets. Step S109 determines whether all K subsets obtained in step S104 have been used as validation sets and trained on the model. If all subsets have been used as validation sets and training has been completed, proceed to step S110; if any subsets have not been used as validation sets, proceed to step S105. This step ensures that every sample in the original dataset has been used as a validation set, improving model stability and preventing overfitting of the model to a particular subset.
[0138] S110: The average AUC value of all models in the prediction model pool Pool is used as the final performance evaluation value of the combined model. The model parameters and the final performance evaluation AUC value are stored in the optimal model pool Poolbest.
[0139] S111: Determine whether all hyperparameter combinations have been used to construct prediction models. Step S111 determines whether prediction models have been constructed for all algorithms and corresponding hyperparameter combinations obtained in step S102. If all combinations have been used to construct models, step S112 is executed. If any combination has not been used to construct models, step S103 is executed.
[0140] S113 , selecting the model with the largest AUC value from the model set Poolbest obtained in step S112 as the final prediction model for breast cancer recurrence risk diagnosis.
[0141] S114, repeat all the above steps until all combinations of markers are modeled.
[0142] By executing the above model-building steps, the optimal models for all combinations of markers were obtained. To compare the performance of the models under these different marker combinations, the ROC method was used to evaluate the AUC values of these models in the training set. The results are shown in Table 4.
[0143] Table 4. Comparison of the area under the ROC curve of the models constructed with different marker combinations in the training group
[0144]
[0145] Table 4 ranks the markers in Table 2 according to the larger VIP value and the smaller FDR value. Starting from the top-ranked marker, the markers are selected in sequence for combination. From 2MP to 9MP, the AUC value, accuracy, and sensitivity of the detection increase with the increase of markers. However, when more markers are added on the basis of 9MP, the diagnostic performance of the constructed model does not continue to increase. The diagnostic efficacy of 9MP is similar to that of 10MP, and 9MP requires fewer markers and is lower in cost. Therefore, the most preferred model is the model constructed by the combination of 9 markers (FBLN5+MUC16+ORM1+ADH1B+CD36+KRT19+DCD+DEFA3+RNASE1).
[0146] 2.2 Optimization of model parameters
[0147] For the optimal marker combination FBLN5+MUC16+ORM1+ADH1B+CD36+KRT19+ DCD+DEFA3+RNASE1, the models constructed under 9 different combinations of glmnet algorithm hyperparameters were analyzed based on this marker combination, and the model performance was evaluated by the AUC value (AUC was calculated using a 10-fold cross-validation method during the modeling process). The results are shown in Tables 5 and Figure 3 shown.
[0148] Table 5. AUC of the constructed model under different hyperparameter combinations of the glmnet algorithm
[0149]
[0150] As can be seen from Table 5, when the hyperparameter combination of the glmnet algorithm is alpha = 0.10, lambda = 0.0005, the AUC reaches a maximum value of 0.954.
[0151] The equation for building a model based on the optimal hyperparameter combination is:
[0152]
[0153] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 9, corresponding to the previous composition), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant of 6.3464595; the coefficients of the 9 biomarkers are:
[0154] Table 6. Coefficients of the 9 biomarkers in the model
[0155]
[0156] The complete model equation is:
[0157] Y=7.180FBLN5+6.227MUC16+1.460ORM1+4.025ADH1B+1.584CD36+8.109KRT19+3.016DCD+6.542DEFA3+9.213RNASE1+6.3464595
[0158] Determination of diagnostic threshold of prediction model:
[0159] ## Setting levels: control = case, case = control
[0160] ## Setting direction: controls < case
[0161] The ROC curve was drawn using the predicted values from the training group, and the optimal diagnostic cutoff value of 0.513363 was set based on the Youden index. That is, when the prediction model score is ≤0.513363, the patient is considered to have a low risk of breast cancer recurrence; when the model prediction score is >0.513363, the patient is considered to have a high risk of breast cancer recurrence. Figure 4 .
[0162] from Figure 4 It can be seen that the AUC of the prediction model in the training group was 0.954, the sensitivity was 0.953, and the specificity was 0.965.
[0163] 2.3 Validation of the prediction model
[0164] The optimal model constructed was verified in the test group and the ROC curve was drawn as follows Figure 5 shown.
[0165] from Figure 5 It is known that the AUC of the prediction model in the test group is 0.932, the sensitivity is 0.941, and the specificity is 0.968, which are very close to the diagnostic effect in the training group.
[0166] In summary, the breast cancer recurrence risk prediction model constructed using 9 protein markers has good predictive performance and accuracy and the best diagnostic efficacy.
[0167] Example 3 Construction and verification of a three-category prediction model
[0168] This example attempts to construct a three-category joint prediction model for distinguishing the breast cancer prognosis non-recurrence group, the breast cancer prognosis in situ recurrence group, and the breast cancer prognosis metastasis group. The specific process includes the following: (1) construction and screening of the optimal prediction model; (2) validation of the effectiveness of the optimal prediction model. The specific screening process and results are as follows (in the present invention, the two-classification model in Example 2 uses the AUC value as the evaluation indicator; when a three-classification model is constructed, since multiple categories are involved, the AUC value is generally not applicable. In this example, the diagnostic efficacy of the model is measured using indicators such as sensitivity, specificity, accuracy, and consistency):
[0169] 3.1 Construction and screening of prediction models
[0170] For the test cohort of 300 breast cancer patients, all enrolled patients signed informed consent. Among them, 64 breast cancer patients relapsed within three years after surgery (30 cases were local recurrence in situ, and 34 cases had metastasis, of which those with both in situ recurrence and metastasis were also classified as metastasis). They were divided into two groups: 236 patient samples without recurrence were classified as low-risk group for breast cancer recurrence, and 64 patients with recurrence or metastasis were classified as high-risk group for breast cancer recurrence. They were randomly divided into training group and test group. The training group included 118 patients with low-risk group for breast cancer recurrence and 32 patients with high-risk group for breast cancer recurrence, including 15 patients with local recurrence in situ and 17 patients with metastasis; the test group included 118 patients with low-risk group for breast cancer recurrence and 32 patients with high-risk group for breast cancer recurrence, including 15 patients with local recurrence in situ and 17 patients with metastasis. This example aims to further construct a three-category detection model based on the marker combination of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1 screened in Example 2 to effectively distinguish between the breast cancer prognosis non-recurrence group (low-risk group), the breast cancer prognosis in situ recurrence group (in situ recurrence group), and the breast cancer prognosis metastasis group (metastasis group, in which both in situ recurrence and metastasis are also included in the metastasis group). All enrolled patients signed an informed consent form. Among them, breast cancer patients were patients diagnosed by pathological histology. Inclusion criteria: (a) no history of other malignant tumors; (b) no patients with other malignant tumors or autoimmune diseases.
[0171] In this embodiment, LC-MS / MS data acquisition and detection were performed on the collected serum samples to obtain the concentrations of nine protein markers, namely FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3 and RNASE1. The Shapiro Wilk test was used to evaluate the normal distribution, and the non-parametric Wilcoxon test was used to analyze the differences in blood marker concentrations between the breast cancer prognosis non-recurrence group (low-risk group), the breast cancer prognosis in situ recurrence group (in situ recurrence group) and the breast cancer prognosis distant metastasis group (metastasis group). A three-classification joint prediction model of 9 markers was constructed by combining machine learning methods. The predicted probability value was used to estimate the area under the receiver operator characteristic (ROC) curve (AUC) with a 95% confidence interval (CI) to evaluate the discriminative ability of the multivariate prediction model. Using the training group, the Youden index (YI) was calculated to determine the predicted probability cut-off value for distinguishing the low-risk group, the in situ recurrence group and the metastasis group. In addition, the ROC of single markers and different subgroups was constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV) and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R3.6.1, and a p value of less than 0.05 was considered statistically significant.
[0172] In this embodiment, in order to construct the optimal three-class joint prediction model, after comparing the six algorithms of gradient boosting, naive Bayes, support vector machine, neural network, generalized linear, and discriminant analysis, the gradient boosting method was selected as the best supervised classification algorithm for constructing the prediction model. The grid search range for hyperparameter optimization of the gradient boosting method model is shown in Table 7 below.
[0173] Table 7. Parameter grid search range of gradient boosting method
[0174]
[0175] Through optimization screening in terms of accuracy, consistency, sensitivity, specificity, etc., the optimal parameter combination mode was determined to be: interaction.depth 1, n.trees 100, shrinkage 0.1, n.minobsinnode 10.
[0176] The training and test groups used two completely different sample batches. Only the training group was used for marker screening and model construction; the test group samples were used solely to validate the diagnostic efficacy of the model. The specific results are shown in Table 8.
[0177] Table 8. Three-category diagnostic efficacy
[0178]
[0179] As shown in Table 8, the gradient boosting model constructed based on nine protein markers, FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, and RNASE1, can be used to predict whether breast cancer patients will experience no recurrence, in situ recurrence, or metastasis after surgery. This also demonstrates that the protein markers screened by the present invention can be used to distinguish the risk of recurrence after surgical treatment of breast cancer patients and, when the risk of recurrence is high, to distinguish between in situ recurrence and metastasis (patients with both in situ recurrence and metastasis are also included in the metastasis group).
[0180] 3.2 Combined Performance of the Three-Classification Joint Prediction Model
[0181] To further enhance the diagnostic value of three-category prediction models (gradient boosting) constructed using different protein combinations of biomarkers, this example compared the performance of prediction models constructed using different protein combinations of biomarkers in a training set based on the 10 protein markers screened in Example 1. The specific combinations of the different models are shown in Table 9.
[0182] Table 9. Combinations of different prediction models
[0183]
[0184] The results are shown in Table 10. Table 10 is a comparison of the performance indicators of different prediction models constructed by the 10 biomarkers screened in Example 1 for the three categories. The calculation method of the minimum value, first quartile, median, mean, third quartile, and maximum value of accuracy and consistency is as follows: (1) Sort the accuracy or consistency values from small to large; (2) Minimum value: the first value after sorting; (3) First quartile (Q1): multiply the number of data by 0.25. If the result is an integer, take the average of the values at this position and the next position; if it is not an integer, round up to get the position, and the value at this position is Q1; (4) Median: If the number of data is odd, the median is the middle value; if it is even, it is the average of the two middle values; (5) Mean: the sum of all values divided by the number of data; (6) Third quartile (Q3): multiply the number of data by 0.75 and process it in the same way as Q1; (7) Maximum value: the last value after sorting. The minimum and maximum values reflect data extremes, demonstrating the worst and best possible model performance. Quartiles help understand the data's distribution and dispersion. Q1 and below indicate lower performance, while Q3 and above indicate higher performance. The median reflects intermediate performance, and the mean comprehensively reflects the overall average performance. By combining these statistical values, we can gain a comprehensive understanding of the overall performance, distribution characteristics, and stability of the model, providing a strong basis for model selection and optimization.
[0185] Table 10. Performance comparison of prediction models based on different protein combination biomarkers
[0186]
[0187] As can be seen from Table 10, for the three-category prediction model, the ten-item joint detection model (10MP) composed of ten markers has the best performance. This also clearly shows that the constructed 10MP model has a very significant improvement in the diagnostic efficacy of distinguishing breast cancer patients with no recurrence, in situ recurrence, or metastasis after surgery. Therefore, the three-category gradient boosting model constructed using these ten protein markers (FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, KLHL22) is preferably used as the best joint prediction model.
[0188] 3.3. Diagnostic Performance Measurement and Validation of the Three-Classification Joint Prediction Model (10MP)
[0189] 3.3.1. Diagnostic Performance Measurement of the Three-Classification Joint Prediction Model (10MP)
[0190] In order to more accurately determine the diagnostic performance and threshold of the model constructed in this embodiment in different disease classifications, a multi-classification model of the gradient boosting (GBM) algorithm in the model group was used to perform predictive analysis in the training group, and the predicted results were calculated as the predicted probability values of the three categories (low-risk group, in situ recurrence group, and metastasis group). The category with the largest predicted probability value was the final prediction result of the system.
[0191] The meaning and calculation method of each indicator are as follows:
[0192] Results: The three-class combined prediction model (9MP) achieved an accuracy of 0.88 and a consistency of 0.89 in the training group. Its diagnostic sensitivity for the low-risk group was 92.5% and specificity was 95.9%. Its diagnostic sensitivity for the primary recurrence group was 85.8% and specificity was 86.3%. Its diagnostic sensitivity for the metastasis group was 86.2% and specificity was 85.6%.
[0193] It should be noted that the three-class joint prediction model (10MP) constructed by gradient boosting is a model built by machine learning and cannot fit a specific equation formula like a generalized linear model.
[0194] 3.3.2 Verification of the Three-Classification Joint Prediction Model (9MP)
[0195] The prediction performance of the model built based on the training group was verified in the test group. The specific results are as follows:
[0196] The accuracy was 0.88 and the consistency was 0.87. The diagnostic sensitivity for the low-risk group was 92.8% and the specificity was 94.8%. The diagnostic sensitivity for the primary recurrence group was 85.6% and the specificity was 86.8%. The diagnostic sensitivity for the metastasis group was 85.1% and the specificity was 84.5%.
[0197] In summary, the three-category joint prediction model containing 10 protein markers constructed in this example can not only accurately determine the risk of breast cancer recurrence, but also better predict the three categories of low-risk group, in situ recurrence group and metastasis group.
[0198] All patents and publications cited in this specification are intended to indicate that they are state of the art and that the present invention may be used. All patents and publications cited herein are incorporated by reference in their entirety, as if each publication were specifically incorporated by reference. The invention described herein may be practiced in the absence of any element or elements, limitation or limitations, unless otherwise specified. For example, in each instance, the terms "comprising," "consisting essentially of," and "consisting of" may be replaced with either of the other two terms. The term "a" or "an" herein simply means "one" and does not exclude the inclusion of only one or more. The terms and expressions used herein are intended to be descriptive, not limiting, and are not intended to exclude any equivalent features. However, it is understood that any suitable changes or modifications may be made within the scope of the present invention and the appended claims. It is understood that the embodiments described herein are preferred embodiments and features, and that modifications and variations can be made by persons of ordinary skill in the art based on the spirit of the present invention. Such modifications and variations are considered to be within the scope of the present invention and the scope of the independent and appended claims.
Claims
1. Use of a reagent for detecting a marker for preparing a reagent for predicting whether a breast cancer patient will not relapse, relapse in situ, or metastasis within three years after surgery, characterized in that: The markers consist of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.
2. A kit for predicting whether a breast cancer patient will not relapse, relapse in situ, or metastasize within three years after surgery, characterized in that: The kit comprises a detection reagent for the biomarker for use according to claim 1.
3. A biomarker combination for predicting whether a breast cancer patient will not relapse, relapse in situ, or metastasize within three years after surgery, characterized in that: The panel consists of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1, and KLHL22.
4. A system for predicting whether a breast cancer patient will not relapse, relapse in situ, or metastasize within three years after surgery, characterized in that: The system includes a data analysis module for analyzing the detection values of markers, wherein the markers consist of FBLN5, MUC16, ORM1, ADH1B, CD36, KRT19, DCD, DEFA3, RNASE1 and KLHL22.
5. The system according to claim 4, wherein: The data analysis module uses the detection values of markers of known samples as a training set, divides breast cancer patients into a non-recurrence group, an in situ recurrence group, and a distant metastasis group according to their post-operative conditions, analyzes the relationship between the detection values of the non-recurrence group, the in situ recurrence group, and the distant metastasis group, and constructs a model; the model is constructed based on a gradient boosting algorithm; the system also includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output prediction results.
Citation Information
Patent Citations
Biomarker for diagnosing breast cancer of to-be-detected person
CN116759070A
Biomarkers for breast cancer
US20090227692A1