Biomarker for predicting recurrence or metastasis risk of bladder cancer and application of biomarker

By screening and constructing a proteomic-based bladder cancer recurrence risk prediction model, using biomarkers such as IGFBP1 and machine learning algorithms, the problem of non-invasive high-sensitivity bladder cancer recurrence risk diagnosis is solved, and diagnostic accuracy and treatment effect are improved.

CN120254285AActive Publication Date: 2025-07-04HANGZHOU GUANGKE ANDE BIOTECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510725542.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The lack of high sensitivity non-invasive biomarkers in the prior art is used to predict the risk of recurrence or metastasis of bladder cancer, especially the diagnosis of recurrence risk of non-muscular invasive bladder cancer, resulting in high recurrence rates and poor treatment effects in patients after surgery.

Method used

Biomarkers such as IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1 and FBLN5 were screened through proteomics to construct a risk prediction model for bladder cancer recurrence or metastasis, and blood samples were analyzed using high-performance liquid chromatography-tandem mass spectrometry technology, and prediction was carried out in combination with machine learning algorithms.

Benefits of technology

It achieves non-invasive and accurate prediction of the risk of recurrence or metastasis of bladder cancer, improves the accuracy and sensitivity of diagnosis, reduces the rate of misdiagnosis, provides personalized treatment suggestions, and improves the treatment effect and quality of life of patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120254285A_ABST
    Figure CN120254285A_ABST
Patent Text Reader

Abstract

The invention provides a biomarker for predicting the recurrence or metastasis risk of bladder cancer and application of the biomarker. According to the application of the substance for detecting the biomarker in preparing the product for predicting the recurrence or metastasis risk of the bladder cancer, through a proteomics method, the significant difference of proteins in blood of patients with recurrence and non-recurrence of the bladder cancer is analyzed, the biomarker for predicting the recurrence risk of the bladder cancer is screened out, and the application of the biomarker in predicting the recurrence or metastasis risk of the bladder cancer is realized. And a bladder cancer recurrence risk prediction model is further constructed. The biomarker and the bladder cancer recurrence risk prediction model can realize simple, accurate and rapid diagnosis of the bladder cancer recurrence risk, improve the diagnosis level, and are suitable for large-scale popularization and application. The biomarker is derived from serum and has the advantages of convenience, instantaneity and non-invasiveness in sampling, so that the biomarker has an important clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical technologies, and particularly to a biomarker for predicting the risk of recurrence or metastasis of bladder cancer and its applications. Background Art

[0002] Bladder cancer is a common malignant tumor of the urinary system, with a high recurrence rate and risk of progression. Its development process usually starts from non-muscle invasive bladder cancer (NMIBC), and some patients may progress to muscle invasive bladder cancer (MIBC) or even distant metastasis. The development of bladder cancer involves multiple biological processes, including cell proliferation, apoptosis, invasion, and metastasis, etc., and these processes are regulated by multiple genes and proteins.

[0003] Currently, the diagnosis of bladder cancer mainly relies on urine cytology and cystoscopy, but these methods have certain limitations. In recent years, a variety of new biomarkers and detection methods have been proposed to improve the accuracy and sensitivity of diagnosis: Urine biomarkers: Such as BCLA-1, BCLA-4, AURKA, APE / Ref-1, etc. These biomarkers can be detected in urine and are not interfered by factors such as infection and smoking. For example, AURKA can be used to distinguish normal urine from low-grade bladder cancer. Multi-biomarker combination: Such as the detection of a combination of multiple urine biomarkers (such as collagen α-1 (I), uromodulin, etc.), which can improve the sensitivity and specificity of diagnosis. Gene detection: Such as Uromonitor-V2 for detecting KRAS hotspot mutations, which has high sensitivity and specificity. EpiCheck has high sensitivity for the recurrence monitoring of bladder cancer by detecting DNA methylation biomarkers in urine.

[0004] Early-stage bladder cancer, such as non-muscle invasive bladder cancer (NMIBC), generally has less distant metastasis because the tumor has not invaded the smooth muscle layer of the bladder, and the tumor can be removed surgically to achieve the purpose of cure. However, about 10%-30% of NMIBC patients still relapse after surgical treatment. If patients who do not benefit from surgical treatment or have progression can be subjected to risk prediction and the treatment plan can be adjusted in a timely manner (such as adjuvant radiotherapy and chemotherapy, secondary surgical resection, targeted therapy, immunotherapy, etc.), the overall survival rate and quality of life of patients can be significantly improved. There is a lack of biomarkers for the diagnosis of NMIBC recurrence risk in clinical practice, especially the discovery of high-sensitivity proteomic biomarkers for the diagnosis of NMIBC recurrence risk is of great significance.

[0005] Proteomics is the science that studies the protein composition, localization, changes, and their interaction rules in cells, tissues, or organisms, including the study of protein expression patterns and proteome function patterns. With the development of mass spectrometry technology, liquid chromatography-tandem mass spectrometry (LC-MS / MS) has become the most important tool in proteomics research. The development of proteomics is of great significance for finding disease diagnostic markers, screening drug targets, toxicology research, etc., and has therefore been widely applied in medical research.

[0006] Therefore, it is urgent to find a more effective and non-invasive biomarker for the diagnosis of the recurrence risk of bladder cancer. Summary of the Invention

[0007] Aiming at the problems existing in the prior art, the present invention provides a biomarker for predicting the recurrence or metastasis risk of bladder cancer and its application. By using proteomics methods, proteins with significantly different abundance levels in the blood of two groups of people with and without recurrence of bladder cancer are analyzed, and biomarkers for predicting the recurrence or metastasis risk of bladder cancer are screened out. Furthermore, a prediction model for the recurrence or metastasis risk of bladder cancer is constructed, which can accurately, non-invasively, and efficiently predict the recurrence or metastasis risk of bladder cancer and meet the clinical needs.

[0008] The first aspect of the present invention provides the use of substances for detecting biomarkers in the preparation of products for predicting the recurrence or metastasis risk of bladder cancer, and the biomarkers include one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5.

[0009] The biomarkers obtained by proteomics in the present invention can accurately predict the recurrence or metastasis risk of bladder cancer, which is beneficial for doctors to judge the severity of the patient's condition and whether treatment plans need to be adjusted, so as to provide more accurate treatment suggestions for patients and truly benefit bladder cancer patients.

[0010] The present invention uses proteomics methods to collect blood samples from patients with and without recurrence of bladder cancer, analyzes different samples by high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS), and based on orthogonal partial least squares discriminant analysis and significance analysis methods, first screens proteins with significant differences between patients with and without recurrence of bladder cancer, and 11 differential proteins with obvious correlation with the recurrence risk of bladder cancer are screened out; these 11 proteins can be used to distinguish whether bladder cancer patients will relapse after surgical treatment and have a certain diagnostic efficacy.

[0011] Among them, the IGFBP1 is a protein or amino acid sequence with the UniProt database number P08833; the SFTPB is a protein or amino acid sequence with the UniProt database number P07988; the MGAM2 is a protein or amino acid sequence with the UniProt database number Q2M2H8; the MUC16 is a protein or amino acid sequence with the UniProt database number Q8WXI7; the SELL is a protein or amino acid sequence with the UniProt database number P14151; the RNASE1 is a protein or amino acid sequence with the UniProt database number P07998; the S100B is a protein or amino acid sequence with the UniProt database number P04271; the TFF1 is a protein or amino acid sequence with the UniProt database number P04155; the CD74 is a protein or amino acid sequence with the UniProt database number P04233; the ORM1 is a protein or amino acid sequence with the UniProt database number P02763; the FBLN5 is a protein or amino acid sequence with the UniProt database number Q9UBX5.

[0012] The present invention also surprisingly discovers that some of the protein markers obtained through proteomic screening are known markers that can be used for other cancers. For example, SFTPB has been reported to be used for the prediction of lung cancer, and TFF1 has been reported to be used for the prediction of gastric cancer. However, this screening discovers that these markers can also be used for the prediction of the risk of recurrence or metastasis of bladder cancer. It can be seen that for many different tumors, their protein markers are not completely separated or irrelevant. In fact, there are many cross - relationships or influences. Many protein markers can be used for the prediction of early - stage cancer, and also for the prognostic diagnosis of cancer. Even many protein markers can be used for the prediction and diagnosis of different stages of multiple different cancers. Therefore, there are still many brand - new functions in the field of proteomics waiting to be explored, and the market prospect is very broad.

[0013] In some embodiments, the biomarker can be 1 marker or a combination of several markers, such as a combination of 2 markers, a combination of 3 markers, a combination of 4 markers, a combination of 5 markers, a combination of 6 markers, a combination of 7 markers, a combination of 8 markers, a combination of 9 markers, a combination of 10 markers.

[0014] In some specific embodiments, the biomarker comprises a combination of more than 2 markers, such as a combination of more than 3 markers.

[0015] In some more specific embodiments, the biomarker comprises a combination of 2 biomarkers, such as IGFBP1 + SFTPB (2MP). In some more specific embodiments, the biomarker comprises a combination of 3 biomarkers, such as IGFBP1 + SFTPB + MGAM2 (3MP). In some embodiments, the biomarker comprises a combination of 4 biomarkers, such as IGFBP1 + SFTPB + MGAM2 + MUC16 (4MP). In some embodiments, the biomarker comprises a combination of 5 biomarkers, such as IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL (5MP). In some embodiments, the biomarker comprises a combination of 6 biomarkers, such as IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 (6MP). In some embodiments, the biomarker comprises a combination of 7 biomarkers, such as IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B (7MP). In some embodiments, the biomarker comprises a combination of 8 biomarkers, such as IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B + TFF1 (8MP). In some embodiments, the biomarker comprises a combination of 9 biomarkers, such as IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B + TFF1 + ORM1 (9MP). In some embodiments, the biomarker comprises a combination of 10 biomarkers, such as IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B + TFF1 + ORM1 + FBLN5 (10MP). In some embodiments, the biomarker comprises a combination of 11 biomarkers, such as IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B + TFF1 + CD74 + ORM1 + FBLN5 (11MP). The combinations include but are not limited to the foregoing. When a risk prediction model for bladder cancer recurrence or metastasis is constructed using the combination of markers, the AUC value is 0.801 - 0.948, the sensitivity is 82.9% - 95.4%, and the specificity is 83.3% - 95.1%.

[0016] Furthermore, to study the diagnostic efficacy of the research department for the recurrence risk of bladder cancer, it is also necessary to combine different differential proteins to construct a diagnostic model, sort them according to the importance obtained by screening, and select different numbers of differential proteins with higher rankings for combination respectively to obtain 10 protein markers. After verification, it is found that constructing a model based on 8 protein markers has good risk prediction ability in the diagnosis of the recurrence risk of bladder cancer. The data of detecting samples of the recurrence risk of bladder cancer show that just using these 8 biomarkers to predict the recurrence risk of bladder cancer, the AUC value can reach 0.948, and the diagnostic performance is good.

[0017] In some embodiments, the product is selected from one or more of a reagent, a kit, a chip, a probe, or a membrane strip. The product is a product with the above-mentioned biomarker as the detection target, and includes biological reagents and kits suitable for the detection of the biomarker, such as sample pretreatment reagents, antigens, or antibodies; it can also be developed into a standardized reagent or kit, a chip, a probe, or a membrane strip suitable for the biomarker.

[0018] In some embodiments, the product is used for predicting patients at risk of recurrence or metastasis of bladder cancer.

[0019] Furthermore, the risk of recurrence or metastasis refers to recurrence or metastasis within three years after the treatment of bladder cancer.

[0020] The recurrence or metastasis of bladder cancer includes in-situ or adjacent area recurrence of bladder cancer and metastasis of bladder cancer.

[0021] Furthermore, the bladder cancer includes non-muscle-invasive bladder cancer.

[0022] In some embodiments, the product is used for detecting the detection amount of the biomarker in a biological sample.

[0023] In some specific embodiments, the biological sample is selected from one or more of saliva, blood, urine, plasma, serum, and spinal fluid.

[0024] In some specific embodiments, the detection amount includes the presence or absence, relative abundance, or concentration of the biomarker.

[0025] The present invention screens biomarkers for predicting the recurrence risk of bladder cancer from blood. These biomarkers have significant differences in the blood of the population with and without recurrence of bladder cancer. By collecting blood samples, it is possible to predict or assist in diagnosing whether an individual has recurrence or non-recurrence of bladder cancer, or to predict or assist in diagnosing whether an individual has no recurrence, in-situ recurrence, and distant recurrence after bladder cancer surgery by detecting these biomarkers in the individual's blood.

[0026] Furthermore, the detection methods generally include radiological methods, immunological methods, fluorescence methods, flow fluorescence methods, latex turbidimetry, biochemical methods, enzymatic methods, hybridization methods, gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), chromatography, chemiluminescence methods, magnetoelectric methods, or photoelectric conversion methods.

[0027] The presence, absence, or high or low content of the biomarker here is a relative concept. For example, when comparing the recurrence group and non-recurrence group of bladder cancer, the content of these specific biomarkers is compared with the recurrence group or non-recurrence group of bladder cancer as a reference. There may be some biomarkers with a relatively higher content in bladder cancer recurrence compared to the non-recurrence group of bladder cancer, and this increase is statistically significant, such as a significant or highly significant increase. Therefore, when judging these biomarkers, if it is a single biomarker, if the probability of a certain risk occurrence increases and the content of the biomarker changes, this change may be a relative increase or may also be a relative decrease, and the difference in this relative increase or decrease is significantly different, and of course, it can also be highly significantly different. Therefore, no matter what means are used for detection, a pre-specified value (cut-off value) can be used as a standard. If it is higher than this value, it is considered that the content has changed, and such a result can be used for prediction or diagnostic value.

[0028] Therefore, in some aspects, the biomarker of the present invention can be obtained by detecting the biomarker content in a sample by any known method, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strips, immunoreaction chips, capillary electrophoresis, infrared spectroscopy, etc. As long as it can be used to detect the protein biomarker content in a sample, it can be used for the diagnosis of bladder cancer recurrence and non-recurrence. As long as the protein biomarker content in the sample can be detected, it can be used to predict or diagnose the probability of occurrence of a certain disease. It can be understood that the detection here is for the individual sample, and then compared with the pre-set standard, and the comparison result is used to judge or predict the occurrence status of the disease. For example, it can be used to predict the probability of bladder cancer recurrence risk. Such prediction or diagnosis is whether it will occur within a certain time. Of course, such detection can be continuous detection, and the progress of the disease can be inferred as the content of certain substances changes.

[0029] In some ways, the relative abundance is the peak area of the biomarker in the detection map obtained by high performance liquid chromatography - tandem mass spectrometry. For example, if the average peak area of a certain biomarker measured in the control sample is 300 and the average peak area measured in the bladder cancer recurrence sample is 1800, then the abundance of this biomarker in the sample is considered to be 6 times that of the control sample.

[0030] The second aspect of the present invention provides the use of a protein as a biomarker in the preparation of a product for predicting the risk of bladder cancer recurrence or metastasis. The protein is selected from one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5. The genes of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5 can be used for the auxiliary judgment of the risk of bladder cancer recurrence or metastasis, the evaluation of drug efficacy, etc. The inventors of the present invention have found that IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, ORM1, and FBLN5 are closely related to the risk of bladder cancer recurrence or metastasis.

[0031] The third aspect of the present invention provides a product for predicting the risk of bladder cancer recurrence or metastasis, including a substance for detecting a biomarker, and the biomarker includes one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5.

[0032] The fourth aspect of the present invention provides a biomarker combination for predicting the risk of bladder cancer recurrence or metastasis, including IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1.

[0033] The fifth aspect of the present invention provides a method for constructing a prediction model for the risk of bladder cancer recurrence or metastasis for non - disease diagnosis purposes, including the following:

[0034] 1) Construct a sample data set based on the detection amount of biomarkers in biological samples, and the biomarkers include one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5;

[0035] 2) Divide the data set into a test set and a training set, and construct and train the prediction model for the risk of bladder cancer recurrence or metastasis through machine learning methods.

[0036] In some embodiments, the biomarkers are IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1.

[0037] In some embodiments, the machine learning method is selected from at least one of gradient boosting algorithm, random forest algorithm, support vector machine algorithm, decision tree algorithm, K-nearest neighbor algorithm, logistic regression algorithm, and neural network algorithm. For example, it can be selected from the gradient boosting algorithm. Different from generalized linear models that can output model formulas and cut-off values, all calculations in the gradient boosting algorithm are directly completed by machine learning. By directly inputting the detection values into the software system, the prediction results can be directly obtained.

[0038] In some embodiments, multiple machine learning methods are used to construct a prediction model for the risk of bladder cancer recurrence or metastasis. It is preliminarily confirmed that for any one of the newly identified biomarkers alone, the change in its concentration can be used to distinguish between the population with bladder cancer recurrence and the non-recurrence population, indicating that these biomarkers have extremely high diagnostic value.

[0039] The present invention discovers that by detecting the detected amount of biomarkers in a biological sample, then inputting the detected amount into the formula of the prediction model for the risk of bladder cancer recurrence or metastasis to obtain the prediction score Y of the prediction model, and comparing this prediction score Y with the threshold (cut off) defined by the Youden index. If the prediction score Y > the threshold (cut off), it is determined that the patient has bladder cancer recurrence; if the prediction score Y ≤ the threshold (cut off), it is determined that the patient does not have bladder cancer recurrence.

[0040] In some embodiments, the equation of the prediction model for the risk of bladder cancer recurrence is:

[0041]

[0042] Where Y is the prediction score, i represents the i-th biomarker, m represents the number of biomarkers (m = 8), Xi represents the detected value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker (see Table 6), and b is a constant 3.554513.

[0043] Table 6 Coefficients of 8 biomarkers in the model

[0044]

[0045] When Y ≤ the cut-off value, the subject is determined not to have bladder cancer recurrence; when Y > the cut-off value, the subject is determined to have bladder cancer recurrence. Specifically, the cut-off value is 0.4720534.

[0046] In some embodiments, it further includes step 3) testing the bladder cancer recurrence or metastasis risk prediction model using a test set. After training is completed, the trained prediction model is verified using the test set. Meanwhile, the AUC value, specificity, and sensitivity are used as evaluation indicators to evaluate the effect of the prediction model.

[0047] The Receiver Operating Characteristic Curve (ROC curve) is a curve plotted with the true positive rate (sensitivity) as the ordinate and the false positive rate (1 - specificity) as the abscissa according to a series of different binary classification methods (cutoff values). The Area Under Curve is defined as the area under the ROC curve. The AUC value is often used to evaluate the diagnostic efficacy of a prediction model. The larger the AUC value, the better the diagnostic efficacy of the corresponding prediction model; conversely, the worse the diagnostic efficacy of the corresponding prediction model.

[0048] The prediction model constructed based on the combination of 8MP in the present invention can distinguish between recurrence and non - recurrence after bladder cancer surgery. Its AUC in the training group is 0.948, the sensitivity is 0.914, and the specificity is 0.878; the AUC in the test group is 0.928, the sensitivity is 0.886, and the specificity is 0.830.

[0049] The sixth aspect of the present invention provides a device for predicting the risk of bladder cancer recurrence or metastasis, including a data acquisition unit and a calculation unit;

[0050] The data acquisition unit is used to obtain the detection amount data of the biomarkers in the biological sample of the subject as described above. The biomarkers include one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5;

[0051] The calculation unit is used to calculate and output a prediction score for the risk of bladder cancer recurrence or metastasis of the subject based on the detection amount data and make a judgment according to a threshold.

[0052] In some embodiments, the biomarkers include IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1.

[0053] In some embodiments, the prediction device further includes a data storage unit and a data output unit; the data storage unit is used to store the detected amount (or detection value) of the biomarker; the data input interface is used to input the detection value of the biomarker, and the data output unit is used to output the prediction result.

[0054] Further, the detection value is the presence or absence, relative abundance or concentration value of each biomarker.

[0055] The seventh aspect of the present invention provides a device, including a processor and a memory, the memory is used to store a computer program, and it is characterized in that the processor is used to execute the computer program stored in the memory, so that the device executes the prediction method or the construction method as described above.

[0056] The eighth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is processed, it executes the prediction method or the construction method as described above.

[0057] In some embodiments, the computer-readable storage medium includes: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs.

[0058] In some embodiments, the being processed means being executed by one or more processors.

[0059] The ninth aspect of the present invention provides a method for predicting the risk of bladder cancer recurrence or metastasis, including the following steps:

[0060] S1. Obtain the detected amount data of the biomarker in the biological sample of the subject, and the biomarker includes one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1 and FBLN5;

[0061] S2. Process the detected amount data by using the bladder cancer recurrence risk prediction model obtained by the construction method as described above to output the prediction result of the risk of bladder cancer recurrence or metastasis.

[0062] In some embodiments, the biomarker includes IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B and TFF1.

[0063] In another aspect, the present invention provides a system for predicting the recurrence risk of bladder cancer. The system includes a data analysis module, which is used to analyze the detection values of biomarkers, and the biomarkers include IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1.

[0064] Further, the data analysis module uses the detection values of biomarkers of known samples as a training set, and divides them into a bladder cancer recurrence group and a non-recurrence group according to whether bladder cancer recurs, analyzes the relationship between the detection values of the bladder cancer recurrence group and the non-recurrence group, and constructs a model.

[0065] In some embodiments, a combined diagnostic model for predicting the recurrence risk of bladder cancer is constructed by combining multiple machine learning methods, and it is preliminarily confirmed that the change in the concentration of any one of the selected novel biomarkers alone can be used to distinguish high-risk and low-risk populations of bladder cancer recurrence, indicating that these biomarkers have extremely high diagnostic value.

[0066] In some embodiments, the equation of the constructed model is:

[0067]

[0068] Where Y is the predicted score, i represents the i-th biomarker, m represents the number of biomarkers (m = 8), Xi represents the detection value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker, and b is a constant 3.554513; the coefficients of the 8 biomarkers are:

[0069]

[0070] When Y ≤ the cut-off value, the tested person has no recurrence of bladder cancer; when Y > the cut-off value, the tested person has recurrence of bladder cancer. Specifically, the cut-off value is 0.4720534.

[0071] Further, the system also includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction result.

[0072] In another aspect, the present invention provides the use of a biomarker for preparing a reagent for predicting non-recurrence, in-situ recurrence or distant metastasis in patients with bladder cancer after surgery. The biomarker includes one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5;

[0073] In another aspect, the present invention provides a kit for predicting non - recurrence, in - situ recurrence or distant metastasis in bladder cancer patients after surgery, and the kit includes detection reagents for the biomarkers as described above for the uses.

[0074] In another aspect, the present invention provides a biomarker combination for predicting non - recurrence, in - situ recurrence or distant metastasis in bladder cancer patients after surgery, and the combination includes any one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1 and FBLN5.

[0075] In another aspect, the present invention provides a system for predicting non - recurrence, in - situ recurrence or distant metastasis after bladder cancer surgery, and the system includes a data analysis module for analyzing the detection values of biomarkers, and the biomarkers include any one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1 and FBLN5.

[0076] Furthermore, the distant metastasis also includes simultaneous in - situ recurrence and distant metastasis, and as long as distant metastasis occurs, it is classified into the distant metastasis group.

[0077] Furthermore, the data analysis module uses the detection values of biomarkers of known samples as a training set, and according to the situation of liver cancer patients after surgery, it is divided into a non - recurrence group, an in - situ recurrence group and a distant metastasis group, analyzes the relationship between the detection values of the non - recurrence group, the in - situ recurrence group and the distant metastasis group, and constructs a model.

[0078] Furthermore, the model is constructed based on the gradient boosting algorithm.

[0079] The gradient boosting algorithm does not output a model formula and a cut - off value like generalized linear models. All calculations are directly completed by machine learning. The detection values can be directly input into the software system, and the prediction results can be directly obtained.

[0080] Furthermore, the system further includes a data storage module, a data input interface and a data output interface; the data storage module is used for storing the detection values of biomarkers; the data input interface is used for inputting the detection values of biomarkers, and the data output interface is used for outputting the prediction results.

[0081] The beneficial effects of the present invention are:

[0082] 1. The present invention has screened 11 brand-new biomarkers that can predict the risk of bladder cancer recurrence or metastasis, developed a new combination of protein biomarkers, which can effectively evaluate and diagnose the risk of bladder cancer recurrence or metastasis, effectively distinguish between recurrent and non-recurrent bladder cancer, and effectively distinguish in-situ recurrence and distant metastasis of bladder cancer. This not only improves the diagnostic accuracy of bladder cancer recurrence and metastasis, but also provides an important biomarker basis for personalized treatment and prognostic monitoring of bladder cancer patients. In addition, the discovery and application of these biomarkers are expected to improve the treatment effect and quality of life of bladder cancer patients, and provide a scientific basis for the early diagnosis and treatment strategy selection of bladder cancer. Compared with traditional detection methods, it reduces the risk of misdiagnosis and missed diagnosis, and provides strong support for the early detection and intervention of the disease.

[0083] 2. The combined differential diagnosis model of 8 biomarkers constructed by the present invention is convenient, fast, and the detection results are highly consistent with the clinical gold standard detection results. At the same time, it significantly reduces the cost of predicting the risk of bladder cancer recurrence or metastasis, and has good application prospects. The combination of biomarkers of the present invention is superior to the diagnosis of broad-spectrum tumor markers or broad-spectrum tumor marker combination models.

[0084] 3. Based on the biomarkers screened for predicting the risk of bladder cancer recurrence or metastasis, a three-classification model that can simultaneously distinguish between non-recurrent group, in-situ recurrence group and distant metastasis group is further constructed, providing a more effective and accurate prediction and diagnosis mode.

[0085] Detailed description

[0086] (1) Diagnosis or detection

[0087] Here, the diagnosis or detection refers to the detection or assay of biomarkers in a sample, or the content of the target biomarker, such as absolute content or relative content, and then it is explained whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease, based on the presence or quantity of the target biomarker. The meanings of diagnosis and detection here can be interchanged. The result of this detection or the result of the diagnosis cannot be directly used as the direct result of having a disease, but is an intermediate result. If a direct result is to be obtained, other auxiliary means such as pathology or anatomy are required to confirm the presence of a certain disease. For example, the present invention provides a variety of new biomarkers related to the risk of bladder cancer recurrence or metastasis, and the change in the content of these biomarkers is directly related to whether the patient belongs to the population at risk of bladder cancer recurrence or metastasis.

[0088] (2) The connection between biomarkers or biological markers or differential proteins and the risk of bladder cancer recurrence or metastasis

[0089] The terms "marker", "biomarker" and "differential protein" have the same meaning in the present invention. The connection here refers to the direct relevance between the appearance or content change of a certain biomarker in a sample and a specific disease. For example, a relative increase or decrease in content indicates a relatively higher likelihood of having this disease compared to the healthy population.

[0090] If multiple different markers in a sample appear simultaneously or show relative changes in content, it also indicates a relatively higher likelihood of having this disease compared to the healthy population. That is to say, among the types of markers, some markers have a strong correlation with the disease, some have a weak correlation, and some may even have no correlation with a specific disease. One or more of those markers with a strong correlation can be used as markers for diagnosing the disease, and those with a weak correlation can be combined with the strong ones to diagnose a certain disease, increasing the accuracy of the test results.

[0091] Regarding the numerous biomarkers found in the serum in the present invention, these biomarkers can all be used to distinguish between recurrent and non-recurrent bladder cancer, as well as non-recurrent, in-situ recurrence and distant metastasis. These markers can be directly detected or diagnosed as individual markers alone. Selecting such a marker indicates that the relative change in the content of this marker has a strong correlation with the risk of bladder cancer recurrence or metastasis. Of course, it can be understood that simultaneous detection of one or more markers with a strong correlation with the risk of bladder cancer recurrence or metastasis can be selected. Normally understood, in some ways, selecting biomarkers with a strong correlation for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90% or 95% accuracy. This means that these markers can obtain an intermediate value for diagnosing a certain disease, but it does not mean that it can directly confirm the presence of a certain disease.

[0092] Of course, differential proteins with a larger ROC value can also be selected as diagnostic markers. The so-called strength or weakness is generally calculated and confirmed through some algorithms, such as the contribution rate or weight analysis of the marker to the prediction of the risk of bladder cancer recurrence or metastasis. Such calculation methods can be significance analysis (p-value or FDR value) and fold change, and multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA) and orthogonal partial least squares discriminant analysis (OPLS-DA). Of course, other methods are also included, such as ROC analysis, etc. Of course, other model prediction methods are also possible. When specifically selecting biomarkers, the differential proteins disclosed in the present invention can be selected, or other existing well-known biomarker combinations can be selected or combined for prediction through model methods.

[0093] (3) Definition of disease terms

[0094] Bladder cancer: It refers to a malignant tumor that occurs on the bladder mucosa. It is the most common malignant tumor in the urinary system and one of the top ten common tumors in the whole body. It ranks first in the incidence of urogenital tumors in China, and in the West, its incidence is second only to prostate cancer, ranking second. In 2012, the incidence of bladder cancer in the national cancer registration areas was 6.61 / 100,000, ranking ninth in the incidence of malignant tumors. Bladder cancer can occur at any age, even in children. Its incidence increases with age, and the high-incidence age is 50-70 years old. The incidence of bladder cancer in men is 3-4 times that of women.

[0095] Non-muscle-invasive bladder cancer: It originates from the bladder mucosa layer and is limited within the mucosa and the lamina propria, and has not invaded the muscle layer of the bladder. The pathological types are mainly bladder urothelial cell carcinoma, squamous cell carcinoma, adenocarcinoma, etc. The clinical manifestation is mainly painless hematuria.

[0096] Recurrence of bladder cancer: It refers to the reappearance of bladder cancer or its spread to other parts of the body after treatment. On the one hand, the recurrent bladder cancer may occur in the primary site or near the bladder, which is called local recurrence (in-situ recurrence). Local recurrence means that the cancer has not spread widely but reappears in the bladder or its directly adjacent tissues. On the other hand, bladder cancer may also spread to other distant parts of the body, and this situation is called distant recurrence or metastasis. Most recurrences of bladder cancer (about 80%) occur within 2-3 years after surgery, and the risk of recurrence is significantly reduced after 5 years. The recurrence of bladder cancer is closely related to the tumor stage, the standardization of treatment, and postoperative management. The prognosis of bladder cancer (prognosis refers to the expected survival situation after disease treatment) is not good because bladder cancer is an aggressive malignant tumor with non-obvious early symptoms and is only discovered at an advanced stage, which leads to relatively poor treatment effects and prognosis.

[0097] Metastasis of bladder cancer: Distant metastasis refers to the spread of tumor cells to parts outside the bladder through the blood or lymphatic system, such as the lungs, liver, bones, etc. Among the patients with non-muscle-invasive bladder cancer after radical cystectomy, about 10-30% will experience recurrence or metastasis. The risk of distant metastasis is closely related to the tumor stage and grade, and high-grade and advanced tumors are more likely to have distant metastasis.

[0098] Therefore, the early detection and intervention of bladder cancer recurrence or metastasis are crucial for improving the treatment effect and increasing the survival rate. The biomarker of the present invention has an accurate and specific prediction effect on the recurrence or metastasis of bladder cancer, which is beneficial to improving the treatment effect of patients and has important significance for improving the prognosis of patients.

[0099] (4) Gold standard for bladder cancer diagnosis: pathological diagnosis (i.e., pathological confirmation of needle biopsy or surgical resection specimens), observing the presence of cancer cells under a microscope, and clarifying the nature of the tumor by combining immunohistochemistry or molecular detection. Description of the Drawings

[0100] Figure 1 It is a volcano plot for the differential analysis of protein markers in the high-risk and low-risk groups of bladder cancer recurrence in Example 1.

[0101] Figure 2 It is a graph of the ROC and OPLS-DA analysis results in the high-risk and low-risk groups of bladder cancer recurrence in Example 1.

[0102] Figure 3 AUC result graph of models constructed with different hyperparameters in Example 2.

[0103] Figure 4 ROC curve graph of the bladder cancer recurrence risk prediction model in the training group in Example 2.

[0104] Figure 5 ROC curve graph of the bladder cancer recurrence risk prediction model in the test group in Example 2. Detailed Implementation Modes

[0105] The present invention will be further described in detail below with reference to the drawings and examples. It should be noted that the following examples are intended to facilitate the understanding of the present invention and do not limit it in any way. The reagents used in this example are all known products and are obtained by purchasing commercially available products.

[0106] Example 1 Screening Biomarkers for Bladder Cancer Recurrence Risk Using Proteomics

[0107] By collecting plasma samples from bladder cancer patients after radical surgery, enriching low-abundance proteins based on the method of removing high-abundance proteins by immunoaffinity chromatography, detecting the protein abundance in the samples by high-performance liquid chromatography-tandem mass spectrometry equipment, and screening for differential proteins by analyzing the differences in their abundances between bladder cancer recurrence and non-recurrence patients. The specific steps are as follows:

[0108] 1.1 Sample Collection

[0109] A total of 134 blood samples were collected from non-muscle-invasive bladder cancer patients (Ta / T1 stage patients) 2 weeks after surgical resection. All bladder cancer patients were confirmed by pathology of the living tissue. Approximately 2 ml of peripheral blood samples from the subjects were collected, placed in a vacuum tube containing EDTA anticoagulant, mixed well, centrifuged at 120 g for 10 minutes at room temperature, and the supernatant was taken, repeated twice. Then centrifuged at 360 g for 20 minutes. Then the platelet samples were collected in centrifuge tubes and stored at -80 °C for later use. The 132 patients included 94 males and 38 females, with an average age of 47 years (25 - 69 years). All enrolled patients signed the informed consent form. Among them, all bladder cancers were patients diagnosed by pathological histology. Inclusion criteria: (a) No history of other malignancies; (b) No patients with other malignancies or autoimmune diseases combined.

[0110] During the following three years, the patients were followed up every 6 months, with regular CYFRA 21-1 and imaging examinations. Once recurrence or metastasis was found, it needed to be confirmed by pathological histology. A total of 32 patients had recurrence within three years after surgery, including 14 cases of local in-situ recurrence and 18 cases of distant metastasis (including 4 cases of liver metastasis, 8 cases of lung metastasis, 4 cases of bone metastasis, and 2 cases of pelvic lymph node metastasis). Then the samples stored at -80 °C were taken out, and the samples were divided into two groups: the samples of 102 patients without recurrence were classified into the low-risk group of bladder cancer recurrence, and 32 cases with recurrence or metastasis were classified into the high-risk group of bladder cancer recurrence, for the screening of proteomic biomarkers.

[0111] 1.2 Sample processing and enzymatic digestion

[0112] 1) First, the plasma samples were centrifuged on a centrifuge for 15 minutes (15,000 g), and the supernatant was taken, filtered, and then 14 high-abundance proteins were removed by immunoaffinity chromatography.

[0113] 2) Then, the low-abundance was concentrated to 350 μL using a concentrator tube with a cut-off molecular weight of 3 kDa on a centrifuge (4000 g, 1 hour).

[0114] 3) The concentrated solution was recovered, and buffer exchange was carried out using a desalting column with a cut-off molecular weight of 7 kDa on a centrifuge (1000 g, 2 minutes). The replacement solution was AEX-A (20 mM Tris, 4 M urea, 3% isopropanol, pH 8.0).

[0115] 4) Using AEX-A as the blank, the protein concentration in the samples was determined by the Bicinchoninic Acid (BCA) method. According to the sample grouping, 25 μL of TCEP was added to the samples, and the samples were incubated at 37 °C for 30 minutes for protein reduction. Then, the corresponding TMT16-plex reagent was added, and the samples were incubated in the dark at room temperature for 1 hour for the TMT labeling reaction. Subsequently, the samples were buffer-exchanged using Zeba columns, and the exchange buffer was AEX-A. After mixing the TMT 16-plex-labeled samples, 2 mL of AEX-A was added to the mixed samples, and the final volume was 5.5 mL.

[0116] 5) The samples were filtered using a 0.22 m filter and the TMT 16-plex-labeled samples were separated using a 2D-HPLC system. The collected fractions were freeze-dried, and finally, a mixture of Trypsin-Lysin C enzymes was added. The samples were incubated at 37 °C for 5 hours for enzymatic digestion, and 5 μL of 10% TFA was added to terminate the enzymatic digestion reaction.

[0117] 6) A total of 60 enzymatically digested 2D-HPLC fractions were used for nanoLC-MS / MS analysis.

[0118] 1.3 LC-MS / MS Data Acquisition

[0119] Each sample obtained in Step 1.2 was separated using a nanoflow liquid chromatography system, Easy nLC-1200, and detected online by coupling with a high-resolution mass spectrometer, Q Exactive HF-X. The details are as follows:

[0120] 1) Separation: Mobile phase - A was an aqueous solution of 0.1% formic acid, and mobile phase - B was an aqueous solution of 0.1% formic acid in acetonitrile (80% acetonitrile and 20% water). The chromatographic column consisted of a trapping column and an analytical column and was equilibrated with 100% mobile phase - A. The samples were loaded onto the trapping column (100 μm ID × 4 cm L, C18, 3 μm, 100 Å) by an autosampler and separated by the analytical column (75 μm ID × 25 cm L, C18, 3 μm, 100 Å) at a flow rate of 300 nL / min.

[0121] 2) The sample after chromatographic separation was subjected to mass spectrometry analysis using a Q Exactive HF-X mass spectrometer. The detection mode was positive ion, the parent ion scanning range was 350 - 1800 m / z, the resolution of the first-order mass spectrometry was 120,000 at 200 m / z, the AGC (Automatic gain control) target was 3e6, the Maximum IT was 50 ms, and the dynamic exclusion time was 40 s. The mass-to-charge ratios of polypeptides and polypeptide fragments were collected according to the data-dependent acquisition (DDA) method: 20 secondary spectra (MS / MS, MS2 scan) were collected after each full scan (first-order mass spectrometry). The MS2 ActivationType was HCD, the Isolation window was 0.7 m / z, the resolution of the second-order mass spectrometry was 30,000 at 200 m / z, the AGC target was 1e5, the Maximum IT was 65 ms, the Fixed first mass was 110.0 m / z, the Normalized Collision Energy was 32 ev, the Minimum AGC target was 2.00e4, the Charge exclusion was 1, 6 - 8, >8, the Multiple charge states was one charge state only, the Peptide match was preferred, and the Exclude isotopes was on.

[0122] 1.4 Data preprocessing

[0123] The MS / MS data was searched using Maxquant (v1.6.15.0). The data type was DIA proteomics data based on MS / MS reporter ion quantification. For the MS / MS spectra used for quantification, it was required that the proportion of precursor ions in the MS spectra was greater than 75%. The database source was the Homo_sapiens_9606_proteome_gene in the Uniprot database (release: 2021-10-14, sequence: 20,437), and a common contaminant library was added to the database. Contaminant proteins were removed during data analysis. The digestion method was set to Trypsin / P; the maximum number of missed cleavage sites was set to 2; the mass error tolerances for precursor ions in the First search and Main search were set to 20 ppm and 5 ppm respectively, and the mass error tolerance for MS / MS fragment ions was 20 ppm. The fixed modification was cysteine alkylation, and the variable modifications were methionine oxidation and protein N-terminal acetylation. The false discovery rate (FDR) for protein identification and PSM identification was set to 1%.

[0124] 1.5 Differential analysis

[0125] Univariate analysis and multivariate statistical analysis were combined to screen for differential proteins and transcripts. Univariate analysis mainly included the significance analysis (p-value or FDR value) and fold change of characteristic molecules in different groups, and multivariate statistical analysis mainly included receiver operating characteristic curve (ROC) analysis and Boruta feature selection based on the random forest algorithm. All statistical analyses were completed using R. The specific R-related information is shown in Table 1.

[0126] Table 1. R and its related information

[0127]

[0128] The variable importance for the projection (VIP) was calculated to measure the influence intensity and interpretability of the expression patterns of each protein on the classification and discrimination of each group of samples. Further, the Wilcoxon rank-sum test was performed to obtain the corrected p-value (FDR). According to the conditions of FDR < 0.01 and Fold change > 2, 70 downregulated proteins and 72 upregulated proteins were screened (see details in Figure 1 ).

[0129] To evaluate the role of each protein biomarker in the diagnosis and prediction of the recurrence risk of bladder cancer, the ROC and Boruta analysis methods were used to evaluate each protein biomarker. The results are shown in Figure 2The abscissa is the AUC obtained from the ROC analysis, the ordinate is the -log10 (FDR) calculated by the Wilcoxon test, and the size of the points represents the VIP value obtained from the Boruta analysis. Further screening was carried out according to VIP > 3 and AUC > 0.6, and a total of 11 more significant candidate protein markers were found, as shown in Table 2 for details.

[0130] Table 2. Differential markers for the recurrence risk of bladder cancer

[0131]

[0132] Among them, the smaller the FDR value and / or the larger the VIP value, to a certain extent, it indicates that the difference in the protein between the high-risk and low-risk bladder cancer recurrences is more significant, and at the same time, it also indicates that the protein may have higher diagnostic value.

[0133] Example 2. Construction and validation of a prediction model for the recurrence risk of bladder cancer

[0134] In this example, based on the 11 protein markers screened in Example 1, a combination was made, and a recurrence risk prediction model was constructed for research.

[0135] Although a single biomarker can also distinguish the recurrence risk of bladder cancer patients after surgery, generally speaking, combining multiple biomarkers has higher accuracy in discrimination or prediction. However, for a single biomarker with higher accuracy in predicting the recurrence risk of bladder cancer, its role in the combination with one or more other biomarkers may not necessarily be greater, and at the same time, it is not the case that the more the number of biomarkers, the higher the prediction accuracy (AUC value) of the combination. Therefore, a large number of verification experiments are still needed.

[0136] 2.1 Obtaining data

[0137] A total of 240 blood samples were collected from non-muscle-invasive bladder cancer patients (Ta / T1 stage patients) 2 weeks after surgical treatment. All enrolled patients signed informed consent forms. Among them, 42 patients had recurrence within three years after surgery (12 cases had local in-situ recurrence and 30 cases had metastasis). Thus, they were divided into two groups: 198 samples from non-recurrent patients were classified into the low-risk group for bladder cancer recurrence, and 42 samples from patients with recurrence or metastasis were classified into the high-risk group for bladder cancer recurrence. They were randomly divided into a test group and a validation group. The test group included 99 samples from the low-risk group for bladder cancer recurrence and 21 samples from the high-risk group for bladder cancer recurrence. The validation group included 99 samples from the low-risk group for bladder cancer recurrence and 21 samples from the high-risk group for bladder cancer recurrence.

[0138] The concentrations of the 11 protein markers in Table 2 in the serum were obtained from the 240 samples using the same steps as in steps 1.2 and 1.3 in Example 1.

[0139] 2.2 Data statistical analysis

[0140] In the training group, a combined diagnostic model of multiple bladder cancer recurrence risk markers was constructed by combining multiple machine learning methods. The area under the receiver operating characteristic (ROC) curve (AUC) was estimated using the predicted probability value with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate diagnostic model.

[0141] Using the test group, the Youden index (YI) was calculated to determine the cut-off value for predicting the probability of distinguishing between the high-risk group and low-risk group of bladder cancer recurrence. In addition, the ROCs of the models formed by combining different markers were constructed and compared. Standard descriptive statistical data such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R 3.6.1, and a p-value less than 0.05 was considered statistically significant.

[0142] 2.3 Construction of the prediction model

[0143] The steps for constructing the bladder cancer recurrence risk prediction model are as follows:

[0144] S101. From the 11 protein markers in Table 2 in the samples of the training group, randomly select the concentration matrix of 2 to 11 markers as the original training data set.

[0145] S102. Select the generalized linear model (glmnet) algorithm for constructing the prediction model and the grid search range during the hyperparameter optimization process of the algorithm. The grid search range for hyperparameter optimization of each algorithm is shown in Table 3.

[0146] Table 3 Parameter grid of glmnet algorithm

[0147]

[0148] S103. According to the algorithm and hyperparameter setting range set in step S102, select one combination of hyperparameters as the parameters for constructing the prediction model.

[0149] S104. Split the original data set into K subsets according to the K-fold cross-validation mechanism. To ensure that the proportion of majority-class samples and minority-class samples in each subset is the same as that in the original data set, the stratified K-fold cross-validation mechanism needs to be used for data splitting.

[0150] S105. From the K training data subsets obtained by splitting according to step S104, select one subset as the validation set Ddev.

[0151] S106. Combine the training data subsets not selected in step S105 to form the training data pool Dtrainl.

[0152] S107. Based on the training dataset Dtrain obtained in step S106, construct a prediction model based on the selected supervised classification algorithm and hyperparameters.

[0153] S108. According to the prediction model obtained in step S107, evaluate it on the validation set Ddev to obtain the AUC value, and store the current prognosis prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 is to evaluate according to the prediction model obtained in step S107 on the validation set determined in the current iteration, and store both the model and the evaluation result in the prediction model pool for later use in selecting the base prediction model. The evaluation mentioned in this step can be the AUC value or other reasonable metrics for evaluating the model performance.

[0154] S109. Determine whether each subset has been used as the validation set. Step S109 is to determine whether the K subsets obtained in step S104 have all been used as the validation set for model training. If all subsets have been used as the validation set and the training is completed, execute step S110; if there is a subset that has not been used as the validation set, execute step S105. This step ensures that each sample in the original dataset has been used as the validation set, improves the model stability, and prevents the model from overfitting to a certain subset.

[0155] S110. Take the average AUC value of all models in the obtained prediction model pool Pool as the final performance evaluation value of the model in this combination method. And store the model parameters and the final performance evaluation AUC value in the optimal model pool Poolbest.

[0156] S111. Determine whether prediction models have been constructed for each hyperparameter combination method. Step S111 is to determine whether prediction models have been constructed for all the algorithms and the corresponding hyperparameter combination methods obtained in step S102. If all combination methods have completed the model construction, execute step S112; if there is a combination method that has not completed the model construction, execute step S103.

[0157] S112. After completing step S111, it is necessary to check whether all hyperparameter combination methods have been used to construct the prediction model. If it is found that there are still hyperparameter combinations for which the model has not been constructed, then it is necessary to return to step S103 to continue constructing the remaining models. If models have been constructed for all hyperparameter combinations, then step S113 can be continued to select the best model from the model pool.

[0158] S113. From the model set Poolbest obtained in step S112, select the model with the largest AUC value as the final prediction model for the risk of bladder cancer recurrence or metastasis.

[0159] S114. Repeat all the above steps until modeling is completed for all combination forms of the markers.

[0160] By performing the above prediction model construction steps, the optimal models constructed from all combination forms of the markers are obtained. To compare the performance of the models under these different marker combination forms, the ROC method was used to evaluate the AUC values of these prediction models in the test group, and the results are shown in Table 4.

[0161] Table 4. Comparison of the areas under the ROC curves of models constructed with different marker combinations in the test group

[0162]

[0163] Table 4 is sorted according to the markers with larger VIP values and smaller FDR values in Table 2. Starting from the markers ranked higher, markers are selected for combination in turn, from 2MP to 8MP. The detected AUC values, accuracy, and sensitivity all increase with the increase in the number of markers. However, when ORM1 is added on the basis of 8MP, the diagnostic performance of the constructed model does not continue to increase. It is found that the model constructed by the combination of 8 markers (8MP, IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B + TFF1) has good AUC, sensitivity, and specificity. It is difficult to continue to increase the diagnostic performance by adding more markers. Therefore, the most preferred model is 8MP, with fewer markers and better performance.

[0164] 2.4 Optimization of model parameters

[0165] For the optimal marker combination form IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B + TFF1 in step 2.3, based on this marker combination, models constructed under 9 different combinations of glmnet algorithm hyperparameters were analyzed, and the performance of the models was evaluated by the AUC value (the AUC was calculated using the 10-fold cross-validation method during the modeling process), and the results are shown in Table 5.

[0166] Table 5. AUC of the model constructed under different combinations of glmnet algorithm hyperparameters

[0167]

[0168] As can be seen from Table 5, when the hyperparameter combination of the glmnet algorithm is alpha = 1 and lambda = 0.0005, the AUC reaches the maximum value of 0.948.

[0169] The equation of the model constructed based on the optimal hyperparameter combination is:

[0170]

[0171] Among them, Y is the predicted score, i represents the i-th biomarker, m represents the number of biomarkers (m = 8), Xi represents the detected value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker (see Table 6), and b is the constant 3.554513.

[0172] Table 6. Coefficients of 8 biomarkers in the model

[0173]

[0174] The complete model equation is:

[0175] Y = 4.845IGFBP1 + 7.675SFTPB1 + 1.827MGAM2 + 4.638MUC16 + 6.259SELL + 4.924RNASE1 + 8.936S100B + 8.727TFF1 + 3.554513

[0176] Determination of the diagnostic threshold of the bladder cancer recurrence risk prediction model:

[0177] ## Setting levels: control = case, case = control

[0178] ## Setting direction: controls < cases

[0179] The ROC curve was plotted with the predicted scores in the training group, and the optimal prediction cut-off value of 0.4720534 was set according to the Youden index value. The results are shown in Figure 4 .

[0180] When the predicted score of the prediction model ≤ 0.4720534, it is considered that the bladder cancer recurrence risk of the subject to be tested is low.

[0181] When the prediction score of the prediction model > 0.4720534, it is considered that the subject has a high risk of bladder cancer recurrence.

[0182] From Figure 4 It can be seen that the AUC of the prediction model in the training group is 0.948, the sensitivity is 0.914, and the specificity is 0.878.

[0183] 2.6 Validation of the Bladder Cancer Recurrence Risk Prediction Model

[0184] The optimal model constructed was verified in the test group, and the ROC curve was plotted as Figure 5 shown.

[0185] From Figure 5 it can be seen that the AUC of the model in the test group is 0.928, the sensitivity is 0.886, and the specificity is 0.830.

[0186] In summary, the bladder cancer recurrence risk prediction model constructed with 8 protein markers has good prediction performance and accuracy, and has the best diagnostic efficacy.

[0187] Example 3: Construction and Validation of a Three - Classification Diagnostic Model

[0188] In this example, an attempt was made to construct a three - classification combined diagnostic model for distinguishing the non - recurrence group, in - situ recurrence group, and distant metastasis group of bladder cancer prognosis, which specifically includes the following processes: (1) construction and screening of the optimal diagnostic model; (2) verification of the effect of the optimal diagnostic model. The specific screening process and results are as follows (in the present invention, the binary classification model in Example 2 uses the AUC value as the evaluation index; while when constructing a three - classification model, since multiple categories are involved, the AUC value is usually not applicable, and in this example, indicators such as sensitivity, specificity, accuracy, and consistency are used to measure the diagnostic efficacy of the model):

[0189] 3.1 Construction and Screening of the Three - Classification Diagnostic Model

[0190] For the test cohort of 280 bladder cancer patients, all enrolled patients signed informed consent forms. Among them, 48 patients relapsed within two years after surgery (20 cases had in - situ local recurrence and 28 cases had metastases). Thus, they were divided into two groups: 232 patient samples without recurrence were classified into the low - risk group of bladder cancer recurrence, and 48 cases with recurrence or metastasis were classified into the high - risk group of bladder cancer recurrence, and were randomly divided into a test group and a validation group. The training group included 116 cases in the low - risk group of bladder cancer recurrence and 24 cases in the high - risk group of bladder cancer recurrence (including 10 cases of in - situ recurrence and 14 cases with metastases); the test group included 116 cases in the low - risk group of bladder cancer recurrence and 24 cases in the high - risk group of bladder cancer recurrence (including 10 cases of in - situ recurrence and 14 cases with metastases).

[0191] Based on the biomarker combination form IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B + TFF1 screened in Example 2, this example further constructs a three-classification detection model that can effectively distinguish the non-recurrence group (low-risk group), in-situ recurrence group (in-situ recurrence group), and distant metastasis group (metastasis group) of bladder cancer prognosis. All enrolled patients signed informed consent forms. Among them, all bladder cancer patients were pathologically and histologically diagnosed, and the inclusion criteria were: (a) no history of other malignant tumors; (b) no patients with combined other malignant tumors or autoimmune diseases.

[0192] LC-MS / MS data collection and detection were performed on the collected serum samples to obtain the concentrations of eight protein biomarkers, namely IGFBP1, SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B + TFF1, respectively.

[0193] The Shapiro Wilk test was used to evaluate the normal distribution, and the non-parametric test Wilcoxon test was used to analyze the differences in blood biomarker concentrations between the non-recurrence group (low-risk group), in-situ recurrence group (in-situ recurrence group), and distant metastasis group (metastasis group) of bladder cancer prognosis. A three-classification combined diagnostic model of 8 biomarkers was constructed by combining machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using the predicted probability value with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate diagnostic model. Using the test group, the Youden index (YI) was calculated to determine the predicted probability cut-off value for distinguishing the non-recurrence group (low-risk group), in-situ recurrence group (in-situ recurrence group), and distant metastasis group (metastasis group) of bladder cancer prognosis. In addition, the ROCs of models constructed with different biomarker combinations were constructed and compared. Standard descriptive statistical data, such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD), were calculated to describe the experimental results of the study population. Statistical analysis was performed using R 3.6.1, and a p-value less than 0.05 was considered statistically significant.

[0194] In this example, in order to construct an optimal three-classification combined diagnostic model, after comparing the models constructed by 6 algorithms, namely gradient boosting, naive Bayes, support vector machine, neural network, generalized linear, and discriminant analysis, the gradient boosting method was selected as the best supervised classification algorithm for constructing the prediction model. The grid search range for hyperparameter optimization of the gradient boosting method is shown in Table 7 below.

[0195] Table 7. Parameter grid search range of the gradient boosting method

[0196]

[0197] Through optimization and screening in terms of accuracy, consistency, sensitivity, specificity, etc., the optimal parameter combination mode was determined as: interaction.depth 1, n.trees 100, shrinkage 0.1, n.minobsinnode 10.

[0198] Completely different two batches of samples were used for the training group and the test group. In this embodiment, only the screening of biomarkers and the construction of the model were carried out from the training group; the samples of the test group were only used to verify the diagnostic efficacy of the model. The specific results are shown in Table 8.

[0199] Table 8. Performance evaluation table for constructing a model to distinguish three classifications by gradient boosting method

[0200]

[0201] It can be seen from Table 8 that the gradient boosting model constructed based on eight protein biomarkers of IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B + TFF1 can be used to predict whether bladder cancer patients will have no recurrence, in-situ recurrence or distant metastasis after surgery. At the same time, it also shows that the protein biomarkers screened in the present invention can be used to distinguish the recurrence risk of bladder cancer patients after surgical treatment, and can also be used to distinguish whether it is in-situ recurrence or distant metastasis when the recurrence risk is high (when there is both in-situ recurrence and metastasis, it is also classified into the metastasis group).

[0202] 3.2 Combined performance of the three-classification combined diagnosis model

[0203] In order to further improve the diagnostic value of the three-classification diagnosis model (gradient boosting) constructed by biomarker combinations of different proteins, in this embodiment, based on the 11 protein biomarkers screened in Example 1, the performance of the diagnostic models constructed by different protein combination biomarkers was compared in the test group. The specific combination forms of different models are shown in Table 9.

[0204] Table 9. Combination forms of different diagnostic models

[0205]

[0206] Comparison results of performance indicators of different diagnostic models constructed using the 11 biomarkers screened in Example 1. The calculation methods for the minimum, first quartile, median, mean, third quartile, and maximum of accuracy and consistency are as follows: (1) Sort the values of accuracy or consistency from smallest to largest; (2) Minimum value: The first value after sorting; (3) First quartile (Q1): Multiply the number of data by 0.25. If the result is an integer, take the average of the values at this position and the next position; if not, round up to get the position, and the value at this position is Q1; (4) Median: If the number of data is odd, the median is the middle value; if even, it is the average of the two middle values; (5) Mean: The sum of all values divided by the number of data; (6) Third quartile (Q3): Multiply the number of data by 0.75, and the processing method is the same as Q1; (7) Maximum value: The last value after sorting. Among them, the minimum and maximum values can reflect the extreme situations of the data and show the worst and best performances that the model may exhibit; the quartiles can help understand the distribution range and dispersion degree of the data; below Q1 represents a lower performance level, and above Q3 represents a higher performance level; the median can reflect the performance at the middle level; the mean comprehensively reflects the overall average performance. Combining the above statistical values can comprehensively understand the overall situation, distribution characteristics, and stability of the model performance, thus providing a strong basis for model selection and optimization.

[0207] Table 10. Performance comparison of diagnostic models constructed based on different protein combination biomarkers

[0208]

[0209] As can be seen from Table 10, for the three-classification diagnostic model, the combined detection model composed of 10 biomarkers has the best performance. This also clearly shows that on the basis of the eight biomarkers IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B + TFF1, continuing to add the biomarkers ORM1 and FBLN5 has a very obvious improvement effect on the diagnostic efficacy of distinguishing whether the prognosis of bladder cancer surgery will recur, whether it is in-situ recurrence or distant metastasis. However, when continuing to add biomarkers to 11, it is difficult to further improve the accuracy and consistency. Therefore, the three-classification gradient boosting model constructed using these 10 protein biomarkers (IGFBP1 + SFTPB + MGAM2 + MUC16 + SELL + RNASE1 + S100B + TFF1 + ORM1 + FBLN5) is used as the best combined diagnostic model.

[0210] 3.2 Determination and verification of the diagnostic performance of the three-classification combined diagnostic model

[0211] 1) Determination of the diagnostic performance of the three-classification combined diagnostic model

[0212] To more precisely determine the diagnostic performance and thresholds of the model constructed in this embodiment for different disease classifications, a multi-classification model based on the gradient boosting (gbm) algorithm was used to perform predictive analysis in the training group. The predicted probability values for the three-class classification (no recurrence, in-situ recurrence, or distant metastasis after bladder cancer surgery) were calculated, and the classification with the largest predicted probability value was the final prediction result of the system.

[0213] The meanings and calculation methods of each index are as follows:

[0214] Calculation results: The accuracy of the three-class combined diagnostic model in the training group was 0.92, and the consistency was 0.91. The diagnostic sensitivity for the non-recurrence group after bladder cancer surgery was 93.2%, and the specificity was 94.6%; the diagnostic sensitivity for the in-situ recurrence group after bladder cancer surgery was 86.5%, and the specificity was 89.6%; the diagnostic sensitivity for the distant metastasis group after bladder cancer surgery was 89.7%, and the specificity was 90.1%.

[0215] It should be noted that the three-class combined diagnostic model constructed by gradient boosting is a model constructed by machine learning and cannot fit a specific equation formula like a generalized linear model.

[0216] 2) Verification of the three-class combined diagnostic model

[0217] Based on the model constructed for the test group, the predictive performance was verified in the training group. The specific results are as follows:

[0218] The accuracy was 0.90, and the consistency was 0.89. The diagnostic sensitivity for the non-recurrence group after bladder cancer surgery was 93.1%, and the specificity was 93.8%; the diagnostic sensitivity for the in-situ recurrence group after bladder cancer surgery was 91.6%, and the specificity was 91.5%; the diagnostic sensitivity for the distant metastasis group after bladder cancer surgery was 88.1%, and the specificity was 89.4%.

[0219] In summary, the three-class combined diagnostic model constructed in this embodiment, which includes 10 protein markers, has good diagnostic value for the three classifications of the non-recurrence group after bladder cancer surgery, the in-situ recurrence group after bladder cancer surgery, and the distant metastasis group after bladder cancer surgery.

[0220] All patents and publications mentioned in the specification of the present invention indicate that these are publicly available technologies in the art and can be used in the present invention. All patents and publications cited herein are equally listed in the references, as if each publication were specifically individually referenced. The present invention described herein can be implemented in the absence of any one or more elements, one or more limitations, where such limitations are not specifically stated. For example, in each instance herein, the terms "comprising," "consisting essentially of," and "consisting of" can be replaced by either of the remaining two terms. The so-called "a" herein merely means "one," and does not exclude including only one, nor does it exclude including more than two. The terms and expressions used herein are for the purpose of description and are not limiting, and there is no intention to indicate that the terms and explanations described herein exclude any equivalent features, but it is understood that any suitable changes or modifications can be made within the scope of the present invention and the claims. It is understood that the embodiments described in the present invention are preferred embodiments and features, and any person of ordinary skill in the art can make some changes and variations based on the essence described in the present invention, and such changes and variations are also considered to be within the scope of the present invention and the scope limited by the independent claims and the dependent claims.

Claims

1. Use of a biomarker in the preparation of a reagent for predicting the risk of recurrence or metastasis of bladder cancer, characterized in that, The biomarker(s) include(s) one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5.

2. The use according to claim 1, characterized in that, The biomarker(s) include(s) IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1.

3. The use according to claim 2, wherein The reagent is used to detect the content of the biomarker in a body fluid sample; the body fluid sample includes any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid; and / or, the reagent is used to detect the presence, absence, relative abundance, or concentration of the biomarker.

4. The use according to claim 2, characterized in that, The risk of recurrence or metastasis refers to recurrence or metastasis within three years after the treatment of bladder cancer, and / or the bladder cancer includes non-muscle-invasive bladder cancer.

5. A product for predicting the risk of recurrence or metastasis of bladder cancer, characterized in that, It includes a substance for detecting a biomarker, and the biomarker includes one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5.

6. A biomarker combination for predicting the risk of recurrence or metastasis of bladder cancer, characterized in that, It includes IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1.

7. A method for constructing a risk prediction model for bladder cancer recurrence or metastasis for non-diagnostic purposes, characterized in that, It includes the following: 1) Construct a data set based on the detected amount of the biomarker in a biological sample, and the biomarker includes IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1; 2) Divide the data set into a test set and a training set, and construct and train the bladder cancer recurrence or metastasis risk prediction model through a machine learning method.

8. A device for predicting the risk of recurrence or metastasis of bladder cancer, characterized in that, It includes a data acquisition unit and a calculation unit; The data acquisition unit is used to obtain the detected amount data of the biomarker in the biological sample of the subject, and the biomarker includes one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, ORM1, and FBLN5; The calculation unit is used to calculate and output a predicted score of the risk of recurrence or metastasis of bladder cancer for the subject based on the detected amount data, and make a judgment according to the cut-off value.

9. An apparatus, comprising a processor and a memory, the memory being configured to store a computer program, characterized in that, The processor is used to execute the computer program stored in the memory, so that the device executes the construction method as described in claim 7.

10. A system for predicting the recurrence risk of bladder cancer, characterized in that, The system includes a data analysis module, and the data analysis module is used to analyze the detection value of the biomarker, and the biomarker includes IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1.

Citation Information

Patent Citations

  • Chromogranin a as a marker for bladder cancer

    CN108780093A

  • Biomarker combination and application thereof in prediction and / or diagnosis of colorectal cancer

    CN117587128A

  • Biomarker combination and application thereof in predicting bladder cancer risk

    CN119410773A

  • Circulating biomarkers for cancer

    US20140228233A1

  • Methods for treating prostate and lung cancer

    US20250154510A1