A biomarker for predicting the risk of bladder cancer recurrence or metastasis and its application
By screening and constructing a proteomic-based biomarker model, the problem of lack of high-sensitivity bladder cancer risk prediction in the prior art is solved, and non-invasive and accurate bladder cancer risk prediction is achieved, which improves diagnostic efficacy and therapeutic effect.
Patent Information
- Application Number
- CN202510725542.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The lack of high sensitivity biomarkers in the prior art is used to predict the risk of recurrence of non-muscular invasive bladder cancer (NMIBC), resulting in a high recurrence rate after surgical treatment, and the inability to timely adjust the treatment plan to improve survival and quality of life.
Biomarkers such as IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1 and FBLN5 were screened through proteomics to construct a risk prediction model for bladder cancer recurrence or metastasis, and blood samples were analyzed using high-performance liquid chromatography-tandem mass spectrometry technology, and predictive models were constructed in combination with machine learning algorithms.
It achieves non-invasive and accurate prediction of the risk of recurrence or metastasis of bladder cancer, improves the accuracy and sensitivity of diagnosis, reduces the risks of misdiagnosis and missed diagnosis, provides personalized treatment suggestions, and improves the treatment effect and quality of life of patients.
Smart Images

Figure CN120254285B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedicine technology, and in particular to a biomarker for predicting the risk of bladder cancer recurrence or metastasis and its application. Background Art
[0002] Bladder cancer is a common urinary tract malignancy with a high risk of recurrence and progression. Its progression typically begins with non-muscle invasive bladder cancer (NMIBC), but some patients may progress to muscle invasive bladder cancer (MIBC) or even distant metastasis. The development of bladder cancer involves multiple biological processes, including cell proliferation, apoptosis, invasion, and metastasis, which are regulated by numerous genes and proteins.
[0003] Currently, the diagnosis of bladder cancer relies primarily on urine cytology and cystoscopy, but these methods have limitations. In recent years, a variety of new biomarkers and detection methods have been proposed to improve diagnostic accuracy and sensitivity: Urinary biomarkers, such as BCLA-1, BCLA-4, AURKA, and APE / Ref-1, are detectable in urine and are not affected by factors such as infection and smoking. For example, AURKA can be used to distinguish between normal urine and low-grade bladder cancer. Multi-marker combinations, such as combining multiple urine biomarkers (such as collagen α-1 (I) and uromodulin), can improve diagnostic sensitivity and specificity. Genetic testing, such as Uromonitor-V2, detects KRAS hotspot mutations and has high sensitivity and specificity. EpiCheck, which detects DNA methylation markers in urine, has high sensitivity for monitoring bladder cancer recurrence.
[0004] Early-stage bladder cancer, such as non-muscle invasive bladder cancer (NMIBC), rarely metastasizes to distant sites because the tumor has not infiltrated the smooth muscle layer of the bladder. Surgical removal of the tumor can achieve a cure. However, approximately 10%-30% of NMIBC patients still experience recurrence after surgical treatment. If patients who do not benefit from surgical treatment or whose disease progresses can have their risk predicted and their treatment plans adjusted in a timely manner (such as adjuvant chemoradiotherapy, secondary surgical resection, targeted therapy, or immunotherapy, etc.), the overall survival rate and quality of life of the patients can be significantly improved. There is a lack of clinical biomarkers for diagnosing the risk of recurrence of NMIBC, and the discovery of highly sensitive proteomic biomarkers for diagnosing the risk of recurrence of NMIBC is of great significance.
[0005] Proteomics is the study of protein composition, localization, changes, and interactions within cells, tissues, or organisms. This includes the study of protein expression patterns and proteome functional patterns. With the advancement of mass spectrometry, liquid chromatography coupled to mass spectrometry (LC-MS / MS) has become the primary tool in proteomics research. The development of proteomics is crucial for identifying disease diagnostic markers, screening drug targets, and conducting toxicology studies, leading to its widespread application in medical research.
[0006] Therefore, it is urgent to find a more effective and non-invasive biomarker for diagnosing the recurrence risk of bladder cancer. Summary of the Invention
[0007] In response to the problems existing in the prior art, the present invention provides a biomarker for predicting the risk of bladder cancer recurrence or metastasis and its application. By using proteomics methods, by analyzing proteins with significantly different abundance levels in the blood of two groups of people with bladder cancer recurrence and those without bladder cancer recurrence, biomarkers that can be used to predict the risk of bladder cancer recurrence or metastasis are screened out, and a bladder cancer recurrence or metastasis risk prediction model is further constructed. This can achieve accurate, non-invasive and efficient prediction of bladder cancer recurrence or metastasis risk to meet clinical needs.
[0008] A first aspect of the present invention provides use of a substance for detecting biomarkers in the preparation of a product for predicting the risk of bladder cancer recurrence or metastasis, wherein the biomarkers include one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1 and FBLN5.
[0009] The present invention provides biomarkers obtained by proteomics that can accurately predict the risk of bladder cancer recurrence or metastasis, which helps doctors judge the severity of the patient's condition and whether the treatment plan needs to be adjusted, thereby providing patients with more accurate treatment recommendations and truly benefiting bladder cancer patients.
[0010] The present invention utilizes proteomics methods to collect blood samples from patients with and without recurrent bladder cancer. The different samples are analyzed using high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS). Based on orthogonal partial least squares discriminant analysis and significance analysis methods, proteins with significant differences between recurrent and non-recurrent bladder cancer are first screened. Eleven differentially expressed proteins with a clear correlation with the risk of bladder cancer recurrence are identified. These 11 proteins can be used to distinguish whether bladder cancer patients will relapse after surgical treatment, demonstrating a certain diagnostic efficacy.
[0011] Among them, the IGFBP1 is the protein or amino acid sequence with the UniProt database number P08833; SFTPB is the protein or amino acid sequence with the UniProt database number P07988; MGAM2 is the protein or amino acid sequence with the UniProt database number Q2M2H8; MUC16 is the protein or amino acid sequence with the UniProt database number Q8WXI7; SELL is the protein or amino acid sequence with the UniProt database number P14151; RNASE1 is the UniProt database number The protein or amino acid sequence of the database number is P07998; S100B is the protein or amino acid sequence of the UniProt database number is P04271; TFF1 is the protein or amino acid sequence of the UniProt database number is P04155; CD74 is the protein or amino acid sequence of the UniProt database number is P04233; ORM1 is the protein or amino acid sequence of the UniProt database number is P02763; FBLN5 is the protein or amino acid sequence of the UniProt database number is Q9UBX5.
[0012] The present inventors surprisingly discovered that some of the protein markers obtained through proteomic screening are already known markers for other cancers. For example, SFTPB has been reported to predict lung cancer, and TFF1 has been reported to predict gastric cancer. However, this screening revealed that these markers can also be used to predict the risk of bladder cancer recurrence or metastasis. This indicates that the protein markers of many different tumors are not completely isolated or unrelated; in fact, there are many cross-relationships or influences. Many protein markers can be used for both early cancer prediction and prognosis, and even for prediction and diagnosis of different cancers at different stages. Therefore, the field of proteomics has many new applications that need to be explored, and the market prospects are very broad.
[0013] In some embodiments, the biomarker can be one marker or a combination of several markers, such as a combination of two markers, a combination of three markers, a combination of four markers, a combination of five markers, a combination of six markers, a combination of seven markers, a combination of eight markers, a combination of nine markers, or a combination of ten markers.
[0014] In some specific embodiments, the biomarker comprises a combination of two or more markers, such as a combination of three or more markers.
[0015] In some more specific embodiments, the biomarker comprises a combination of two biomarkers, such as IGFBP1+SFTPB (2MP). In some more specific embodiments, the biomarker comprises a combination of three biomarkers, such as IGFBP1+SFTPB+MGAM2 (3MP). In some embodiments, the biomarker comprises a combination of four biomarkers, such as IGFBP1+SFTPB+MGAM2+MUC16 (4MP). In some embodiments, the biomarker comprises a combination of five biomarkers, such as IGFBP1+SFTPB+MGAM2+MUC16+SELL (5MP). In some embodiments, the biomarker comprises a combination of six biomarkers, such as IGFBP1+SFTPB+MGAM2+MUC16+SELL+RNASE1 (6MP). In some embodiments, the biomarkers comprise a combination of 7 biomarkers, such as IGFBP1+SFTPB+MGAM2+MUC16+SELL+RNASE1+S100B (7MP). In some embodiments, the biomarkers comprise a combination of 8 biomarkers, such as IGFBP1+SFTPB+MGAM2+MUC16+SELL+RNASE1+S100B+TFF1 (8MP). In some embodiments, the biomarkers comprise a combination of 9 biomarkers, such as IGFBP1+SFTPB+MGAM2+MUC16+SELL+RNASE1+S100B+TFF1+ORM1 (9MP). In some embodiments, the biomarkers comprise a combination of 10 biomarkers, such as IGFBP1+SFTPB+MGAM2+MUC16+SELL+RNASE1+S100B+TFF1+ORM1+FBLN5 (10MP). In some embodiments, the biomarkers comprise a combination of 11 biomarkers, such as IGFBP1+SFTPB+MGAM2+MUC16+SELL+RNASE1+S100B+TFF1+CD74+ORM1+FBLN5 (11MP). This combination includes, but is not limited to, the aforementioned combinations. When this combination of biomarkers is used to construct a bladder cancer recurrence or metastasis risk prediction model, the AUC value is 0.801-0.948, with a sensitivity of 82.9%-95.4% and a specificity of 83.3%-95.1%.
[0016] Furthermore, to further investigate the diagnostic efficacy of the model for bladder cancer recurrence risk, it was necessary to combine different differentially expressed proteins to construct a diagnostic model. The identified differentially expressed proteins were ranked according to their importance, and different numbers of proteins with the highest rankings were selected for combination, resulting in 10 protein markers. After verification, it was found that the model based on eight protein markers had good risk prediction capabilities in diagnosing bladder cancer recurrence risk. Data from bladder cancer recurrence risk samples showed that using only these eight biomarkers to predict bladder cancer recurrence risk achieved an AUC value of 0.948, demonstrating good diagnostic performance.
[0017] In some embodiments, the product is selected from one or more of a reagent, a kit, a chip, a probe, or a membrane strip. The product is a product that detects the biomarkers described above, and includes, for example, sample pretreatment reagents, antigens, or antibodies, and other biological reagents and kits suitable for detecting the biomarkers. The product can also be developed into standardized reagents or kits, chips, probes, or membrane strips suitable for detecting the biomarkers.
[0018] In some embodiments, the product is used to predict the risk of bladder cancer recurrence or metastasis in patients.
[0019] Furthermore, the risk of recurrence or metastasis refers to recurrence or metastasis within three years after bladder cancer treatment.
[0020] The bladder cancer recurrence or metastasis includes bladder cancer recurrence in situ or adjacent areas, and bladder cancer metastasis.
[0021] Furthermore, the bladder cancer includes non-muscle invasive bladder cancer.
[0022] In some embodiments, the product is used to detect the amount of a biomarker in a biological sample.
[0023] In some specific embodiments, the biological sample is selected from one or more of saliva, blood, urine, plasma, serum and cerebrospinal fluid.
[0024] In some specific embodiments, the detected amount includes the presence or relative abundance or concentration of a biomarker.
[0025] The present invention uses blood screening to identify biomarkers that predict the risk of bladder cancer recurrence. These biomarkers show significant differences in the blood of people with and without bladder cancer recurrence. By collecting blood samples, these biomarkers can be detected in the individual's blood to predict or assist in diagnosing whether the individual has bladder cancer recurrence or not, or to predict or assist in diagnosing whether the individual has not recurred after bladder cancer surgery, has recurred in situ, or has distant recurrence.
[0026] Furthermore, the detection method generally includes a radiometric method, an immunological method, a fluorescence method, a flow cytometry, a latex turbidimetry method, a biochemical method, an enzymatic method, a hybridization method, a gas chromatography-mass spectrometry method, a liquid chromatography-mass spectrometry method, a chromatography method, a chemiluminescence method, a magnetoelectric method or a photoelectric conversion method.
[0027] The presence or absence of biomarkers, or the level of these biomarkers, is a relative concept. For example, when comparing a bladder cancer recurrence group with a non-recurrence group, the levels of these specific biomarkers are compared relative to the baseline of the recurrence or non-recurrence group. It's possible that certain biomarkers may be higher in recurrence than in the non-recurrence group, and this increase can be statistically significant, such as a significant or highly significant increase. Therefore, when assessing the presence of these biomarkers, if a single biomarker indicates an increased probability of a certain risk, its level may change. This change can be a relative increase or decrease, and this relative increase or decrease can be considered significant, or even highly significant. Therefore, regardless of the testing method, a predetermined cut-off value can be used as a standard. A value above this cut-off value is considered a change in the level, and such a result can be used for prognostic or diagnostic purposes.
[0028] Therefore, in some aspects, the biomarkers described herein can be obtained by detecting the marker content in a sample using any known method, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strips, immunoreaction chips, capillary electrophoresis, infrared spectroscopy, and the like. As long as the protein marker content in a sample can be detected, it can be used to predict or diagnose the probability of a disease. It is understood that the detection here involves testing an individual sample and then comparing it with a pre-set standard. The comparison results are used to determine or predict the disease state. For example, it can be used to predict the probability of bladder cancer recurrence risk. This prediction or diagnosis is based on whether the disease will recur within a certain period of time. Of course, such detection can be continuous, and the progression of the disease can be inferred based on the changes in the content of certain substances.
[0029] In some embodiments, the relative abundance is the peak area of the biomarker in a detection spectrum obtained by high-performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a biomarker measured in a control sample is 300 and the average peak area measured in a bladder cancer recurrence sample is 1800, then the abundance of the biomarker in the sample is considered to be 6 times that in the control sample.
[0030] A second aspect of the present invention provides the use of proteins as biomarkers in the preparation of products for predicting the risk of bladder cancer recurrence or metastasis. The proteins are selected from one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5. The IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5 genes can be used to assist in determining the risk of bladder cancer recurrence or metastasis, assess drug efficacy, and more. The inventors have discovered that IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, ORM1, and FBLN5 are closely associated with the risk of bladder cancer recurrence or metastasis.
[0031] A third aspect of the present invention provides a product for predicting the risk of bladder cancer recurrence or metastasis, comprising a substance for detecting biomarkers, wherein the biomarkers include one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1 and FBLN5.
[0032] A fourth aspect of the present invention provides a biomarker combination for predicting the risk of bladder cancer recurrence or metastasis, comprising IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1.
[0033] A fifth aspect of the present invention provides a method for constructing a bladder cancer recurrence or metastasis risk prediction model for purposes other than disease diagnosis, comprising the following steps:
[0034] 1) constructing a sample dataset based on the detected amount of biomarkers in the biological sample, wherein the biomarkers include one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5;
[0035] 2) dividing the data set into a test set and a training set, and constructing and training the bladder cancer recurrence or metastasis risk prediction model through machine learning methods.
[0036] In some embodiments, the biomarkers are IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1.
[0037] In some embodiments, the machine learning method is selected from at least one of a gradient boosting algorithm, a random forest algorithm, a support vector machine algorithm, a decision tree algorithm, a K-nearest neighbor algorithm, a logistic regression algorithm, and a neural network algorithm. For example, the gradient boosting algorithm can be selected. Unlike generalized linear regression, which outputs a model formula and cutoff value, the gradient boosting algorithm performs all calculations directly through machine learning. Test values can be directly input into the software system to directly obtain prediction results.
[0038] In some embodiments, multiple machine learning methods are used to construct a bladder cancer recurrence or metastasis risk prediction model, and it is preliminarily confirmed that the concentration changes of any one of the screened new biomarkers alone can be used to distinguish between people with bladder cancer recurrence and those without recurrence, indicating that these biomarkers have extremely high diagnostic value.
[0039] The present invention discovered that by detecting the detection amount of biomarkers in biological samples and then inputting the detection amount into the formula of a bladder cancer recurrence or metastasis risk prediction model, a prediction score Y of the prediction model is obtained, and the prediction score Y is compared with a threshold (cut off) defined by the Youden Index. If the prediction score Y is greater than the threshold (cut off), it is judged that the bladder cancer has recurred; if the prediction score y is less than or equal to the threshold (cut off), it is judged that the bladder cancer has not recurred.
[0040] In some embodiments, the equation of the constructed bladder cancer recurrence risk prediction model is:
[0041]
[0042] Where Y is the prediction score, i represents the i-th biomarker, m represents the number of biomarkers (m = 8), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker (see Table 6), and b is a constant of 3.554513.
[0043] Table 6 Coefficients of 8 biomarkers in the model
[0044]
[0045] When Y≤cutoff value, the patient has no bladder cancer recurrence; when Y>cutoff value, the patient has bladder cancer recurrence. Specifically, the cutoff value is 0.4720534.
[0046] In some embodiments, the method further includes step 3) testing the bladder cancer recurrence or metastasis risk prediction model using a test set. After training, the trained prediction model is validated using the test set, and the effectiveness of the prediction model is evaluated using AUC, specificity, and sensitivity as evaluation indicators.
[0047] The Receiver Operating Characteristic Curve (ROC Curve) is a curve drawn based on a series of different binary classification methods (cut-off values), with the true positive rate (sensitivity) as the vertical axis and the false positive rate (1-specificity) as the horizontal axis. The area under the receiver operating characteristic curve (AUC) is defined as the area under the ROC curve. The AUC value is often used to evaluate the diagnostic efficacy of a prediction model. The larger the AUC value, the better the diagnostic efficacy of the corresponding prediction model; conversely, the lower the AUC value, the worse the diagnostic efficacy of the corresponding prediction model.
[0048] The prediction model constructed based on the combination of 8MP in the present invention can distinguish between recurrence and non-recurrence of bladder cancer after surgery. In the training group, the AUC was 0.948, the sensitivity was 0.914, and the specificity was 0.878; in the test group, the AUC was 0.928, the sensitivity was 0.886, and the specificity was 0.830.
[0049] A sixth aspect of the present invention provides a device for predicting the risk of bladder cancer recurrence or metastasis, comprising a data acquisition unit and a calculation unit;
[0050] The data acquisition unit is used to obtain the detection amount data of the biomarkers as described above in the biological sample of the subject, wherein the biomarkers include one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1 and FBLN5;
[0051] The calculation unit is used to calculate and output a predicted score for the subject's risk of bladder cancer recurrence or metastasis based on the detection amount data, and make a judgment based on a threshold.
[0052] In some embodiments, the biomarkers include IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1.
[0053] In some embodiments, the prediction device further includes a data storage unit and a data output unit; the data storage unit is used to store the detection amount (or detection value) of the biomarker; the data input interface is used to input the detection value of the biomarker, and the data output unit is used to output the prediction result.
[0054] Furthermore, the detection value is the presence or absence, relative abundance or concentration value of each biomarker.
[0055] The seventh aspect of the present invention provides a device comprising a processor and a memory, wherein the memory is used to store a computer program, and is characterized in that the processor is used to execute the computer program stored in the memory so that the device performs the prediction method as described above or the construction method as described above.
[0056] An eighth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is processed, the prediction method as described above or the construction method as described above is executed.
[0057] In some embodiments, the computer-readable storage medium includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0058] In some implementations, being processed refers to being executed by one or more processors.
[0059] A ninth aspect of the present invention provides a method for predicting the risk of bladder cancer recurrence or metastasis, comprising the following steps:
[0060] S1. Obtaining detection amount data of biomarkers in a biological sample of a subject, wherein the biomarkers include one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5;
[0061] S2. Applying the bladder cancer recurrence risk prediction model obtained by the construction method described above to process the detection data to output a bladder cancer recurrence or metastasis risk prediction result.
[0062] In some embodiments, the biomarkers include IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1.
[0063] In another aspect, the present invention provides a system for predicting the risk of bladder cancer recurrence, the system comprising a data analysis module for analyzing the detection values of markers, wherein the markers include IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B and TFF1.
[0064] Furthermore, the data analysis module uses the detection values of markers of known samples as a training set, divides the samples into a bladder cancer recurrence group and a bladder cancer non-recurrence group according to whether the bladder cancer recurs, analyzes the relationship between the detection values of the bladder cancer recurrence group and the bladder cancer non-recurrence group, and constructs a model.
[0065] In some embodiments, a combined diagnostic model for predicting the risk of bladder cancer recurrence is constructed by combining multiple machine learning methods, and it is preliminarily confirmed that the concentration changes of any one of the screened new biomarkers alone can be used to distinguish between people at high risk and low risk of bladder cancer recurrence, indicating that these biomarkers have extremely high diagnostic value.
[0066] In some embodiments, the equation of the constructed model is:
[0067]
[0068] Where Y is the prediction score, i represents the i-th biomarker, m represents the number of biomarkers (m=8), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant of 3.554513; the coefficients of the eight biomarkers are:
[0069]
[0070] When Y≤cutoff value, the patient has no bladder cancer recurrence; when Y>cutoff value, the patient has bladder cancer recurrence. Specifically, the cutoff value is 0.4720534.
[0071] Furthermore, the system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0072] In another aspect, the present invention provides a use of a marker for preparing a reagent for predicting whether a bladder cancer patient will not relapse after surgery, will relapse in situ, or will have distant metastasis, wherein the marker comprises one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5;
[0073] In another aspect, the present invention provides a kit for predicting whether a bladder cancer patient will not relapse after surgery, will relapse in situ, or will metastasize to a distant site. The kit comprises a detection reagent for the biomarker for the purpose described above.
[0074] In another aspect, the present invention provides a biomarker combination for predicting whether a bladder cancer patient will not relapse after surgery, relapse in situ, or have distant metastasis, wherein the combination comprises any one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5.
[0075] In another aspect, the present invention provides a system for predicting whether bladder cancer will not recur after surgery, will recur in situ, or will metastasize to a distant site. The system includes a data analysis module for analyzing the detection values of markers, wherein the markers include any one or more of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, CD74, ORM1, and FBLN5.
[0076] Furthermore, the distant metastasis also includes simultaneous in situ recurrence and distant metastasis. As long as distant metastasis occurs, it will be classified into the distant metastasis group.
[0077] Furthermore, the data analysis module uses the detection values of markers of known samples as a training set, and divides liver cancer patients into a non-recurrence group, an in situ recurrence group, and a distant metastasis group according to their post-operative conditions. The relationship between the detection values of the non-recurrence group, the in situ recurrence group, and the distant metastasis group is analyzed to construct a model.
[0078] Furthermore, the model is constructed based on the gradient boosting algorithm.
[0079] Unlike generalized linear regression, the gradient boosting algorithm cannot output model formulas and cutoff values. All calculations are done directly by machine learning. The test values can be directly input into the software system to obtain the prediction results.
[0080] Furthermore, the system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0081] The beneficial effects of the present invention are:
[0082] 1. The present invention has screened 11 new biomarkers that can predict the risk of bladder cancer recurrence or metastasis, and developed a new protein marker combination that can effectively assess and diagnose the risk of bladder cancer recurrence or metastasis, effectively distinguish between bladder cancer recurrence and non-recurrence, and effectively distinguish between in situ recurrence and distant metastasis of bladder cancer. This not only improves the diagnostic accuracy of bladder cancer recurrence and metastasis, but also provides an important biomarker basis for personalized treatment and prognosis monitoring of bladder cancer patients. In addition, the discovery and application of these markers are expected to improve the treatment effect and quality of life of bladder cancer patients, and provide a scientific basis for the early diagnosis and treatment strategy selection of bladder cancer. Compared with traditional detection methods, it reduces the risk of misdiagnosis and missed diagnosis, providing strong support for early detection and intervention of the disease.
[0083] 2. The eight-biomarker combined differential diagnosis model constructed by this invention is convenient and rapid, and its test results are highly consistent with the clinical gold standard test results. It also significantly reduces the cost of predicting the risk of bladder cancer recurrence or metastasis, and has promising application prospects. The combination of markers of this invention is superior to the diagnosis of broad-spectrum tumor markers or broad-spectrum tumor marker combination models.
[0084] 3. Based on the biomarkers screened for predicting the risk of bladder cancer recurrence or metastasis, a three-classification model was further constructed that can simultaneously distinguish between the non-recurrence group, the in situ recurrence group, and the distant metastasis group, providing a more effective and accurate predictive diagnostic model.
[0085] Detailed description
[0086] (1) Diagnosis or testing
[0087] The term "diagnosis" or "detection" here refers to the detection or testing of biomarkers in a sample, or the content of a target biomarker, such as the absolute content or relative content, and then the presence or amount of the target marker is used to indicate whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease. The terms "diagnosis" and "detection" are interchangeable. The results of such a test or diagnosis cannot be directly used as a direct result of the disease, but rather an intermediate result. If a direct result is obtained, other auxiliary means such as pathology or anatomy are required to confirm the presence of a certain disease. For example, the present invention provides a variety of new biomarkers associated with the risk of bladder cancer recurrence or metastasis. Changes in the content of these markers are directly correlated with whether the patient belongs to the group at risk of bladder cancer recurrence or metastasis.
[0088] (2) Association between markers, biomarkers, or differentially expressed proteins and the risk of bladder cancer recurrence or metastasis
[0089] The terms "marker," "biomarker," and "differential protein" have the same meaning in this invention. Association here refers to a direct correlation between the presence or change in the level of a biomarker in a sample and a specific disease. For example, a relative increase or decrease in the level indicates a higher likelihood of the individual having the disease compared to a healthy population.
[0090] The simultaneous presence of multiple markers in a sample, or the relative changes in their levels, indicate a higher likelihood of the individual having the disease compared to healthy individuals. This means that among marker types, some are strongly associated with disease, while others are weakly associated, or even unrelated to a particular disease. One or more markers with strong correlations can be used as diagnostic markers, while markers with weaker correlations can be combined with stronger markers to diagnose a disease, increasing the accuracy of test results.
[0091] The numerous serum biomarkers discovered by the present invention can be used to distinguish between recurrent and non-recurrent bladder cancer, as well as between non-recurrent, primary recurrence, and distant metastasis. These markers can be used alone as single markers for direct detection or diagnosis. The selection of such markers indicates that the relative change in their level is strongly correlated with the risk of bladder cancer recurrence or metastasis. Of course, it is understood that one or more markers with strong correlations with the risk of bladder cancer recurrence or metastasis can be selected for simultaneous detection. It is generally understood that, in some approaches, selecting biomarkers with strong correlations for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90%, or 95%. This indicates that these markers can achieve intermediate values for diagnosing a certain disease, but does not directly confirm the presence of a specific disease.
[0092] Of course, the differential protein with the larger ROC value can also be selected as a diagnostic marker. The so-called strength is generally calculated and confirmed by some algorithms, such as the contribution rate or weight analysis of the marker to the risk prediction of bladder cancer recurrence or metastasis. Such calculation methods can be significance analysis (p value or FDR value) and fold change (Fold change). Multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA) and orthogonal partial least squares discriminant analysis (OPLS-DA), and of course other methods, such as ROC analysis, etc. Of course, other model prediction methods are also possible. When specifically selecting biomarkers, the differential proteins disclosed in the present invention can be selected, or other existing well-known marker combinations can be selected or combined to make predictions through model methods.
[0093] (3) Definition of disease terms
[0094] Bladder cancer refers to a malignant tumor that develops in the bladder mucosa. It is the most common malignant tumor of the urinary system and one of the top ten most common tumors in the body. It ranks first in incidence among genitourinary tumors in my country and second in incidence in the West, second only to prostate cancer. In 2012, the incidence of bladder cancer in the National Cancer Registry was 6.61 per 100,000 people, ranking ninth among malignant tumors. Bladder cancer can occur at any age, even in children. Its incidence increases with age, peaking in those aged 50 to 70. The incidence of bladder cancer in men is three to four times higher than in women.
[0095] Non-muscle-invasive bladder cancer (NIBC) is a bladder malignancy that originates in the bladder mucosa and is confined to the mucosa and lamina propria, without invading the muscularis. Pathological types include urothelial carcinoma, squamous cell carcinoma, and adenocarcinoma. Clinical manifestations include painless hematuria.
[0096] Bladder cancer recurrence occurs when bladder cancer reappears or spreads to other parts of the body after treatment. Recurrence can occur at or near the original site of the bladder, known as a local recurrence (in situ recurrence). Local recurrence means the cancer has not spread widely but reappears in the bladder or its immediate surrounding tissues. Alternatively, bladder cancer can spread to distant sites in the body, known as distant recurrence or metastasis. Most bladder cancer recurrences (approximately 80%) occur within 2-3 years after surgery, with the risk of recurrence significantly decreasing after 5 years. Bladder cancer recurrence is closely related to tumor stage, treatment compliance, and postoperative management. The prognosis (expected survival after treatment) for bladder cancer is poor because it is a highly aggressive malignancy with few early symptoms and is often discovered in the late stages. This results in relatively poor treatment outcomes and a poor prognosis.
[0097] Bladder cancer metastasis: Distant metastasis refers to the spread of tumor cells through the blood or lymphatic system to sites beyond the bladder, such as the lungs, liver, and bones. Approximately 10% to 30% of patients with non-muscle invasive bladder cancer who undergo radical cystectomy will experience recurrence or metastasis. The risk of distant metastasis is closely related to tumor stage and grade, with high-grade and advanced tumors being more prone to distant metastasis.
[0098] Therefore, early detection and intervention of bladder cancer recurrence or metastasis are crucial for improving treatment efficacy and increasing survival rates. The biomarkers of the present invention have accurate and specific predictive efficacy for bladder cancer recurrence or metastasis, which is beneficial to improving patients' treatment efficacy and has important significance for improving patients' prognosis.
[0099] (4) The gold standard for diagnosing bladder cancer is pathological diagnosis (i.e., pathological confirmation of a biopsy or surgical resection specimen), which involves observing the presence of cancer cells under a microscope and combining immunohistochemistry or molecular testing to clarify the nature of the tumor. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] Figure 1 This is a volcano plot of the differential analysis of protein markers in the high-risk group and low-risk group for bladder cancer recurrence in Example 1.
[0101] Figure 2 Graphs showing the ROC and OPLS-DA analysis results for the high-risk and low-risk groups for bladder cancer recurrence in Example 1.
[0102] Figure 3 AUC results of the models constructed with different hyperparameters in Example 2.
[0103] Figure 4 ROC curve of the bladder cancer recurrence risk prediction model in Example 2 in the training group.
[0104] Figure 5 ROC curve of the bladder cancer recurrence risk prediction model in Example 2 in the test group. DETAILED DESCRIPTION
[0105] The present invention will be described in further detail below in conjunction with the accompanying drawings and Examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not serve to limit the present invention in any way. The reagents used in this example are all known products and were obtained by purchasing commercially available products.
[0106] Example 1 Screening for biomarkers of bladder cancer recurrence risk using proteomics
[0107] Plasma samples were collected from patients undergoing radical surgery for bladder cancer. Low-abundance proteins were enriched using immunoaffinity chromatography to remove high-abundance proteins. Protein abundance in the samples was detected using a tandem high-performance liquid chromatography-mass spectrometry device. The differences in protein abundance between patients with and without bladder cancer recurrence were analyzed to identify differentially expressed proteins. The specific steps are as follows:
[0108] 1.1 Sample Collection
[0109] Blood samples were collected from 134 patients with non-muscle-invasive bladder cancer (stage Ta / T1) 2 weeks after surgical resection. Bladder cancer was confirmed by biopsy. Approximately 2 ml of peripheral blood was collected from each patient, placed in a vacuum tube containing EDTA anticoagulant, mixed, and centrifuged twice at 120 g for 10 minutes at room temperature. The supernatant was removed and the mixture was centrifuged at 360 g for 20 minutes. Platelet samples were then collected in centrifuge tubes and stored at -80°C until further use. The 132 patients included 94 males and 38 females, with an average age of 47 years (range, 25-69 years). All enrolled patients provided written informed consent. Bladder cancer was confirmed by histopathology. Inclusion criteria included: (a) no history of other malignancies; and (b) no concurrent malignancies or autoimmune diseases.
[0110] For the next three years, patients were followed up every six months with regular CYFRA 21-1 and imaging examinations. Any recurrence or metastasis was confirmed by pathological histology. A total of 32 patients relapsed within three years of surgery, including 14 with local recurrence in situ and 18 with distant metastasis (including 4 liver metastases, 8 lung metastases, 4 bone metastases, and 2 pelvic lymph node metastases). Samples were then removed from the -80°C storage and divided into two groups: 102 patients without recurrence were classified as low-risk for bladder cancer recurrence, and 32 patients with recurrence or metastasis were classified as high-risk for bladder cancer recurrence for proteomic biomarker screening.
[0111] 1.2 Sample processing and enzymatic hydrolysis
[0112] 1) First, plasma samples were centrifuged for 15 minutes (15,000 g), and the supernatant was filtered and then subjected to immunoaffinity chromatography to remove 14 highly abundant proteins.
[0113] 2) The low-abundance fraction was concentrated to 350 μL using a 3 kDa cut-off concentration tube in a centrifuge (4000 g, 1 hour).
[0114] 3) The concentrate was recovered and buffer exchanged using a 7 kDa cutoff desalting column in a centrifuge (1000 g, 2 minutes) using AEX-A (20 mM Tris, 4 M urea, 3% isopropanol, pH 8.0).
[0115] 4) Determine the protein concentration in the sample using the biuret assay (BCA) with AEX-A as a blank. According to the sample grouping, add 25 μL of TCEP to the sample and incubate at 37°C for 30 minutes to reduce the protein. Then, add the corresponding TMT 16-plex reagent and incubate at room temperature in the dark for 1 hour to perform the TMT labeling reaction. Then, perform buffer exchange of the sample using Zeba columns with AEX-A. After mixing the TMT 16-plex-labeled sample, add 2 mL of AEX-A to the mixed sample to a final volume of 5.5 mL.
[0116] 5) Filter the sample through a 0.22 µm filter and separate the TMT 16-plex-labeled sample using a 2D-HPLC system. Freeze-dry the collected fractions and add Trypsin-Lysin C enzyme mix. Incubate the sample at 37°C for 5 hours to digest the sample. Terminate the digestion reaction by adding 5 µL of 10% TFA.
[0117] 6) A total of 60 2D-HPLC fractions after enzymatic digestion were used for nanoLC-MS / MS analysis.
[0118] 1.3 LC-MS / MS Data Acquisition
[0119] Each sample obtained in step 1.2 was separated using the nanoliter flow rate liquid chromatography system Easy nLC-1200 and detected online using a high-resolution mass spectrometer Q Exactive HF-X. The details are as follows:
[0120] 1) Separation: Mobile phase A consisted of 0.1% formic acid in water, and mobile phase B consisted of 0.1% formic acid in acetonitrile (80% acetonitrile, 20% water). The chromatographic column consisted of an enrichment column and an analytical column equilibrated with 100% mobile phase A. The sample was loaded via an autosampler onto the enrichment column (100 μm ID × 4 cmL, C18, 3 μm, 100 Å) and separated on the analytical column (75 μm ID × 25 cmL, C18, 3 μm, 100 Å) at a flow rate of 300 nL / min.
[0121] 2) After chromatographic separation, samples were analyzed by mass spectrometry using a Q Exactive HF-X mass spectrometer. Positive ion detection was used, with a parent ion scan range of 350–1800 m / z, a primary MS resolution of 120,000 at 200 m / z, an AGC (Automatic Gain Control) target of 3e6, a maximum IT of 50 ms, and a dynamic exclusion time of 40 s. The mass-to-charge ratios of peptides and peptide fragments were acquired using data-dependent acquisition (DDA): 20 secondary MS / MS spectra (MS2 scans) were collected after each full scan (MS2 scan). The MS2 Activation Type is HCD, the Isolation window is 0.7 m / z, the secondary mass spectrometry resolution is 30,000 at 200 m / z, the AGC target is 1e5, the Maximum IT is 65 ms, the Fixed first mass is 110.0 m / z, the Normalized Collision Energy is 32 eV, the Minimum AGC target is 2.00e4, the Charge exclusion is 1, 6-8, >8, the Multiple charge states is one charge state only, the Peptidematch is preferred, and the Exclude isotopes is on.
[0122] 1.4 Data Preprocessing
[0123] Secondary mass spectrometry data were retrieved using Maxquant (v1.6.15.0). The data type is DIA proteomics data based on secondary reporter ion quantification. The secondary spectrum used for quantification requires that the parent ion accounts for more than 75% in the primary spectrum. The database comes from the Homo_sapiens_9606_proteome_gene (release: 2021-10-14, sequence: 20,437) of the Uniprot database, and a common contamination library is added to the database. Contaminating proteins are deleted during data analysis; the enzyme cleavage method is set to Trypsin / P; the number of missed cleavage sites is set to 2; the parent ion mass error tolerance of the First search and Main search is set to 20 ppm and 5 ppm, respectively, and the mass error tolerance of the secondary fragment ion is 20 ppm. The fixed modification is cysteine alkylation, and the variable modification is methionine oxidation and acetylation of the protein N-terminus. The FDR for protein identification and PSM identification is set to 1%.
[0124] 1.5 Gap Analysis
[0125] We screened for differentially expressed proteins and transcripts using a combination of univariate and multivariate statistical analyses. Univariate analysis primarily included significance analysis (p-value or FDR value) and fold change analysis of signature molecules across different groups. Multivariate statistical analysis primarily included receiver operating characteristic (ROC) curve analysis and Boruta signature screening based on the random forest algorithm. All statistical analyses were performed using R. Specific R information is provided in Table 1.
[0126] Table 1. R and related information
[0127]
[0128] The variable importance for the projection (VIP) was calculated to measure the influence and explanatory power of each protein expression pattern on the classification and discrimination of each group of samples. The Wilcoxon rank sum test was further performed to obtain the corrected p value (FDR). According to the conditions of FDR < 0.01 and Fold change > 2, 70 down-regulated proteins and 72 up-regulated proteins were screened (see Figure 1 ).
[0129] In order to evaluate the role of each protein marker in the diagnosis and prediction of bladder cancer recurrence risk, ROC and Boruta analysis methods were used to evaluate each protein marker. The results are shown in Figure 2The horizontal axis represents the AUC obtained from ROC analysis, and the vertical axis represents the -log10 (FDR) calculated by the Wilcoxon test. The size of the dots represents the VIP value obtained from Boruta analysis. Further screening based on VIP > 3 and AUC > 0.6 identified 11 more significant candidate protein biomarkers, as detailed in Table 2.
[0130] Table 2. Differential markers of bladder cancer recurrence risk
[0131]
[0132] Among them, the smaller the FDR value and / or the larger the VIP value, to a certain extent, it indicates that the difference between the protein at high risk of bladder cancer recurrence and low risk of bladder cancer recurrence is more significant, and it also indicates that the protein may have a higher diagnostic value.
[0133] Example 2: Construction and validation of a bladder cancer recurrence risk prediction model
[0134] In this example, the 11 protein markers screened and obtained in Example 1 were combined to construct a recurrence risk prediction model for research.
[0135] While a single biomarker can differentiate post-operative bladder cancer recurrence risk in patients, combining multiple biomarkers generally offers greater accuracy in differentiation or prediction. However, a single biomarker with a higher accuracy in predicting bladder cancer recurrence risk may not necessarily be more effective when combined with one or more other biomarkers. Furthermore, a greater number of biomarkers does not necessarily equate to a higher predictive accuracy (AUC value) in the combination. Therefore, extensive validation experiments are still needed.
[0136] 2.1 Get Data
[0137] A total of 240 patients with non-muscle invasive bladder cancer (Ta / T1 stage patients) were collected with blood samples 2 weeks after surgical treatment. All enrolled patients signed informed consent forms, of which 42 patients relapsed within three years after surgery (12 were local recurrences in situ and 30 were metastatic). They were divided into two groups: 198 patient samples without recurrence were classified as a low-risk group for bladder cancer recurrence, and 42 patients with recurrence or metastasis were classified as a high-risk group for bladder cancer recurrence. They were randomly divided into a test group and a validation group. The test group included 99 patients in a low-risk group for bladder cancer recurrence and 21 patients in a high-risk group for bladder cancer recurrence, and the validation group included 99 patients in a low-risk group for bladder cancer recurrence and 21 patients in a high-risk group for bladder cancer recurrence.
[0138] The same steps as steps 1.2 and 1.3 in Example 1 were used to obtain the concentrations of 11 protein markers in serum as shown in Table 2 for 240 samples.
[0139] 2.2 Data Statistical Analysis
[0140] In the training group, a combined diagnostic model for multiple bladder cancer recurrence risk markers was constructed using a combination of multiple machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using the predicted probability values with 95% confidence intervals (CI) to assess the discriminative ability of the multivariate diagnostic model.
[0141] Using the test group, the Youden index (YI) was calculated to determine the cut-off value for distinguishing the high-risk group for bladder cancer recurrence from the low-risk group for bladder cancer recurrence. In addition, the ROC of the model formed by the combination of different markers was constructed and compared. Standard descriptive statistics, such as frequency, mean, median, positive predictive score (PPV), negative predictive score (NPV) and standard deviation (SD) were calculated to describe the experimental results of the study group. Statistical analysis was performed using R3.6.1, and p values less than 0.05 were considered to be statistically significant.
[0142] 2.3 Construction of prediction model
[0143] The specific steps for constructing a bladder cancer recurrence risk prediction model are as follows:
[0144] S101. Randomly select 2 to 11 marker concentration matrices from the 11 protein markers in Table 2 in the samples of the training group as the original training data set.
[0145] S102: Selecting the generalized linear model (glmnet) algorithm for constructing a prediction model and the grid search range during the algorithm hyperparameter optimization process. The grid search range for model hyperparameter optimization is set for each algorithm as shown in Table 3.
[0146] Table 3 Parameter grid of the glmnet algorithm
[0147]
[0148] S103. According to the algorithm and hyperparameter setting range set in step S102, one of the hyperparameter combinations is selected as a parameter for constructing a prediction model.
[0149] S104: Split the original dataset into K subsets using a K-fold cross validation mechanism. To ensure that the ratio of majority class samples to minority class samples in each subset is the same as in the original dataset, a Stratified K-Folds cross validation mechanism is used for data segmentation.
[0150] S105 , selecting one of the K training data subsets obtained by segmentation in step S104 as a validation set Ddev.
[0151] S106: Combine the training data subsets not selected in step S105 to form a training data pool Dtrain1.
[0152] S107 . According to the training data set Dtrain obtained in step S106 , a prediction model is constructed based on the selected supervised classification algorithm and hyperparameters.
[0153] S108: Evaluate the prediction model obtained in step S107 on the validation set Ddev to obtain an AUC value, and store the current prognosis prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 involves evaluating the prediction model obtained in step S107 on the validation set determined in the current iteration, and storing both the model and the evaluation results in the prediction model pool for future selection and use by the base prediction model. The evaluation mentioned in this step can be an AUC value or other reasonable metric for evaluating model performance.
[0154] S109: Determine whether all subsets have been used as validation sets. Step S109 determines whether all K subsets obtained in step S104 have been used as validation sets and trained on the model. If all subsets have been used as validation sets and training has been completed, proceed to step S110; if any subsets have not been used as validation sets, proceed to step S105. This step ensures that every sample in the original dataset has been used as a validation set, improving model stability and preventing the model from overfitting to a particular subset.
[0155] S110: The average AUC value of all models in the prediction model pool Pool is used as the final performance evaluation value of the combined model. The model parameters and the final performance evaluation AUC value are stored in the optimal model pool Poolbest.
[0156] S111: Determine whether all hyperparameter combinations have been used to construct prediction models. Step S111 determines whether all algorithms and corresponding hyperparameter combinations obtained in step S102 have been used to construct prediction models. If all combinations have been used to construct models, then step S112 is executed; if any combination has not been used to construct models, then step S103 is executed.
[0157] S112. After completing step S111, it is necessary to check whether all hyperparameter combinations have been used to build prediction models. If it is found that there are still hyperparameter combinations for which models have not been built, it is necessary to return to step S103 and continue to build the remaining models. If all hyperparameter combinations have been used to build models, it is possible to continue to step S113 and select the best model from the model pool.
[0158] S113 . From the model set Poolbest obtained in step S112 , select the model with the largest AUC value as the final prediction model for the risk of bladder cancer recurrence or metastasis.
[0159] S114. Repeat all the above steps until all combinations of markers are modeled.
[0160] By executing the above prediction model construction steps, the optimal models for all combinations of markers were obtained. To compare the performance of the models under these different marker combinations, the ROC method was used to evaluate the AUC values of these prediction models in the test group. The results are shown in Table 4.
[0161] Table 4. Comparison of the area under the ROC curve of the models constructed by different marker combinations in the test group
[0162]
[0163] Table 4 shows the ranking of markers according to the larger VIP value and the smaller FDR value in Table 2. Starting from the top-ranked markers, markers are selected in sequence for combination. From 2MP to 8MP, the AUC value, accuracy, and sensitivity of the detection increase with the increase of markers. However, when ORM1 is added on the basis of 8MP, the diagnostic performance of the constructed model does not continue to improve. It is found that the model constructed by the combination of 8 markers (8MP, IGFBP1+SFTPB+MGAM2+MUC16+SELL+RNASE1+S100B+ TFF1) has good AUC, sensitivity, and specificity. It is difficult to further improve the diagnostic performance by adding more markers. Therefore, the most preferred model is 8MP, which has fewer markers and better performance.
[0164] 2.4 Optimization of model parameters
[0165] For the optimal marker combination of IGFBP1+SFTPB+MGAM2+MUC16+SELL+RNASE1+S100B+TFF1 in step 2.3, the models constructed under 9 different combinations of glmnet algorithm hyperparameters were analyzed based on this marker combination, and the model performance was evaluated by the AUC value (AUC was calculated using a 10-fold cross-validation method during the modeling process). The results are shown in Table 5.
[0166] Table 5. AUC of the constructed model under different combinations of glmnet algorithm hyperparameters
[0167]
[0168] As can be seen from Table 5, when the hyperparameter combination of the glmnet algorithm is alpha = 1, lambda = 0.0005, the AUC reaches a maximum value of 0.948.
[0169] The equation for building a model based on the optimal hyperparameter combination is:
[0170]
[0171] Where Y is the prediction score, i represents the i-th biomarker, m represents the number of biomarkers (m = 8), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker (see Table 6), and b is a constant of 3.554513.
[0172] Table 6. Coefficients of the eight biomarkers in the model
[0173]
[0174] The complete model equation is:
[0175] Y=4.845IGFBP1+7.675SFTPB1+1.827MGAM2+4.638MUC16+6.259SELL+4.924RNASE1+8.936S100B+8.727TFF1+3.554513
[0176] Determination of diagnostic threshold for bladder cancer recurrence risk prediction model:
[0177] ## Setting levels: control = case, case = control
[0178] ## Setting direction: controls < cases
[0179] The ROC curve was drawn with the prediction scores in the training group, and the optimal prediction cutoff value was set to 0.4720534 according to the Youden index value. The results are shown in Figure 4 .
[0180] When the prediction score of the prediction model is ≤0.4720534, the subject is considered to have a low risk of bladder cancer recurrence.
[0181] When the prediction score of the prediction model is greater than 0.4720534, the subject is considered to have a high risk of bladder cancer recurrence.
[0182] from Figure 4 It can be seen that the prediction model has an AUC of 0.948, a sensitivity of 0.914, and a specificity of 0.878 in the training group.
[0183] 2.6 Validation of the Bladder Cancer Recurrence Risk Prediction Model
[0184] The optimal model constructed was verified in the test group and the ROC curve was drawn as follows Figure 5 shown.
[0185] from Figure 5 It can be seen that the AUC of the model in the test group is 0.928, the sensitivity is 0.886, and the specificity is 0.830.
[0186] In summary, the bladder cancer recurrence risk prediction model constructed using 8 protein markers has good predictive performance and accuracy and the best diagnostic efficacy.
[0187] Example 3: Construction and verification of a three-category diagnostic model
[0188] This example attempts to construct a three-category combined diagnostic model for distinguishing between the bladder cancer prognosis non-recurrence group, the bladder cancer prognosis in situ recurrence group, and the bladder cancer prognosis distant metastasis group. The specific process includes the following: (1) construction and screening of the optimal diagnostic model; (2) verification of the effectiveness of the optimal diagnostic model. The specific screening process and results are as follows (in the present invention, the two-category model in Example 2 uses the AUC value as the evaluation indicator; when a three-category model is constructed, since multiple categories are involved, the AUC value is generally not applicable. In this example, indicators such as sensitivity, specificity, accuracy, and consistency are used to measure the diagnostic efficacy of the model):
[0189] 3.1 Construction and screening of the three-category diagnostic model
[0190] For the test cohort of 280 bladder cancer patients, all enrolled patients signed informed consent. Among them, 48 patients relapsed within two years after surgery (20 of them had local recurrence in situ and 28 had metastasis). They were divided into two groups: 232 patient samples without recurrence were classified as the low-risk group for bladder cancer recurrence, and 48 patients with recurrence or metastasis were classified as the high-risk group for bladder cancer recurrence. They were randomly divided into a test group and a validation group. The training group included 116 patients in the low-risk group for bladder cancer recurrence and 24 patients in the high-risk group for bladder cancer recurrence (including 10 patients with in situ recurrence and 14 patients with metastasis); the test group included 116 patients in the low-risk group for bladder cancer recurrence and 24 patients in the high-risk group for bladder cancer recurrence (including 10 patients with in situ recurrence and 14 patients with metastasis).
[0191] This example further constructs a three-category detection model based on the marker combination of IGFBP1+SFTPB+MGAM2+MUC16+SELL+RNASE1+S100B+TFF1 screened in Example 2, capable of effectively distinguishing bladder cancer patients with no recurrence (low-risk group), those with in situ recurrence (in situ recurrence group), and those with distant metastasis (metastasis group). All enrolled patients provided informed consent. Bladder cancer patients were diagnosed by histopathology, and inclusion criteria included: (a) no history of other malignancies; and (b) no concurrent malignancies or autoimmune diseases.
[0192] The collected serum samples were subjected to LC-MS / MS data acquisition and detection to obtain the concentrations of eight protein markers, including IGFBP1, SFTPB+MGAM2+MUC16+SELL+RNASE1+S100B+TFF1.
[0193] The Shapiro-Wilk test was used to assess normal distribution, and the nonparametric Wilcoxon test was used to analyze differences in blood marker concentrations between the bladder cancer prognosis non-recurrence group (low-risk group), the bladder cancer prognosis in situ recurrence group (in situ recurrence group), and the bladder cancer prognosis distant metastasis group (metastasis group). An eight-marker three-classification combined diagnostic model was constructed using a combination of machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using the predicted probability values with 95% confidence intervals (CI) to assess the discriminatory ability of the multivariate diagnostic model. Using the test set, the Youden index (YI) was calculated to determine the predicted probability cutoff value for distinguishing the bladder cancer prognosis non-recurrence group (low-risk group), the bladder cancer prognosis in situ recurrence group (in situ recurrence group), and the bladder cancer prognosis distant metastasis group (metastasis group). In addition, the ROC values of the models constructed with different marker combinations were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV) and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R3.6.1, and a p value of less than 0.05 was considered statistically significant.
[0194] In this embodiment, in order to construct the optimal three-class joint diagnosis model, after comparing the six algorithms of gradient boosting, naive Bayes, support vector machine, neural network, generalized linear, and discriminant analysis, the gradient boosting method was selected as the best supervised classification algorithm for constructing the prediction model. The grid search range for hyperparameter optimization of the gradient boosting method model is shown in Table 7 below.
[0195] Table 7. Parameter grid search range of gradient boosting method
[0196]
[0197] Through optimization screening in terms of accuracy, consistency, sensitivity, specificity, etc., the optimal parameter combination mode was determined to be: interaction.depth 1, n.trees 100, shrinkage 0.1, n.minobsinnode 10.
[0198] The training and test groups used two completely different batches of samples. This example only screened markers and constructed models from the training group; the test group samples were only used to verify the diagnostic efficacy of the model. The specific results are shown in Table 8.
[0199] Table 8. Performance evaluation table of the gradient boosting method to build a model to distinguish three categories
[0200]
[0201] Table 8 shows that the gradient boosting model constructed based on eight protein markers (IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF1) can be used to predict whether bladder cancer patients will experience no recurrence, in situ recurrence, or distant metastasis after surgery. This also demonstrates that the protein markers screened by this invention can be used to distinguish the risk of recurrence in bladder cancer patients after surgical treatment and, when the risk of recurrence is high, to distinguish between in situ recurrence and distant metastasis (patients with both in situ recurrence and metastasis are also included in the metastasis group).
[0202] 3.2 Combined performance of the three-class joint diagnosis model
[0203] To further enhance the diagnostic value of the three-category diagnostic model (gradient boosting) constructed using different protein combinations of biomarkers, this example compared the performance of diagnostic models constructed using different protein combinations of biomarkers in a test group based on the 11 protein markers screened in Example 1. The specific combinations of the different models are shown in Table 9.
[0204] Table 9. Combinations of different diagnostic models
[0205]
[0206] The performance index comparison results of different diagnostic models constructed by the 11 biomarkers screened in Example 1 are shown. The calculation method of the minimum value, first quartile, median, mean, third quartile and maximum value of accuracy and consistency is as follows: (1) Sort the accuracy or consistency values from small to large; (2) Minimum value: the first value after sorting; (3) First quartile (Q1): multiply the number of data by 0.25. If the result is an integer, take the average of the values at this position and the next position; if it is not an integer, round up to get the position, and the value at this position is Q1; (4) Median: if the number of data is odd, the median is the middle value; if it is even, it is the average of the two middle values; (5) Mean: the sum of all values divided by the number of data; (6) Third quartile (Q3): multiply the number of data by 0.75 and process it in the same way as Q1; (7) Maximum value: the last value after sorting. The minimum and maximum values reflect data extremes, demonstrating the worst and best possible model performance. Quartiles help understand the data's distribution and dispersion. Q1 and below indicate lower performance, while Q3 and above indicate higher performance. The median reflects intermediate performance, and the mean comprehensively reflects the overall average performance. By combining these statistical values, we can gain a comprehensive understanding of the overall performance, distribution characteristics, and stability of the model, providing a strong basis for model selection and optimization.
[0207] Table 10. Performance comparison of diagnostic models based on different protein combination biomarkers
[0208]
[0209] As shown in Table 10, the combined 10-marker model performed best for the three-category diagnostic model. This clearly demonstrates that, in addition to the eight markers (IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, and TFF11), the addition of ORM1 and FBLN5 significantly improved the diagnostic efficacy of differentiating postoperative bladder cancer recurrence, in situ recurrence, and distant metastasis. However, further improvement in accuracy and consistency was difficult when the number of markers was increased to 11. Therefore, the three-category gradient boosting model constructed using these 10 protein markers (IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, ORM1, and FBLN5) was selected as the optimal combined diagnostic model.
[0210] 3.2 Diagnostic Performance Measurement and Verification of the Three-Classification Joint Diagnosis Model
[0211] 1) Diagnostic performance measurement of the three-classification combined diagnosis model
[0212] In order to more accurately determine the diagnostic performance and threshold of the model constructed in this embodiment for different disease classifications, a multi-classification model of the gradient boosting (GBM) algorithm was used to perform predictive analysis in the training group, and the predicted results were calculated as the predicted probability values of three categories (bladder cancer surgical prognosis: no recurrence, in situ recurrence, or distant metastasis). The category with the largest predicted probability value was the final prediction result of the system.
[0213] The meaning and calculation method of each indicator are as follows:
[0214] Results: The three-class combined diagnosis model achieved an accuracy of 0.92 and a consistency of 0.91 in the training group. The diagnostic sensitivity for patients with non-recurrence after bladder cancer surgery was 93.2%, and the specificity was 94.6%. The diagnostic sensitivity for patients with in situ recurrence after bladder cancer surgery was 86.5%, and the specificity was 89.6%. The diagnostic sensitivity for patients with distant metastasis after bladder cancer surgery was 89.7%, and the specificity was 90.1%.
[0215] It should be noted that the three-classification joint diagnosis model constructed by gradient boosting is a model constructed by machine learning and cannot fit a specific equation formula like a generalized linear model.
[0216] 2) Validation of the three-classification combined diagnosis model
[0217] The prediction performance of the model built based on the test group was verified in the training group. The specific results are as follows:
[0218] The accuracy was 0.90 and the consistency was 0.89. The sensitivity for diagnosing non-recurrence after bladder cancer surgery was 93.1% and the specificity was 93.8%; the sensitivity for diagnosing in situ recurrence after bladder cancer surgery was 91.6% and the specificity was 91.5%; and the sensitivity for diagnosing distant metastasis after bladder cancer surgery was 88.1% and the specificity was 89.4%.
[0219] In summary, the three-category combined diagnostic model containing 10 protein markers constructed in this example has good diagnostic value for the three categories of bladder cancer non-recurrence group after surgery, bladder cancer in situ recurrence group after surgery, and bladder cancer distant metastasis group after surgery.
[0220] All patents and publications cited in this specification are intended to indicate that they are state of the art and that the present invention may be used. All patents and publications cited herein are incorporated by reference in their entirety, as if each publication were specifically incorporated by reference. The invention described herein may be practiced in the absence of any element or elements, limitation or limitations, unless otherwise specified. For example, in each instance, the terms "comprising," "consisting essentially of," and "consisting of" may be replaced with either of the other two terms. The term "a" or "an" herein simply means "one" and does not exclude the inclusion of only one or more. The terms and expressions used herein are intended to be descriptive, not limiting, and are not intended to exclude any equivalent features. However, it is understood that any suitable changes or modifications may be made within the scope of the present invention and the appended claims. It is understood that the embodiments described herein are preferred embodiments and features, and that modifications and variations can be made by persons of ordinary skill in the art based on the spirit of the present invention. Such modifications and variations are considered to be within the scope of the present invention and the scope of the independent and appended claims.
Claims
1. Use of a reagent for detecting a marker for preparing a reagent for predicting whether a bladder cancer patient will not relapse within two years after surgery, whether the patient will relapse in situ or have distant metastasis, characterized in that: The markers consist of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, ORM1 and FBLN5.
2. A kit for predicting whether a bladder cancer patient will not relapse, relapse in situ, or metastasize within two years after surgery, characterized in that: The kit comprises a detection reagent for the biomarker for use as claimed in claim 1.
3. A biomarker combination for predicting whether a bladder cancer patient will not relapse, relapse in situ, or metastasize within two years after surgery, characterized in that: The panel consists of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, ORM1, and FBLN5.
4. A system for predicting whether bladder cancer will not recur, recur in situ, or metastasize within two years after surgery, characterized in that: The system includes a data analysis module for analyzing the detection values of markers, wherein the markers consist of IGFBP1, SFTPB, MGAM2, MUC16, SELL, RNASE1, S100B, TFF1, ORM1 and FBLN5.
5. The system according to claim 4, wherein: The data analysis module uses the detection values of markers of known samples as a training set, divides bladder cancer patients into a non-recurrence group, an in situ recurrence group, and a distant metastasis group according to their post-operative conditions, analyzes the relationship between the detection values of the non-recurrence group, the in situ recurrence group, and the distant metastasis group, and constructs a model; the model is constructed based on a gradient boosting algorithm; the system also includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output prediction results.
Citation Information
Patent Citations
Biomarker combination and application thereof in prediction and / or diagnosis of colorectal cancer
CN117587128A
Biomarker combination and application thereof in predicting bladder cancer risk
CN119410773A