Biomarkers for detecting ovarian cancer and their applications
By screening out biomarkers such as CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, and FABP4, an ovarian cancer detection system was constructed, which solved the problem of insufficient sensitivity and specificity in the early diagnosis of ovarian cancer in existing technologies, and achieved efficient and accurate prediction and staging diagnosis of ovarian cancer.
Patent Information
- Application Number
- CN202510970187.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2025-06-03
- Filing Date
- 2025-07-15
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-07-15
AI Technical Summary
Existing ovarian cancer screening tools, such as CA125 and transvaginal ultrasound, have insufficient sensitivity and specificity in early diagnosis, making it difficult to effectively distinguish between ovarian cancer and benign ovarian diseases. This leads to difficulties in early diagnosis of ovarian cancer and a high mortality rate.
A group of biomarkers, including CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, and FABP4, were screened using proteomics technology to construct an ovarian cancer detection system. The expression levels of these biomarkers were detected through blood samples, and predictions were made using a data analysis module. A model equation was established to perform non-invasive, convenient, and efficient prediction of ovarian cancer.
It achieves early, non-invasive, convenient, and efficient prediction of ovarian cancer. The model has high sensitivity and specificity, can effectively distinguish between benign and malignant ovarian cancer, and can differentiate between different clinical stages, with high accuracy and discriminative ability.
Smart Images

Figure CN120761643B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical testing technology, specifically involving the use of proteomics to screen biomarkers for ovarian cancer, for the detection of ovarian cancer, and to differentiate between benign and malignant ovarian tumors. Background Technology
[0002] Proteomics is the science that studies the composition, location, changes, and interactions of proteins in cells, tissues, or organisms, including the study of protein expression patterns and proteome functional patterns. With the development of mass spectrometry, liquid chromatography-mass spectrometry (LC-MS / MS) has become the most important tool in proteomics research. The advancements in proteomics are significant for identifying diagnostic markers for diseases, screening drug targets, and conducting toxicological studies, and are therefore widely used in medical research.
[0003] Ovarian cancer (OC) is a leading cause of cancer death in women and has the worst prognosis among gynecological malignancies. The mortality rate for advanced ovarian cancer is over three-quarters. Currently, the tumor is confined to the ovary and can be cured with cytoreductive surgery and chemotherapy; however, due to the lack of obvious early symptoms, diagnosis of early-stage localized cancer is extremely difficult.
[0004] The most widely used screening tool currently is a combination of cancer antigen 125 blood testing and transvaginal ultrasound. However, neither multimodal screening nor transvaginal ultrasound screening has significantly reduced ovarian cancer mortality.
[0005] While the protein biomarker CA125 exhibits high sensitivity (approximately 79%), its specificity is insufficient for early diagnosis, failing to accurately differentiate ovarian cancer from benign ovarian diseases. HE4, a protease inhibitor, can improve diagnostic accuracy when used in combination with CA125, demonstrating higher specificity, particularly in early ovarian cancer diagnosis; however, its use alone still has limitations. p53 is overexpressed in over 96% of high-grade serous ovarian cancers and is an important prognostic marker, but its sensitivity as a diagnostic marker is low. Several other biomarkers have been disclosed in the prior art, showing some value in the diagnosis and prognosis of ovarian cancer, but all suffer from insufficient sensitivity or specificity.
[0006] Therefore, finding new biomarkers and their combinations to detect emerging diseases as early as possible, detect ovarian tumors and differentiate between benign and malignant tumors, and construct predictive models for benign and malignant ovarian cancer have important clinical value. Summary of the Invention
[0007] To address the problems existing in the prior art, this invention provides a biomarker, detection reagent, detection kit, and ovarian cancer risk prediction system for ovarian cancer detection. This invention screens out a series of novel biomarkers that can predict the risk of ovarian cancer at an early stage and can distinguish between benign and malignant ovarian tumors.
[0008] The technical solution adopted in this invention is:
[0009] The application of biomarkers in the preparation of ovarian cancer detection products, wherein the biomarkers are selected from one or more of the following: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4.
[0010] Preferably, the biomarker is one of the following combinations:
[0011] (a) Combinations of FABP4, ORM2, CTSG, TALDO1, and TFF1:
[0012] (ii) Combinations of TFF1, FABP4, ORM2, TALDO1, PGM5, and SFRP1:
[0013] (iii) Combinations of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, and SFRP1.
[0014] More preferably, it is a combination of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1 and FABP4.
[0015] The ovarian cancer detection product can be an ovarian cancer detection reagent or an ovarian cancer detection kit.
[0016] The present invention also provides a kit for ovarian cancer detection, the kit comprising reagents for detecting biomarkers, the biomarkers being selected from one or more of the following: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4;
[0017] Preferably, the biomarker is one of the following combinations:
[0018] (a) Combinations of FABP4, ORM2, CTSG, TALDO1, and TFF1:
[0019] (ii) Combinations of TFF1, FABP4, ORM2, TALDO1, PGM5, and SFRP1:
[0020] (iii) Combinations of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, and SFRP1.
[0021] More preferably, it is a combination of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1 and FABP4.
[0022] Furthermore, the reagents for detecting biomarkers can detect the expression level of biomarkers. The reagents for detecting biomarkers can be sample pretreatment reagents, antigens or antibodies, or other biological reagents and kits suitable for the detection of the biomarkers; they can also be developed into standardized reagents or kits suitable for the detection of the biomarkers by LC-UV or LC-MS.
[0023] In some methods, the reagent for detecting the biomarker is an antibody that is a biomarker as described above, and the antibody is a monoclonal antibody.
[0024] Furthermore, the CTSG is cathepsin G, which is a protein or amino acid sequence with the UniProt database number P08311;
[0025] PGM5 is a phosphoglucoside mutase 5, and is a protein or amino acid sequence with the UniProt database number Q15124.
[0026] The ORM2 is α-1-acidic glycoprotein 2, which is the protein or amino acid sequence with the UniProt database number P19652;
[0027] The TFF1 is Trefoil factor 1, which is a protein or amino acid sequence with the UniProt database number P04155;
[0028] The SFRP1 is a secreted Fritz-associated protein 1, which is a protein or amino acid sequence with the UniProt database number Q8N474;
[0029] The TALDO1 is transaldolase 1, which is the protein or amino acid sequence with UniProt database number P37837;
[0030] The FABP4 is fatty acid binding protein 4, which is a protein or amino acid sequence with the UniProt database number P15090.
[0031] Furthermore, the reagent is used to detect biomarkers in body fluid samples, which include any one of blood, urine, saliva, and sweat.
[0032] In some preferred embodiments, the biomarkers of the present invention are obtained through screening of blood samples, and are particularly suitable for development into blood test reagents or kits for ovarian cancer prediction.
[0033] Furthermore, biomarker detection refers to detecting the presence, relative abundance, or concentration of biomarkers in an individual's bodily fluid samples.
[0034] In some methods, relative abundance is preferred, which is the peak area of the biomarker in the detection chromatogram obtained by high performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a biomarker measured in a control sample (an individual without ovarian cancer) is 500, and the average peak area measured in an ovarian cancer sample is 3000, then the abundance of the biomarker in the ovarian cancer sample is considered to be 6 times that in the control sample.
[0035] This invention also provides a system for detecting ovarian cancer, the system including a data analysis module for analyzing the detection values of biomarkers in samples from patients to be tested, wherein the biomarkers are selected from one or more of the following: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4; preferably, the biomarkers are one of the following combinations:
[0036] (a) Combinations of FABP4, ORM2, CTSG, TALDO1, and TFF1:
[0037] (ii) Combinations of TFF1, FABP4, ORM2, TALDO1, PGM5, and SFRP1:
[0038] (iii) Combinations of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, and SFRP1.
[0039] More preferably, it is a combination of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1 and FABP4.
[0040] The data analysis module includes an analysis model equation, which is as follows:
[0041]
[0042] logit(Y) = log(Y / 1 - Y), the formula for calculating Y is as follows:
[0043]
[0044] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m=7), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant -4.536.
[0045] The coefficients of Ki are shown in Table 6 below:
[0046] Table 6: Coefficients of 7 biomarkers in the model
[0047]
[0048] The data analysis module calculates a predictive value for whether the patient under test has ovarian cancer based on the detection value of the biomarker using the analysis model equation, and determines whether the patient under test has ovarian cancer based on the predicted value.
[0049] The determination criteria are as follows:
[0050] When the predicted value is less than or equal to a preset threshold, the patient to be tested is determined not to be an ovarian cancer patient.
[0051] When the predicted value is greater than the preset threshold, the patient to be tested is determined to be an ovarian cancer patient;
[0052] The preset threshold is 0.524.
[0053] Furthermore, the system also includes a data detection module, a data input module, and a data output module; the data detection module is used to detect biomarkers in the sample and obtain detection values; the data input module is used to input the detection values of the biomarkers, and after the data analysis module analyzes the detection values, the data output interface is used to analyze whether the patient under test has ovarian cancer.
[0054] The detection value is generally obtained by performing enzyme-linked immunosorbent assay (ELISA) on the sample to obtain the concentration of biomarkers in the sample, which is used as the detection value in μg / mL.
[0055] The present invention also provides the application of biomarkers in the preparation of clinical staging diagnostic products for ovarian cancer, wherein the biomarkers are a combination of TFF1, FABP4, ORM2, TALDO1, PGM5, and SFRP1 or a combination of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, and SFRP1.
[0056] The technical solution of this application has the following beneficial effects:
[0057] (1) This invention employs proteomics technology to conduct in-depth screening of differentially expressed proteins in blood samples from ovarian cancer patients and healthy controls, identifying a set of key biomarkers that can indicate tumor risk in the early stages of ovarian cancer. Based on these biomarkers, an ovarian cancer detection system is constructed, enabling non-invasive, convenient, and efficient prediction of ovarian cancer in patients to be tested. This system not only meets the clinical need for early diagnosis but also has broad prospects for widespread application.
[0058] (2) The ovarian cancer detection model of the present invention has an AUC of 0.953284, a sensitivity of 0.876, a specificity of 0.9, an AUC of 0.92706 in the test group, an accuracy of 0.846, a sensitivity of 0.816, and a specificity of 0.876. It has high accuracy and discrimination ability and can more efficiently predict whether an individual has ovarian cancer.
[0059] (3) The combination of biomarkers of the present invention can construct a clinical staging prediction model for ovarian cancer, which can effectively distinguish between stage I, II, III and IV of ovarian cancer and has the potential to diagnose the clinical staging of ovarian cancer. Attached Figure Description
[0060] Figure 1 Volcano plot for differential analysis of benign and malignant ovarian tumors.
[0061] Figure 2 The graph shows the ROC and OPLS-DA analysis results for benign and malignant ovarian tumors.
[0062] Figure 3 A bar chart comparing the performance AUC of models built for different combinations of markers.
[0063] Figure 4 AUC results for models with different hyperparameters.
[0064] Figure 5 The ROC curve of the combined diagnostic model in the model group.
[0065] Figure 6 The ROC curve of the joint diagnostic model in the test group. Detailed Implementation
[0066] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate understanding of the present invention and are not intended to limit it in any way. The reagents used in this embodiment are all known products and were obtained by purchasing commercially available products.
[0067] It should be noted that:
[0068] (1) Diagnosis or testing
[0069] Diagnosis or testing here refers to the detection or analysis of biomarkers in a sample, or the determination of the content of a target biomarker, such as its absolute or relative content. The presence or quantity of the target biomarker then indicates whether the individual providing the sample may have or suffer from a certain disease, or the likelihood of having a certain disease. The meanings of diagnosis and testing are interchangeable here. The results of such testing or diagnosis cannot be directly considered as a direct result of disease; rather, they are intermediate results. To obtain a direct result, further auxiliary methods such as pathology or anatomy are needed to confirm the presence of a certain disease. For example, this invention provides several novel biomarkers associated with ovarian cancer, and changes in the levels of these biomarkers are directly related to the presence or absence of ovarian cancer.
[0070] (2) The link between biomarkers or markers and ovarian cancer
[0071] In this invention, "marker" and "biomarker" have the same meaning. Here, "linkage" refers to a direct correlation between the presence or change in the level of a certain biomarker in a sample and a specific disease; for example, a relative increase or decrease in the level indicates a higher likelihood of having the disease compared to healthy individuals.
[0072] If multiple different biomarkers are present simultaneously in a sample, or if their relative levels change, it indicates a higher likelihood of having the disease compared to healthy individuals. In other words, among biomarkers, some are strongly associated with the disease, some are weakly associated, and some may not even be associated with a specific disease. One or more of the strongly associated biomarkers can be used as diagnostic markers, while weakly associated biomarkers can be combined with strong biomarkers to diagnose a disease, increasing the accuracy of test results.
[0073] The numerous serum biomarkers discovered in this invention can be used to differentiate ovarian cancer from healthy individuals. These biomarkers can be used individually for direct detection or diagnosis, indicating a strong correlation between the relative change in their levels and ovarian cancer. Of course, it is understood that one or more biomarkers strongly associated with ovarian cancer can be detected simultaneously. It is generally understood that in some methods, selecting highly correlated biomarkers for detection or diagnosis can achieve a certain level of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90%, or 95%. This indicates that these biomarkers can provide intermediate values for diagnosing a certain disease, but does not necessarily mean a direct diagnosis of that disease.
[0074] Of course, differentially expressed proteins with higher ROC values can also be selected as diagnostic biomarkers. The terms "strong" and "weak" are generally determined using algorithms, such as biomarker contribution rate to ovarian cancer or weighted analysis. Such calculation methods can include significance analysis (p-value or FDR value) and fold change. Multivariate statistical analyses mainly include principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA), and orthogonal partial least squares discriminant analysis (OPLS-DA), as well as other methods such as ROC analysis. Other model prediction methods are also possible. When specifically selecting biomarkers, the differentially expressed proteins disclosed in this invention can be chosen, or combinations of other existing known biomarkers can be selected and predicted using model methods.
[0075] Example 1: Screening of Ovarian Cancer Biomarkers
[0076] 1. Experimental Design
[0077] This experiment was designed to collect plasma samples from newly diagnosed ovarian tumor patients, enrich low-abundance proteins using immunoaffinity chromatography to remove high-abundance proteins, and detect the protein abundance in the samples using a high-performance liquid chromatography-mass spectrometry tandem device. The differences in protein abundance between patients with benign and malignant ovarian tumors were analyzed, and the diagnostic performance was also analyzed.
[0078] 2. Sample Collection
[0079] Blood samples from 100 patients with benign and malignant ovarian tumors were included. All ovarian tumor patients underwent biopsy and pathological confirmation. Approximately 2 ml of peripheral blood was collected from each subject before chemotherapy, radiotherapy, and surgery. The blood was mixed in a vacuum tube containing EDTA anticoagulant and centrifuged at 120 g at room temperature for 10 minutes. The supernatant was collected, and this process was repeated twice. Then, the centrifuge was performed at 360 g for 20 minutes. Platelet samples were then collected in centrifuge tubes and stored at -80°C for later use.
[0080] 3. Protein sample processing and enzymatic digestion
[0081] First, plasma samples were centrifuged for 15 minutes (15,000g), and the supernatant was collected, filtered, and subjected to immunoaffinity chromatography to remove 14 high-abundance proteins. Then, low-abundance proteins were concentrated to 350 μL using a 3 kDa cutoff concentration tube at 4000 g for 1 hour. The concentrate was recovered, and buffer exchange was performed using a 7 kDa cutoff desalting column at 1000 g for 2 minutes. The replacement buffer was AEX-A (20 mM Tris, 4 M Urea, 3% isopropanol, pH 8.0). Using AEX-A as a blank, protein concentration in the samples was determined using the BCA method. According to sample grouping, 25 mL of TCEP was added to the samples, and incubation was performed at 37°C for 30 minutes for protein reduction. Then, the corresponding TMT 16-plex reagent was added, and incubation was performed at room temperature in the dark for 1 hour for TMT labeling. Subsequently, the sample was buffer-replaced using a Zeba column with AEX-A as the replacement buffer. After mixing the TMT 16-plex labeled sample, 2 mL of AEX-A was added to the mixture, resulting in a final volume of 5.5 mL. The sample was filtered using a 0.22 μm filter, and the TMT 16-plex labeled sample was separated using a 2D-HPLC system. The collected fractions were lyophilized, and finally, a Trypsin-LysinC mixed enzyme was added, and the sample was incubated at 37°C for 5 hours to digest the enzymes. 5 μL of 10% TFA was added to terminate the digestion reaction. A total of 60 digested 2D-HPLC fractions were used for nanoLC-MS / MS analysis.
[0082] 4. LC-MS / MS Data Acquisition
[0083] Each sample was separated using an Easy nLC-1200 high-resolution liquid chromatography system with a flow rate of nanoliters and coupled online to a Q Exactive HF-X mass spectrometer. Mobile phase A was a 0.1% formic acid aqueous solution, and mobile phase B was a 0.1% formic acid-acetonitrile aqueous solution (acetonitrile 80%, water 20%). The chromatographic column consisted of an enrichment column and an analytical column, equilibrated with 100% mobile phase A. Samples were loaded onto the enrichment column (specifications: 100 μm inner diameter (ID), 4 cm length (L), C18 packing material, 3 μm particle size, 100 Å pore size) via an autosampler, and then separated on the analytical column (75 μm inner diameter, 25 cm length, C18 packing material, 3 μm particle size, 100 Å pore size) at a flow rate of 300 nL / min. After chromatographic separation, the samples were analyzed by mass spectrometry using a Q Exactive HF-X mass spectrometer. The detection method is positive ion, with a precursor ion scan range of 350-1800 m / z, a first-order mass spectrometry resolution of 120,000 at 200 m / z, and an AGC target of 3 × 10⁻⁶. 6 The maximum injection time (Maximum IT) was 50 ms, and the dynamic exclusion time (Dynamic exclusion) was 40 s. The mass-charge ratio of peptides and peptide fragments was acquired using a data-dependent acquisition (DDA) method: 20 secondary spectra (MS / MS, MS2 scans) were acquired after each full scan (primary mass spectrometry). The MS2 activation type was high-energy collisional dissociation (HCD), the selection window was 0.7 m / z, the secondary mass spectrometry resolution was 30,000 at 200 m / z, and the automatic gain control target was set to 1 × 10⁻⁶. 5 The maximum injection time is 65 ms, the fixed first term mass is 110.0 m / z, the normalized collision energy is 32 eV, and the minimum automatic gain control target is 2.00 × 10⁻⁶. 4 It excludes ions with 1-valent, 6–8-valent, and >8-valent charges, allowing only single charge states, sets peptide matching as a priority, and enables the exclusion of isotopes.
[0084] 5. Data Preprocessing
[0085] Secondary mass spectrometry (PMS) data were retrieved using Maxquant (v1.6.15.0). The data type was DIA proteomics data based on secondary reporter ion quantification; the secondary spectra used for quantification required the precursor ion to account for more than 75% of the primary spectrum. The database source was the Uniprot database's Homo_sapiens_9606_proteome_gene (release: 2021-10-14, sequence: 20,437), and common contamination libraries were added, with contaminating proteins removed during data analysis. Enzyme digestion was set to Trypsin / P; the number of missed cleavage sites was set to 2; the precursor ion mass error tolerance for First Search and Main Search was set to 20 ppm and 5 ppm, respectively, and the secondary fragment ion mass error tolerance was 20 ppm. Fixed modification was cysteine alkylation, and variable modifications included methionine oxidation and N-terminal acetylation of proteins. The free-dip ion ratio (FDR) for protein identification and PSM identification was set to 1%.
[0086] 6. Difference Analysis
[0087] A combination of univariate and multivariate statistical analyses was used to screen for differentially expressed proteins and transcripts. Univariate analysis primarily included significance analysis (p-value or FDR value) and fold change of characteristic molecules in different groups. Multivariate statistical analyses mainly included receiver operating characteristic (ROC) curve analysis and Boruta feature selection based on the random forest algorithm. All statistical analyses were performed using R; specific R-related information is shown in Table 1.
[0088] Table 1: R used in this invention and related information
[0089]
[0090] Variable Importance for the Projection (VIP) was calculated to measure the influence and explanatory power of each protein's expression pattern on the classification of each group of samples. A Wilcoxon rank-sum test was then performed to obtain the corrected p-value (FDR). Based on the criteria of FDR < 0.01 and Fold change > 2, 59 downregulated and 76 upregulated proteins were obtained (see details). Figure 1 ).
[0091] To evaluate the role of each biomarker in the prediction of ovarian cancer diagnosis, we used ROC and Boruta analyses to evaluate each biomarker. The results are shown in the figure below. Figure 2The x-axis represents the AUC obtained from ROC analysis, and the y-axis represents -log10(FDR) calculated by the Wilcoxon test. The size of the point represents the VIP value obtained from Boruta analysis. Further screening based on VIP>3 and AUC>0.6 identified a total of 7 more significant candidate biomarkers, as detailed in Table 2.
[0092] Table 2: Differential markers for benign and malignant ovarian tumors
[0093]
[0094] In Table 2, a smaller FDR value and / or a larger VIP value indicate, to some extent, a more significant difference in protein between the two groups, and also suggest that the protein may have higher diagnostic value.
[0095] Example 2: A classification model for identifying benign and malignant ovarian tumors using a combination of seven differentially expressed proteins and its establishment.
[0096] While a single biomarker can differentiate between benign and malignant ovarian tumors in serum samples or predict ovarian cancer, generally, combining multiple biomarkers results in higher accuracy in differentiation or prediction. However, a single biomarker with higher predictive accuracy for ovarian cancer may not necessarily have a greater effect in a combination with one or more other biomarkers. Furthermore, a higher number of biomarkers does not necessarily lead to higher predictive accuracy (AUC value) for the combined combination; therefore, extensive validation experiments are still needed. This example studies a model constructed using seven protein biomarkers from serum: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, and FABP4.
[0097] 1. Obtaining data
[0098] Study population:
[0099] A total of 1000 samples were collected, including 500 blood samples from patients with benign ovarian tumors and 500 blood samples from patients with malignant ovarian tumors. All samples were obtained from biopsies and confirmed by pathology. Participants were divided into a model group and a test group at a ratio of 8:2.
[0100] Inclusion criteria for ovarian cancer patients: (a) no history of other malignant tumors, and (b) surgical treatment performed within one month after blood collection, with postoperative pathology confirming ovarian cancer. All collected serum samples were stored in a serum bank at -80°C after obtaining informed consent.
[0101] In this embodiment, the collected serum samples were subjected to enzyme-linked immunosorbent assay (ELISA) to obtain the concentrations of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, and FABP4 in the serum.
[0102] 2. Statistical Analysis of Experimental Data
[0103] In the model cohort, a combined diagnostic model for multiple ovarian cancer biomarkers was constructed using a combination of machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using predicted probability values with a 95% confidence interval (CI) to evaluate the discriminative power of the multivariate diagnostic model. In the test cohort, the Youden index (YI) was calculated to determine the cut-off value for differentiating between primary and metastatic ovarian cancer patients. Furthermore, ROCs for individual biomarkers and different subgroups were constructed and compared. Standard descriptive statistics, such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD), were calculated to describe the experimental results in the study population. Statistical analysis was performed using R 3.6.1, and a p-value less than 0.05 was considered statistically significant.
[0104] 3. Steps for constructing a joint diagnostic model
[0105] S101, randomly select the concentration matrix of 3 to 7 protein biomarkers (CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4) from the samples in the model group as the original training dataset.
[0106] S102, Select the generalized linear model (glmnet) algorithm to construct the prediction model, and define the grid search range during the hyperparameter optimization process. In this step, the grid search range for hyperparameter optimization of each algorithm is set as shown in Table 3.
[0107] Table 3: Parameter grid search range of the glmnet algorithm
[0108]
[0109] S103. Based on the algorithm and hyperparameter setting range set in step S102, select one of the hyperparameter combination methods as the parameters for constructing the prediction model.
[0110] S104. The original dataset is split into K subsets using a K-fold cross-validation mechanism. To ensure that the proportion of majority and minority class samples in each subset is the same as in the original dataset, a stratified K-fold cross-validation mechanism is used for data splitting.
[0111] S105. Based on the K training data subsets obtained from step S104, select one subset as the validation set Ddev.
[0112] S106, Combine the unselected subset of training data from step S105 to form the training data pool Dtrain. S107, Based on the training dataset Dtrain obtained in step S106, construct a prediction model using the selected supervised classification algorithm and hyperparameters.
[0113] S108: Based on the prediction model obtained in step S107, evaluate it on the validation set Ddev to obtain the AUC value, and store the current prognostic prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 involves evaluating the prediction model obtained in step S107 on the validation set determined in the current iteration, and storing both the model and the evaluation result in the prediction model pool for future use in selecting the baseline prediction model. The evaluation mentioned in this step can be the AUC value or other reasonable metrics for evaluating model performance.
[0114] S109, determine if each subset has been used as a validation set. Step S109 checks if all K subsets obtained in step S104 have been used as validation sets for model training. If all subsets have been used as validation sets and training has been completed, proceed to step S110; otherwise, proceed to step S105. This step ensures that every sample in the original dataset has been used as a validation set, improving model stability and preventing the model from overfitting to a particular subset.
[0115] S110: The average AUC of all models in the prediction model pool (Pool) is used as the final performance evaluation value of the model for this combination method. The model parameters and the final performance evaluation AUC value are then stored in the optimal model pool (Poolbest).
[0116] S111: Determine whether all hyperparameter combinations have constructed a prediction model. Step S111 involves determining whether all algorithms and corresponding hyperparameter combinations obtained in step S102 have been used to construct prediction models. If all combinations have completed model construction, proceed to step S112; otherwise, if any combination has not completed model construction, proceed to step S103.
[0117] S113. From the model set Poolbest obtained in step S112, select the model with the largest AUC value as the final predictive model for ovarian cancer diagnosis.
[0118] S114, Repeat all the above steps until all combinations of the markers have been modeled.
[0119] 4. Determining the optimal combination of markers
[0120] By executing the model building steps described above, we obtained the optimal model for all combinations of markers. To compare the performance of these models under different marker combinations, we used the ROC method to evaluate the AUC values of these models in the test group. See Table 4 below. Figure 3 As shown:
[0121] Table 4: Comparison of the area under the ROC curve for models constructed with different combinations of biomarkers
[0122]
[0123] As shown in Table 4, the AUCs of 5MP, 6MP, and 7MP are all greater than 0.7, and the maximum AUC is greater than 0.8, demonstrating good performance. Among them, the model composed of TALDO1+PGM5+TFF1+CTSG+ORM2+FABP4+SFRP1 (7MP) has the highest AUC, which is higher than the mean of the models of other biomarker combinations.
[0124] 5. Optimization results of 7MP model parameters
[0125] Based on the above analysis, we obtained the optimal combination of markers as TALDO1+PGM5+TFF1+CTSG+ORM2+FABP4+SFRP1. Based on this marker combination, we analyzed the models constructed under nine different combinations of glmnet algorithm hyperparameters. Figure 4 The model performance was evaluated using AUC values, as shown in Table 5. Figure 4 As shown: When the hyperparameter combination of the glmnet algorithm is alpha=0.1 and lambda=0.0005, the AUC reaches the maximum value of 0.882 (the AUC is calculated using the 10x cross-validation method during the modeling process).
[0126] Table 5: AUC of the model built under different hyperparameter combinations of the glmnet algorithm
[0127]
[0128] The equations for the model constructed based on the optimal hyperparameter combination are:
[0129]
[0130] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m=7), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker (Table 6), and b is a constant -4.536.
[0131] And logit(Y) = log(Y / 1 - Y), therefore the formula for calculating Y is as follows:
[0132]
[0133] Table 6: Coefficients of 7 biomarkers in the model
[0134]
[0135] 6. Determination of diagnostic thresholds for the combined diagnostic model (7MP) for ovarian cancer
[0136] ROC curves were plotted using the predicted values from the model group, and the optimal diagnostic cutoff value of 0.524 was set based on the Youden index. That is, when the model's predicted value was ≤0.524, the subject was considered not to have ovarian cancer; when the model's predicted value was >0.524, the subject was considered to have ovarian cancer. The results are as follows: Figure 5 As shown, the model's AUC was 0.953284, sensitivity was 0.876, and specificity was 0.9 in the model group.
[0137] 7. Validation of the combined diagnostic model for ovarian cancer (7MP)
[0138] Plot the ROC curve using the predicted values from the test group, such as... Figure 6 As shown, the AUC was 0.92706. The optimal diagnostic cutoff value was set to 0.524 based on the Youden index. That is, when the model's predicted value was ≤0.524, the subject was considered not to have ovarian cancer; when the model's predicted value was >0.524, the subject was considered to have ovarian cancer. The results showed that the model's accuracy was 0.846, sensitivity was 0.816, and specificity was 0.876 in the test group.
[0139] Example 3: Comparison of the diagnostic value of different ovarian cancer diagnostic models
[0140] Table 7: Comparison of Area Under the ROC Curve for Different Diagnostic Models
[0141]
[0142] As shown in Table 7, the AUC of our model (7MP) was 0.334 and 0.349 higher than that of traditional single biomarkers, respectively. The DeLong's test, a method for testing the significance of AUC differences, showed that the diagnostic value of our model (7MP) was significantly higher (p<0.05) than that of traditional biomarkers or combinations of traditional biomarkers.
[0143] Example 4: Constructing a Clinical Staging Prediction Model for Ovarian Cancer
[0144] This embodiment attempts to construct a clinical staging diagnostic model for ovarian cancer based on seven protein biomarkers in serum: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, and FABP4. The aim is to use these biomarkers to detect and diagnose the clinical staging of ovarian cancer.
[0145] One hundred patients clinically diagnosed with stage I, II, III, and IV ovarian cancer were collected, with 45, 51, 63, and 41 cases in stage I, II, III, and IV, respectively. All samples were obtained from pathologically confirmed biopsies. The participants were divided into a model group and a test group at an 8:2 ratio.
[0146] The data from the model group is used to construct a classification model using machine learning algorithms, and the model satisfies:
[0147]
[0148] i represents the i-th biomarker, m represents the number of biomarkers (m=7), Xi represents the detection value of the i-th biomarker (μg / mL), and b and Ki are parameters to be optimized.
[0149] P represents the probability of any clinical stage of ovarian cancer, namely stage I, II, III, or IV.
[0150] The model is trained using maximum likelihood estimation or gradient descent, logistic regression or neural network, to obtain the classification model parameters b and Ki for each clinical stage, thereby obtaining the prediction model for each clinical stage.
[0151] A confusion matrix is constructed using the prediction model for each clinical stage, a classification report is generated, an ROC curve is plotted, and an optimal diagnostic cutoff value is set based on the Youden index. That is, when the diagnostic model's predicted value P ≤ the cutoff value, the patient is determined not to belong to that clinical stage; when the model's predicted value > the cutoff value, the patient is determined to belong to that clinical stage.
[0152] Test set data is input into the model for each clinical stage for validation, evaluation, and parameter optimization.
[0153] If a certain metric (such as AUC or accuracy) fails to meet the preset requirements, the model will proceed to iteration.
[0154] Based on the optimal combinations selected in Table 4, the maximum AUC area under the ROC curve of the predictive model for clinical staging constructed for different biomarker combinations is shown in Table 8.
[0155] Table 8: Comparison of the maximum AUC area under the ROC curve for predictive models of clinical staging constructed with different biomarker combinations
[0156]
[0157] Table 8 shows that both 6MP and 7MP have a maximum AUC greater than 0.8, indicating good performance. 6MP showed the best predictive performance in clinical stage I, while 7MP performed best in clinical stages II, III, and IV. Therefore, 6MP was chosen to construct the clinical stage I classification model, and 7MP was chosen to construct the clinical stage II, III, and IV classification models.
[0158] The test set data was input into the models of the four clinical staging processes mentioned above. The AUC area, accuracy, sensitivity, and specificity under the ROC curve are shown in Table 9 below.
[0159] Table 9 Performance Verification Results of the Test Set
[0160]
[0161] This invention utilizes six biomarkers to establish a clinical stage I classification model and seven biomarkers to construct clinical stage II, III, and IV classification models. The AUC is greater than 0.8, and the accuracy, sensitivity, and specificity are all greater than 70%, indicating that the clinical staging model for ovarian cancer of this invention can effectively distinguish between stages I, II, III, and IV of ovarian cancer, and therefore can be used to monitor the progression of ovarian cancer.
Claims
1. The application of biomarkers in the preparation of ovarian cancer detection products, characterized in that... The biomarkers are combinations of TFF1, FABP4, ORM2, TALDO1, PGM5, and SFRP1, or combinations of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, and SFRP1.
2. A reagent kit for ovarian cancer detection, characterized in that... The kit includes reagents for detecting biomarkers, which are combinations of TFF1, FABP4, ORM2, TALDO1, PGM5, and SFRP1, or combinations of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, and SFRP1.
3. A system for detecting ovarian cancer, characterized in that... The system includes a data analysis module, which is used to analyze the detection values of biomarkers in the samples of the patients to be tested. The biomarkers are a combination of TFF1, FABP4, ORM2, TALDO1, PGM5, and SFRP1 or a combination of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, and SFRP1.
4. The system for detecting ovarian cancer as described in claim 3, characterized in that... The biomarkers are a combination of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, and FABP4.
5. The system for detecting ovarian cancer as described in claim 4, characterized in that... The data analysis module includes an analysis model equation, which is as follows: logit(Y) = log(Y / 1 - Y), the formula for calculating Y is as follows: Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers, m=7, Xi represents the detection value of the i-th biomarker in μg / mL, Ki represents the coefficient of the i-th biomarker, and b is a constant -4.536; The coefficients for each biomarker are as follows: The coefficients for TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, and SFRP1 are 1.407, 1.748, 7.999, 8.674, 6.031, 1.735, and 7.516 respectively.
6. The system for detecting ovarian cancer as described in claim 5, characterized in that... The data analysis module calculates a predictive value for whether the patient under test has ovarian cancer based on the detection value of the biomarker using the analysis model equation, and determines whether the patient under test has ovarian cancer based on the predicted value. The judgment criteria are: When the predicted value is less than or equal to a preset threshold, the patient to be tested is determined not to be an ovarian cancer patient. When the predicted value is greater than the preset threshold, the patient to be tested is determined to be an ovarian cancer patient; The preset threshold is 0.
524.
7. The system for detecting ovarian cancer as described in claim 5, characterized in that... The system includes a data detection module, a data input module, and a data output module. The data detection module is used to detect biomarkers in samples and obtain detection values. The data input module is used to input the detection values of biomarkers. After the data analysis module analyzes the detection values, the data output interface is used to output the analysis results of whether the patient under test has ovarian cancer.
8. The application of biomarkers in the preparation of clinical staging diagnostic products for ovarian cancer, characterized in that... The biomarkers are a combination of TFF1, FABP4, ORM2, TALDO1, PGM5, and SFRP1, or a combination of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, and SFRP1.
Citation Information
Patent Citations
Device for predicting chemosensitivity of ovarian cancer patient based on protein markers, construction method of protein marker classifier and application
CN118315068A
Composition for predicting recurrence rate of cancer or survival rate in ovarian cancer patients
KR102316178B1