Biomarker for detecting ovarian cancer and application thereof

By screening the combination of biomarkers such as CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4, etc., an ovarian cancer detection system was built, which solved the problem of insufficient early diagnosis sensitivity and specificity of ovarian cancer in the existing technology, and achieved efficient and accurate early prediction of ovarian cancer.

CN120254286AInactive Publication Date: 2025-07-04HANGZHOU GUANGKEANDE BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510725700.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The lack of sensitivity and specificity in the prior art has led to difficulties in early diagnosis of ovarian cancer, and existing screening tools have failed to significantly reduce mortality.

Method used

The biomarker combinations of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4 and other biomarker combinations were used to screen differentially expressed proteins through LC-MS/MS technology to construct an ovarian cancer detection system, and the biomarker concentration was obtained using enzyme-linked immunosorbent assay (ELISA) and predicted in combination with the data analysis module.

Benefits of technology

A non-invasive, convenient and efficient early prediction of ovarian cancer was achieved. The detection model AUC was 0.953284, sensitivity was 0.876, specificity was 0.9, and accuracy was 0.846, which significantly improved the accuracy of early diagnosis of ovarian cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120254286A_ABST
    Figure CN120254286A_ABST
Patent Text Reader

Abstract

The invention discloses an application of a biomarker in preparation of an ovarian cancer detection product, the biomarker is selected from one or more of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1 and FABP4, and discloses a system and a kit for detecting ovarian cancer, which comprise the biomarker. The biomarkers capable of predicting the tumor occurrence risk in the early stage of ovarian cancer are screened out, an ovarian cancer detection system is constructed based on the markers, the benign and malignant states of the ovarian tumor of a patient to be detected can be noninvasively, conveniently and efficiently predicted, high accuracy, sensitivity and specificity are achieved, clinical detection requirements are met, and the application prospect is wide. And the method has a relatively great application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical detection, and specifically relates to screening biomarkers for ovarian cancer by proteomics for detecting ovarian cancer and distinguishing between benign and malignant ovarian tumors. Background Art

[0002] Proteomics is the science that studies the protein composition, localization, changes and their interaction rules in cells, tissues or organisms, including the study of protein expression patterns and proteome function patterns. With the development of mass spectrometry technology, liquid chromatography-tandem mass spectrometry (LC-MS / MS) has become the most important tool in proteomics research. The development of proteomics is of great significance for finding disease diagnostic markers, screening drug targets, toxicology research, etc., and has thus been widely applied in medical research.

[0003] Ovarian cancer (OC) is the main cause of death among female cancers and also has the worst prognosis among gynecological malignancies. The mortality rate of patients with advanced ovarian cancer is as high as over three-quarters. Currently, the tumor is still confined to the ovaries and can be cured by debulking surgery and chemotherapy, but due to the lack of obvious early symptoms, the diagnosis of local early cancer is very difficult.

[0004] Currently, the most widely used screening tools are the combination of cancer antigen 125 blood test and transvaginal ultrasound examination. However, neither multimodal screening nor transvaginal ultrasound screening methods have significantly reduced the mortality rate of ovarian cancer.

[0005] Although the protein biomarker CA125 has a relatively high sensitivity (about 79%), it lacks specificity in early diagnosis and cannot accurately distinguish ovarian cancer from benign ovarian diseases. HE4 is a protease inhibitor, and its combined use with CA125 can improve the diagnostic accuracy, especially showing higher specificity in the diagnosis of early ovarian cancer, but there are still certain limitations when used alone. p53 is overexpressed in more than 96% of high-grade serous ovarian cancers and is an important prognostic biomarker, but its sensitivity as a diagnostic biomarker is low. The prior art has also disclosed some other biomarkers, which have certain value in the diagnosis and prognosis of ovarian cancer, but all have problems of insufficient sensitivity or specificity.

[0006] Therefore, finding new biomarkers and their combinations to detect new diseases as early as possible, detect ovarian tumors and distinguish between benign and malignant, and construct a prediction model for benign and malignant ovarian cancer has important clinical value. Summary of the Invention

[0007] In view of the problems existing in the prior art, the present invention provides a biomarker, a detection reagent, a detection kit and an ovarian cancer risk prediction system for ovarian cancer detection. The present invention screens out a series of brand-new biomarkers that can early predict the risk of ovarian cancer occurrence and can distinguish between benign and malignant ovarian tumors.

[0008] The technical solution adopted by the present invention is as follows: The application of a biomarker in the preparation of an ovarian cancer detection product, wherein the biomarker is selected from one or more of the following: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4.

[0009] Preferably, the biomarker is one of the following combinations: (1) The combination of FABP4, ORM2, CTSG, TALDO1, TFF1: (2) The combination of TFF1, FABP4, ORM2, TALDO1, PGM5, SFRP1: (3) The combination of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, SFRP1.

[0010] More preferably, it is the combination of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1 and FABP4.

[0011] The ovarian cancer detection product can be an ovarian cancer detection reagent or an ovarian cancer detection kit.

[0012] The present invention also provides a kit for ovarian cancer detection, the kit includes a reagent for detecting a biomarker, and the biomarker is selected from one or more of the following: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4; Preferably, the biomarker is one of the following combinations: (1) The combination of FABP4, ORM2, CTSG, TALDO1, TFF1: (2) The combination of TFF1, FABP4, ORM2, TALDO1, PGM5, SFRP1: (3) The combination of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, SFRP1.

[0013] More preferably, it is the combination of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1 and FABP4.

[0014] Furthermore, the reagent for detecting the biomarker can detect the expression level of the biomarker. The reagent for detecting the biomarker can be a sample pretreatment reagent, an antigen, an antibody, or other biological reagents and kits suitable for the detection of the biomarker; it can also be developed into a standardized reagent or kit suitable for the LC-UV or LC-MS detection of the biomarker, etc.

[0015] In some embodiments, the reagent for detecting the biomarker is an antibody against the biomarker as described above, and the antibody is a monoclonal antibody.

[0016] Furthermore, the CTSG is cathepsin G, which is a protein or amino acid sequence with the UniProt database number P08311; PGM5 is phosphoglucomutase-like 5, which is a protein or amino acid sequence with the UniProt database number Q15124; The ORM2 is α-1-acid glycoprotein 2, which is a protein or amino acid sequence with the UniProt database number P19652; The TFF1 is trefoil factor 1, which is a protein or amino acid sequence with the UniProt database number P04155; The SFRP1 is secreted frizzled-related protein 1, which is a protein or amino acid sequence with the UniProt database number Q8N474; The TALDO1 is transaldolase 1, which is a protein or amino acid sequence with the UniProt database number P37837; The FABP4 is fatty acid-binding protein 4, which is a protein or amino acid sequence with the UniProt database number P15090.

[0017] Furthermore, the reagent is used to detect the biomarker in a body fluid sample, and the body fluid sample includes any one of blood, urine, saliva, and sweat.

[0018] In some preferred embodiments, the biomarker of the present invention is obtained by screening a blood sample, and is particularly suitable for being developed into a blood detection reagent or kit for ovarian cancer prediction, etc.

[0019] Furthermore, detecting the biomarker refers to detecting the presence, relative abundance, or concentration of the biomarker in an individual's body fluid sample.

[0020] In some ways, it is preferably expressed by relative abundance, and the relative abundance is the peak area of the biomarker in the detection spectrum obtained by high performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a certain biomarker measured in a control sample (an individual without ovarian cancer) is 500, and the average peak area measured in an ovarian cancer sample is 3000, then the abundance of this biomarker in the ovarian cancer sample is considered to be 6 times that in the control sample.

[0021] The present invention also provides a system for detecting ovarian cancer, the system includes a data analysis module, and the data analysis module is used to analyze the detection values of biomarkers in the samples of the patient to be tested. The biomarkers are selected from one or more of the following: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4; preferably, the biomarker is one of the following combinations: (1) The combination of FABP4, ORM2, CTSG, TALDO1, TFF1: (2) The combination of TFF1, FABP4, ORM2, TALDO1, PGM5, SFRP1: (3) The combination of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, SFRP1.

[0022] More preferably, it is the combination of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1 and FABP4.

[0023] The data analysis module includes an analysis model equation, and the equation is as follows:

[0024] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 7), Xi represents the detection value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker, and b is a constant 4.536.

[0025] The coefficients of Ki are shown in Table 6 below: Table 6: Coefficients of 7 Biomarkers in the Model

[0026] The data analysis module calculates the predicted value of whether the patient to be tested has ovarian cancer by using the analysis model equation according to the detection values of the biomarkers, and determines whether the patient to be tested has ovarian cancer based on the predicted value.

[0027] The determination condition is: When the predicted value ≤ the preset threshold, it is determined that the patient to be tested is not an ovarian cancer patient; When the predicted value > the preset threshold, it is determined that the patient to be tested has ovarian cancer; Among them, the preset threshold is 0.524.

[0028] Furthermore, the system further includes a data detection module, a data input module, and a data output module; the data detection module is used to detect the biomarker in the sample to obtain a detection value; the data input module is used to input the detection value of the biomarker. After the detection value is analyzed by the data analysis module, the data output interface is used for the analysis result of whether the patient to be tested has ovarian cancer.

[0029] The detection value is generally obtained by performing an enzyme-linked immunosorbent assay (ELISA) on the sample to obtain the concentration of the biomarker in the sample as the detection value, with the unit of μg / mL.

[0030] The technical solution of this application has the following beneficial effects: (1) The present invention uses proteomics technology to deeply screen the differentially expressed proteins in the blood samples of ovarian cancer patients and healthy controls, and screens out a group of key biomarkers that can indicate tumor risk in the early stage of ovarian cancer; based on these biomarkers, an ovarian cancer detection system is constructed to achieve non-invasive, convenient and efficient prediction of ovarian cancer in patients to be tested. This system not only meets the clinical needs for early diagnosis, but also has a broad application prospect for promotion.

[0031] (2) The ovarian cancer detection model of the present invention has an AUC of 0.953284, a sensitivity of 0.876, and a specificity of 0.9. In the test group, the AUC is 0.92706, the accuracy is 0.846, the sensitivity is 0.816, and the specificity is 0.876, which has high accuracy and discrimination ability and can more efficiently predict whether an individual has ovarian cancer. Description of the Drawings

[0032] Figure 1 It is a volcano plot for the differential analysis of the benign and malignant of ovarian tumors.

[0033] Figure 2 It is a graph of the ROC and OPLS-DA analysis results of the benign and malignant of ovarian tumors.

[0034] Figure 3 It is a columnar comparison chart of the performance AUC of the models constructed by different biomarker combinations.

[0035] Figure 4 It is a graph of the AUC results of the models constructed with different hyperparameters.

[0036] Figure 5 It is an ROC curve graph of the combined diagnosis model in the model group.

[0037] Figure 6ROC curve of the combined diagnostic model in the test group. Specific implementation mode

[0038] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not limit it in any way. The reagents used in this embodiment are all known products and are obtained by purchasing commercially available products.

[0039] It should be noted that: (1) Diagnosis or detection The diagnosis or detection here refers to the detection or assay of biomarkers in a sample, or the content of the target biomarker, such as the absolute content or relative content, and then it is explained whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease, based on the presence or quantity of the target biomarker. The meanings of diagnosis and detection here can be interchanged. The result of this detection or the result of the diagnosis cannot be directly used as the direct result of having a disease, but is an intermediate result. If a direct result is to be obtained, other auxiliary means such as pathology or anatomy are still required to confirm having a certain disease. For example, the present invention provides a variety of new biomarkers associated with ovarian cancer, and the changes in the content of these biomarkers are directly related to whether one has ovarian cancer.

[0040] (2) The connection between the biomarker or biomarker and ovarian cancer The biomarker and the biomarker have the same meaning in the present invention. The connection here means that the appearance or the change in the content of a certain biomarker in a sample is directly related to a specific disease. For example, a relative increase or decrease in the content indicates that the possibility of having this disease is relatively higher than that of healthy individuals.

[0041] If multiple different biomarkers in a sample appear simultaneously or there are relative changes in their contents, it also indicates that the possibility of having this disease is relatively higher than that of healthy individuals. That is to say, among the types of biomarkers, some biomarkers have a strong correlation with the disease, some biomarkers have a weak correlation with the disease, or some are even not related to a certain specific disease. One or more of those biomarkers with a strong correlation can be used as biomarkers for diagnosing the disease, and those biomarkers with a weak correlation can be combined with the strong ones to diagnose a certain disease, increasing the accuracy of the detection result.

[0042] For the numerous biomarkers found in the serum of the present invention, these biomarkers can all be used to distinguish ovarian cancer from healthy individuals. These biomarkers can be directly detected or diagnosed as individual biomarkers alone. Selecting such a biomarker indicates that the relative change in the content of this biomarker has a strong correlation with ovarian cancer. Of course, it can be understood that one or more biomarkers with strong correlation with ovarian cancer can be selected for simultaneous detection. Normally understood, in some ways, selecting biomarkers with strong correlation for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90% or 95% accuracy. This indicates that these biomarkers can obtain an intermediate value for diagnosing a certain disease, but it does not mean that it can directly confirm the presence of a certain disease.

[0043] Of course, differential proteins with larger ROC values can also be selected as diagnostic biomarkers. The so-called strong or weak is generally calculated and confirmed through some algorithms, such as the contribution rate or weight analysis of the biomarker to ovarian cancer. Such calculation methods can be significance analysis (p-value or FDR value) and fold change, and multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA) and orthogonal partial least squares discriminant analysis (OPLS-DA). Of course, other methods are also included, such as ROC analysis, etc. Of course, other model prediction methods are also possible. When specifically selecting biomarkers, the differential proteins disclosed in the present invention can be selected, or other existing well-known biomarker combinations can be selected or combined for prediction through model methods.

[0044] Example 1 Screening of Ovarian Cancer Biomarkers 1. Experimental Design This experimental design is to collect plasma samples from patients with newly diagnosed ovarian tumors, enrich low-abundance proteins based on the method of removing high-abundance proteins by immunoaffinity chromatography, detect the protein abundance in the samples by high-performance liquid chromatography-tandem mass spectrometry equipment, analyze the differences in its abundance between patients with ovarian benign tumors and ovarian malignant tumors, and analyze its diagnostic performance.

[0045] 2. Sample Collection Blood samples from 100 patients with ovarian benign tumors and ovarian malignant tumors were included. All ovarian tumor patients were confirmed by pathology of the living tissue. Approximately 2 ml of peripheral blood samples of the subjects were collected before chemotherapy, radiotherapy and surgical treatment, placed in a vacuum tube containing EDTA anticoagulant, mixed well, centrifuged at 120 g for 10 minutes at room temperature, the supernatant was taken, and this was repeated twice. Then centrifuged at 360 g for 20 minutes. Then the platelet samples were collected in centrifuge tubes and stored at -80 °C for later use.

[0046] 3. Processing and Enzymolysis of Protein Samples First, the plasma sample is centrifuged on a centrifuge for 15 minutes (15,000 g), and the supernatant is taken, filtered, and then 14 high-abundance proteins are removed by immunoaffinity chromatography. Then, a concentrator with a cut-off molecular weight of 3 kDa is used to concentrate the low abundance to 350 μL on a centrifuge (4000 g, 1 hour). The concentrated solution is recovered, and a desalting column with a cut-off molecular weight of 7 kDa is used to perform buffer exchange on a centrifuge (1000 g, 2 minutes), and the exchange solution is AEX-A (20 mM Tris, 4 M urea, 3% isopropanol, pH 8.0). Using AEX-A as a blank, the protein concentration in the sample is determined by the BCA method. According to the sample grouping, 25 μL of TCEP is added to the sample, and the sample is incubated at 37 °C for 30 minutes for protein reduction. Then, the corresponding TMT 16-plex reagent is added, and the sample is incubated in the dark at room temperature for 1 hour for the TMT labeling reaction. Then, the sample is subjected to buffer exchange using a Zeba column, and the exchange solution is AEX-A. After mixing the samples labeled with TMT 16-plex, 2 mL of AEX-A is added to the mixed sample, and the final volume is 5.5 mL. The sample is filtered using a 0.22 μm filter and the TMT 16-plex labeled sample is separated using a 2D-HPLC system. The collected fractions are freeze-dried, and finally, a mixture of Trypsin-LysinC enzymes is added, and the sample is incubated at 37 °C for 5 hours for enzymatic digestion of the sample. 5 μL of 10% TFA is added to terminate the enzymatic digestion reaction. A total of 60 enzymatically digested 2D-HPLC fractions are used for nanoLC-MS / MS analysis.

[0047] 4. LC-MS / MS Data Acquisition Each sample was separated using a nano-flow liquid chromatography system, Easy nLC-1200, and online coupled to a high-resolution mass spectrometer, Q Exactive HF-X. Mobile phase - A was an aqueous solution of 0.1% formic acid, and mobile phase - B was an aqueous solution of 0.1% formic acid in acetonitrile (80% acetonitrile and 20% water). The chromatographic column consisted of a trapping column and an analytical column and was equilibrated with 100% mobile phase - A. The sample was loaded onto the trapping column (specification: inner diameter (ID) 100 μm, length (L) 4 cm, C18 packing material, particle size 3 μm, pore size 100 Å) by an autosampler and separated by the analytical column (inner diameter 75 μm, length 25 cm, C18 packing material, particle size 3 μm, pore size 100 Å) at a flow rate of 300 nL / min. After chromatographic separation, the sample was subjected to mass spectrometry analysis using the Q Exactive HF-X mass spectrometer. The detection mode was positive ion, the parent ion scan range was 350 - 1800 m / z, the resolution of the first-stage mass spectrometry was 120,000 at 200 m / z, the AGC target (automatic gain control target) was 3×10 6 , the maximum injection time (Maximum IT) was 50 ms, and the dynamic exclusion time was 40 s. The mass-to-charge ratios of polypeptides and polypeptide fragments were acquired according to the data-dependent acquisition (DDA) method: 20 MS / MS (MS2 scan) spectra were acquired after each full scan (first-stage mass spectrometry). The MS2 activation type was high-energy collision dissociation (HCD), the selection window was 0.7 m / z, the resolution of the second-stage mass spectrometry was 30,000 at 200 m / z, and the automatic gain control target was set to 1×10 5 , the maximum injection time was 65 ms, the fixed first mass was 110.0 m / z, the normalized collision energy was 32 eV, and the minimum automatic gain control target was 2.00×10 4 , ions with a valence of 1, 6 - 8, and >8 were excluded, only single charge states were allowed, peptide match was set to be prioritized, and the isotope exclusion function was enabled.

[0048] 5. Data preprocessing The MS / MS data was searched using Maxquant (v1.6.15.0). The data type was DIA proteomics data based on MS / MS reporter ion quantification. For the MS / MS spectra used for quantification, it was required that the proportion of precursor ions in the MS1 spectra was greater than 75%. The database source was Homo_sapiens_9606_proteome_gene from the Uniprot database (release: 2021-10-14, sequence: 20,437), and a common contaminant library was added to the database. Contaminant proteins were removed during data analysis. The digestion method was set to Trypsin / P; the maximum number of missed cleavage sites was set to 2; the mass error tolerances for precursor ions in the first search and main search were set to 20 ppm and 5 ppm respectively, and the mass error tolerance for MS / MS fragment ions was 20 ppm. The fixed modification was cysteine alkylation, and the variable modifications were methionine oxidation and protein N-terminal acetylation. The false discovery rate (FDR) for protein identification and PSM identification was set to 1%.

[0049] 6. Differential analysis A combination of univariate analysis and multivariate statistical analysis was used to screen for differential proteins and transcripts. Univariate analysis mainly included the significance analysis (p-value or FDR value) and fold change of characteristic molecules in different groups, and multivariate statistical analysis mainly included receiver operating characteristic curve (ROC) analysis and Boruta feature selection based on the random forest algorithm. All statistical analyses were performed using R. The specific R-related information is shown in Table 1.

[0050] Table 1: R and its related information used in the present invention

[0051] The variable importance for the projection (VIP) was calculated to measure the influence intensity and interpretability of the expression patterns of each protein on the classification and discrimination of each group of samples. Further, the Wilcoxon rank-sum test was performed to obtain the corrected p-value (FDR). 59 downregulated and 76 upregulated proteins were screened according to the conditions of FDR < 0.01 and Fold change > 2 (see details in Figure 1 ).

[0052] To evaluate the role of each biomarker in the diagnosis and prediction of ovarian cancer, we used ROC and Boruta analysis methods to evaluate each biomarker. The result graphs are shown in Figure 2The abscissa is the AUC obtained from the ROC analysis, the ordinate is the -log10(FDR) calculated by the Wilcoxon test, and the size of the points represents the VIP value obtained from the Boruta analysis. Further screening was carried out according to VIP>3 and AUC>0.6, and a total of 7 more significant candidate markers were found, as shown in Table 2 for details.

[0053] Table 2: Differential markers between benign and malignant ovarian tumors

[0054] The smaller the FDR value and / or the larger the VIP value in Table 2, to a certain extent, it indicates that the difference of the protein between the two groups is more significant, and at the same time, it also indicates that the protein may have higher diagnostic value.

[0055] Example 2: Classification model for differentiating benign and malignant ovarian tumors by combining 7 differential proteins and its establishment Although a single biomarker can also distinguish serum samples of benign and malignant ovarian tumors or predict ovarian cancer, generally speaking, combining multiple biomarkers can achieve higher accuracy in differentiation or prediction. However, for a single biomarker with higher accuracy in predicting ovarian cancer, its role in the combination with one or more other biomarkers may not necessarily be greater. At the same time, it is not that the more the number of biomarkers, the higher the prediction accuracy (AUC value) of the combination. Therefore, a large number of verification experiments are still needed. In this example, a model constructed by 7 protein markers, namely CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, and FABP4 in serum, was studied.

[0056] 1. Obtain data Study population: A total of 1000 samples were collected, including 500 blood samples from patients with benign ovarian tumors and 500 blood samples from patients with malignant ovarian tumors. All sample source patients were confirmed by pathological examination of living tissues. The enrolled subjects were divided into a model group and a test group according to a ratio of 8:2.

[0057] Inclusion criteria for ovarian cancer patients: (a) no history of other malignancies, (b) surgical treatment within one month after blood collection, and confirmed as ovarian cancer by postoperative pathology. After obtaining informed consent, all collected serum samples were stored in a serum bank at -80°C.

[0058] In this example, the collected serum samples were detected by enzyme-linked immunosorbent assay (ELISA) to obtain the concentrations of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, and FABP4 in the serum.

[0059] 2. Statistical analysis of experimental data In the model group, a combined diagnostic model for multiple ovarian cancer markers was constructed using a combination of multiple machine learning methods. The predicted probability values were used to estimate the area under the receiver operating characteristic (ROC) curve with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate diagnostic model. Using the test group, the Youden index (YI) was calculated to determine the cut-off value of the predicted probability for distinguishing primary ovarian cancer patients from metastases. In addition, the ROCs of individual markers and different subgroups were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R 3.6.1, and a p-value less than 0.05 was considered statistically significant.

[0060] 3. Steps for constructing the combined diagnostic model S101: Among the 7 protein markers CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, and FABP4 in the samples of the model group, randomly select the concentration matrices of 3 to 7 markers as the original training data set.

[0061] S102: Select the generalized linear model (glmnet) algorithm for constructing the prediction model and the grid search range during the hyperparameter optimization process of the algorithm. In this step, the grid search range for hyperparameter optimization of each algorithm is set as shown in Table 3.

[0062] Table 3: Parameter grid search range of the glmnet algorithm

[0063] S103: According to the algorithm and the hyperparameter setting range set in step S102, select one of the hyperparameter combination methods as the parameters for constructing the prediction model.

[0064] S104: Divide the original data set into K subsets according to the K-fold cross-validation mechanism. To ensure that in each subset, the ratio of majority-class samples to minority-class samples is the same as that of the original data set, the stratified K-fold cross-validation mechanism needs to be used for data segmentation.

[0065] S105: According to the K training data subsets obtained by segmentation in step S104, select one of the subsets as the validation set Ddev.

[0066] S106. Combine the unselected subsets of the training data in step S105 to form a training data pool Dtrain. S107. Based on the training data set Dtrain obtained in step S106, construct a prediction model based on the selected supervised classification algorithm and hyperparameters.

[0067] S108. According to the prediction model obtained in step S107, evaluate it on the validation set Ddev to obtain the AUC value, and store the current prognosis prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 is to evaluate the prediction model obtained in step S107 on the validation set determined in the current iteration, and store both the model and the evaluation results in the prediction model pool for use in future base prediction model selection. The evaluation mentioned in this step can be the AUC value or other reasonable metrics for evaluating the model performance.

[0068] S109. Determine whether each subset has been used as the validation set. Step S109 is to determine whether the K subsets obtained in step S104 have all been used as the validation set for model training. If all subsets have been used as the validation set and the training is completed, execute step S110; if there is a subset that has not been used as the validation set, execute step S105. This step ensures that every sample in the original data set has been used as the validation set, improving the model stability and preventing the model from overfitting to a certain subset.

[0069] S110. Take the average AUC value of all the models in the obtained prediction model pool Pool as the final performance evaluation value of the model for this combination method. And store the model parameters and the final performance evaluation AUC value in the optimal model pool Poolbest.

[0070] S111. Determine whether prediction models have been constructed for all combinations of hyperparameters. Step S111 is to determine whether prediction models have been constructed for all the algorithms and the corresponding combinations of hyperparameters obtained in step S102. If all combinations have completed the model construction, execute step S112; if there is a combination that has not completed the model construction, execute step S103.

[0071] S113. From the model set Poolbest obtained in step S112, select the model with the largest AUC value as the final prediction model for ovarian cancer diagnosis.

[0072] S114. Repeat all the above steps until all combinations of the markers have completed the modeling.

[0073] 4. Determination of the Optimal Marker Combination By performing the above model construction steps, we obtained the optimal model constructed from all combinations of the biomarkers. To compare the performance of the models under these different biomarker combinations, we used the ROC method to evaluate the AUC values of these models in the test group. As shown in Table 4 below and Figure 3 as follows: Table 4: Comparison of the areas under the ROC curves of the models constructed from different biomarker combinations

[0074] As shown in Table 4, the AUC values of 5MP, 6MP, and 7MP are all greater than 0.7, and the maximum AUC is greater than 0.8, showing good performance. Among them, the model (7MP) composed of TALDO1 + PGM5 + TFF1 + CTSG + ORM2 + FABP4 + SFRP1 has the highest AUC, which is higher than the average value of the models with other biomarker combinations 5. Optimization results of the 7MP model parameters Through the above analysis, we obtained the optimal combination form of the biomarkers as TALDO1 + PGM5 + TFF1 + CTSG + ORM2 + FABP4 + SFRP1. Based on this biomarker combination, we analyzed the models constructed under 9 different combinations of glmnet algorithm hyperparameters ( Figure 4 ), and evaluated the model performance through the AUC value. As shown in Table 5 and Figure 4 as follows: When the combination of glmnet algorithm hyperparameters is alpha = 0.1 and lambda = 0.0005, the AUC reaches the maximum value of 0.882 (the AUC is calculated using the 10-fold cross-validation method during the modeling process).

[0075] Table 5: AUC of the models constructed under different combinations of glmnet algorithm hyperparameters

[0076] The equation of the model constructed based on the optimal hyperparameter combination is:

[0077] where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 7), Xi represents the measured value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker (Table 6), and b is the constant 4.536.

[0078] Table 6: Coefficients of the 7 biomarkers in the model

[0079] 6. Determination of the diagnostic threshold of the ovarian cancer combined diagnosis model (7MP) The ROC curve was plotted using the predicted values in the model group, and the optimal diagnostic cut-off value of 0.524 was set according to the Youden index value. That is, when the predicted value of the diagnostic model ≤ 0.524, the subject to be tested was considered not to be an ovarian cancer patient; when the predicted value of the model > 0.524, the subject to be tested was considered to be an ovarian cancer patient. The results are as Figure 5 shown: The AUC of the model in the model group was 0.953284, the sensitivity was 0.876, and the specificity was 0.9.

[0080] 7. Validation of the combined ovarian cancer diagnostic model (7MP) The ROC curve was plotted using the predicted values in the test group, as Figure 6 shown, and the AUC was 0.92706. The optimal diagnostic cut-off value was set to 0.524 according to the Youden index value. That is, when the predicted value of the diagnostic model ≤ 0.524, the subject to be tested was considered not to be an ovarian cancer patient; when the predicted value of the model > 0.524, the subject to be tested was considered to be an ovarian cancer patient. The results showed that the accuracy of the model in the test group was 0.846, the sensitivity was 0.816, and the specificity was 0.876.

[0081] Example 3: Comparison of the diagnostic values of different ovarian cancer diagnostic models Table 7: Comparison of the areas under the ROC curves of different diagnostic models

[0082] As shown in Table 7, the AUC of our model (7MP) was 0.334 and 0.349 higher than that of the traditional single marker respectively. The results of the DeLong's test for the significance of the AUC difference showed that the diagnostic value of our model (7MP) was significantly (p < 0.05) higher than that of the traditional marker or the combined model of traditional markers.

Claims

1. Use of a biomarker in the preparation of an ovarian cancer detection product, characterized in that The biomarker is selected from one or more of the following: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4.

2. The application according to claim 1, characterized in that The biomarker is one of the following combinations: (1) The combination of FABP4, ORM2, CTSG, TALDO1, TFF1; (2) The combination of TFF1, FABP4, ORM2, TALDO1, PGM5, SFRP1; (3) The combination of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, SFRP1.

3. A kit for ovarian cancer detection, characterized in that The kit includes reagents for detecting the biomarker, and the biomarker is selected from one or more of the following: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4.

4. The kit for ovarian cancer detection according to claim 3, characterized in that The kit includes reagents for detecting the biomarker, and the biomarker is one of the following combinations: (1) The combination of FABP4, ORM2, CTSG, TALDO1, TFF1; (2) The combination of TFF1, FABP4, ORM2, TALDO1, PGM5, SFRP1; (3) The combination of TALDO1, PGM5, TFF1, CTSG, ORM2, FABP4, SFRP1.

5. A system for detecting ovarian cancer, characterized in that The system includes a data analysis module, and the data analysis module is used to analyze the detection values of the biomarker in the sample of the patient to be tested. The biomarker is selected from one or more of the following: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, FABP4.

6. The system for detecting ovarian cancer according to claim 5, wherein The biomarker is the combination of CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1 and FABP4.

7. The system for detecting ovarian cancer according to claim 6, wherein The data analysis module includes an analysis model equation, and the equation is as follows: Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers, m = 7, Xi represents the detection value of the i-th biomarker, the unit is μg / mL, Ki represents the coefficient of the i-th biomarker, and b is a constant 4.

536.

8. The system for detecting ovarian cancer according to claim 7, characterized in that The coefficients of the Ki are shown in the following table: 。 9. The system for detecting ovarian cancer according to claim 5 or 6, characterized in that The data analysis module calculates the predicted value of whether the patient to be tested has ovarian cancer by using the analysis model equation according to the detection values of the biomarker, and determines whether the patient to be tested has ovarian cancer based on the predicted value; The determination condition is: When the predicted value ≤ the preset threshold, it is determined that the patient to be tested is not an ovarian cancer patient; When the predicted value > the preset threshold, it is determined that the patient to be tested is an ovarian cancer patient; Wherein, the preset threshold is 0.

524.

10. The system for detecting ovarian cancer according to claim 5 or 6, characterized in that The system includes a data detection module, a data input module, and a data output module; the data detection module is used to detect the biomarker in the sample to obtain the detection value; the data input module is used to input the detection value of the biomarker. After the data analysis module analyzes the detection value, the data output interface is used to output the analysis result of whether the patient to be tested has ovarian cancer.