Device for predicting chemotherapy sensitivity of ovarian cancer patients based on protein markers, construction method and application of protein marker classifier

By constructing a classifier based on plasma protein markers in ovarian cancer patients, the problem of low accuracy in predicting chemotherapy sensitivity in the prior art was solved, and the accuracy of predicting recurrence within one year after chemotherapy was achieved before surgery was achieved, and the targeted treatment plan was improved.

CN118315068BActive Publication Date: 2025-08-08WESTLAKE LAB OF LIFE SCI & BIOMEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410600639.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2025-08-08
Estimated Expiration
2044-05-15

AI Technical Summary

Technical Problem

The prior art lacks accurate commercial tools to predict chemotherapy sensitivity in ovarian cancer patients, and tissue samples are invasive and mRNA unstable, resulting in low accuracy in predicting recurrence after chemotherapy in ovarian cancer patients.

Method used

A classifier based on protein markers was constructed, and chemotherapy sensitivity prediction was predicted using plasma protein markers before surgery in ovarian cancer patients, such as SERPINA3, GPLD1, SAA4, etc. combined with extreme gradient enhancement algorithms and other machine learning models.

Benefits of technology

It improves the accuracy of predicting chemotherapy sensitivity in ovarian cancer patients, can predict whether chemotherapy recurs within one year after chemotherapy before surgery, and guides personalized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118315068B_ABST
    Figure CN118315068B_ABST
Patent Text Reader

Abstract

This application provides a device for predicting chemotherapy sensitivity in ovarian cancer patients based on protein markers, a method for constructing a protein marker classifier, and a protein marker. The device includes a processor configured to obtain a first characteristic parameter of a first protein marker associated with chemotherapy sensitivity in ovarian cancer patients. The first protein marker is a plasma protein obtained from the patient's preoperative plasma. Based on the first characteristic parameter, a pre-constructed protein marker classifier is used to predict chemotherapy sensitivity in ovarian cancer patients and output a prediction result. In this way, based on the prediction result, it is possible to accurately determine whether an ovarian cancer patient will relapse within one year after chemotherapy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of biomedical technology, and in particular to a device for predicting chemotherapy sensitivity of ovarian cancer patients based on protein markers, and a method for constructing and applying a protein marker classifier. Background Art

[0002] Ovarian cancer is a common gynecological malignancy. First-line treatment research focuses on optimizing conventional platinum / taxane conjugate chemotherapy (such as dose intensity, dose density, the addition of different agents, and intraperitoneal administration) and prolonging the maintenance of remission with cytotoxic chemotherapy. Generally, patients with a recurrence-free survival (RFS) of less than 6 months after platinum-based chemotherapy are considered platinum-resistant, those with an RFS between 6 and 12 months are considered relatively resistant, and those with an RFS greater than 12 months are considered platinum-sensitive.

[0003] Gene expression-based tools for predicting patient prognosis after chemotherapy are currently available for certain cancers, but there are currently no commercially available predictive kits specifically for predicting chemotherapy sensitivity in ovarian cancer patients. Currently reported biomarkers for predicting ovarian cancer prognosis primarily come from tissue samples, which require invasive procedures. Furthermore, most biomarkers are mRNA, which is unstable and cannot directly reflect biological activity. Furthermore, these biomarkers often lack independent validation sets, leading to potential overfitting or lack of universal applicability, making them inaccurate for predicting recurrence within one year of chemotherapy. Summary of the Invention

[0004] The present application addresses the above-mentioned technical problems existing in the prior art. This application aims to provide a device for predicting chemotherapy sensitivity in ovarian cancer patients based on protein markers, a method for constructing a protein marker classifier, and its application, which can predict chemotherapy sensitivity in ovarian cancer patients based on protein markers and improve the prediction accuracy of chemotherapy sensitivity in ovarian cancer patients.

[0005] According to a first embodiment of the present application, a device for predicting chemotherapy sensitivity of ovarian cancer patients based on protein markers is provided. The device includes a processor, wherein the processor is configured to: obtain a first characteristic parameter of a first protein marker related to the chemotherapy sensitivity of ovarian cancer patients, wherein the first protein marker is a plasma protein obtained based on the plasma of ovarian cancer patients, and the first protein marker includes SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU, and C4BPA; and predict the chemotherapy sensitivity of ovarian cancer patients based on the first characteristic parameter using a pre-constructed protein marker classifier, and output a prediction result.

[0006] According to the second scheme of the present application, a method for constructing a protein marker classifier for predicting chemotherapy sensitivity of ovarian cancer patients is provided, the construction method comprising: constructing an initial prediction classification model using an extreme gradient boosting algorithm; obtaining a first characteristic parameter of a first protein marker related to chemotherapy sensitivity of ovarian cancer patients, and a second characteristic parameter representing a first clinical factor, and optimizing the training parameters of the initial prediction classification model based on the first characteristic parameter and the second characteristic parameter to obtain a first prediction classification model, wherein the first protein marker comprises SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU and C4BPA; the first representative clinical factors include diagnosis age, tumor residual size, FIGO stage, lymph node metastasis, CA125 before treatment, HE4 level, and CA125 in the last chemotherapy cycle; based on the first predictive classification model, a third characteristic parameter of the second protein marker for predicting chemotherapy sensitivity of ovarian cancer patients and a fourth characteristic parameter of the second representative clinical factor are obtained; based on the third characteristic parameter and the fourth characteristic parameter, the training parameters of the first predictive classification model are optimized to obtain a second predictive classification model. When the accuracy of the second predictive classification model in predicting chemotherapy sensitivity of ovarian cancer patients meets the preset accuracy requirements, the second predictive classification model is used as a protein marker classifier.

[0007] According to the third embodiment of the present application, a protein marker is provided for predicting chemotherapy sensitivity in ovarian cancer patients. The protein markers include SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU and C4BPA.

[0008] According to the fourth embodiment of the present application, a kit for predicting chemotherapy sensitivity in ovarian cancer patients is provided, wherein the kit comprises protein markers, and the protein markers include SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU and C4BPA.

[0009] According to a fifth embodiment of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program. When the computer program is executed by a processing device, the computer program performs the steps performed by the processor of the device for predicting chemotherapy sensitivity of ovarian cancer patients based on protein markers as described in each embodiment of the present application, or the method for constructing a protein marker classifier for predicting chemotherapy sensitivity of ovarian cancer patients as described in each embodiment of the present application.

[0010] Compared with the prior art, the embodiments of the present application have the following advantages:

[0011] The apparatus for predicting chemotherapy sensitivity in ovarian cancer patients based on protein markers provided in the embodiments of the present application can predict chemotherapy sensitivity in ovarian cancer patients based on a first protein marker using a protein marker classifier, and can improve the accuracy of the prediction. The first protein marker is a plasma protein obtained from the preoperative plasma of ovarian cancer patients. On the one hand, obtaining the first protein marker based on plasma has good clinical operability. On the other hand, the first protein marker obtained from the preoperative plasma sample can accurately predict the chemotherapy sensitivity of ovarian cancer patients, thereby predicting whether ovarian cancer patients are likely to relapse within one year after chemotherapy.

[0012] Based on the device provided in the embodiment of the present application, the prognosis of ovarian cancer patients within one year after chemotherapy can be predicted before surgery, which helps doctors to make targeted and refined adjustments to the treatment plan for ovarian cancer patients before surgery.

[0013] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above description and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In the drawings, which are not necessarily drawn to scale, the same reference numerals may describe similar components in different views. Similar reference numerals with letter suffixes or different letter suffixes may represent different examples of similar components. The accompanying drawings generally illustrate various embodiments by way of example and not by way of limitation, and together with the description and claims, serve to illustrate the disclosed embodiments. Such embodiments are illustrative and exemplary and are not intended to be exhaustive or exclusive embodiments of the present method, apparatus, system, or non-transitory computer-readable medium having instructions for implementing the method.

[0015] Figure 1 A schematic diagram of a device for predicting chemotherapy sensitivity of ovarian cancer patients based on protein markers according to an embodiment of the present application is shown.

[0016] Figure 2 A schematic flow chart illustrating a method for constructing a protein marker classifier for predicting chemotherapy sensitivity of ovarian cancer patients according to an embodiment of the present application is shown.

[0017] Figure 3 Shown are the SHAP values of the prediction model based on preoperative plasma proteins.

[0018] Figure 4 A schematic diagram showing the prediction results of the prognosis of ovarian cancer patients within one year after chemotherapy based on the test set and the independent validation set. DETAILED DESCRIPTION

[0019] In order to enable those skilled in the art to better understand the technical solution of the present application, the present application is described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and specific embodiments, but are not intended to limit the present application.

[0020] The words "first", "second" and similar terms used in this application do not indicate any order, quantity or importance, but are only used to distinguish. The words "include" or "comprises" and similar terms used in this application mean that the elements before the word include the elements listed after the word, and do not exclude the possibility of covering other elements. In this application, the arrows shown in the figures of each step are only examples of the execution order, not limitations. The technical solution of this application is not limited to the execution order described in the embodiments. The steps in the execution order can be combined, decomposed, or swapped, as long as the logical relationship of the execution content is not affected.

[0021] All terms (including technical or scientific terms) used in this application have the same meaning as those understood by ordinary technicians in the field to which this application belongs, unless otherwise specifically defined. It should also be understood that terms defined in general dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an idealized or extremely formal sense unless explicitly defined as such here. Techniques and equipment known to ordinary technicians in the relevant field may not be discussed in detail, but where appropriate, the techniques and equipment should be considered as part of the specification.

[0022] Figure 1 A schematic diagram of a device for predicting chemotherapy sensitivity of ovarian cancer patients based on protein markers according to an embodiment of the present application is shown. The device 101 for predicting chemotherapy sensitivity of ovarian cancer patients based on protein markers includes a processor 102, which may be a processing device including one or more general-purpose processing devices, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), and the like. More specifically, the processor 102 may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor that runs other instruction sets, or a processor that runs a combination of instruction sets. The processor 102 may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a system on a chip (SoC), and the like.

[0023] The processor 102 is configured to obtain a first characteristic parameter of a first protein marker associated with chemotherapy sensitivity of an ovarian cancer patient, wherein the first protein marker is a plasma protein obtained based on the preoperative plasma of the ovarian cancer patient, and the first protein marker includes SERPINA3 (Uniprot ID: P01011), GPLD1 (Uniprot ID: P80108), SAA4 (Uniprot ID: P35542), SAA1 (Uniprot ID: P0DJI8), SERPINC1 (Uniprot ID: P01008), GSN (Uniprot ID: P06396), CFB (Uniprot ID: P00751), RBP4 (Uniprot ID: P02753), CRP Uniprot ID: P02741), C9 (Uniprot ID: P02748), APOD (Uniprot ID: P05090), C2 (Uniprot ID: P06681), TTR (Uniprot ID: P06682), and VEGF (Uniprot ID: P01097). ID: P02766), C5 (Uniprot ID: P01031), CFP (Uniprot ID: P27918), CFI (Uniprot ID: P05156), FCN3 (Uniprot ID: O75636), ITIH3 (Uniprot ID: Q06033), C8G (Uniprot ID: P07360), ATRN (Uniprot ID: O75882), SERPINA1 (Uniprot ID: P01009), ITIH2 (Uniprot ID: P19823), ORM1 (Uniprot ID: P02763), ORM2 (Uniprot ID: P19652), BCHE (Uniprot ID: P06276), HP (Uniprot ID: P00738), SERPINA4 (UniprotID: P29622), SERPINA5 (Uniprot ID: P05154), LYZ (Uniprot ID: P61626), LRG1 (Uniprot ID: P02750), CD14 (Uniprot ID: P08571), F13A1 (Uniprot ID: P00488), CALU (Uniprot ID: O43852) and C4BPA (Uniprot ID: P04003).Uniprot (Universal Protein) is a protein database that contains protein sequences, functional information, and research paper indexes. The Uniprot ID refers to the address of the protein in Uniprot. Each protein ID is unique, and information about each protein can be obtained through the protein's Uniprot ID. The first characteristic parameter can be the protein expression level of plasma proteins in the human body. The protein expression level can be expressed as the number or relative abundance of proteins. For example, mass spectrometry and immunoassay methods are used to measure protein expression levels to provide relevant information such as the relative or absolute number of proteins in the human body.

[0024] The first protein marker is obtained based on the preoperative plasma of ovarian cancer patients. The first characteristic parameter of the first protein marker can be directly obtained from a database or storage device, or can be obtained by analyzing the preoperative plasma of ovarian cancer patients who are about to undergo chemotherapy.

[0025] For example, the database or storage device may pre-store the first characteristic parameters of various plasma proteins of multiple ovarian cancer patients before surgery, and the processor 102 may read the first characteristic parameter of the first protein marker from the database or storage device.

[0026] Alternatively, before chemotherapy surgery is performed on an ovarian cancer patient, a small amount of plasma can be collected from the patient, and the plasma can be processed by pretreatment and fractionation methods to extract plasma proteins such as SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU and C4BPA as first protein markers, and the protein expression level (e.g., protein abundance) of each plasma protein in the ovarian cancer patient can be obtained.

[0027] This is merely an example and does not constitute a specific limitation on obtaining the first characteristic parameter.

[0028] Specifically, the protein marker classifier can be trained based on classifiers such as logistic regression, decision tree, random forest, and naive Bayes. For example, the first characteristic parameters of plasma proteins such as SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU, and C4BPA are used as input features, and the chemotherapy sensitivity of ovarian cancer patients is used as the target label. Classifiers such as logistic regression, decision tree, random forest, and naive Bayes are trained, and the training parameters of the classifiers are adjusted during the training process until the trained classifier has a high prediction accuracy verified by the test set, and then the classifier can be used as a protein marker classifier. This is only an example and does not constitute a limitation to the specific solution.

[0029] That is, after obtaining the first characteristic parameter of the first protein marker, the first characteristic parameter is used as an input feature of the protein marker classifier to predict the chemotherapy sensitivity of ovarian cancer patients and output a prediction result. Therefore, the processor 102 is configured to predict the chemotherapy sensitivity of ovarian cancer patients based on the first characteristic parameter using the pre-built protein marker classifier and output a prediction result.

[0030] On the one hand, plasma proteins such as SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU and C4BPA are significantly correlated with chemotherapy sensitivity in ovarian cancer patients. For example, in order to determine the first protein marker, plasma samples can be collected from multiple ovarian cancer patients before surgery. After the preoperative plasma samples of the ovarian cancer patients are collected, chemotherapy is performed on the ovarian cancer patients, and the prognosis within one year after chemotherapy is tracked to divide the plasma samples into an experimental group with recurrence within one year and a control group with recurrence after one year. Plasma samples from each ovarian cancer patient were processed and analyzed to generate global proteomic data. Univariate Cox regression analysis and T-tests were then used to identify prognostic proteins that met prognostic criteria. These criteria included a univariate Cox regression p-value < 0.05, or a T-test p-value < 0.05 between patients with ovarian cancer who relapsed within one year and those who relapsed after one year. There are no specific prognostic criteria and they can be set by physicians or other researchers.

[0031] Then, multiple reaction monitoring (MRM) was used to quantify the prognostic proteins for further screening. The screened prognostic proteins were then verified. The plasma proteins that met the prognostic criteria included SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU, and C4BPA. Among them, the use of MRM to further screen the prognostic proteins has higher stability than other methods and improves the accuracy of the first protein marker screened.

[0032] On the other hand, predictions based on the first characteristic parameter of the first protein marker can improve prediction accuracy. Based on the prediction results, doctors can predict whether ovarian cancer patients will relapse within one year after chemotherapy before surgery. For example, if the prediction results indicate that the ovarian cancer patient has high chemotherapy sensitivity, they can predict that the ovarian cancer patient is unlikely to relapse within one year after chemotherapy. If the prediction results indicate that the ovarian cancer patient has low chemotherapy sensitivity, they can predict that the ovarian cancer patient is likely to relapse within one year after chemotherapy. In this way, the prognosis of ovarian cancer patients within one year after chemotherapy can be predicted in advance, and treatment plans can be scientifically adjusted based on the chemotherapy sensitivity characteristics of each ovarian cancer patient, achieving targeted treatment for each ovarian cancer patient.

[0033] In some embodiments of the present application, the processor 102 is further configured to: obtain a second characteristic parameter representing a first clinical factor, and input the second characteristic parameter and the first characteristic parameter into a protein marker classifier for prediction to obtain a prediction result, wherein the first representative clinical factor includes age at diagnosis, residual tumor size, FIGO stage, lymph node metastasis, CA125 before treatment, HE4 level and CA125 in the last chemotherapy cycle. Specifically, the age at diagnosis can be understood as the age of the ovarian cancer patient at the time of tumor diagnosis. The residual tumor size can be understood as the size of the residual lesion after chemotherapy surgery for ovarian cancer patients. The FIGO stage refers to the standard for staging gynecological malignancies (such as cervical cancer, ovarian cancer, endometrial cancer, etc.) formulated by the International Federation of Gynecology and Obstetrics (FIGO). The lymph node metastasis can be understood as the degree of metastasis of the tumor in the lymph nodes. CA125 is a protein called cancer antigen 125, commonly used as a marker for ovarian cancer. In ovarian cancer patients, CA125 levels in the blood are often elevated. Pre-treatment CA125 refers to the CA125 level in the blood of ovarian cancer patients before chemotherapy, and final chemotherapy cycle CA125 refers to the CA125 level in the blood after chemotherapy. HE4 (Human Epididymal Protein 4) is a protein that increases in certain cancers, particularly ovarian cancer. HE4 levels can be understood as the expression level of the HE4 protein in the blood of ovarian cancer patients before chemotherapy.

[0034] The second characteristic parameter may be age, residual lesion size, tumor stage, lymph node metastasis degree, protein content, or other data parameters used to characterize various first representative clinical factors, which are not specifically limited.

[0035] In this embodiment, the first characteristic parameter and the second characteristic parameter are used together as input features to predict the chemotherapy sensitivity of ovarian cancer patients, thereby improving the prediction accuracy.

[0036] In some embodiments of the present application, the processor 102 is further configured to: obtain a third characteristic parameter of a second protein marker associated with chemotherapy sensitivity in ovarian cancer patients, and a fourth characteristic parameter representing a second clinical factor, and input the third characteristic parameter and the fourth characteristic parameter together into a protein marker classifier for prediction to obtain a prediction result; wherein the second protein marker is CFP, LYSC, APOD, C4BPA, CO8G, A1AT, and F13A; and the second representative clinical factor is age of diagnosis, pre-treatment CA125, HE4 level, and CA125 in the last chemotherapy cycle. Specifically, the second protein marker has a higher correlation with chemotherapy sensitivity in ovarian cancer patients. By using the third characteristic parameter and the fourth characteristic parameter as input features of the protein marker classifier, the accuracy of predicting chemotherapy sensitivity in ovarian cancer patients can be further improved.

[0037] In some embodiments of the present application, the method for constructing the protein marker classifier specifically includes: using the Extreme Gradient Boosting (XGBoost) algorithm to construct an initial prediction classification model. For example, when constructing the initial prediction classification model, the XGBoost algorithm can be used to consider XGBoost classifier, LightGBM, CatBoost, Gradient Boosting Machines (GBM). These models are all based on gradient boosting tree algorithms, which gradually improve the prediction performance of the model by iteratively constructing decision trees during the training process. When selecting a model, a trade-off can be made based on factors such as the specific data set size, feature type, training efficiency, and model performance, and there is no limitation on the specific initial prediction classification model.

[0038] The training parameters of the initial prediction classification model are optimized based on the first and second feature parameters, or based on the third and fourth feature parameters, and the prediction classification model whose accuracy in predicting chemotherapy sensitivity of ovarian cancer patients meets the preset accuracy requirements after the optimization is used as a protein marker classifier. In other words, the first and second feature parameters can be used as input features of the initial prediction classification model, or the third and fourth feature parameters can be used as input features of the initial prediction classification model to perform the prediction process of chemotherapy sensitivity of ovarian cancer patients. During the prediction process of the initial prediction classification model, the training parameters, such as the number of trees in the random forest, the kernel function type of the support vector machine, and other related parameters, are adjusted to optimize the training parameters of the initial prediction classification model.

[0039] The optimized predictive classification model can be used to validate its predictions using a test set. For example, it can predict the chemotherapy sensitivity of each ovarian cancer patient in the test set and, based on the predictions, determine whether the ovarian cancer patient will relapse within one year after chemotherapy. If the optimized predictive classification model predicts that multiple ovarian cancer patients are chemotherapy-sensitive, and these patients actually relapse one year after chemotherapy, the prediction accuracy of the optimized predictive classification model can be considered to meet the preset accuracy requirements and can be used as a protein marker classifier. The preset accuracy requirements are specifically defined and can be set by the user based on the test set.

[0040] Figure 2 A method for constructing a protein marker classifier for predicting chemotherapy sensitivity in ovarian cancer patients is shown. In step S201, an initial prediction classification model is constructed using the extreme gradient boosting algorithm. The example of constructing the initial prediction classification model is as described above and will not be repeated here. In step S202, a first characteristic parameter of a first protein marker associated with chemotherapy sensitivity of ovarian cancer patients and a second characteristic parameter of a first representative clinical factor are obtained, and training parameters of the initial prediction classification model are optimized based on the first characteristic parameter and the second characteristic parameter to obtain a first prediction classification model, wherein the first protein marker includes SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU, and C4BPA; and the first representative clinical factor includes age of diagnosis, residual tumor size, FIGO stage, lymph node metastasis, pre-treatment CA125, HE4 level, and CA125 in the last chemotherapy cycle.

[0041] Among them, the first protein marker has a high correlation with chemotherapy sensitivity in ovarian cancer patients. Combining it with the first representative clinical factor is beneficial for improving the optimization effect of the training parameters of the initial predictive classification model. Specifically, XGBoost can be used to use the first feature parameter and the second feature parameter as input features, perform 100 60% undersampling iterations on the training set, and optimize the two parameters of subsampling and learning rate to obtain the first predictive classification model.

[0042] The first predictive classification model can then rank the importance of each feature. For example, feature importance can be used to assess the contribution of each feature to the splitting of a decision tree node and the features can be ranked according to importance. That is, in step S203, based on the first predictive classification model, a third characteristic parameter of the second protein marker for predicting chemotherapy sensitivity in ovarian cancer patients and a fourth characteristic parameter representing the second clinical factor are obtained.

[0043] By further screening the first protein marker and the first representative clinical factor, a feature with a higher correlation with predicting the chemotherapy sensitivity of ovarian cancer can be determined, namely the second protein marker and the second representative clinical factor. Then, in step S204, based on the third characteristic parameter of the second protein marker and the fourth characteristic parameter of the second representative clinical factor, the training parameters of the first prediction classification model are optimized to obtain a second prediction classification model. When the accuracy of the second prediction classification model in predicting the chemotherapy sensitivity of ovarian cancer patients meets the preset accuracy requirement, the second prediction classification model is used as a protein marker classifier. For example, the Gamma parameter, the maximum depth (max_depth), the colsamp_bytree parameter, and the minimum child node weight (min_child_weight) can be optimized during the training process. The protein marker classifier can accurately predict the chemotherapy sensitivity of ovarian cancer patients, which helps doctors to judge in advance whether ovarian cancer patients will relapse within one year after chemotherapy based on the prediction results of the protein marker classifier and adjust their treatment plan.

[0044] Furthermore, the second protein markers are CFP, LYZ, APOD, C4BPA, C8G, SERPINA1 and F13A1; the second representative clinical factors are diagnosis age, pre-treatment CA125, HE4 levels and CA125 in the last chemotherapy cycle.

[0045] In some embodiments of the present application, the construction method further includes obtaining whole proteomic data based on preoperative plasma sample data from ovarian cancer patients, wherein the plasma sample data is obtained before chemotherapy surgery. For example, before chemotherapy surgery is performed on an ovarian cancer patient, preoperative plasma sample data can be collected from the patient, and then processed and analyzed. The amount of plasma required for each patient is 3-10 μl.

[0046] To build a classification model, preoperative plasma samples from several ovarian cancer patients can be collected. After chemotherapy, these patients are tracked for recurrence within one year. Based on their prognosis within one year, the plasma samples are divided into samples that relapsed within one year and samples that relapsed after one year. The preoperative plasma sample data is divided into a training set and a test set. The initial predictive classification model is trained based on the training set, and the training parameters are adjusted. The optimized predictive classification model is then validated using the test set.

[0047] The construction method includes obtaining prognostic protein sample data that meets the prognostic criteria based on the whole proteomics data, wherein the prognostic criteria can be set by the user or determined based on the results of clinical trials. For example, the expression level of the prognostic protein may be related to the patient's survival rate or survival period, and the prognostic criteria can be implemented by monitoring the expression level of the protein within a specific period, so that the protein that meets the prognostic criteria is used as the prognostic protein. This is only used as an example and does not constitute a limitation to the specific scheme.

[0048] After identifying the prognostic protein, the protein expression levels of some of the prognostic protein sample data in the human body can be monitored using the multiple reaction monitoring (MRM) method to screen for first protein markers whose protein expression levels exceed a preset level. For example, based on the protein expression levels obtained using MRM, proteins whose protein expression levels exceed a preset level can be used as first protein markers. MRM has the advantages of a wide linear dynamic range, high sensitivity, high accuracy, and good reproducibility.

[0049] The first characteristic parameter of the first protein marker and the second characteristic parameter of the first representative clinical factor are input into the initial prediction classification model, and the training parameters are optimized based on the training set to obtain a first prediction classification model.

[0050] The first predictive classification model further screened out features that were more correlated with predicting chemotherapy sensitivity in ovarian cancer patients, namely the second protein marker and the second representative clinical factor, and continued to optimize the first predictive classification model using the third characteristic parameter and the fourth characteristic parameter as input features to obtain the second predictive classification model.

[0051] Furthermore, based on the test set, the prediction results of the second predictive classification model are verified using the third and fourth feature parameters as input features. The second predictive classification model whose accuracy in predicting chemotherapy sensitivity in ovarian cancer patients meets the preset accuracy requirements is used as a protein marker classifier. Verifying the prediction results of the second predictive classification model using the test set further verifies the important role of the second protein marker in predicting chemotherapy sensitivity in ovarian cancer patients. After verification, if the second predictive classification model predicts chemotherapy sensitivity in ovarian cancer patients with high accuracy, meeting the preset accuracy requirements, the second predictive classification model can be used as a protein marker classifier.

[0052] In this embodiment, a second protein marker can be obtained based on a trace amount of plasma sample, and the chemotherapy sensitivity of ovarian cancer patients can be accurately predicted based at least on the second protein marker, thereby guiding doctors to adjust the treatment plan for ovarian cancer patients.

[0053] In some embodiments of the present application, a protein marker is provided for predicting chemotherapy sensitivity in ovarian cancer patients. The protein markers include SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU, and C4BPA. Using the protein markers as input features, a predictive classification model can be used to predict the chemotherapy sensitivity of ovarian cancer patients and output a prediction result with high accuracy. Based on the prediction results, doctors can determine in advance whether the ovarian cancer patient will relapse within one year after chemotherapy and make scientific adjustments to the specific treatment plan.

[0054] In some embodiments of the present application, a kit for predicting chemotherapy sensitivity in ovarian cancer patients is provided, the kit comprising protein markers, wherein the protein markers include SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU and C4BPA.

[0055] Example:

[0056] This example is only one implementation method of a method for constructing a protein marker classifier for predicting chemotherapy sensitivity in ovarian cancer patients, and the details are as follows:

[0057] Plasma sample data acquisition: All plasma sample data were collected from ovarian cancer patients before chemotherapy. Patients with ovarian cancer received at least six cycles of platinum-based chemotherapy. The discovery cohort included 69 patients who relapsed within 12 months of the last adjuvant chemotherapy (experimental group) and 72 patients who relapsed after 12 months (control group). An independent validation set comprised 30 samples (10 in the experimental group and 20 in the control group) for proteomic analysis.

[0058] The specific proteomic analysis methods are as follows:

[0059] Peptide extraction from plasma samples: 5 μL of plasma sample was diluted in 400 μL of PBS and purified by human affinity chromatography kit (Thermo Fisher Scientific TM , San Jose, USA) were used to remove the 14 most abundant proteins from the blood. The sample was then concentrated to 50 μL by centrifugation for 30 minutes using a 3K MWCO ultrafiltration tube (Thermo Fisher Scientific, San Jose, USA). The concentrated sample was mixed with 500 μL of 8 M urea (Sigma) and further concentrated to 50 μL by centrifugation for 40 minutes. The sample was transferred to a 1.5 mL EP tube and reduced with 10 mM tris(2-carboxyethyl)phosphine (TCEP, Sigma) and alkylated with 40 mM iodoacetamide (IAA) for 40 minutes each in a horizontal shaker at 800 rpm and 32°C in the dark. Proteins were then digested with two trypsin steps (enzyme to protein ratio: 1:60; Wallis Technology Co., Ltd., Beijing), with incubations at 600 rpm and 32°C for 4 and 12 hours, respectively. Digestion was then stopped by acidification to pH 2-3 with 1% trifluoroacetic acid (TFA) (Thermo Fisher), and peptides were desalted by SOLAμHRP (Thermo Fisher).

[0060] Proteomics of peptide samples based on TMTpro labeling:

[0061] 5 μg of peptide was quantified per sample, dried, and reconstituted with 3 μL of 100 mM TEAB solution. 2 μL of 20 μg / μL TMTpro labeling reagent was added, vortexed, and incubated at room temperature on a horizontal shaker at 1500 rpm for 60 minutes. The reaction was terminated by adding 0.5 μL of 5% hydroxylamine solution, mixing thoroughly, and incubating at room temperature on a horizontal shaker at 1500 rpm for 15 minutes.

[0062] 0.1 μL of each sample from the same TMT batch was mixed, dried, and then labeled by mass spectrometry. After confirming that both N-terminal and K-terminal labeling efficiencies reached 95%, all samples from the same batch were combined for each labeling channel. Each mixed sample batch was desalted using a Waters C18 (50 mg) desalting column: 400 μL of 100% (vol / vol) methanol was added to activate the desalting column, and this step was repeated once. The filtrate was discarded, and the C18 column was treated with 400 μL of 80% (vol / vol) ACN / 0.1% (vol / vol) TFA (vol / vol). This step was repeated once, and the filtrate was discarded. Each C18 column was equilibrated with 400 μL of 2% (vol / vol) ACN / 0.1% (vol / vol) TFA, and this step was repeated twice, and the filtrate was discarded. The peptide supernatant (~160 μL) of the acidified peptides was loaded onto each C18 column and centrifuged at 60 g for 2 minutes. The filtrate was reloaded and centrifuged once, and the filtrate was discarded. Wash each C18 column with 800 μL of 2% (vol / vol) ACN / 0.1% (vol / vol) TFA. Repeat this step three times. Replace a new set of 1.5 mL tubes, label them, and place each C18 column in the corresponding tube. The desalted peptides were eluted with 300 μL of 60% (vol / vol) ACN / 0.1% (vol / vol) TFA. All solvents used for desalting were analytical grade. The desalted and spin-dried peptides were re-dissolved in 110 μL of 10 mM ammonia (pH = 10.0) and fractionated using a Thermo Ultimate Dionex 3000 (Thermo Fisher Scientific, USA). The 60 fractions were combined into 30 fractions in the order of 1+31, 2+32, ..., 30+60, spin-dried, and re-dissolved.

[0063] By UltiMate TM 3000 RSLCnano system (Thermo Fisher Scientific) and a homemade 15 cm × 75 μm, filled with 1.9 μm Separation was performed on a C18-Aqua analytical column. The mobile phases consisted of buffer A (2% ACN, 0.1% formic acid) and buffer B (98% ACN, 0.1% formic acid). Peptides were separated using a 60-min effective gradient at a flow rate of 300 nL / min. The mobile phases consisted of mass buffer A and mass buffer B (98% ACN, 0.1% formic acid), respectively. Buffer B varied from 7% to 30% over the 60-min effective gradient. The mass spectrometer was operated in positive mode using a FAIMS Pro Orbitrap Exploris 480 mass spectrometer in data-dependent analysis (DDA) mode (Thermo Fisher Scientific, USA). The optimal offset voltage (CV) was set at -48 V and -68 V, with a cycle time of 1 s. The MS1 resolution was set to 60,000, with a normalized AGC target of 300%. The mass range was set to 375–1800. The dynamic exclusion mode was set to custom exclusion mode, with an exclusion duration of 40 s. The MS2 resolution was set to 30,000, the normalized AGC target to 200%, the isolation window to 0.7 m / z, the normalized HCD collision energy to 38%, and the Turbo-TMT function enabled.

[0064] Proteomic data analysis of TMT labeling:

[0065] The acquired mass spectrometry data were analyzed using Proteome Discoverer (version 2.5.0.400, Thermo Fisher Scientific). The protein database consisted of a human database downloaded from UniProtKB on February 9, 2018 (containing 20,415 reviewed protein sequences). The enzyme was set to include a maximum of two sites missed by trypsin. Static modifications included aminomethylation of cysteine (+57.021464), lysine residues, and TMTpro at the N-terminus of the peptide (+304.207145); variable modifications included methionine oxidation (+15.994915) and acetylation of the N-terminus of the peptide (+42.010565). The precursor ion mass tolerance was set to 10 ppm, and the product ion mass tolerance was set to 0.02 Da. Peptide spectrum matching allowed a target false discovery rate (FDR) of 1% (strict) and a target FDR of 5% (relaxed). Normalization was performed on the total peptide amount. Other parameters followed the default settings.

[0066] Model building and test set validation based on plasma proteome data:

[0067] To predict recurrence one year after the last chemotherapy, we first identified prognostic proteins in the discovery cohort's full proteomics data and validated these proteins using targeted proteomics and machine learning optimization models. Univariate Cox analysis and Student's t test were used to identify prognostic proteins. A total of 241 proteins had either a univariate Cox analysis p-value < 0.05 or a Student's t test p-value < 0.05 between patients who relapsed within 1 year and those who relapsed after 1 year.

[0068] Multiple reaction monitoring (MRM) was then used to quantify 51 of the 241 prognostic proteins. Retention time calibration was performed using commonly used internal retention time (CiRT) standard peptides. Twelve peptides were selected from a published blood chromatogram library and analyzed on a Jasper™ HPLC system (SCIEX, CA, USA) in buffer B (buffer A: 0.1% formic acid in water; buffer B: 0.1% formic acid in acetonitrile) ranging from 10% to 40%. Ionized peptides were transferred to a TRIPLE QUAD™ 4500MD (SCIEX, CA, USA). A total of 389 ion transitions representing 101 peptides were analyzed from the plasma sample using timed acquisition within a ±1 minute time window with a target scan time of 1.7 seconds.

[0069] Thirty-four prognostic proteins were validated using MRM. These 34 proteins were SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU, and C4BPA. Subsequently, the eXtreme Gradient Boosting (XGBoost) algorithm was used to construct the predictive classification models A and B. Seven clinical factors, including age at diagnosis, residual tumor size, FIGO stage, lymph node metastasis, pretreatment CA125 and HE4 levels, and CA125 during the last chemotherapy cycle, were used, along with a 34-protein prognostic feature screened by MRM quantitative validation, to optimize predictive classification models A and B. Predictive classification model A was based solely on clinical factors, while predictive classification model B was based on both clinical factors and the 34 protein prognostic features.

[0070] XGBoost was used to perform 100 iterations of 60% undersampling on the training set in the discovery cohort, and the two parameters of subsample (step size of 0.05, from 0.5 to 1) and learning rate (step size of 0.04, from 0.1 to 0.3) were optimized. The features in the prediction classification models A and B were ranked by importance, and the top 5-15 features were selected and trained using XGBoost for the entire training set. The other four parameters, namely gamma (from 0 to 0.2, step size of 0.05), max_depth (from 3 to 10, step size of 1), colsamp_bytree (from 0.1 to 1, step size of 0.1), and min_child_weight (from 1 to 5, step size of 1) were all optimized. Finally, the prediction classification model with the highest prediction accuracy on the training set and the internal test set was selected (such as Figure 3 , 4).

[0071] An independent validation set was used to evaluate the predictive utility of the final predictive classification model.

[0072] Finally, it was found that the prediction classification model A based on five clinical factors (diagnosis age, pre-treatment CA125, HE4 levels, the last chemotherapy cycle CA125 and FIGO stage) could not distinguish the prognostic differences between the experimental group and the control group. Figure 3 As shown in the figure, plasma proteins CFP, LYZ, APOD, C4BPA, C8G, SERPINA1 and F13A1 and four clinical factors (diagnosis age, pre-treatment CA125, HE4 level, and last chemotherapy cycle CA125) have a greater impact on the output of the prediction classification model. When these seven plasma proteins and four clinical factors are included as input features of the prediction classification model, it can be predicted that the prognosis of the two groups of patients who relapsed one year later and those who relapsed within one year in the independent validation set are significantly different (e.g. Figure 4 ).

[0073] In summary, CFP, LYZ, APOD, C4BPA, C8G, SERPINA1 and F13A1 can be used to effectively predict whether ovarian cancer patients will relapse within one year after surgery and adjuvant chemotherapy.

[0074] This application describes various operations or functions that can be implemented as software code or instructions or defined as software code or instructions. Such content can be source code or differential code ("incremental" or "patch" code) that can be directly executed ("object" or "executable" form). Software code or instructions can be stored in a computer-readable storage medium and, when executed, can cause a machine to perform the described functions or operations, and include any mechanism for storing information in a form accessible to a machine (e.g., a computing device, an electronic system, etc.), such as recordable or non-recordable media (e.g., read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, etc.).

[0075] The exemplary methods described herein can be at least partially implemented by a machine or computer. In some embodiments, a computer-readable storage medium stores computer program instructions that, when executed by a processing device, cause the processing device to execute the steps performed by a processor of the apparatus for predicting chemotherapy sensitivity of ovarian cancer patients based on protein markers as described in various embodiments of this application, or the method for constructing a protein marker classifier for predicting chemotherapy sensitivity of ovarian cancer patients as described in various embodiments of this application.

[0076] Implementations of such methods may include software code, such as microcode, assembly language code, high-level language code, and the like. Various software programming techniques may be used to create various programs or program modules. For example, program portions or program modules may be designed using or with the aid of Java, Python, C, C++, assembly language, or any other known programming language. One or more of such software portions or modules may be integrated into a computer system and / or computer-readable media. Such software code may include computer-readable instructions for performing the various methods. The software code may form part of a computer program product or computer program module. Furthermore, in an example, the software code may be tangibly stored on one or more volatile, non-transitory, or non-volatile tangible computer-readable media, such as during execution or at other times. Examples of such tangible computer-readable media may include, but are not limited to, hard disks, removable disks, removable optical disks (e.g., optical disks and digital video disks), magnetic cassettes, memory cards or sticks, random access memory (RAM), read-only memory (ROM), and the like.

[0077] In addition, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present application with equivalent elements, modifications, omissions, combinations (e.g., solutions that intersect various embodiments), adaptations, or changes. The elements in the claims are to be interpreted broadly based on the language employed in the claims and are not limited to the examples described in this specification or during the prosecution of this application, which examples are to be interpreted as non-exclusive. Therefore, this specification and examples are intended to be considered as examples only, with the true scope and spirit being indicated by the following claims and the full scope of their equivalents.

[0078] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of their solutions) can be used in combination with each other. For example, those of ordinary skill in the art may use other embodiments when reading the above description. In addition, in the above-mentioned specific embodiments, various features can be grouped together to simplify the application. This should not be interpreted as an intention that a disclosed feature that is not required to be protected is necessary for any claim. On the contrary, the subject matter of the present application may be less than all the features of a specific disclosed embodiment. Thus, the claims are incorporated into the specific embodiments as examples or embodiments, wherein each claim is independently a separate embodiment, and it is considered that these embodiments can be combined with each other in various combinations or arrangements. The scope of this application should be determined with reference to the appended claims and the full scope of equivalents to which these claims are entitled.

[0079] The above embodiments are merely exemplary embodiments of the present application and are not intended to limit the scope of the present application. The scope of protection of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present application within the essence and scope of protection of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the scope of protection of the present application.

Claims

1. A device for predicting chemotherapy sensitivity of ovarian cancer patients based on protein markers, characterized in that: The apparatus comprises a processor configured to: Obtaining a first characteristic parameter of a first protein marker associated with chemotherapy sensitivity in ovarian cancer patients, where the first protein marker is a plasma protein obtained based on preoperative plasma of the ovarian cancer patient, and the first protein marker is SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU, and C4BPA; Obtaining a second characteristic parameter of the first representative clinical factor, wherein the first representative clinical factor is diagnosis age, tumor residual size, FIGO stage, lymph node metastasis, pre-treatment CA125, HE4 level, and CA125 of the last chemotherapy cycle; Obtaining a third characteristic parameter of a second protein marker associated with chemotherapy sensitivity in ovarian cancer patients, and a fourth characteristic parameter of a second representative clinical factor, wherein the second protein marker is CFP, LYZ, APOD, C4BPA, C8G, SERPINA1, and F13A1; and the second representative clinical factor is age at diagnosis, pre-treatment CA125, HE4 levels, and CA125 during the last chemotherapy cycle; Based on the first characteristic parameter, the second characteristic parameter, the third characteristic parameter and the fourth characteristic parameter, a pre-constructed protein marker classifier is used to predict the chemotherapy sensitivity of ovarian cancer patients and output a prediction result.

2. The device according to claim 1, characterized in that The method for constructing the protein marker classifier specifically includes: Use extreme gradient boosting algorithm to build initial prediction classification model; The training parameters of the initial prediction classification model are optimized based on the first and second characteristic parameters, or based on the third and fourth characteristic parameters, and the prediction classification model whose accuracy in predicting chemotherapy sensitivity of ovarian cancer patients meets the preset accuracy requirements is used as a protein marker classifier.

3. A method for constructing a protein marker classifier for predicting chemotherapy sensitivity in ovarian cancer patients, characterized in that: The construction method comprises: Use extreme gradient boosting algorithm to build initial prediction classification model; Obtaining a first characteristic parameter of a first protein marker associated with chemotherapy sensitivity in ovarian cancer patients and a second characteristic parameter of a first representative clinical factor, and optimizing the training parameters of the initial prediction classification model based on the first characteristic parameter and the second characteristic parameter to obtain a first prediction classification model, wherein the first protein marker is SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU, and C4BPA; and the first representative clinical factor is age of diagnosis, residual tumor size, FIGO stage, lymph node metastasis, pre-treatment CA125, HE4 level, and CA125 in the last chemotherapy cycle; Based on the first prediction and classification model, a third characteristic parameter of a second protein marker for predicting chemotherapy sensitivity in ovarian cancer patients and a fourth characteristic parameter of a second representative clinical factor are obtained, where the second protein marker is CFP, LYZ, APOD, C4BPA, C8G, SERPINA1, and F13A1; and the second representative clinical factor is age at diagnosis, pre-treatment CA125 and HE4 levels, and CA125 during the last chemotherapy cycle; Based on the third characteristic parameter and the fourth characteristic parameter, the training parameters of the first prediction classification model are optimized to obtain a second prediction classification model. When the accuracy of the second prediction classification model in predicting chemotherapy sensitivity of ovarian cancer patients meets the preset accuracy requirements, the second prediction classification model is used as a protein marker classifier.

4. The construction method according to claim 3, characterized in that The construction method further comprises: Based on preoperative plasma sample data from ovarian cancer patients, full proteomic data were obtained; The patients' preoperative plasma sample data were divided into training set and test set; Obtaining prognostic protein sample data that meets prognostic criteria based on the whole proteomics data; The protein expression levels of some prognostic protein sample data in the human body are monitored using a multiple reaction monitoring method to screen out the first protein marker whose protein expression level exceeds a preset level; The first characteristic parameter of the first protein marker and the second characteristic parameter of the first representative clinical factor are input into the initial prediction classification model, and the training parameters are optimized based on the training set to obtain a first prediction classification model.

5. The construction method according to claim 4, characterized in that The construction method further comprises: Based on the test set, the third characteristic parameter and the fourth characteristic parameter are used as feature inputs to verify the prediction results of the second prediction classification model, and the second prediction classification model whose accuracy in predicting chemotherapy sensitivity of ovarian cancer patients meets the preset accuracy requirement is used as the protein marker classifier; The amount of plasma required by each patient is 3-10ul.

6. A kit for predicting chemotherapy sensitivity of ovarian cancer patients, characterized in that: The kit includes protein markers, and the protein markers are SERPINA3, GPLD1, SAA4, SAA1, SERPINC1, GSN, CFB, RBP4, CRP, C9, APOD, C2, TTR, C5, CFP, CFI, FCN3, ITIH3, C8G, ATRN, SERPINA1, ITIH2, ORM1, ORM2, BCHE, HP, SERPINA4, SERPINA5, LYZ, LRG1, CD14, F13A1, CALU and C4BPA.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processing device, performs the steps performed by a processor of the device for predicting chemotherapy sensitivity of ovarian cancer patients based on protein markers according to claim 1 or 2, or the method for constructing a protein marker classifier for predicting chemotherapy sensitivity of ovarian cancer patients according to any one of claims 3 to 5.

Citation Information

Patent Citations

  • Establishment method of ovarian malignancy and junctional tumor diagnosis model

    CN117133439A

  • Training method and device for prediction model of chemosensitivity of ovarian cancer patient

    CN118039152A