Method for predicting preeclampsia by using clinical laboratory data and use thereof
The preeclampsia risk prediction model, which combines maternal clinical indicators and laboratory data, solves the problem of predicting preeclampsia in early pregnancy, achieving efficient and low-cost prediction results, and is suitable for early clinical prevention and treatment.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2026-03-12
AI Technical Summary
Existing technologies are insufficient to effectively predict preeclampsia in early pregnancy, and existing methods require complex maternal information or multiple biomarkers, resulting in high testing costs and insufficient accuracy.
By combining maternal clinical indicators, clinical laboratory data, and IGFBP1 and/or PLGF, a preeclampsia risk prediction model is constructed using machine learning methods. Clinical laboratory data such as white blood cell count and urine protein are used to simplify the detection indicators and improve the prediction accuracy.
It enables timely prediction of preeclampsia risk between 11 and 15+6 weeks of gestation, reduces testing costs, and improves the predictive specificity and sensitivity of early and late-onset preeclampsia, making it suitable for early clinical prevention and treatment.
Smart Images

Figure CN2024117462_12032026_PF_FP_ABST
Abstract
Description
Method for predicting preeclampsia using clinical laboratory data and applications thereof TECHNICAL FIELD
[0001] The present application relates to the field of preeclampsia analysis, in particular to a method for constructing a preeclampsia risk prediction model based on clinical laboratory data, a prediction product, a prediction system and related applications thereof. BACKGROUND
[0002] In recent years, there are many articles and patents about preeclampsia prediction. The use of protein markers in blood and urine to predict preeclampsia has been commercialized, and international and domestic companies have related products on the market. In addition to protein markers, the use of other types of markers such as metabolites, cfRNA, cfDNA, etc. is still in the research stage.
[0003] Currently, the methods for predicting preeclampsia using protein markers include: "Method and means for excluding the onset of pre-eclampsia in a certain period of time using sFlt-1 / PlGF or Endoglin / PlGF ratio" (CN104412107B), which relates to a method for diagnosing whether a pregnant subject is at risk of pre-eclampsia within a short time window, which involves sFlt-1, Endoglin and PlGF, comparing the above protein markers with a reference to diagnose the subject at risk of developing pre-eclampsia within a short period, wherein the reference allows a diagnosis with a negative predictive value of at least about 98%. "Dynamic ratio of SFlt-1 or Endoglin / PlGF used as an indicator of critical pre-eclampsia and / or HELLP syndrome" (CN103917875), which determines the amount of sFlt-1 or Endoglin and the amount of P1GF, and determines whether the subject is at risk of pre-eclampsia within a short period by comparing the amounts. "Method for determining the risk of prenatal complications" (CN102216468B), which relates to a method, medical profile, kit and device for determining the risk of developing pre-eclampsia in a pregnant individual, based on placental growth factor (PlGF) and pregnancy-associated plasma protein A (PAPP-A), combined with blood pressure and maternal factors, determining the likelihood ratio by multivariate Gaussian analysis to determine the risk of pre-eclampsia in the individual with a false positive rate of 10% and a detection rate greater than 65%. "A method for screening early preeclampsia" (CN115116570A), which discloses a method for screening early preeclampsia, involving the serum biochemical indicators of pregnant women, pregnancy-associated protein A and placental growth factor, double-arm mean arterial pressure and uterine artery pulsatility index, to predict the risk of preeclampsia in pregnant women. "Preeclampsia biomarkers and related systems and methods" (CN111094988A), which is based on PlGF, sFlt1, KIM1 and CLEC4A, and LOESS corrected for gestational age, applies a machine learning algorithm classifier to the expression profile of the above proteins to determine whether to avoid unnecessary treatment for pre-eclampsia. "Method for predicting hypertensive disorders in pregnancy in early pregnancy by combining MAP, PlGF and PAPP-A in a model for pregnant women" (CN112466460), which is based on MAP detection, serum PlGF and PAPP-A levels, and uses a risk calculation model constructed by calibrating MoM values combined with weight and gestational age for screening. "Method for predicting preeclampsia preterm birth using metabolic biomarkers and protein biomarkers" (CN112105931A), which discloses a computer-implemented method for early prediction of the risk of pregnancy outcome in pregnant women, comprising the following steps: inputting values selected from a plurality of preeclampsia-specific biomarkers (including a plurality of protein markers and metabolites) into a calculation model, calculating the predicted risk of the selected pregnancy outcome; and outputting the predicted risk of pregnancy outcome in pregnant women.
[0004] But in the above research results, CN104412107 and CN103917875 can only predict the risk of preeclampsia in the short term or in the third trimester of pregnancy, and cannot make predictions in the first trimester of pregnancy. Preeclampsia usually occurs in the third trimester of pregnancy, and domestic and foreign guidelines point out that taking low-dose aspirin before 16 weeks can effectively prevent the risk of preeclampsia, so timely prediction in the first trimester of pregnancy is crucial for guiding doctors to take medicine. The maternal information required by CN102216468 and CN115116570 includes uterine artery pulsatility index, which requires high operation of the physician and needs professional training. The pulsatility of this index has a greater impact on the prediction result. The prediction accuracy of CN102216468, CN115116570, CN111094988 and CN112466460 for early-onset preeclampsia is not high enough. The proteins and metabolic indicators required by CN112105931 will increase the experimental detection cost in actual application, and more markers will require more biological sample size. Therefore, the development of a more efficient and easy-to-implement preeclampsia risk prediction method is of great significance.
[0005] SUMMARY
[0006] The above existing prediction model mainly uses protein markers and metabolic indicators (metabolic markers) combined with several maternal clinical information to construct the model, and the present application adds maternal clinical laboratory data to the protein markers and maternal clinical indicators (clinical information) to jointly construct the preeclampsia prediction model, thereby providing a new method for constructing a preeclampsia risk prediction model.
[0007] In a first aspect, the present application provides a method for constructing a preeclampsia risk prediction model based on clinical laboratory data, which comprises constructing the above-mentioned preeclampsia risk prediction model based on several clinical indicators of pregnant women, several clinical laboratory data and insulin growth binding factor binding protein 1 (IGFBP1) and / or placental growth factor (PLGF), wherein the pregnant women include both diseased pregnant women with early-onset preeclampsia or late-onset preeclampsia and healthy pregnant women.
[0008] Specifically, the above-mentioned construction method comprises:
[0009] Obtaining the clinical indicators of pregnant women, several clinical laboratory data and IGFBP1 and / or PLGF as related modeling factors, and constructing a training sample set of the modeling factors and the corresponding pregnant women, wherein the pregnant women include both diseased pregnant women with early-onset preeclampsia or late-onset preeclampsia and healthy pregnant women;
[0010] Based on the above-mentioned training sample set, several types of model training are carried out, and the several types of trained models are evaluated, and the best prediction model is confirmed according to the model evaluation index.
[0011] In one embodiment, the prediction model is constructed by machine learning method using the clinical indicators, the clinical laboratory data, and the IGFBP1 and / or PLGF of the pregnant woman as input, and using whether the pregnant woman has preeclampsia as output.
[0012] The clinical laboratory data includes white blood cell count and / or urine protein. Further, the clinical laboratory data also includes one or more of the following: platelet count, monocyte ratio, aspartate aminotransferase / alanine aminotransferase ratio, glutamyl transpeptidase, uric acid, and high-density lipoprotein.
[0013] In a second aspect, the present application provides an application of the clinical laboratory data, i.e. white blood cell count and / or urine protein, in constructing a prediction model for preeclampsia risk.
[0014] Further, the clinical laboratory data also includes one or more of the following: platelet count, monocyte ratio, aspartate aminotransferase / alanine aminotransferase ratio, glutamyl transpeptidase, uric acid, and high-density lipoprotein.
[0015] In a third aspect, based on the prediction model constructed by the method of the first aspect, the present application provides a method for predicting preeclampsia risk, which comprises:
[0016] obtaining the clinical indicators, the clinical laboratory data, and the IGFBP1 and / or PLGF of the pregnant woman;
[0017] inputting the obtained clinical indicators, the clinical laboratory data, and the IGFBP1 and / or PLGF of the pregnant woman into the prediction model for processing, and obtaining the preeclampsia risk value of the pregnant woman, wherein if the risk value exceeds a threshold value, it is determined that the pregnant woman has a high risk of preeclampsia.
[0018] In a fourth aspect, based on the prediction model constructed by the method of the first aspect, the present application provides a system for predicting preeclampsia risk, which comprises:
[0019] a device for obtaining the clinical indicators, the clinical laboratory data, and the IGFBP1 and / or PLGF of the pregnant woman;
[0020] a device for processing the clinical indicators, the clinical laboratory data, and the IGFBP1 and / or PLGF of the pregnant woman by the prediction model;
[0021] a device for outputting the prediction result.
[0022] Fifthly, based on the prediction model obtained by the construction method of the first aspect and the preeclampsia risk prediction method provided in the third aspect, this invention discloses a preeclampsia risk prediction product, which includes:
[0023] Memory, used to store programs;
[0024] A processor for implementing the prediction method provided in the third aspect by executing the program stored in the aforementioned memory.
[0025] In a sixth aspect, based on the prediction model constructed by the construction method of the first aspect, the present invention provides a computer-readable storage medium storing a program that can be executed by a processor to implement the preeclampsia risk prediction method provided in the third aspect.
[0026] In a seventh aspect, based on the construction method of the first aspect above, the present invention provides a computer-readable storage medium storing a program and storing the prediction model constructed in the first aspect above.
[0027] The beneficial effects of this invention are as follows: the preeclampsia risk prediction model constructed by combining maternal clinical indicators, clinical laboratory data, and IGFBP1 and / or PLGF data provides a reliable means for the early prevention and treatment of preeclampsia. Firstly, the detection range of this invention is gestational age 11-15. +6 Firstly, this invention allows for timely and effective medication guidance from doctors before the recommended 16 weeks of gestation for aspirin use. Secondly, the maternal clinical indicators used in this invention include body mass index, mean arterial pressure, presence or absence of IVF, and adverse pregnancy history, excluding uterine artery pulsatility index. All required indicators are easily measurable. Furthermore, the clinical laboratory data used in this invention are routine and readily available tests performed in hospitals during early pregnancy. Additionally, the protein predictive biomarkers required by this invention are only IGFBP1 and PlGF, resulting in low testing costs and facilitating commercial use. Moreover, the predictive model of this invention exhibits a specificity of 85.04% and a sensitivity of 88.0% for early-onset preeclampsia in a fully independent validation set, and a sensitivity of 62.5% for late-onset preeclampsia with a specificity of 85.04%, exceeding the performance of currently invented detection methods. Attached Figure Description
[0028] Figure 1 shows the change of PLGF concentration with gestational week in an embodiment of the present invention;
[0029] Figure 2 shows the change of IGFBP1 concentration with gestational week in an embodiment of the present invention;
[0030] Figure 3 shows the difference of MoM values of IGFBP1 and PLGF in early-onset preeclampsia and healthy control samples in different data sets, wherein Figure 3a is IGFBP1 and Figure 3b is PLGF;
[0031] Figure 4 shows the difference of MoM values of IGFBP1 and PLGF in late-onset preeclampsia and healthy control samples in different data sets, wherein Figure 4a is IGFBP1 and Figure 4b is PLGF;
[0032] Figure 5 shows the ROC graph of the optimal model of early-onset preeclampsia in the embodiment of the present application.
[0033] Figure 6 shows the ROC graph of the optimal model of late-onset preeclampsia in the embodiment of the present application. DETAILED DESCRIPTION
[0034] As introduced in the background, several clinical indicators and biomarkers (such as various protein markers and metabolites) of pregnant women are generally used in the prior art to make risk assessment and prediction of preeclampsia. However, through a large number of studies on clinical laboratory data, the present application has found a group of clinical laboratory data which has special guiding significance for the risk prediction of preeclampsia. Specifically, in the present application, the laboratory examination information of the corresponding pregnant women samples in the early stage of pregnancy is collected as the source of clinical laboratory data, including blood routine test in 0-17 +6 weeks, liver and kidney function test in 11-15 +6 weeks, urine routine test in 11-15 +6 weeks, and hepatitis B two pairs of half detection within one year before and after pregnancy, a total of 102 indicators, see Table 1.
[0035] Table 1 Clinical laboratory data - laboratory examination included items
[0036] It should be understood that, in addition to the clinical indicators of pregnant women used in the construction of the prediction model of preeclampsia in the prior art, the present application classifies the laboratory examination items of pregnant women in Table 1 as clinical laboratory data for better understanding and description of the construction of the prediction model of the present application. In the relevant clinical laboratory data in Table 1, the present application finds that white blood cell count (WBC) and urine protein (NCG PRO) have a special effect on other clinical laboratory data, which can effectively improve the sensitivity and specificity of the prediction model of preeclampsia risk. In addition to the above-mentioned WBC and NCG PRO, the combination of one or more of platelet count (PLT), monocyte percentage (MO%), aspartate / alanine aminotransferase ratio (AS / AL), glutamyl transpeptidase (GGT), uric acid (UA) and high-density lipoprotein (HDL-C) together with WBC and / or NCG PRO can also effectively improve the prediction quality of the model.
[0037] Based on the above inventive concept, the present application develops a new method for constructing a prediction model of preeclampsia risk by using the above-mentioned clinical laboratory data, maternal clinical indicators of pregnant women (including BMI, MAP, PMH, IVF) and related data of protein markers PLGF and / or IGFBP1 as modeling factors. In the maternal clinical indicators of pregnant women: body mass index (Body mass index, BMI) is obtained by the height and weight of the pregnant woman; mean arterial pressure is calculated by measuring the blood pressure of both arms of the individual; whether the pregnant woman is an IVF (in vitro fertilization, IVF) or not; adverse past medical history (PMH) includes previous gestational diabetes, preeclampsia history, chronic hypertension history, systemic lupus erythematosus and antiphospholipid syndrome. By using the above-mentioned modeling factors of pregnant women and the corresponding conditions of pregnant women (such as early-onset preeclampsia, late-onset preeclampsia and healthy controls) to establish a training set and train the prediction model, and to verify and evaluate the obtained prediction model, the best model is finally obtained.
[0038] In a specific embodiment of the present application, the prediction model is constructed by using the clinical indicators, clinical laboratory data and IGFBP1 and / or PLGF of the pregnant woman as inputs and whether the pregnant woman has preeclampsia as outputs through machine learning methods. The machine learning methods suitable for use in the present application include but are not limited to random forest, naive Bayes, logistic regression, support vector machine, AdaBoost, K-nearest neighbor algorithm, neural network, passive attack algorithm, stochastic gradient descent or xgboost. In the embodiments of the present application, the machine learning method actually used is random forest. For better understanding of the present application, the data processing of the modeling factors involved in the construction method of the prediction model is described as follows.
[0039] 1. Acquisition of data related to modeling factors
[0040] 1.1. Collect clinical information of pregnant women, including their height, gestational age (11-15 days). +6 Zhou's weight, adverse medical history (including history of gestational diabetes, preeclampsia, chronic hypertension, systemic lupus erythematosus, and antiphospholipid syndrome), and whether or not she had undergone in vitro fertilization. The information data of the above clinical indicators are converted. In one specific embodiment of the present invention, the specific conversion method is as follows:
[0041] BMI - BMI is calculated from height and weight;
[0042] PMH - 1 is awarded if the individual has a history of gestational diabetes, preeclampsia, chronic hypertension, systemic lupus erythematosus, or antiphospholipid syndrome; otherwise, 0 is awarded.
[0043] IVF - If the pregnancy was conceived through IVF, the value is 1; otherwise, it is 0.
[0044] In addition, obtaining the MAP (partial maternal motion) clinical indicator in pregnant women includes measuring the pregnant woman's MAP at 11-15 minutes. +6 The mean arterial pressure (MAP) is calculated from the diastolic blood pressure (SBP) and systolic blood pressure (DBP) of both arms using the following formula 1:
[0045] 1.2. Select 11-15 pregnant women. +6The sample used in the embodiments of the present application is plasma, but the sample source can also be other body fluids of the human body, including but not limited to whole blood, serum, urine, saliva, amniotic fluid, cerebrospinal fluid, nipple aspirate, etc., and it should be understood that the prediction model can be constructed and its effectiveness can be evaluated according to different sample sources. At the same time, other similar polypeptide sequences of the protein prediction marker IGFBP1, various modified forms of polypeptides, and various variant forms of polypeptides can also be suitable for the present application, and the prediction model can be constructed and its effectiveness can be evaluated according to different forms. The protein markers PLGF and IGFBP1 can be detected by various methods, and the protein detection technology used in the embodiments of the present application is ELISA enzyme-linked immunoassay technology. Those skilled in the art should understand that other protein detection technologies, such as protein arrays, proteomics, expression proteomics, mass spectrometry (for example, liquid chromatography-mass spectrometry (LC-MS), multiple reaction monitoring (MRM), selected reaction monitoring (SRM), scheduled MRM, scheduled SRM), 2D PAGE, 3D PAGE, electrophoresis, proteome chips, proteome microarrays, Edman degradation, direct or indirect ELISA, immunosorbent assay, immuno-PCR, proximity extension assay, Luminex analysis or homogeneous analysis, time-resolved fluorescence (TRF), fluorescence oxygen channel immunoassay (FOCI), or luminescence oxygen channel immunoassay, target protein detection technology of liquid chromatograph-mass spectrometer, chemiluminescence immunoassay protein detection technology, fluorescence immunoassay protein detection technology, etc. can also have similar detection results.
[0046] 1.3. Collecting pregnant women 11-15 +6 Week laboratory test information, including blood routine indexes: white blood cell count (WBC), platelet count (PLT), and monocyte ratio (MO%); liver and kidney function indexes: aspartate / alanine aminotransferase ratio (AS / AL), glutamyl transpeptidase (GGT), uric acid (UA), and high-density lipoprotein (HDL-C); urine routine indexes: urine protein (NCG_PRO).
[0047] 2. Data MoM processing: processing the relevant data of the above modeling factors obtained
[0048] 2.1. Fixed median correction of BMI in clinical indicators, i.e., MoM (Multiple Of Median) processing, the calculation formula is formula 2:
[0049] wherein, BMI is calculated from height and weight, Median BMIMedian of population MAP, in one embodiment of the present application, the fixed median is 20.69, MoM BMI Median calibrated BMI MoM value.
[0050] 2.2. MAP is also corrected by fixed median, the calculation formula is formula 3:
[0051] Wherein, MAP is the raw value of mean arterial pressure of pregnant women according to formula 1, Median MAP Median of population MAP, in one embodiment of the present application, the median of population MAP is assigned as 82.98, MoM MAP Median calibrated MAP MoM value.
[0052] 2.3. Two protein marker concentrations are corrected by median corresponding to different gestational weeks, and the median protein concentration values in different gestational weeks are derived from the expected values after linear regression fitting of healthy population, the fitting equation of insulin growth binding factor binding protein 1 (IGFBP1) is shown in formula 4: E(Median) IGFBP1,GA =1.94×10 5 -1.19×10 3 ×GA Formula 4
[0053] The fitting equation of placental growth factor (PlGF) is shown in formula 5: E(Median) PLGF,GA =-10.3+0.672×GA Formula 5
[0054] Wherein, GA is gestational age, E(Median) IGFBP1,GA And E(Median) PLGF,GA Is the expected median of IGFBP1 and PLGF protein concentration of each gestational week. By correcting the protein concentration value by the median corresponding to the gestational week, the calculation formula is formula 6:
[0055] Concentration protein Is the raw value of IGFBP1 and PLGF protein concentration of each pregnant woman, E(Median)protein, GA is the expected median of protein concentration corresponding to the gestational week, MoM protein Gestational age calibrated protein concentration MoM value.
[0056] 2.4 Median correction of clinical laboratory data
[0057] In the present application, the median correction of each clinical laboratory data is performed by the following formula 7, wherein lab is the detection original value of each clinical laboratory data, medianhospital.lab is the fixed median of the corresponding clinical laboratory data, MoM hospital,lab is the MoM value of the clinical laboratory data
[0058] 3. Construction and application of prediction model
[0059] 3.1. In one specific embodiment of the present application, a random forest machine learning method is used to construct the prediction model. The MoM processed clinical indicators, clinical laboratory data and protein marker concentrations are used to train the early-onset and late-onset preeclampsia models. Specifically, the specific parameter features used include: MoM value of body mass index (BMI), MoM value of mean arterial pressure (MAP), in vitro fertilization (IVF), adverse past medical history (PMH), MoM value of IGFBP1 and / or MoM value of PLGF, MoM value of white blood cell count (WBC), MoM value of platelet count (PLT), MoM value of monocyte ratio (MO%), MoM value of aspartate aminotransferase / alanine aminotransferase ratio (AS / AL), MoM value of glutamyl transpeptidase (GGT), MoM value of uric acid (UA), MoM value of high-density lipoprotein (HDL-C) and MoM value of urine protein (NCG_PRO). When the prediction model is constructed, the output is whether the pregnant woman has preeclampsia (if yes, it can be further determined whether the pregnant woman belongs to early-onset preeclampsia or late-onset preeclampsia).
[0060] For the early-onset preeclampsia prediction model, the clinical indicators include BMI, MAP, IVF and PMH, and the protein markers include IGFBP1 and PLGF; the clinical laboratory data preferably include WBC, and in addition, one or more of PLT, MO%, HDL-C, NCG_PRO, AS / AL, GGT and UA can be further included.
[0061] For the late-onset preeclampsia prediction model, the clinical indicators include BMI, MAP, IVF and PMH, and the protein markers preferably include PLGF or only IGFBP1; the clinical laboratory data preferably include NCG_PRO, and in addition, one or more of MO%, HDL-C, WBC and AS / AL can be further included.
[0062] In one specific embodiment of the present application, the preferred input feature combination of the early-onset preeclampsia prediction model is:
[0063] BMI + IVF + PMH + MAP + IGFBP1 * PLGF + WBC + MO% + PLT;
[0064] In one specific embodiment of the present application, the input feature combination of the preferred prediction model for late-onset preeclampsia is:
[0065] BMI+IVF+PMH+MAP+PLGF+MO%+AS / AL+HDL-C+NCG_PRO.
[0066] The specificity of the above prediction model for early-onset preeclampsia is 85.04%, and the sensitivity is 88.0%. The specificity of the prediction model for late-onset preeclampsia is 85.04%, and the sensitivity is 62.5%.
[0067] It should be understood that a plurality of prediction models can be constructed by the above method. Corresponding to different input feature combinations, a plurality of different ROC curves (Receiver operating characteristic curve, ROC) are also obtained, and according to different ROC curves, the sensitivity, specificity, positive predictive value and negative predictive value of each model for predicting early-onset preeclampsia and late-onset preeclampsia can be obtained respectively, and the optimal model can be selected according to the above parameters. At the same time, the ROC curve is constructed according to the probability of early-onset preeclampsia or late-onset preeclampsia of the pregnant woman calculated by the random forest model, and the optimal cutoff value and the area under the curve (Area under curve, AUC) are determined according to the ROC curve, and then the risk threshold is determined. The risk threshold is used as a standard for judging the risk value output by the model when predicting the risk of preeclampsia of the to-be-tested pregnant woman using the model. Specifically, if the risk threshold of early-onset preeclampsia or late-onset preeclampsia is exceeded, it is determined that the to-be-tested pregnant woman has a high risk of early-onset preeclampsia or late-onset preeclampsia.
[0068] 3.2. Compare the calculated early-onset preeclampsia risk value with the early-onset model threshold value, and compare the late-onset preeclampsia risk value with the late-onset model threshold value. If it is higher than the threshold value, it is determined to be high risk, and if it is lower than the threshold value, it is determined to be low risk. Specifically, the threshold value can be set to the threshold value when the specificity of the model is 90%, 85%, or 80%.
[0069] It should be understood that after the prediction model is constructed by the above method, the prediction models for early-onset preeclampsia and late-onset preeclampsia can be used independently when predicting the risk of preeclampsia, that is, the same pregnant woman is independently predicted for early-onset preeclampsia and late-onset preeclampsia. At this time, the data input of the relevant pregnant woman can be independent, that is, different protein data is input into the prediction model of different categories. Specifically, according to the feature input requirements of the early-onset preeclampsia prediction model, the data of each clinical index, clinical laboratory data and IGFBP1 and / or PLGF of the pregnant woman to be tested is input into the early-onset preeclampsia model; according to the feature input requirements of the late-onset preeclampsia prediction model, the data of each clinical index, clinical laboratory data and PLGF of the pregnant woman to be tested is input into the late-onset preeclampsia model to predict the early-onset and late-onset risk of the pregnant woman. It should be understood that the two models can also be combined into one, and after inputting the two protein data and each clinical index and clinical laboratory data of the pregnant woman, the prediction of early-onset preeclampsia and late-onset preeclampsia is performed by the prediction model through data recognition and retrieval.
[0070] The preeclampsia prediction model obtained by the present application can be used to predict at 11-15 weeks of pregnancy, which can provide timely medication guidance for clinicians and effectively reduce the risk of disease. It does not depend on the uterine artery pulsatility index, and clinical information and laboratory test information are easier to obtain. The prediction model of the above-mentioned clinical indicators, clinical laboratory data and protein markers improves the prediction accuracy of the current existing model, and only one or two protein markers are used, so that the detection cost required for prediction is lower. Overall, it is more convenient for clinical promotion, reduces medical expenditure, effectively improves the outcome of mother and infant, and has important significance for reducing the maternal and infant mortality rate. +6 Weeks of pregnancy can provide timely medication guidance for clinicians and effectively reduce the risk of disease. It does not depend on the uterine artery pulsatility index, and clinical information and laboratory test information are easier to obtain. The prediction model of the above-mentioned clinical indicators, clinical laboratory data and protein markers improves the prediction accuracy of the current existing model, and only one or two protein markers are used, so that the detection cost required for prediction is lower. Overall, it is more convenient for clinical promotion, reduces medical expenditure, effectively improves the outcome of mother and infant, and has important significance for reducing the maternal and infant mortality rate.
[0071] In addition, the prediction effect of the random forest machine learning algorithm is evaluated in the embodiments of the present application, and other machine learning, deep learning, reinforcement learning and the like can be applicable to the present application.
[0072] The present application will be further described in detail below by specific embodiments in conjunction with the accompanying drawings.
[0073] Embodiment:
[0074] 1. Sample collection
[0075] 1.1. Inclusion criteria:
[0076] The inclusion criteria for preeclampsia are as follows: the systolic blood pressure of the pregnant woman is ≥ 140 mmHg and / or the diastolic blood pressure is ≥ 90 mmHg after 20 weeks of pregnancy, accompanied by any one of the following: urine protein quantitative ≥ 0.3 g / 24 h, or urine protein / creatinine ratio ≥ 0.3, or random urine protein ≥ (+) (when the examination method without conditional protein quantification is carried out); without proteinuria but with any one of the following organ or system involvement: heart, lung, liver, kidney and other important organs, or abnormal changes in blood system, digestive system, nervous system, placenta-fetus involvement, etc. The inclusion criteria for early-onset preeclampsia are as follows: the blood pressure of the pregnant woman with preeclampsia is abnormal or the 24-hour proteinuria is abnormal, and the diagnosis is made at ≤ 34 weeks of gestation; the blood pressure of the pregnant woman with preeclampsia is abnormal or the 24-hour proteinuria is abnormal, and the diagnosis is made at > 34 weeks of gestation.
[0077] The inclusion criteria for healthy controls are as follows: the sample is a full-term pregnancy without pregnancy complications, the fetus is healthy at birth, and there are no obstetric, medical or surgical complications during pregnancy. The exclusion criteria are as follows: ① complicated by other pregnancy complications; ② severe heart, liver and kidney dysfunction; ③ patients with autoimmune diseases, malignant diseases; and ③ abnormal pregnant women caused by chromosomal abnormalities, congenital abnormalities, preterm birth and multiple pregnancy.
[0078] 1.2. Participants:
[0079] According to the above inclusion criteria, the remaining plasma samples of pregnant women who underwent NIPT screening in the early pregnancy were retrieved, and samples detected at 11-15 +6 weeks of gestation were screened. At the same time, the clinical information of the corresponding samples was collected, including the height of the pregnant woman, the body weight at the time of NIPT (Noninvasive Prenatal Testing) detection, the systolic / diastolic blood pressure at the time of NIPT detection, adverse medical history (including previous gestational diabetes, preeclampsia history, chronic hypertension history, systemic lupus erythematosus and antiphospholipid syndrome), assisted reproduction and other clinical data. In addition, laboratory examination information of the corresponding samples in the early pregnancy was also collected, including blood routine tests recorded in Table 1 at 0-17 +6 weeks, liver and kidney function tests at 11-15 +6 weeks, urine routine tests at 11-15 +6 weeks, and hepatitis B two pairs of half detection within one year before and after pregnancy, a total of 102 indicators. The data set in this case is shown in Table 2.
[0080] Table 2. Sample quantity
[0081] 1.3. Determination of mean arterial pressure:
[0082] The pregnant woman is comfortably seated, back supported, with legs uncrossed, and rests for 5 minutes. Then, the blood pressure is measured at least twice on both arms using an electronic sphygmomanometer. The difference between the systolic blood pressure (SBP) of the left and right arm should be <= 10 mmHg, and the difference between the diastolic blood pressure (DBP) of the left and right arm should be <= 5 mmHg. If not, wait 1 minute from the moment the cuff is deflated and repeat the set of measurements.
[0083] The measured systolic and diastolic blood pressures are averaged and introduced into the formula for calculating the mean arterial pressure (MAP):
[0084] 2. Detection of protein markers:
[0085] Using the above samples, the content of IGFBP1 protein in the plasma samples was detected using the ab233618 Human IGFBP1 SimpleStep Kit, and the content of PLGF protein was detected using the ab260056 Human PIGF SimpleStep Kit.
[0086] 2.1. Detection principle:
[0087] The two protein markers play an important role in the process of pregnancy (see Table 3), and their concentrations can be measured from blood samples. The content of IGFBP1 protein in serum samples was detected using the ab233618 Human IGFBP1 SimpleStep Kit, and the content of PLGF protein was detected using the ab260056 Human PIGF SimpleStep Kit. This kit uses a double antibody sandwich enzyme-linked immunosorbent assay technique. The specific anti-human IGFBP1 or PLGF antibody is pre-coated on a high-affinity enzyme-labeled plate. Standard and sample are added to the enzyme-labeled plate, incubated, and the IGFBP-1 or PLGF in the sample binds to the solid-phase antibody. After washing to remove unbound substances, the detection antibody is added for incubation. After washing, the color developing substrate TMB is added, and the color develops in the dark. The color reaction is proportional to the concentration of IGFBP-1 or PLGF in the sample, and the stop solution is added to stop the reaction. The absorbance value is measured at 450 nm wavelength. The data is analyzed by bioinformatics software to produce protein expression values.
[0088] Table 3. Detection indicators and clinical significance of protein markers
[0089] 2.2. Quality control standards:
[0090] In this experiment, the following quality control standards for ELISA detection were set: 1) linear regression r of the standard was greater than 0.99; 2) OD value of blank well was less than 0.1; 3) signal-to-noise ratio = OD value of the lowest concentration standard well / OD value of the blank well, signal-to-noise ratio > 1.2; 4) the OD value of the test sample should be within the range of the maximum and minimum OD values of the standard.
[0091] 3. Data Information MOM Processing:
[0092] 3.1. Maternal clinical indicators:
[0093] Clinical information collected from pregnant women includes their height and gestational age (11-15 days). +6 Zhou's weight, adverse medical history (including history of gestational diabetes, preeclampsia, chronic hypertension, systemic lupus erythematosus, and antiphospholipid syndrome), and whether she had undergone in vitro fertilization. The clinical indicators were converted as described above and will not be repeated here.
[0094] The processed data, along with mean arterial pressure, were compared between the training and validation sets for early-onset, late-onset eclampsia and healthy controls. The Wilcoxon rank sum test was used to statistically analyze continuous variables, and the chi-square test was used to analyze categorical variables. Results showed that the proportions of body mass index (BMI), mean arterial pressure (MAP), and in vitro fertilization (IVF) in pregnant women with early-onset and late-onset eclampsia in the training set were significantly higher than those in the control group (P < 0.05, Table 4). The proportion of adverse pre-eclampsia history (PMH) in early-onset eclampsia was also significantly higher in the training set (P < 0.001, Table 4). These results indicate that higher BMI, MAP, PMH, and IVF rates are high-risk factors for preeclampsia during pregnancy, significantly increasing the risk of early-onset and late-onset preeclampsia.
[0095] Table 4. Comparison of clinical characteristics of pregnant women with early-onset eclampsia and healthy controls in the training and test sets.
[0096] For continuous variables with a normal distribution, the median (standard deviation) is given; for skewed distributions, the median (25th percentile, 75th percentile) is given. For binary variables, the number of values of 1 (proportion) is given. EPE: early-onset preeclampsia; LPE: late-onset preeclampsia. *P<0.05, **P<0.01, ***P<0.001, Wilcoxon rank-sum test is used for continuous variables; chi-square test is used for binary variables.
[0097] After processing the pregnant woman's BMI and MAP as described above, a fixed median correction is applied, calculated using the following formula:
[0098] where the corresponding fixed median is:
[0099] Table 5. Fixed median values of BMI and MAP
[0100] 3.2. Clinical laboratory data:
[0101] For the laboratory examination information collected in the early pregnancy, the median correction was made for the continuous experimental information variables:
[0102] For the corrected indicators, it was found that the white blood cell count (WBC), platelet count (PLT) and monocyte ratio (MO%) in the routine blood test; uric acid (UA), high-density lipoprotein (HDL-C), glutamyl transpeptidase (GGT) and glutamic acid / alanine transaminase ratio (AS / AL) in liver and kidney function test; urine protein (NCG_PRO) in urine routine were the examination indicators with significant differences or important significance for the onset of preeclampsia.
[0103] Table 6. Comparison of laboratory examination information of early-onset, late-onset preeclampsia and healthy control pregnant women in training set and test set
[0104] For continuous variables with normal distribution, the median (standard deviation) is given; for skewed distribution, the median (25% quantile value, 75% quantile value) is given. For binary variables, the number of values with 1 (proportion) is given. EPE: early-onset preeclampsia; LPE: late-onset preeclampsia. *P<0.05, **P<0.01, ***P<0.001, Wilcoxon rank-sum test for continuous variables; chi-square test for binary variables.
[0105] 3.3. Protein markers
[0106] After the absolute quantitative concentration of IGFBP1 and PLGF was obtained by ELISA experiment for all data set samples, it was found that the protein marker concentration changed with gestational age, and the PLGF concentration significantly increased with gestational age (correlation coefficient 0.44, P value <0.01), as shown in Figure 1; the IGFBP1 concentration significantly decreased with gestational age (correlation coefficient 0.18, P value <0.01), as shown in Figure 2. For 1653 cases of healthy control samples, linear fitting was made for the relationship between the protein concentrations of IGFBP1 and PLGF and gestational days, and the corresponding expected median values of gestational days were obtained, and the fitting equation of IGFBP1 protein is as follows: IGFBP1,GA = -2.21 x 10 4 + 1.62 x 10 3 x GA
[0107] The fitting equation of PLGF protein is as follows: E(Median) PLGF,GA = -64.4 + 1.28 x GA
[0108] wherein GA is gestational age.
[0109] After obtaining the expected median value of IGFBP1 and PLGF protein corresponding to gestational age, the concentration of the protein is median-corrected:
[0110] After median correction of IGFBP1 and PLGF protein concentration, IGFBP1 in the training set of early preeclampsia samples was significantly higher than that in the healthy control samples (p value <0.05, see Figure 3a). PLGF in the validation set of early preeclampsia samples was significantly higher than that in the healthy control samples (p value <0.05, see Figure 3b). IGFBP1 in the training set of late preeclampsia samples was significantly higher than that in the healthy control samples (p value <0.05, see Figure 4a). PLGF in the validation set of late preeclampsia samples was significantly higher than that in the healthy control samples (p value <0.01, see Figure 4b). It is shown that IGFBP1 and PLGF have the value of being used as a preeclampsia disease prediction marker.
[0111] 4. Constructing early and late preeclampsia prediction model
[0112] 4.1. Model training
[0113] The clinical indicators, clinical laboratory data and protein marker concentrations processed by MoM are used to train the early and late preeclampsia model, and the input features used include: MoM value of body mass index (BMI), MoM value of mean arterial pressure (MAP), in vitro fertilization (IVF), adverse medical history (PMH), MoM value of IGFBP1 and / or MoM value of PLGF, MoM value of white blood cell count (WBC), MoM value of platelet count (PLT), MoM value of monocyte ratio (MO%), MoM value of uric acid (UA), MoM value of high-density lipoprotein (HDL-C), MoM value of glutamyl transpeptidase (GGT), MoM value of aspartate / alanine aminotransferase ratio (AS / AL), and urine protein in urine routine (NCG_PRO). In this example, a random forest machine learning model is used, and the samples in the training set in Table 2 are used to train the model, and the validation set is used as an external evaluation data set for the algorithm.
[0114] The hyperparameters Criterion, max_depth and n_estimators of the random forest model are tuned using the sklearn package of python, and the optimal hyperparameter combination is obtained.
[0115] 4.2. Model prediction
[0116] After the optimal combination of hyperparameters is obtained in the previous step, the trained random forest machine learning model is obtained. The early-onset preeclampsia and late-onset preeclampsia models use different combinations of clinical information and protein prediction markers for prediction. The prediction index is AUC (Area Under Curve), and the prediction results are shown in Tables 7 and 8 as follows:
[0117] Table 7. Prediction results of early-onset preeclampsia model
[0118] * AUC of the data set after selecting fixed hyperparameters for 10 times of model prediction, average value ± 1 / 2 range
[0119] Table 8. Prediction results of late-onset preeclampsia model
[0120] * AUC of the data set after selecting fixed hyperparameters for 10 times of model prediction, average value ± 1 / 2 range
[0121] From the data in Table 7 above, it can be found that after adding clinical laboratory data, especially the prediction model containing WBC, the validation set AUC data of the prediction model is significantly better than that of the prediction model not containing WBC, and the validation set AUC of part of the prediction model containing WBC reaches more than 0.9; similarly, in Table 8, the data of the prediction model containing NCG PRO is better than that of the prediction model not containing NCG PRO, and the validation set AUC is significantly improved. It should be understood that in the art, the above-mentioned related AUC data is very difficult to improve, especially from 0.8 to 0.9, therefore, the prediction model provided by the present application has high application value.
[0122] 4.3. Model evaluation
[0123] From the model prediction results of the previous step, select the optimal feature combination of early-onset and late-onset preeclampsia models. The optimal combination of the early-onset preeclampsia model is: BMI + IVF + PMH + MAP + IGFBP1*PLGF + WBC + MO% + PLT, and the optimal model is the random forest model. The optimal combination of the late-onset preeclampsia model is: BMI + IVF + PMH + MAP + PLGF + MO% + AS / AL + HDL-C + NCG PRO. After fixing the optimal hyperparameters, the model prediction effect still has a small fluctuation. Select the model with the optimal sensitivity under the condition of 85% specificity of the test set in 100 prediction results as the final optimal model result. The data of the relevant preferred early and late-onset preeclampsia prediction models are shown in FIG. 5, FIG. 6 and Table 9.
[0124] These results show that the model constructed in the present application performs better for predicting early-onset preeclampsia in early pregnancy.
[0125] It should be understood that existing research data shows that the complication rate of pregnant women with early-onset preeclampsia is much higher than that of late-onset preeclampsia. In particular, because pregnant women with early-onset preeclampsia often have liver and placental damage, the incidence of adverse outcomes of perinatal infants is also much higher than that of late-onset preeclampsia. Therefore, the significance and role of early-onset preeclampsia prediction are more important.
[0126] It should be understood that the specificity setting can be adjusted according to actual needs.
[0127] Table 9. Evaluation index of the optimal model of early and late onset
[0128] Those skilled in the art can understand that all or part of the functions of the various methods in the above embodiments can be realized by hardware or by a computer program. When all or part of the functions in the above embodiments are realized by a computer program, the program can be stored in a computer readable storage medium, which can include read-only memory, random access memory, magnetic disk, optical disk, hard disk, etc. The above functions are realized by executing the program by a computer. For example, the program is stored in the memory of the device, and when the program in the memory is executed by the processor, the above all or part of the functions are realized. In addition, when all or part of the functions in the above embodiments are realized by a computer program, the program can also be stored in a server, another computer, a storage medium such as a disk, an optical disk, a flash disk or a mobile hard disk, and downloaded or copied into the memory of the local device, or the system of the local device is updated, and when the program in the memory is executed by the processor, the above all or part of the functions in the above embodiments are realized.
[0129] The above application of specific examples to illustrate the present invention, is only used to help understand the present invention, and does not limit the present invention. For the skilled in the art to which the present invention belongs, according to the idea of the present invention, several simple deductions, deformation or replacement can be made.
Claims
1. A method for constructing a pre-eclampsia risk prediction model, characterized by, The preeclampsia risk prediction model is constructed based on a plurality of clinical indicators of pregnant women, a plurality of clinical laboratory data of the pregnant women, and IGFBP1 and / or PLGF, the pregnant women including both diseased pregnant women with early-onset preeclampsia or late-onset preeclampsia and healthy pregnant women.
2. The construction method of claim 1, wherein, The construction method comprises: Obtaining the clinical indicators, the plurality of clinical laboratory data, and IGFBP1 and / or PLGF of the pregnant women as relevant modeling factors, and constructing a training sample set of the modeling factors and the corresponding pregnant women, the pregnant women including both diseased pregnant women with early-onset preeclampsia or late-onset preeclampsia and healthy pregnant women; Based on the training sample set, a plurality of types of model training are performed, and the trained plurality of types of models are evaluated, and the best prediction model is confirmed according to a model evaluation index.
3. The construction method of claim 1, wherein, The construction method comprises: The plurality of clinical indicators, the plurality of clinical laboratory data, and IGFBP1 and / or PLGF are taken as inputs, and whether the pregnant women have preeclampsia is taken as an output, and a model is constructed by a machine learning method to obtain the prediction model.
4. The construction method according to claim 3, characterized in that, The machine learning method is random forest, naive Bayes, logistic regression, support vector machine, AdaBoost, K-nearest neighbor algorithm, neural network, passive attack algorithm, stochastic gradient descent, or xgboost.
5. The construction method according to any one of claims 1 to 4, characterized in that, The plurality of clinical laboratory data includes white blood cell count and / or urine protein.
6. The construction method of claim 5, wherein, The plurality of clinical laboratory data further includes any one or more of platelet count, monocyte ratio, glutamic-oxaloacetic transaminase ratio, glutamyl transpeptidase, uric acid, and high-density lipoprotein.
7. The construction method of claim 6, wherein, For early-onset preeclampsia, the plurality of clinical laboratory data includes one or more of platelet count, monocyte ratio, glutamic-oxaloacetic transaminase ratio, glutamyl transpeptidase, uric acid, high-density lipoprotein, and urine protein, and white blood cell count.
8. The construction method of claim 6, wherein, For late-onset preeclampsia, the plurality of clinical laboratory data includes one or more of platelet count, monocyte ratio, glutamic-oxaloacetic transaminase ratio, glutamyl transpeptidase, uric acid, high-density lipoprotein, and white blood cell count, and urine protein.
9. The construction method of claim 1, wherein, The IGFBP1 and / or PLGF are from the pregnant woman 11-15 +6 Peripheral blood serum / plasma of the week.
10. The construction method of claim 1, wherein, The plurality of clinical indicators include body mass index, mean arterial pressure, adverse medical history, and in-vitro fertilization of the pregnant women, and the adverse medical history includes previous gestational diabetes, preeclampsia history, chronic hypertension history, systemic lupus erythematosus, and anti-phospholipid syndrome.
11. Application of clinical laboratory data white blood cell count and / or urine protein in constructing a preeclampsia risk prediction model.
12. Use according to claim 11, characterized in that, The plurality of clinical laboratory data further includes any one or more of platelet count, monocyte ratio, glutamic-oxaloacetic transaminase ratio, glutamyl transpeptidase, uric acid, and high-density lipoprotein.
13. A method of predicting the risk of pre-eclampsia, characterized by, The prediction method comprises: Obtaining the clinical indicators, the plurality of clinical laboratory data, and IGFBP1 and / or PLGF of the pregnant women as relevant modeling factors, and constructing a training sample set of the modeling factors and the corresponding pregnant women, the pregnant women including both diseased pregnant women with early-onset preeclampsia or late-onset preeclampsia and healthy pregnant women; The plurality of clinical indicators include body mass index, mean arterial pressure, adverse medical history, and in-vitro fertilization of the pregnant women, and the adverse medical history includes previous gestational diabetes, preeclampsia history, chronic hypertension history, systemic lupus erythematosus, and anti-phospholipid syndrome.
11. Application of clinical laboratory data white blood cell count and / or urine protein in constructing a preeclampsia risk prediction model. The plurality of clinical laboratory data further includes any one or more of platelet count, monocyte ratio, glutamic-oxaloacetic transaminase ratio, glutamyl transpeptidase, uric acid, and high-density lipoprotein. The prediction method comprises: Obtaining the clinical indicators, the plurality of clinical laboratory data, and IGFBP1 and / or PLGF of the pregnant women as relevant modeling factors, and constructing a training sample set of the modeling factors and the corresponding pregnant women, the pregnant women including both diseased pregnant women with early-onset preeclampsia or late-onset preeclampsia and healthy pregnant women; The plurality of clinical indicators include body mass index, mean arterial pressure, adverse medical history, and in-vitro fertilization of the pregnant women, and the adverse medical history includes previous gestational diabetes, preeclampsia history, chronic hypertension history, systemic lupus erythematosus, and anti-phospholipid syndrome.
11. Application of clinical laboratory data white blood cell count and / or urine protein in constructing a preeclampsia risk prediction model. The plurality of clinical laboratory data further includes any one or more of platelet count, monocyte ratio, glutamic-oxaloacetic transaminase ratio, glutamyl transpeptidase, uric acid, and high-density lipoprotein. The prediction method comprises: Obtaining the clinical indicators, the plurality of clinical laboratory data, and IGFBP1 and / or PLGF of the pregnant women as relevant modeling factors, and constructing a training sample set of the modeling factors and the corresponding pregnant women, the pregnant women including both diseased pregnant women with early-onset preeclampsia or late-onset preeclampsia and healthy pregnant women; The plurality of clinical indicators include body mass index, mean arterial pressure, adverse medical history, and in-vitro fertilization of the pregnant women, and the adverse medical history includes previous gestational diabetes, preeclampsia history, chronic hypertension history, systemic lupus erythematosus, and anti-phospholipid syndrome.
11. Application of clinical laboratory data white blood cell count and / or urine protein in constructing a preeclampsia risk prediction model. The plurality of clinical laboratory data further includes any one or more of platelet count, monocyte ratio, glutamic-oxaloacetic transaminase ratio, glutamyl transpeptidase, uric acid, and high-density lipoprotein. The prediction model is obtained by the construction method of any one of claims 1-10.
14. A system for predicting the risk of pre-eclampsia, characterized in that, The prediction system comprises: a device for acquiring clinical indicators, clinical laboratory data, and data of IGFBP1 and / or PLGF of a pregnant woman to be tested; a device for performing prediction model processing on the clinical indicators, clinical laboratory data, and data of IGFBP1 and / or PLGF of the pregnant woman; a device for outputting a prediction result; The prediction model is obtained by the construction method of any one of claims 1-10.
15. A product for the prediction of the risk of pre-eclampsia, characterized in that comprise: a memory for storing a program; a processor for executing the program stored in the memory to implement the method of claim 13.
16. A computer readable storage medium characterized by: The medium has a program stored thereon, which can be executed by a processor to implement the method of claim 13.
17. A computer-readable storage medium, characterized in that, The medium has a prediction model stored thereon, which is obtained by the construction method of any one of claims 1-10.
Citation Information
Patent Citations
Methods and kits for providing a preeclampsia assessment and prognosing preterm birth
CN109891239A
Pre-eclampsia risk prediction method based on MLP multi-platform calibration
CN113724873A
Gene combination for predicting risk of premature eclampsia, prediction model and construction method of prediction model
CN115631855A
Preeclampsia risk prediction model based on key enzyme fusion index EHI
CN117747103A
A system and method of generating a model to detect, or predict the risk of, an outcome
US20220005605A1