Construction method and application of preeclampsia prediction model

By combining IGFBP1, PlGF, clinical information and mean arterial pressure, a machine learning algorithm is used to construct a preeclampsia prediction model, which solves the problem that preeclampsia cannot be effectively predicted in early pregnancy, and achieves high-accurate risk prediction and reduced patient risks.

CN120048501APending Publication Date: 2025-05-27BGI GENOMICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311594295.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing preeclampsia prediction methods cannot be effectively predicted in the early stage of pregnancy, cannot guide doctors to conduct drug intervention in the early stage of pregnancy, and require multiple markers to increase experimental costs and sample demand.

Method used

By combining insulin growth binding factor binding protein 1 (IGFBP1) and placental growth factor (PlGF) with the clinical information and mean arterial pressure of the maternal body, a machine learning algorithm was used to construct a preeclampsia prediction model to achieve risk prediction in the early pregnancy.

Benefits of technology

Effectively predict preeclampsia risks in early pregnancy, reduce patient risks, improve prediction accuracy, sensitivity reaches 92.31%, specificity reaches 90.08%, reduce medical costs, and is suitable for clinical promotion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048501A_ABST
    Figure CN120048501A_ABST
Patent Text Reader

Abstract

The invention develops a construction method of a novel preeclampsia prediction model based on insulin growth binding factor binding protein 1 (IGFBP1) and placental growth factor (PlGF) and related application of the novel preeclampsia prediction model. Specifically, according to the method, IGFBP1 and / or PLGF of a pregnant woman, clinical information and MAP data serve as input, whether the pregnant woman suffers from preeclampsia or not serves as output, model construction is conducted through a machine learning method, and a preeclampsia prediction model is obtained. As effective detection can be carried out only by adopting one or two protein predictive markers without depending on the uterine artery pulsation index, the cost is low, and on the whole, clinical popularization is more convenient, the medical expenditure is reduced, the mom and infant outcome is effectively improved, and the method has important significance on reducing the mortality rate of pregnant and lying-in women and infants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of preeclampsia analysis, and particularly to a method for constructing a preeclampsia prediction model, a prediction method, a prediction product, a prediction system and related applications thereof. Background Art

[0002] In recent years, there have been numerous articles and patents on preeclampsia prediction. The use of protein markers in blood and urine to predict preeclampsia has been commercialized, and related products have been launched by international and domestic companies. However, the use of other types of markers such as metabolites, cfRNA, cfDNA, etc. is still in the research stage.

[0003] Protein markers used in the existing related technologies for predicting preeclampsia by protein markers include soluble fms-like tyrosine kinase-1 (sFlt-1), placental growth factor (PlGF), endoglin, pregnancy-associated plasma protein A (PAPP-A), kidney injury molecule-1 (KIM1), C-type lectin domain family 4 member A (CLEC4A), etc.

[0004] Patent CN104412107 mainly uses the ratio of sFlt-1 / PlGF or endoglin / PlGF to exclude the risk of preeclampsia onset during a certain period. Patent CN103917875 dynamically uses the ratio of SFlt-1 or Endoglin / PlGF as an indicator of critical preeclampsia and / or HELLP syndrome to determine the risk of a subject developing preeclampsia in the short term. It should be understood that preeclampsia generally occurs in the late pregnancy, and domestic and international guidelines have pointed out that taking low-dose aspirin before 16 weeks can effectively prevent the risk of preeclampsia onset. Therefore, timely prediction in the early pregnancy is crucial for guiding doctors in medication. These two methods in CN104412107 and CN103917875 can only predict the risk of preeclampsia onset in the short term or in the late pregnancy, and cannot make a prediction in the early pregnancy. Therefore, they cannot effectively guide doctors in medication in the early pregnancy.

[0005] Patent CN102216468 involves measuring blood pressure and maternal factors by combining one or more of placental growth factor (PlGF) and pregnancy-associated plasma protein A (PAPP-A), and performing multivariate Gaussian analysis to determine the likelihood ratio, so as to determine the risk of preeclampsia with a false positive rate of 10% and a detection rate greater than 65% for an individual. Patent CN115116570 utilizes pregnancy-associated protein A and placental growth factor, measures the mean arterial pressure of both arms and the uterine artery pulsatility index, and modifies the risk calculation model by constructing a median equation for local early pregnancy preeclampsia screening markers, and predicts the risk of preeclampsia in pregnant women according to a reasonable cut-off value corresponding to a 15% false positive rate. However, the maternal information required in both patents CN102216468 and CN115116570 includes the uterine artery pulsatility index, and this index requires high requirements for operating physicians and needs to be professionally trained, and the fluctuation of this index has a relatively large impact on the prediction results.

[0006] Patent CN111094988 discloses that by using the measured concentrations or amounts of maternal PlGF, sFlt1, KIM1, and CLEC4A, and combining LOESS correction for gestational age, applying a machine learning algorithm classifier to the protein expression profile to determine whether to avoid unnecessary treatment for preeclampsia. Patent CN112466460 utilizes the combination of MAP, PlGF, and PAPP-A in early pregnancy pregnant women to construct a model to predict hypertensive disorders in pregnancy. However, based on the descriptions in patents CN102216468, CN115116570, CN102216468, and CN115116570, the prediction accuracy for early-onset preeclampsia in these patents is not high enough and it cannot effectively meet the actual needs.

[0007] Patent CN112105931 predicts the risk of preterm preeclampsia using metabolic biomarkers and protein biomarkers, which involves the values of various metabolites and protein markers. In practical applications, the two types of predictive markers will increase the experimental detection cost and more markers will require more biological sample volumes, so its application prospect is not ideal.

[0008] In summary, although the existing protein markers have achieved certain results in the application of preeclampsia, the existing research results or products are not sufficient to meet the actual needs, and it is urgent to develop better preeclampsia prediction methods. Summary of the Invention

[0009] Based on the deficiencies in the prior art, the present invention develops a method for constructing a new preeclampsia prediction model based on insulin-like growth factor-binding protein 1 (IGFBP1) and placental growth factor (PlGF). Specifically, by combining IGFBP1 and / or PlGF with maternal clinical information and mean arterial pressure (MAP) information, the risk of a mother developing preeclampsia can be effectively predicted in the early pregnancy through machine learning algorithms, providing a reliable implementation means for the early prevention and treatment of preeclampsia.

[0010] In the first aspect, the present invention provides a method for constructing a preeclampsia prediction model. Specifically, this method takes the data of a pregnant woman's IGFBP1 and / or PLGF, clinical information, and MAP as inputs, and whether the pregnant woman has preeclampsia as the output, and constructs a model through machine learning methods to obtain a preeclampsia prediction model.

[0011] Preferably, the clinical information includes the pregnant woman's age, height, weight at 11 - 15 +6 weeks of gestation, parity, number of previous deliveries, history of adverse pregnancy, history of adverse previous medical conditions (including previous gestational diabetes, preeclampsia, chronic hypertension, systemic lupus erythematosus, and antiphospholipid syndrome), whether this pregnancy is an in vitro fertilization pregnancy, and multiple pregnancy. The pregnant woman's height and weight at 11 - 15 +6 weeks of gestation are preferably converted to BMI.

[0012] It should be specifically noted that the mean arterial pressure (MAP), which is calculated after measuring the blood pressure of both arms of an individual, has a special importance compared with the general clinical information of pregnant women in the present invention. Therefore, the mean arterial pressure (MAP) is listed separately in the present invention.

[0013] The above machine learning methods may include random forest, naive Bayes, logistic regression, support vector machine, AdaBoost, K-nearest neighbor algorithm, neural network, passive-aggressive algorithm, stochastic gradient descent, or xgboost.

[0014] Preferably, the optimal input combination for the early-onset preeclampsia model is: clinical information + mean arterial pressure + IGFBP1 + PLGF, and the optimal model is a random forest model.

[0015] Preferably, the optimal input combination for the late-onset preeclampsia model is: clinical information + mean arterial pressure + PLGF, and the optimal model is a random forest model.

[0016] After obtaining the above prediction model, a threshold can be set according to the model evaluation index of the prediction model, such as specificity. Specifically, the threshold is set based on the specificity of the prediction model being 80% - 98%.

[0017] Second aspect, based on the above first aspect, the present invention provides the use of protein markers IGFBP1 and / or PLGF in constructing a preeclampsia prediction model.

[0018] Third aspect, the present invention provides a protein marker detection kit related to preeclampsia, and the above protein markers include IGFBP1 and / or PLGF.

[0019] Fourth aspect, based on the prediction model obtained in the above first aspect, the present invention provides a preeclampsia risk assessment prediction method, and this method includes:

[0020] 1) Obtain the data of IGFBP1 and / or PLGF, clinical information and MAP of the pregnant woman to be tested;

[0021] 2) Input the obtained data of IGFBP1 and / or PLGF, clinical information and MAP of the above pregnant woman into the prediction model for processing, obtain the risk value of preeclampsia of the above pregnant woman, and when the risk value is higher than the set threshold, it is determined that the pregnant woman has a high risk of preeclampsia.

[0022] In a specific embodiment of the present invention, for the prediction of early-onset preeclampsia, after inputting the clinical information, MAP, IGFBP1 and PLGF data of the above pregnant woman into the prediction model, the early-onset preeclampsia risk value of the pregnant woman to be tested is obtained.

[0023] In a specific embodiment of the present invention, for the prediction of late-onset preeclampsia, after inputting the clinical information, MAP and PLGF data of the above pregnant woman into the prediction model, the late-onset preeclampsia risk value of the pregnant woman to be tested is obtained.

[0024] Fifth aspect, based on the prediction method in the above fourth aspect, the present invention provides a preeclampsia prediction system, and this system includes:

[0025] 1) A device for obtaining the data of IGFBP1 and / or PLGF, clinical information and MAP of the pregnant woman to be tested;

[0026] 2) A device for performing prediction model processing on the data of IGFBP1 and / or PLGF, clinical information and MAP of the pregnant woman obtained in the above 1);

[0027] 3) A device for outputting the prediction result.

[0028] Sixth aspect, the present invention provides a preeclampsia risk assessment prediction product, and this product includes: a memory and a processor. This memory is used to store programs; this processor is used to implement the preeclampsia risk assessment prediction method as mentioned in the above fifth aspect by executing the programs stored in the above memory.

[0029] Meanwhile, the present invention also provides a computer-readable storage medium, on which a program is stored, and the program can be executed by a processor to implement the preeclampsia risk assessment and prediction method mentioned in the fourth aspect above.

[0030] In addition, the present invention also provides a computer-readable storage medium, on which a prediction model obtained by the construction method of the first aspect is stored.

[0031] The beneficial effects of the present invention are as follows: In the early pregnancy, the present invention combines clinical information, mean arterial pressure, the concentration value of insulin-like growth factor binding protein 1 (IGFBP1) alone or the concentration value of placental growth factor (PlGF) to predict the preeclampsia risk of pregnant women to be examined, without relying on the uterine artery pulsatility index. Clinical information is easier to obtain, and predicting at 11-15 +6 weeks of pregnancy can provide timely medication guidance for clinicians and effectively reduce the risk of disease. The present invention constructs a prediction model by combining newly discovered prediction markers with other indicators, improving the prediction accuracy of the existing models. The sensitivity of early-onset preeclampsia reaches 92.31%, and the specificity reaches 90.08%. At the same time, since only one or two protein prediction markers are used in the present invention for effective detection, the cost is low. Generally speaking, it is more convenient for clinical promotion, reduces medical expenses, effectively improves the maternal and child outcomes, and is of great significance for reducing the maternal and infant mortality rates. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It shows the change of PLGF concentration with gestational age in the examples, and the PLGF concentration increases significantly with the increase of gestational age;

[0033] Figure 2 It shows the change of IGFBP1 concentration with gestational age in the examples, and the IGFBP1 concentration decreases significantly with the increase of gestational age;

[0034] Figure 3 It shows the comparison of the multiples of median (MoM) values of IGFBP1 and PLGF in early-onset preeclampsia and healthy control samples in different datasets - training set (train), validation set (validation) and test set (test) in the examples, where EPE: early-onset preeclampsia, control: healthy control;

[0035] Figure 4 It shows the comparison of the MoM values of IGFBP1 and PLGF in late-onset preeclampsia and healthy control samples in different datasets - training set (train), validation set (validation) and test set (test) in the examples, where LPE: late-onset preeclampsia, control: healthy control;

[0036] Figure 5 ROC curve of the optimal model for early-onset preeclampsia in the embodiment

[0037] Figure 6 ROC curve of the optimal model for late-onset preeclampsia in the embodiment Detailed implementation manners

[0038] Preeclampsia refers to newly-occurred hypertension, proteinuria, and symptoms such as dysfunction of other organs occurring after 20 weeks of pregnancy. As the second leading cause of maternal death after postpartum hemorrhage, it is particularly crucial to predict the risk of preeclampsia outbreak in advance. Since preeclampsia generally occurs in the late pregnancy, current research shows that taking low-dose aspirin before 16 weeks of pregnancy can effectively prevent the risk of preeclampsia. Therefore, it is very crucial to timely predict whether a pregnant woman has preeclampsia in the early pregnancy for guiding doctors in medication.

[0039] Based on the above purpose, the present invention finds that in the early pregnancy, combining clinical information, mean arterial pressure, concentration value of a new protein prediction marker insulin-like growth factor binding protein 1 (IGFBP1), and concentration value of an existing protein prediction marker placental growth factor (PlGF) can effectively predict the preeclampsia risk of the tested pregnant woman. Therefore, the present application provides a method for constructing a prediction model for early-pregnancy preeclampsia based on the above research results. This method involves measuring the concentrations of biochemical markers of insulin-like growth factor binding protein 1 (IGFBP1) and placental growth factor (PlGF) in one or more blood samples from a pregnant woman. At the same time, by measuring the blood pressure of both arms of the pregnant woman, calculating the mean arterial pressure, and collecting other maternal clinical information of the pregnant woman, the above data are used with a machine learning algorithm to construct a prediction model for the risk of the tested person having early-onset and late-onset preeclampsia. Through the constructed prediction model, doctors can be assisted in intervening in the early pregnancy of the tested person and reducing the risk of preeclampsia.

[0040] The detection range of the prediction model constructed by the present invention is from 11 to 15 weeks of pregnancy +6During the first 16 weeks of pregnancy when aspirin is recommended, doctors can be provided with timely and effective medication guidance. Since the maternal clinical information used for model construction and use includes the pregnant woman's age, body mass index (BMI), parity, two or more abortions or induced labors, in vitro fertilization (IVF), previous gestational diabetes, preeclampsia, chronic hypertension, systemic lupus erythematosus, antiphospholipid syndrome, and multiple pregnancies, excluding the uterine artery pulsatility index, the required indicators are all easily measurable and do not rely on the operating experience of physicians, eliminating interference caused by human factors and thus improving the accuracy of the final prediction result. At the same time, the present invention uses a new protein prediction marker - insulin-like growth factor-binding protein 1 (IGFBP1) in the prediction, making the sensitivity of the model used in the test set for predicting early-onset preeclampsia reach 92.31% and the specificity reach 90.08%, significantly higher than the detection performance of the current existing technologies. Moreover, the protein prediction markers required by the present invention are only insulin-like growth factor-binding protein 1 (IGFBP1) and placental growth factor (PlGF), with low detection costs and being conducive to commercial use.

[0041] The model construction method in the present invention includes: obtaining the data of IGFBP1 and / or PLGF, clinical information, and MAP in the training samples and their corresponding annotation results. Among them, the annotation result is a label representing whether the sample has preeclampsia. The training samples include pregnant women with diseases and healthy pregnant women, where the pregnant women with diseases include those with early-onset preeclampsia or late-onset preeclampsia. Input the data of the above IGFBP1 and / or PLGF, clinical information, and MAP and the corresponding annotation results into a machine learning model using machine learning methods. This machine learning model can operate in a training mode. In the training mode, the data of the above IGFBP1 and / or PLGF, clinical information, and MAP and the corresponding annotation results are respectively input as the conditions and results for training. After being trained with multiple training samples, this machine learning model becomes an available model. This available model can operate in a prediction mode. In the prediction mode, this available model can output the risk value of the corresponding prediction sample having preeclampsia based on the data of IGFBP1 and / or PLGF, clinical information, and MAP input as conditions. In other words, this available model is the prediction model obtained after the construction (training) of the present invention is completed.

[0042] In a specific embodiment of the present invention, the model construction method of the present invention specifically includes 1) collection of clinical information, MAP, IGFBP1, and PlGF. Among them, collecting the clinical information of pregnant women includes the pregnant woman's age, height, and gestational weeks from 11 to 15 +6The weight in a week, the number of pregnancies, the number of deliveries, a history of adverse pregnancy, a history of adverse previous conditions (including previous gestational diabetes, preeclampsia, chronic hypertension, systemic lupus erythematosus, and antiphospholipid syndrome), whether this pregnancy is an in vitro fertilization pregnancy, and whether this pregnancy is a multiple pregnancy. Data conversion is performed on the clinical information data collected above. Specifically, BMI is calculated from height and weight; abortion - when the difference between the number of deliveries and the number of pregnancies is greater than or equal to 2, it is 1, otherwise it is 0; previous medical history - if there is any one of previous gestational diabetes, preeclampsia, chronic hypertension, systemic lupus erythematosus, and antiphospholipid syndrome, it is 1, otherwise it is 0; IVF - if this pregnancy is an in vitro fertilization pregnancy, it is 1, otherwise it is 0; multiple pregnancy - if the number of fetuses in this pregnancy is greater than one, it is 1, otherwise it is 0. The data of the above clinical information after conversion is further used for the construction of a prediction model. Obtain the mean arterial pressure of the pregnant woman, and measure the diastolic blood pressure (Systolic Blood Pressure, SBP) and systolic blood pressure (Diastolic Blood Pressure, DBP) of both arms of the pregnant woman at 11 - 15 +6 weeks, and calculate the mean arterial pressure (MAP). The calculation formula (Formula 1) is:

[0043]

[0044] The acquisition of IGFBP1 and PlGF is achieved by extracting the blood samples of the pregnant woman at 11 - 15 +6 weeks, and measuring the concentrations of insulin - like growth factor - binding protein 1 (IGFBP1) and placental growth factor (PlGF) in the plasma / serum.

[0045] Regarding the detection of protein markers IGFBP1 and PlGF, it can be detected by using the ELISA enzyme - linked immunosorbent assay technique as in the embodiments of the present invention, or other protein detection techniques can also be used, such as protein arrays, proteomics, expression proteomics, mass spectrometry (such as liquid chromatography - mass spectrometry (LC - MS), multiple reaction monitoring (MRM), selected reaction monitoring (SRM), scheduled MRM, scheduled SRM), 2D PAGE, 3D PAGE, electrophoresis, protein chips, proteomic microarrays, Edman degradation, direct or indirect ELISA, immunosorbent assay, immunological PCR, proximity extension assay, Luminex assay or homogeneous assay, time - resolved fluorescence (TRF), fluorescence oxygen channel immunoassay (FOCI) or luminescence oxygen channel immunoassay and other targeted protein detection techniques of liquid mass spectrometers, chemiluminescent immunoassay protein detection techniques, and fluorescent immunoassay protein detection techniques may all have similar detection results. It should be understood that the detection of IGFBP1 and PlGF includes, but is not limited to, other similar polypeptide sequences, various modified forms of polypeptides, and various variant forms of polypeptides.

[0046] In addition to detecting protein markers IGFBP1 and PlGF in plasma / serum samples as described above, other body fluids in the human body, such as whole blood, urine, saliva, amniotic fluid, cerebrospinal fluid, nipple aspirate, etc., can also be collected, and the corresponding model parameters can be adjusted according to the specific data of different samples.

[0047] 2) MoM processing of relevant data:

[0048] In the present invention, it is necessary to further perform median processing (MoM processing) on age, BMI, MAP, IGFBP1, and PlGF. Among them, the calculation formulas (Formula 2) for fixed median correction of age and BMI in clinical information are as follows:

[0049]

[0050] In Formula 2, MoM age / BMI is the MoM value of age or BMI after median correction, age / BMI is the original value of pregnant woman's age or BMI, and median age / BMI is the fixed median of age or BMI. In the present invention, the fixed median of age is 29, and the fixed median of BMI is 20.69.

[0051] The calculation formula (Formula 3) for fixed median correction of mean arterial pressure is:

[0052]

[0053] In Formula 3, MoM MAP is the MoM value of MAP after median correction, MAP is the original value of pregnant woman's MAP, and median MAP is the fixed median of mean arterial pressure, and the fixed median value is 82.98.

[0054] The concentrations of two protein markers IGFBP1 and PLGF are corrected by the median corresponding to different gestational weeks. The median protein concentration values at different gestational weeks are the expected values obtained from linear regression fitting of healthy populations. The fitting equation of insulin-like growth factor binding protein 1 (IGFBP1) is shown in Formula 4:

[0055] E(Median) IGFBP1,GA = 1.94×10 5 - 1.19×10 3 ×GA Formula 4

[0056] The fitting equation of placental growth factor (PlGF) is shown in Formula 5:

[0057] E(Median) PLGF,GA = -10.3 + 0.672×GA Formula 5

[0058] where GA is the gestational age, and E(Median) IGFBP1,GA and E(Median) PLGF,GA are the median IGFBP1 and PLGF protein concentrations expected for each gestational day. The protein concentration values are corrected using the median for the corresponding gestational day. The calculation formula (Formula 6) is as follows:

[0059]

[0060] In Formula 6, Concentration protein is the original value of the IGFBP1 or PLGF protein concentration in the pregnant woman, and E(Median) protein,GA is the median IGFBP1 or PLGF protein concentration expected for the corresponding gestational day, and MoM protein is the IGFBP1 or PLGF MoM value calibrated for the gestational age.

[0061] The data of age, BMI, MAP, IGFBP1, and PlGF processed by MoM above are further used for the construction of the prediction model.

[0062] 3) Construction of the prediction model

[0063] In the present invention, 10 machine learning methods are adopted, including Naive Bayes, Logistic Regression, Support Vector Machine, Random Forest, AdaBoost, K-Nearest Neighbor Algorithm, Neural Network, Passive Aggressive Algorithm, Stochastic Gradient Descent, and XGBoost. The clinical information, MAP, and protein marker concentrations after transformation and / or MoM processing are used as inputs, and whether the pregnant woman has preeclampsia is used as the output to train the early-onset and late-onset preeclampsia models. The features used in model training include: the MoM value of IGFBP1 and / or the MoM value of PLGF, as well as the MoM value of age, the MoM value of BMI, the mean arterial pressure (MAP) MoM value, parity, the number of abortions or induced labors more than 2 times, IVF, adverse past history, and multiple fetuses. At the same time, after obtaining the prediction model, the threshold for judging the preeclampsia risk can be set according to model evaluation indicators such as the specificity of the model. This threshold is used as the judgment criterion when using the prediction model for prediction, that is, for the prediction sample, its early-onset preeclampsia risk value calculated by the constructed prediction model is compared with the set early-onset model threshold, and the late-onset preeclampsia risk value is compared with the set late-onset model threshold. If it is higher than the threshold, it is judged as high risk, and if it is lower than the threshold, it is judged as low risk. It should be understood that the setting of the threshold can be the thresholds for different specificities of the model - for example, specificities of 80%, 85%, 90%, 95%, 98%.

[0064] It should be understood that in addition to the 10 machine learning algorithms evaluated in the embodiments of the present invention, those skilled in the art can also adopt other specific algorithms such as machine learning, deep learning, and reinforcement learning to perform model prediction.

[0065] In a specific embodiment of the present application, the preferred feature combination for the early-onset preeclampsia model is: clinical information + mean arterial pressure + IGFBP1 + PLGF, and the model is a random forest model; the preferred input feature combination for the late-onset preeclampsia model is: clinical information + mean arterial pressure + PLGF, and the model is a random forest model. Among them, whether the pregnant woman has early-onset preeclampsia or late-onset preeclampsia. If so (suffering from the disease), the output during model construction is set to 1; if not (normal and not suffering from the disease), the output during model construction is set to 0. After constructing the model using the training set, the validation set is then used to validate the prediction model, and the threshold is adjusted accordingly.

[0066] It should be understood that after obtaining the prediction model through the above method, when predicting the risk of preeclampsia, the IGFBP1 and / or PLGF of the pregnant woman to be tested, as well as the data of various clinical information and MAP, are input into the constructed prediction model after being processed by data conversion and / or MoM as used in the above model construction, and the judgment is made based on the risk value output by the model in combination with the threshold. When the risk value of the pregnant woman to be tested exceeds the set high-risk threshold for early-onset preeclampsia or late-onset preeclampsia, it is determined that the pregnant woman to be tested has a high risk of early-onset preeclampsia or late-onset preeclampsia. Through the above prediction method, the risk of early-onset and / or late-onset preeclampsia during pregnancy can be effectively judged, so as to realize the prediction of the risk of preeclampsia in the early pregnancy stage to guide clinicians to carry out early intervention.

[0067] It should be particularly noted that the constructed prediction models can be combined into one, that is, after uniformly inputting the IGFBP1 and PLGF data of the pregnant woman, as well as the data of various clinical information and MAP, the prediction of early-onset preeclampsia and late-onset preeclampsia is carried out through data recognition and data retrieval models respectively. Or they can be two completely independent separate models, each independently inputting and calculating for judgment.

[0068] The present invention will be further described in detail below in conjunction with the specific embodiments and the accompanying drawings.

[0069] Embodiment

[0070] 1. Sample collection

[0071] 1.1. Inclusion criteria:

[0072] Inclusion criteria for preeclampsia: After 20 weeks of pregnancy, the pregnant woman has a systolic blood pressure ≥ 140 mmHg and / or a diastolic blood pressure ≥ 90 mmHg, accompanied by any one of the following: urinary protein quantification ≥ 0.3 g / 24 h, or urinary protein / creatinine ratio ≥ 0.3, or random urinary protein ≥ (+) (the examination method when protein quantification is not possible); no proteinuria but accompanied by any one of the following organ or system involvements: important organs such as the heart, lungs, liver, and kidneys, or abnormal changes in the hematological system, digestive system, and nervous system, and involvement of the placenta-fetus, etc. Inclusion criteria for early-onset preeclampsia: The blood pressure abnormality of preeclampsia or the abnormality of 24-hour urinary protein is diagnosed at ≤ 34 weeks of pregnancy; late-onset preeclampsia: The blood pressure abnormality of preeclampsia or the abnormality of 24-hour urinary protein is diagnosed > 34 weeks of pregnancy.

[0073] Inclusion criteria for healthy controls: The sample is a full-term pregnancy without pregnancy complications, the fetus grows well at birth, and there are no obstetric, medical, or surgical complications during pregnancy. Exclusion criteria: ① Concurrent with other pregnancy complications; ② Severe heart, liver, and kidney insufficiency; ③ Patients with autoimmune diseases and malignant tumor diseases. Exclude abnormal pregnant women caused by chromosomal, congenital abnormalities, premature birth, and multiple pregnancies.

[0074] 1.2. Participants:

[0075] According to the above inclusion criteria, we retrieved the remaining plasma / serum samples of pregnant women who underwent NIPT screening in the first trimester from 8 cooperative hospitals, and screened the samples detected at 11 - 15 +6 weeks of pregnancy. At the same time, the clinical information corresponding to the samples was retrieved, including the age, height, weight at the time of NIPT (Noninvasive Prenatal Testing), systolic / diastolic blood pressure at the time of NIPT, parity, gravidity, adverse past history (including previous gestational diabetes, preeclampsia, chronic hypertension, systemic lupus erythematosus, and antiphospholipid syndrome), assisted reproduction, multiple pregnancies, and other clinical data. The sample sources are shown in Table 1.

[0076] Table 1. Brief introduction of sample sources

[0077]

[0078] 1.3. Determination of mean arterial pressure:

[0079] After the pregnant woman sits comfortably with back support and legs uncrossed for 5 minutes of rest, the electronic sphygmomanometer is used to measure the blood pressure at least 2 times in both arms of the pregnant woman. The difference in systolic blood pressure (SBP) between the left and right arms measured twice is ≤ 10 mmHg, and the difference in diastolic blood pressure (DBP) is ≤ 5 mmHg. If not met, wait for 1 minute from the moment the cuff is deflated, and then repeat a set of measurements.

[0080] The measured systolic and diastolic blood pressures are averaged and substituted into the mean arterial pressure (MAP) calculation formula:

[0081]

[0082] 2. Detection of protein markers:

[0083] Using the above samples, the ab233618 Human IGFBP1 SimpleStep Kit is used to detect the IGFBP1 protein content in the samples, and the ab260056 Human PIGF SimpleStep Kit is used to detect the PLGF protein content.

[0084] 2.1. Detection principle:

[0085] Two protein predictive markers play important roles during pregnancy (see Table 2), and their concentrations can be measured from blood samples. We use the ab233618 Human IGFBP1 SimpleStep Kit to detect the IGFBP1 protein content in the samples, and the human placental growth factor (PGF) enzyme-linked immunosorbent assay kit to detect the PLGF protein content. This kit uses the double-antibody sandwich enzyme-linked immunosorbent assay technique, and specific anti-human IGFBP1 or PLGF antibodies are pre-coated on high-affinity enzyme-linked immunosorbent assay plates. Standard products and test samples are added to the wells of the enzyme-linked immunosorbent assay plates. After incubation, the IGFBP-1 or PLGF present in the samples binds to the solid-phase antibodies. After washing to remove unbound substances, detection antibodies are added for incubation. After washing, the chromogenic substrate TMB is added and the color is developed in the dark. The intensity of the color reaction is proportional to the concentration of IGFBP-1 or PLGF in the samples. A stop solution is added to terminate the reaction, and the absorbance value is measured at a wavelength of 450 nm. The data output from the instrument is analyzed by bioinformatics software to generate protein expression values.

[0086] Table 2. Detection indicators and clinical significance

[0087]

[0088]

[0089] 2.2. Quality control standards:

[0090] In this experiment, ELISA detection quality control standards are set: 1) The linear regression r of the standard products is greater than 0.99; 2) The OD value of the blank wells is less than 0.1; 3) The signal-to-noise ratio = OD value of the lowest-concentration standard product well / OD value of the blank wells, and the signal-to-noise ratio > 1.2; 4) The OD value of the test samples needs to be within the maximum and minimum OD values of the standard products.

[0091] 3. MoM processing of data information:

[0092] 3.1 Maternal clinical information:

[0093] Collect the clinical information of pregnant women, including the age of the pregnant woman, height, weight at 11 - 15 +6 weeks of gestation, parity, number of deliveries, history of adverse pregnancy, history of adverse previous illnesses (including previous gestational diabetes, preeclampsia, chronic hypertension, systemic lupus erythematosus and antiphospholipid syndrome), whether there is in vitro fertilization (IVF), and multiple pregnancy. Convert the clinical information data:

[0094] BMI - Calculated from height and weight;

[0095] Miscarriage - If the difference between the number of deliveries and parity is greater than or equal to 2, it is 1, otherwise it is 0;

[0096] Previous medical history - If there is any one of previous gestational diabetes, preeclampsia, chronic hypertension, systemic lupus erythematosus and antiphospholipid syndrome, it is 1, otherwise it is 0;

[0097] IVF - If this pregnancy is an IVF pregnancy, it is 1, otherwise it is 0;

[0098] Multiple pregnancy - If the number of fetuses in this pregnancy is greater than 1, it is 1, otherwise it is 0.

[0099] For the processed information, plus the mean arterial pressure, by comparing the differences between the early-onset, late-onset preeclampsia and healthy controls in the training set, validation set and test set, use the Wilcoxon rank sum test to statistically analyze the differences in continuous variables among the three groups of pregnant women, and use the chi-square test to statistically analyze the differences in categorical variables. The results show that the BMI, mean arterial pressure (MAP), and proportion of adverse previous illnesses of pregnant women with early-onset preeclampsia and late-onset preeclampsia are all higher than those of the control group, and the differences are all statistically significant (P values are all < 0.01, Table 4). The age, proportion of primiparas, proportion of more than 2 times of miscarriage or induced abortion, proportion of IVF, and proportion of multiple pregnancies of pregnant women with early-onset preeclampsia and late-onset preeclampsia are higher than those of the control group in different datasets, and the differences are statistically significant (P values < 0.05, Table 4). These results show that higher age, BMI, mean arterial pressure, primiparity, history of more than 2 times of miscarriage or induced abortion, history of adverse previous illnesses, IVF, and multiple pregnancy are all high-risk factors for the occurrence of preeclampsia during pregnancy, significantly increasing the risks of early-onset and late-onset preeclampsia.

[0100] Table 4. Comparison of clinical characteristics of pregnant women with early-onset, late-onset preeclampsia and healthy controls in the training set, validation set and test set

[0101]

[0102] When the data is normally distributed, it is presented as the mean (standard deviation); when it is skewed, it is presented as the median (25th percentile value, 75th percentile value). EPE: early-onset preeclampsia; LPE: late-onset preeclampsia. *P<0.05, **P<0.01, ***P<0.001. The Wilcoxon rank sum test was used for continuous variables; the chi-square test was used for binary variables.

[0103] After the above processing of the maternal age, BMI, and MAP, fixed median correction was performed, and the calculation formula is as follows:

[0104]

[0105] Among them, the corresponding fixed median is:

[0106] Table 4. Fixed median values of age, BMI, and MAP

[0107] Clinical information Median Age 29 BMI 82.89 MAP 20.69

[0108] 3.2. Protein prediction markers

[0109] After obtaining the absolute quantitative concentrations of IGFBP1 and PLGF from all dataset samples through ELISA experiments, we found that the concentrations of protein prediction markers changed with gestational age. The concentration of PLGF increased significantly with increasing gestational age (correlation coefficient 0.18, P value <0.01), see Figure 1 ; the concentration of IGFBP1 decreased significantly with increasing gestational age (correlation coefficient -0.22, P value <0.01), see Figure 2 . We selected 1653 large-scale healthy control samples and performed linear fitting on the relationship between the protein concentrations of IGFBP1 and PLGF and gestational days to obtain the expected median values for the corresponding gestational days. The fitting equation for the IGFBP1 protein is as follows:

[0110] E(Median) IGFBP1,GA = 1.94×10 5 - 1.19×10 3 ×GA

[0111] The fitting equation for the PLGF protein is as follows:

[0112] E(Median) PLGF,GA = -10.3 + 0.672×GA

[0113] where GA is the gestational day.

[0114] After obtaining the expected median values of IGFBP1 and PLGF proteins for the corresponding gestational days, we performed median correction on the protein concentrations:

[0115]

[0116]

[0117] After median correction of the IGFBP1 and PLGF protein concentrations, the early-onset preeclampsia samples were significantly higher than the healthy control samples in all datasets (p values were all < 0.05, see Figure 3 ). IGFBP1 was significantly higher in the late-onset preeclampsia samples of the training set than in the healthy control samples (p value < 0.01, see Figure 4 ). PLGF was significantly higher in the late-onset preeclampsia samples of all datasets than in the healthy control samples (p values were all < 0.01, see Figure 4 ). This indicates that IGFBP1 and PLGF have the value of being predictive markers for preeclampsia.

[0118] 4. Construction of prediction models for early- and late-onset preeclampsia

[0119] 4.1. Model training

[0120] We used the clinical information and protein prediction marker concentrations after MoM processing to train the early-onset and late-onset preeclampsia models. The features used were: the MoM value of IGFBP1 and / or the MoM value of PLGF, as well as the MoM value of age, the MoM value of BMI, the mean arterial pressure (MAP) MoM value, parity, the number of abortions or inductions more than 2 times, IVF, adverse previous history, and multiple pregnancies. We used 10 machine learning methods, including: Naive Bayes, Logistic Regression, Support Vector Machine, Random Forest, AdaBoost, K-Nearest Neighbor Algorithm, Neural Network, Passive Aggressive Algorithm, Stochastic Gradient Descent, and XGBoost. We used the samples in the training set in Table 1 to train the models, and the validation set and test set as the external evaluation datasets for the algorithms. The hyperparameters used by different machine learning models are (refer to the sklearn package and XGBoost package):

[0121] Table 5. Hyperparameters for tuning different machine learning algorithms

[0122]

[0123] After tuning the 10 machine learning algorithms, the optimal hyperparameter combinations were obtained.

[0124] 4.2. Model prediction

[0125] After obtaining the optimal hyperparameter combinations in the previous step, the trained machine learning models were obtained. The early-onset preeclampsia and late-onset preeclampsia models were predicted using different combinations of clinical information and protein prediction markers and different machine learning models respectively. The prediction index was AUC (Area Under Curve), and the prediction results are as follows:

[0126] Table 6. Prediction Results of Early-Onset Preeclampsia Model

[0127]

[0128]

[0129] * Prediction results of 10 times of the model with fixed hyperparameters for the AUC of 3 datasets, mean ± 1 / 2 range

[0130] Table 7. Prediction Results of Late-Onset Preeclampsia Model

[0131]

[0132]

[0133] * Prediction results of 10 times of the model with fixed hyperparameters for the AUC of 3 datasets, mean ± 1 / 2 range

[0134] 4.3. Model Evaluation

[0135] From the model prediction results of the previous step, we selected the optimal feature combinations and machine learning models for the early-onset and late-onset preeclampsia models. The optimal combination for the early-onset preeclampsia model is: clinical information + mean arterial pressure + IGFBP1 + PLGF, and the optimal model is the random forest model. The optimal combination for the late-onset preeclampsia model is: clinical information + mean arterial pressure + PLGF, and the optimal model is the random forest model. After fixing the optimal hyperparameters of the random forest model, there are still slight fluctuations in the model prediction effect. We selected the model with the optimal sensitivity when the specificity of the test set is 90% (the threshold of the early-onset preeclampsia model is 0.6045, and the threshold of the late-onset preeclampsia model is 0.6260) among the 100 prediction results as the final optimal model result.

[0136] Table 8. Evaluation Metrics of the Optimal Models for Early and Late Onsets

[0137]

[0138] The above uses specific examples to illustrate the present invention, which is only used to help understand the present invention and does not limit the present invention. For those skilled in the technical field to which the present invention pertains, based on the idea of the present invention, several simple deductions, deformations or substitutions can also be made.

Claims

1. A method for constructing a preeclampsia prediction model, characterized in that, using the data of IGFBP1 and / or PLGF, clinical information and MAP of pregnant women as input, and whether the pregnant woman has preeclampsia as output, a model is constructed through a machine learning method to obtain a preeclampsia prediction model.

2. The construction method according to claim 1, characterized in that, The clinical information includes the age of the pregnant woman, height, weight at 11-15 +6 weeks of gestation, parity, number of previous deliveries, history of adverse pregnancy, history of adverse previous medical conditions, whether in vitro fertilization or multiple pregnancy exists in this pregnancy; Preferably, the adverse past history includes previous gestational diabetes, preeclampsia, chronic hypertension, systemic lupus erythematosus and antiphospholipid syndrome.

3. The construction method according to claim 1, characterized in that, For the prediction of early-onset preeclampsia, the input includes the data of IGFBP1, PLGF, clinical information and MAP; for the prediction of late-onset preeclampsia, the input includes the data of PLGF, clinical information and MAP.

4. The construction method according to any one of claims 1-3, characterized in that, The machine learning method is random forest, naive Bayes, logistic regression, support vector machine, AdaBoost, K-nearest neighbor algorithm, neural network, passive attack algorithm, stochastic gradient descent or xgboost; Preferably, the machine learning method is random forest.

5. The construction method according to claim 4, characterized in that, After obtaining the prediction model, set the threshold according to the model evaluation index of the prediction model; Preferably, the model evaluation index is the specificity of the prediction model; Preferably, the setting of the threshold is the threshold when the specificity of the prediction model is 80%-98%; Preferably, the threshold is set based on the specificity of the prediction model being 90%.

6. The construction method according to any one of claims 1-4, characterized in that, It includes: Obtain the data of IGFBP1 and / or PLGF, clinical information and MAP in the training samples and their corresponding annotation results; wherein, the annotation result is a label representing whether the sample has preeclampsia; Input the data of IGFBP1 and / or PLGF, clinical information and MAP and the corresponding annotation results into a machine learning model, and the machine learning model can run in the training mode. In the training mode, the data of IGFBP1 and / or PLGF, clinical information and MAP and the corresponding annotation results are respectively used as the conditions and results for training. After being trained by multiple training samples, the machine learning model becomes an available model. The available model can run in the prediction mode. In the prediction mode, it can output the risk value of the corresponding sample having preeclampsia according to the data of IGFBP1 and / or PLGF, clinical information and MAP used as the input conditions.

7. The construction method according to claim 2, characterized in that, Among the clinical information, the height of the pregnant woman and the weight at 11-15 +6 weeks of gestation are converted into BMI, and the age and BMI are corrected by a fixed median. The calculation formula is: The MoM age / BMI is the age or BMI MoM value after median correction, age / BMI is the original value of the pregnant woman's age or BMI, and median age / BMI is the fixed median of age or BMI, where the fixed median of age is 29 and the fixed median of BMI is 20.

69.

8. The construction method according to claim 1, characterized in that, The MAP is corrected by a fixed median, and the fixed median value is 82.98, and the calculation formula is: The MoM mentioned above MAP is the MAP MoM value after median correction, where MAP is the original value of the pregnant woman's MAP, and median MAP is the fixed median of the mean arterial pressure, and the fixed median value is 82.

98.

9. The construction method according to claim 1, characterized in that, The IGFBP1 and PLGF are corrected by the median of the corresponding gestational weeks of pregnant women. The median protein concentration values at different gestational weeks are the expected values obtained from the linear regression fitting of healthy populations. The fitting equation for insulin-like growth factor binding protein 1 (IGFBP1) is: E(Median) IGFBP1,GA = 1.94×10 5 - 1.19×10 3 ×GA The fitting equation for placental growth factor (PlGF) is: E(Median) PLGF,GA = -10.3 + 0.672 × GA where GA is the gestational age, E(Median) IGFBP1,GA and E(Median) PLGF,GA are the median concentrations of IGFBP1 and PLGF proteins expected for each gestational day; The protein concentration value is corrected by the median of the corresponding gestational days, and the calculation formula is: Concentration protein is the original value of the IGFBP1 or PLGF protein concentration in pregnant women, E(Median) protein,GA is the expected median of the IGFBP1 or PLGF protein concentration at the corresponding gestational age, MoM protein is the MoM value of IGFBP1 or PLGF calibrated by gestational age.

10. A preeclampsia prediction system characterized in that the method includes: a device for obtaining IGFBP1 and / or PLGF, clinical information, and MAP data of a pregnant woman to be tested; a device for performing a prediction model process on the obtained IGFBP1 and / or PLGF, clinical information, and MAP data of the pregnant woman; a device for outputting a prediction result; the prediction model is obtained by the construction method described in any one of claims 1-9.

11. A prediction product for preeclampsia risk assessment characterized in that it includes; a memory for storing a program; a processor for implementing the following detection method by executing the program stored in the memory; the method includes: obtaining IGFBP1 and / or PLGF, clinical information, and MAP data of a pregnant woman to be tested; inputting the obtained IGFBP1 and / or PLGF, clinical information, and MAP data of the pregnant woman to be tested into a prediction model for processing, and outputting the risk value of preeclampsia of the pregnant woman to be tested. When the risk value is higher than a set threshold, it is determined that the pregnant woman has a high risk of preeclampsia; the prediction model is obtained by the construction method described in any one of claims 1-9.

12. A computer-readable storage medium characterized in that a program is stored on the medium, and the program can be executed by a processor to implement the following method: the method includes: obtaining IGFBP1 and / or PLGF, clinical information, and MAP data of a pregnant woman to be tested; inputting the obtained IGFBP1 and / or PLGF, clinical information, and MAP data of the pregnant woman to be tested into a prediction model for processing, and outputting the risk value of preeclampsia of the pregnant woman to be tested. When the risk value is higher than a set threshold, it is determined that the pregnant woman has a high risk of preeclampsia; the prediction model is obtained by the construction method described in any one of claims 1-9.

13. A computer-readable storage medium characterized in that the medium stores a prediction model obtained by the construction method described in any one of claims 1-9.