Early prediction method and system for gestational diabetes mellitus and computer readable medium

By screening the basic clinical characteristics and biomarkers of pregnant women before 16 weeks of gestation, an early prediction model for GDM was constructed, which solved the problem of late diagnosis in existing technologies, and enabled early and accurate risk assessment and intervention for GDM, reducing the risk of adverse maternal and infant outcomes.

CN120809171APending Publication Date: 2025-10-17TIANJIN CENT OBSTETRICS & GYNECOLOGY HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511048002.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing gestational diabetes mellitus (GDM) screening methods diagnose late, which makes it difficult to meet the needs of early, efficient and personalized prediction, and lack convenient and accurate early prediction tools.

Method used

By obtaining sample data from pregnant women before 16 weeks of gestation, the most predictive basic clinical characteristics and biomarkers were screened out. A GDM early prediction model was constructed using a binary logistic regression model and machine learning algorithms, including age, BMI, and biomarkers such as IL-4, IL-10, and IL-17, to conduct GDM risk assessment.

Benefits of technology

It enables early identification of high-risk individuals for GDM before 16 weeks of gestation, reduces the risk of adverse maternal and infant outcomes, and provides a convenient and accurate GDM prediction tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120809171A_ABST
    Figure CN120809171A_ABST
Patent Text Reader

Abstract

The invention discloses a gestational diabetes mellitus early prediction method, a gestational diabetes mellitus early prediction system and a computer readable medium, and the method comprises the steps: obtaining pregnant woman sample data used for researching GDM early prediction, the pregnant woman sample data for researching the early prediction of the GDM comprises a plurality of basic clinical characteristics of each pregnant woman before 16 weeks of pregnancy, a plurality of biomarkers related to the pathogenesis of the GDM before 16 weeks of pregnancy, and a diagnosis result whether the biomarkers are GDM or not; screening a plurality of basic clinical features and a plurality of biomarkers in the obtained pregnant woman sample data to obtain basic clinical features and biomarkers with the highest GDM prediction value; and performing GDM prediction on the pregnant woman which is insufficient for 16 weeks of pregnancy by using the basic clinical features and the biomarkers with the highest GDM prediction value, so as to obtain a prediction result that GDM appears or does not appear in 24-28 weeks of pregnancy. According to the method, the GDM high-risk crowd can be predicted in the early stage of pregnancy, the GDM high-risk crowd is intervened, and the occurrence risk of bad outcomes is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of medicine and machine learning, and in particular to a method and system for early prediction of gestational diabetes mellitus (GDM) based on a machine learning algorithm, and a computer readable medium. BACKGROUND

[0002] Gestational diabetes mellitus (GDM) refers to abnormal glucose metabolism that occurs for the first time during pregnancy. The global prevalence of GDM is about 14%, reaching 14.8% in China, and showing an increasing trend year by year. GDM not only increases the risk of complications during pregnancy, such as hypertension, premature birth, and macrosomia, but also significantly increases the long-term risk of maternal and infant diabetes and cardiovascular diseases.

[0003] The current standard screening method for GDM is to perform a 75g oral glucose tolerance test (OGTT) on pregnant women without a history of diabetes at 24-28 weeks of pregnancy. However, this method has limitations such as late diagnosis and delayed intervention opportunity.

[0004] Studies have shown that identifying high-risk groups of GDM in early pregnancy and carrying out interventions can significantly reduce the risk of adverse outcomes. However, there is still a lack of convenient, accurate, and early prediction tools for GDM in clinical practice, making it difficult for existing methods to meet the needs of early, efficient, and personalized GDM prediction. SUMMARY

[0005] The present application provides a method and system for early prediction of gestational diabetes mellitus (GDM), and a computer readable medium to solve the above problems existing in the prior art.

[0006] The present application provides a method for early prediction of gestational diabetes mellitus (GDM), comprising: obtaining pregnant woman sample data for studying early prediction of GDM, wherein the pregnant woman sample data for studying early prediction of GDM includes multiple basic clinical characteristics of each pregnant woman before 16 weeks of pregnancy, multiple biomarkers related to the pathogenesis of GDM before 16 weeks of pregnancy, and a diagnosis result of whether GDM is determined at 24-28 weeks of pregnancy; screening the multiple basic clinical characteristics and multiple biomarkers in the obtained pregnant woman sample data to obtain basic clinical characteristics and biomarkers with the highest GDM prediction value; using the basic clinical characteristics and biomarkers with the highest GDM prediction value to predict GDM for pregnant women less than 16 weeks of pregnancy, and obtaining a prediction result of whether GDM occurs or not in the pregnant women less than 16 weeks of pregnancy at 24-28 weeks of pregnancy.

[0007] Preferably, the acquiring the pregnant woman sample data for studying early prediction of GDM comprises: including pregnant woman sample data of pregnant women who have conducted blood tests during pregnancy and have been diagnosed with GDM at 24-28 weeks of pregnancy; excluding, from the included pregnant woman sample data, pregnant woman sample data of pregnant women with multiple pregnancies, pregnant woman sample data of pregnant women with pregnancy complications, pregnant woman sample data of pregnant women without peripheral blood samples or unavailable peripheral blood samples, pregnant woman sample data of pregnant women with blood sampling gestational weeks after 16 weeks of pregnancy, and pregnant woman sample data of pregnant women diagnosed as obese.

[0008] Preferably, the screening the plurality of basic clinical characteristics and the plurality of biomarkers in the acquired pregnant woman sample data to obtain the basic clinical characteristics and the biomarkers with the highest prediction value for GDM comprises: selecting an optimal penalty parameter of a binary logistic regression model by cross-validation when training the constructed binary logistic regression model by using the acquired pregnant woman sample data; re-fitting the binary logistic regression model by using the optimal penalty parameter to obtain respective coefficients of the plurality of basic clinical characteristics and the plurality of biomarkers in the acquired pregnant woman sample data; and determining the basic clinical characteristics with non-zero coefficients as the basic clinical characteristics with the highest prediction value for GDM and determining the biomarkers with non-zero coefficients as the biomarkers with the highest prediction value for GDM.

[0009] Preferably, the binary logistic regression model adopts a LASSO logistic regression model.

[0010] Preferably, the plurality of basic clinical characteristics at least include age, gestational weeks, and BMI of the pregnant woman, and the plurality of biomarkers include tumor necrosis factor-α, interleukin-4, interleukin-6, interleukin-10, interleukin-17, interleukin-1β, interferon-γ, plasma fatty acid binding protein 4, high molecular weight adiponectin, leptin, insulin-like growth factor, insulin-like growth factor binding protein 2, plasma sex hormone binding globulin, testosterone, and placental growth factor.

[0011] Preferably, the basic clinical characteristics with the highest prediction value for GDM include age and BMI, and the biomarkers with the highest prediction value for GDM include interleukin-4, interleukin-10, interleukin-17, and interleukin-1β.

[0012] Preferably, the GDM prediction of the pregnant women less than 16 weeks of gestation by using the basic clinical characteristics and biomarkers with the highest GDM prediction value obtains the prediction result of the pregnant women less than 16 weeks of gestation that GDM appears or does not appear at 24-28 weeks of gestation, including: taking the basic clinical characteristics and biomarkers with the highest GDM prediction value in the acquired pregnant women samples as model input variables, taking the diagnosis result of whether the pregnant women samples are GDM as a classification label, training and verifying a plurality of machine learning models respectively to obtain an optimal machine learning model as a GDM early prediction model; inputting the basic clinical characteristics and biomarkers with the highest GDM prediction value of the pregnant women less than 16 weeks of gestation into the GDM early prediction model to predict the GDM of the pregnant women less than 16 weeks of gestation, and obtain the prediction result of the pregnant women less than 16 weeks of gestation that GDM appears or does not appear at 24-28 weeks of gestation.

[0013] Preferably, the plurality of machine learning models include Random Forest, XGBoost, KNN, SVM, and Logistic regression model; and the optimal machine learning model is Random Forest.

[0014] The application provides a pregnancy-induced diabetes early prediction system, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and when the computer program is executed by the processor, the steps of the pregnancy-induced diabetes early prediction method are realized.

[0015] The application provides a computer readable medium, which stores a pregnancy-induced diabetes early prediction method program, and when the pregnancy-induced diabetes early prediction method program is executed by a processor, the steps of the pregnancy-induced diabetes early prediction method are realized.

[0016] The application screens the basic clinical characteristics and biomarkers with the highest GDM prediction value, and uses the basic clinical characteristics and biomarkers with the highest GDM prediction value of the individual pregnant women to predict GDM early, which can assist doctors in predicting the GDM high-risk population (i.e., the pregnant women who are likely to have GDM at 24-28 weeks of gestation) from the pregnant women less than 16 weeks of gestation, and timely intervening in the GDM high-risk population to reduce the risk of adverse outcomes. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 is a flowchart of the pregnancy-induced diabetes early prediction method provided by the application;

[0018] Figure 2 is a screening flowchart of the research samples meeting the inclusion criteria in the preliminary study provided by the application;

[0019] Figure 3 is a model feature map screened out by LASSO regression provided by the application;

[0020] Figure 4 is an ROC curve and an AUC value of five machine learning models provided by the application;

[0021] Figure 5 is a confusion matrix of Random Forest in an independent validation set provided by the application;

[0022] Figure 6 is a feature importance ranking graph determined by Gini index in Random Forest provided by the application;

[0023] Figure 7 is a flowchart for predicting GDM of a pregnant woman less than 16 weeks pregnant provided by the application;

[0024] Figure 8 is a structural block diagram of the early prediction system for gestational diabetes mellitus provided by the application. DETAILED DESCRIPTION

[0025] The embodiments of the application will be described in detail below with reference to the accompanying drawings. It should be understood that the following described embodiments are only used to illustrate and explain the application, and are not used to limit the application.

[0026] Referring to Figure 1 , the application provides an early prediction method for gestational diabetes mellitus, which is used for early prediction (before 16 weeks of pregnancy) of GDM occurrence, and the method comprises the following steps:

[0027] Step S101: obtaining pregnant woman sample data for studying early prediction of GDM, wherein the pregnant woman sample data for studying early prediction of GDM comprises a plurality of basic clinical characteristics of each pregnant woman before 16 weeks of pregnancy, a plurality of biomarkers related to the pathogenesis of GDM before 16 weeks of pregnancy, and a diagnosis result of whether GDM is determined at 24-28 weeks of pregnancy.

[0028] In the historical pregnant woman sample data, the pregnant woman sample data of pregnant women who have undergone blood tests during pregnancy and have been diagnosed with GDM at 24-28 weeks of pregnancy are included, and in the included pregnant woman sample data, the pregnant woman sample data of pregnant women with multiple pregnancies, the pregnant woman sample data of pregnant women with pregnancy complications, the pregnant woman sample data of pregnant women without peripheral blood samples or unavailable peripheral blood samples, the pregnant woman sample data of pregnant women with blood sampling after 16 weeks of pregnancy, and the pregnant woman sample data of pregnant women diagnosed as obese are excluded.

[0029] The plurality of basic clinical characteristics include age, BMI, gestational age, and on the basis of the above characteristics, blood pressure, pregnancy history, weight, etc. can also be included.

[0030] The biomarkers are factors related to the pathogenesis of GDM, including but not limited to tumor necrosis factor-α (TNF-α), interleukin-4 (IL-4), interleukin-6 (IL-6), interleukin-10 (IL-10), interleukin-17 (IL-17), interleukin-1β (IL-1β), interferon-γ (IFN-γ), plasma fatty acid-binding protein 4 (FABP4), high-molecular-weight adiponectin (HWMA), leptin, insulin-like growth factor (IGF), insulin-like growth factor binding protein 2 (IGFBP2), plasma sex hormone-binding globulin (SHGB), testosterone, placental growth factor (PLGF), which can be obtained by detecting the peripheral blood of pregnant women collected before 16 weeks of pregnancy.

[0031] Step S102: screening the plurality of basic clinical features and the plurality of biomarkers in the obtained pregnant woman sample data to obtain the basic clinical features and the biomarkers with the highest GDM prediction value.

[0032] The binary logistic regression model is trained by using the obtained pregnant woman sample data, the optimal penalty parameter of the binary logistic regression model is selected by cross-validation, the binary logistic regression model is refitted by using the optimal penalty parameter, the respective coefficients of the plurality of basic clinical features and the plurality of biomarkers are obtained, the basic clinical features with non-zero coefficients are determined as the basic clinical features with the highest GDM prediction value, the biomarkers with non-zero coefficients are determined as the biomarkers with the highest GDM prediction value, the basic clinical features with the coefficients compressed to 0 and the biomarkers with the coefficients compressed to 0 are determined as unimportant basic clinical features and biomarkers, respectively, and are screened out.

[0033] The binary logistic regression model adopts a LASSO logistic regression model.

[0034] Through feature selection, the basic clinical features with the highest GDM prediction value finally obtained include age and BMI, and the biomarkers with the highest GDM prediction value include interleukin-4, interleukin-10, interleukin-17 and interleukin-1β.

[0035] Step S103: Using the basic clinical features with the highest GDM prediction value and the biomarkers, GDM prediction is performed on the pregnant women less than 16 weeks pregnant, and a prediction result of whether the pregnant women less than 16 weeks pregnant will have GDM at 24-28 weeks of pregnancy is obtained.

[0036] The present application takes the values of the basic clinical features with the highest GDM prediction value and the concentrations of the biomarkers in the obtained pregnant woman samples as model input variables, takes the diagnosis results of whether the pregnant woman samples are GDM as classification labels, trains and verifies a plurality of machine learning models including but not limited to Random Forest, XGBoost, KNN, SVM and Logistic regression model, obtains an optimal machine learning model Random Forest as a GDM early prediction model, and inputs the values of the basic clinical features with the highest GDM prediction value and the concentrations of the biomarkers of a pregnant woman less than 16 weeks pregnant into the GDM early prediction model to predict GDM of the pregnant woman, so that a prediction result of whether the pregnant woman less than 16 weeks pregnant will have GDM at 24-28 weeks of pregnancy is obtained.

[0037] The prediction result of whether the pregnant woman less than 16 weeks pregnant will have GDM at 24-28 weeks of pregnancy specifically refers to that the pregnant woman less than 16 weeks pregnant does not have diabetes, but the pregnant woman first has or does not have abnormal glucose metabolism at 24-28 weeks of pregnancy, that is, the probability of GDM occurrence (that is, the risk of GDM occurrence) is high or low. In the present application, the pregnant woman less than 16 weeks pregnant with a GDM occurrence probability higher than a threshold value is identified as a GDM high-risk group, and otherwise is identified as a GDM low-risk group.

[0038] In actual application, the present application can be applied to the following network environment:

[0039] 1. In combination with a smart phone, a tablet computer and the like personal terminal, a user uploads sample detection data and personal clinical feature information through an App, and obtains a GDM risk assessment result in real time, so that convenient individual health management is realized.

[0040] 2. A special server or terminal software program is deployed inside a hospital or a detection institution, and is responsible for data reception, model calculation and result output. Since data is processed locally, privacy security can be guaranteed, and the present application is suitable for medical scenarios with high requirements for data security.

[0041] 3. Use cloud computing platforms to provide model deployment and call services. The test results are uploaded to the cloud via the network and calculated by the cloud server, supporting remote collaboration and data sharing across regions and multiple institutions.

[0042] The present invention combines quantitatively detected biomarkers with an early GDM prediction model (or GDM risk prediction model) based on a machine learning algorithm. Compared with the existing standard GDM screening method (i.e., an oral glucose tolerance test (OGTT) after 24 weeks of gestation), the present invention can predict the risk of GDM in advance from early to mid-pregnancy (before 16 weeks of gestation) (i.e., the probability of GDM occurring after 24 weeks of gestation), which helps to achieve early intervention when a high risk of GDM is predicted, thereby reducing the risk of adverse maternal and infant outcomes.

[0043] The following combination Figures 2 to 6 The present invention will be described in detail.

[0044] 1. Biomarker Screening

[0045] Based on in-depth research of a large number of relevant literature and systematic searches and analyses in conjunction with multiple authoritative biological databases, the inventors have preliminarily identified biomarkers that may play a key role in the development and progression of GDM (see Table 1). These biomarkers play an important role in adipose inflammation, glucose metabolism, and placental regulation, and therefore are considered potential biomarkers for predicting GDM in early pregnancy.

[0046] Table 1. Preliminary screening of biomarkers (candidate biomarkers)

[0047] TNF-a IL-6 IL-1 b IL-10 IL-17A IFN-g IL-4 FABP4 HWMA Leptin IGF IGFBP2 SHGB Testosterone PLGF

[0048] 2. Sample collection and grouping

[0049] The medical records of pregnant women who visited Tianjin Central Obstetrics and Gynecology Hospital were systematically screened to ensure the representativeness of the research subjects and the reliability of the data. Figure 2 As shown in the figure, all available cases were strictly screened according to the inclusion and exclusion criteria to eliminate interfering factors that may affect the research results. The specific processing process is as follows:

[0050] Step 1: Obtain a total of 11,971 pregnant women's samples who had a discharge diagnosis (i.e., GDM diagnosed at 24-28 weeks of gestation) and underwent blood tests at Tianjin Central Obstetrics and Gynecology Hospital from 2022 to 2024.

[0051] Step 2: Exclude samples from pregnant women with multiple pregnancies; exclude samples from pregnant women with other pregnancy complications such as hypertension, eclampsia, polycystic ovary syndrome, and hyperthyroidism. A total of 5707 pregnant women samples remain.

[0052] Step 3: Excluding the cases of no peripheral blood sample, peripheral blood sample has been discarded, and other peripheral blood samples are not available, a total of 381 pregnant women samples were left.

[0053] Step 4: Excluding the samples of pregnant women whose blood was collected after 16 weeks of pregnancy, a total of 176 samples were left.

[0054] Step 5: Since obesity is an important risk factor for GDM, and the potential patients with GDM in the non-obese population are often more hidden, the samples of pregnant women diagnosed as obese were excluded to explore the predictive effect of biomarkers. On the basis of strictly implementing the above screening criteria, a total of 136 qualified pregnant women samples were finally retained as research objects, and the relevant clinical information of these pregnant women was retrospectively collected.

[0055] Step 6: According to whether they were diagnosed as GDM, the 136 pregnant women samples were grouped, among which, a total of 80 samples were in the GDM group, and a total of 56 samples were in the normal pregnancy control group (referred to as the control group). The information of pregnant women samples in the GDM group and the control group is shown in Table 2, in which the gestational age, age and BMI are displayed in the form of "mean ± standard deviation".

[0056] Table 2. Sample information of GDM group and normal control group

[0057] Variables Control group (n=56) GDM group (n=80) p value Age (years) 30.11±4.60 33.17±4.41 0.0002 Gestational age 14.64(14.14,15.43) 14.64(14.00,15.29) 0.3817 Height (m) 164.05±4.67 162.81±4.46 0.1233 Weight (kg) 57.49±6.07 60.49±7.57 0.0119 BMI 21.37±2.14 22.82±2.70 0.0007 Parity 0(0,1) 0(0,1) 0.6182 Normal weight (%) 47(83.9) 53(66.2) 0.0355 Overweight (%) 9(16.1) 27(33.8)

[0058] 3. Maternal plasma marker enzyme-linked immunosorbent assay (ELISA) analysis

[0059] According to the ELISA kit instructions, the protein concentration of the candidate biomarker was detected.

[0060] The createDataPartition method in R language was used to randomly divide the data set into training set and independent validation set in the ratio of 7:3. After random division, the training set had 96 samples and the validation set had 40 samples.

[0061] Since different biomarkers and different basic clinical features have different predictive abilities, in order to select the combination of the most contribution to prediction and the least redundancy from 15 biomarkers and multiple basic clinical features, LASSO regression model is used for feature screening in this case. Specifically, for the ELISA detection results of 15 biomarkers and basic clinical features (gestational age, BMI, age are used in this case, and more clinical features such as family history can be included when the sample is sufficient) collected in the training set, LASSO regression is used for feature screening. The steps include: first, a binary logistic regression model is constructed based on the training set, LASSO regression (alpha = 1) is realized by using the glmnet package, and the regularization parameter λ (lambda) is cross-validated by 10-fold cross-validation by the cv.glmnet function, and the minimum mean squared error (Mean Squared Error) is used as the LASSO regression model goodness standard to determine the optimal penalty parameter (lambda.min). The LASSO model is refitted under the selected optimal λ, and the regression coefficients are automatically shrunk. In the obtained regression coefficients, the variables with non-zero coefficients are considered as the key features that significantly contribute to the LASSO regression model, and after removing the intercept term, the final screened biomarkers and basic clinical features are obtained, which provide a simple, stable and discriminative input variable set for subsequent modeling. In this case, IL-1β, IL-10, IL-17, IL-4, age and BMI are finally selected as basic features.

[0062] To illustrate the differences in the expression levels of IL-1β, IL-10, IL-17, IL-4 and age and BMI between the GDM group and the control group, see Figure 3 In this case, the ggplot2 package is used to draw violin plots and box plots to show the distribution differences and individual jitter points of the two groups. In terms of significance test, Shapiro-Wilk normality test is performed on the data after removing outliers of the two groups; if both groups meet the normal distribution, two independent sample t-test is used to calculate the p value, if any group does not meet the normal distribution, Wilcoxon rank sum test is used to calculate the p value, and finally the p value of the significance test is marked on the corresponding graph. It can be seen that the six basic features selected exist significant differences (p<0.05) between the control group and the GDM group.

[0063] 4. Model training and verification

[0064] The model is constructed based on R language (version 3.6.1), in order to ensure that the data is not leaked, the training set is used for feature screening, the training of the model and the optimization of the parameters, and the determination of the best threshold value, the independent verification set is used to evaluate the generalization ability of the model, to ensure its stability and accuracy on unseen data.

[0065] Five classic machine learning models were selected for supervised training, namely Random Forest, XGBoost, KNN, SVM, and Logistic Regression. After training on the training set and determining the optimal threshold, the independent validation set was validated. Finally, the AUC of the Random Forest model was 0.995 (95% CI: 0.982-1.000), achieving the highest prediction performance, followed by XGBoost, with an AUC of 0.977 (95% CI: 0.933-1.00), KNN with an AUC of 0.941 (95% CI: 0.875-1.000), SVM with an AUC of 0.919 (95% CI: 0.839-0.999), and Logistic Regression with an AUC of 0.865 (95% CI: 0.741-0.988). See Figure 4 The ROC curves of the five models and their corresponding AUC values and 95% confidence intervals are shown. In terms of sensitivity, all models except KNN achieved a sensitivity of 1.00, with KNN achieving a sensitivity of 0.958 (95% confidence interval: 0.789-0.999). Specificity varied by model, with Random Forest having the highest specificity of 0.875 (95% CI: 0.617-0.984), followed by XGBoost with a specificity of 0.813 (95% CI: 0.544-0.960). In terms of positive predictive value (PPV), Random Forest and XGBoost showed better PPV (0.92 and 0.89, respectively). The negative predictive value (NPV) of all models was high, exceeding 0.90. See Table 3 for the AUC, sensitivity, specificity, PPV, and NPV of the five models and their 95% confidence intervals.

[0066] Table 3. Performance of five machine learning models on the independent validation set

[0067]

[0068] The AUC, sensitivity, specificity, negative predictive value, and positive predictive value of the five machine learning models shown in Table 3 were compared, and the optimal model Random Forest was selected from the five machine learning models as the classification model for early prediction of GDM. Random Forest is an ensemble learning method and a representative model of the bagging strategy. The basic idea is to build multiple decision trees and integrate the prediction results of each tree (in classification tasks, majority voting is used, and in regression tasks, the average value is taken) to improve the stability, accuracy, and generalization ability of the overall model. During model training, Random Forest generates multiple sub-samples from the original training data using bootstrap sampling with replacement, and at each tree split node, a subset of features is randomly selected for division, introducing additional randomness and reducing the correlation between decision trees.

[0069] In this example, the specific process of constructing the Random Forest model is as follows: the age, BMI, IL-1β, IL-10, IL-17, and IL-4 concentrations are used as input variables (i.e., independent variables), and the control group and GDM group are used as classification label variables (i.e., dependent variables) to construct a Random Forest classification model. In the RandomForest function, set ntree = 500, i.e., the Random Forest model trained by constructing 500 decision trees contains multiple decision trees, and each decision tree is composed of node splitting rules (such as splitting features, splitting thresholds), leaf node prediction values (class probability or class label), etc. Based on the prediction probability, the ROC curve is drawn in the training set, and the optimal classification threshold is determined using the Youden index. Then apply this threshold in the independent validation set, calculate the predicted label, and evaluate the performance indicators such as accuracy, sensitivity, specificity, AUC, and confidence interval. The confusion matrix of the Random Forest model on the independent validation set and the ranking of the Gini index (representing feature importance) are shown in Tables 2 and 3, respectively. Figure 5 and Figure 6

[0070] For subsequent calling and deployment, the model parameters are saved in a structured file format.

[0071] 5. Model application

[0072] Deploy the model file in a computer or server environment, which can perform early prediction of GDM based on the input basic features and output the early prediction results of GDM, such as the risk degree of GDM, which is high risk or low risk.

[0073] Referring to Figure 7 ​, for a pregnant woman less than 16 weeks of gestation, peripheral blood sample collection is performed on the pregnant woman, plasma is extracted, and a kit is used to detect the protein concentration of four biomarkers IL-1β, IL-10, IL-17, and IL-4. The concentrations of the four biomarkers L-1β, IL-10, IL-17, and IL-4 and the age and BMI of the pregnant woman are input into the model to obtain the risk level of the pregnant woman developing GDM at 24-28 weeks of gestation, i.e. high risk or low risk, thereby achieving efficient and accurate prediction of GDM.

[0074] The four biomarkers IL-1β, IL-10, IL-17, and IL-4 in the example can be detected using existing kits, or the existing kits can be integrated into a multi-factor ELISA detection kit for quantitatively detecting IL-1β, IL-10, IL-17, and IL-4 in peripheral blood samples to facilitate operation.

[0075] The present application can obtain GDM prediction results before 16 weeks of gestation by detecting the concentrations of specific metabolic, inflammatory, and immune-related biomarkers (such as IL-1β, IL-10, IL-17, and IL-4) in combination with the basic clinical characteristics (age, BMI) of pregnant women, thereby achieving early risk assessment of GDM, assisting doctors in early identification and intervention decision-making for high-risk groups of GDM, and effectively improving the clinical management level of GDM.

[0076] Referring to Figure 8 The present application also provides an early prediction system for gestational diabetes mellitus, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the computer program implements the steps of the early prediction method for gestational diabetes mellitus when executed by the processor. The system has wide application prospects, including but not limited to:

[0077] 1. Early screening and risk assessment of clinical GDM

[0078] In obstetric clinics, maternal and child health care hospitals, and general hospitals, as a routine screening tool for pregnant women in early to middle pregnancy, the system assists doctors in early intervention for high-risk pregnant women with GDM, thereby reducing pregnancy complications and adverse events for mothers and infants.

[0079] 2. Health management of community and primary medical institutions

[0080] The system is suitable for primary hospitals, community health service centers, and maternal and child health care stations, realizes rapid screening and remote monitoring of GDM, and promotes the improvement of public health level.

[0081] 3. Individualized health monitoring and self-assessment during pregnancy

[0082] Combined with mobile terminal application program and cloud platform service, the pregnant woman can collect samples at home or community and upload detection data through the intelligent terminal to realize dynamic monitoring and risk prompt of the health state.

[0083] In addition, the application further provides a computer readable medium, which stores a program of the early prediction method of gestational diabetes mellitus, and the program of the early prediction method of gestational diabetes mellitus realizes the steps of the early prediction method of gestational diabetes mellitus as described above when executed by a processor, and the risk of GDM is evaluated before 16 weeks of pregnancy, so that the high-risk population of GDM is early identified and intervention decision is made.

[0084] Although the application has been described in detail above, it is not limited thereto, and those skilled in the art can make various modifications according to the principles of the application. Therefore, any modification made according to the principles of the application should be understood as falling within the scope of the application.

Claims

1. A method for early prediction of gestational diabetes, characterized in that: The method comprises: Obtaining pregnant woman sample data for studying early prediction of gestational diabetes mellitus (GDM), wherein the pregnant woman sample data for studying early prediction of GDM includes multiple basic clinical characteristics of each pregnant woman before 16 weeks of gestation, multiple biomarkers related to the pathogenesis of GDM before 16 weeks of gestation, and a diagnosis of GDM determined at 24-28 weeks of gestation; Screen multiple basic clinical features and multiple biomarkers in the obtained pregnant woman sample data to obtain the basic clinical features and biomarkers with the greatest predictive value for GDM; The basic clinical features and biomarkers with the greatest predictive value for GDM are used to predict GDM in pregnant women who are less than 16 weeks of gestation, and the predicted results of whether the pregnant women who are less than 16 weeks of gestation will develop GDM or not at 24-28 weeks of gestation are obtained.

2. The method according to claim 1, characterized in that The sample data of pregnant women obtained for studying early prediction of GDM includes: The sample data of pregnant women who had undergone blood tests during pregnancy and had been diagnosed with GDM at 24-28 weeks of gestation were included; Among the included pregnant women sample data, sample data from pregnant women with multiple pregnancies, sample data from pregnant women with pregnancy complications, sample data from pregnant women without peripheral blood samples or with unavailable peripheral blood samples, sample data from pregnant women with blood sampling after 16 weeks of gestation, and sample data from pregnant women diagnosed with obesity were excluded.

3. The method according to claim 1, characterized in that The screening of multiple basic clinical features and multiple biomarkers in the obtained pregnant woman sample data to obtain the basic clinical features and biomarkers with the greatest predictive value for GDM includes: When using the obtained pregnant woman sample data to train the constructed binary logistic regression model, the optimal penalty parameter of the binary logistic regression model is selected through cross-validation; Refitting the binary logistic regression model using the optimal penalty parameter to obtain respective coefficients of multiple basic clinical characteristics and multiple biomarkers in the obtained pregnant woman sample data; The basic clinical features with non-zero coefficients were identified as the basic clinical features with the greatest predictive value for GDM, and the biomarkers with non-zero coefficients were identified as the biomarkers with the greatest predictive value for GDM.

4. The method according to claim 3, characterized in that The binary logistic regression model adopts the LASSO logistic regression model.

5. The method according to any one of claims 1 to 4, characterized in that The multiple basic clinical characteristics include at least the age, gestational age, and body mass index (BMI) of the pregnant woman; The multiple biomarkers include tumor necrosis factor-α, interleukin-4, interleukin-6, interleukin-10, interleukin-17, interleukin-1β, interferon-γ, plasma fatty acid binding protein 4, high molecular weight adiponectin, leptin, insulin-like growth factor, insulin-like growth factor binding protein 2, plasma sex hormone binding globulin, testosterone, and placental growth factor.

6. The method according to claim 5, characterized in that The basic clinical characteristics with the greatest predictive value for GDM include: age and BMI; The biomarkers with the greatest predictive value for GDM include interleukin-4, interleukin-10, interleukin-17, and interleukin-1β.

7. The method according to claim 1, characterized in that The method of using the most valuable basic clinical features and biomarkers for predicting GDM to predict GDM in pregnant women less than 16 weeks of gestation, and obtaining the prediction results of whether or not the pregnant women less than 16 weeks of gestation will develop GDM at 24-28 weeks of gestation, includes: The most predictive clinical features and biomarkers for GDM in the obtained pregnant women's samples were used as model input variables, and the diagnosis of GDM in the obtained pregnant women's samples was used as the classification label. Several machine learning models were trained and validated separately to obtain an optimal machine learning model as the early prediction model for GDM. The basic clinical characteristics and biomarkers with the most predictive value for GDM of pregnant women less than 16 weeks of gestation are input into the GDM early prediction model, and GDM is predicted for the pregnant women less than 16 weeks of gestation, to obtain the prediction results of whether the pregnant women less than 16 weeks of gestation will develop GDM at 24-28 weeks of gestation.

8. The method according to claim 7, characterized in that The several machine learning models include RandomForest, XGBoost, KNN, SVM, and Logistic Regression models; the optimal machine learning model is Random Forest.

9. A system for early prediction of gestational diabetes, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is executed by the processor, the steps of the method for early prediction of gestational diabetes mellitus according to any one of claims 1 to 8 are implemented.

10. A computer-readable medium, characterized in that A program of an early prediction method for gestational diabetes mellitus is stored thereon, and when the program of the early prediction method for gestational diabetes mellitus is executed by a processor, the steps of the early prediction method for gestational diabetes mellitus as claimed in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Application method and system of an intervention model for predicting the health of pregnant women with epilepsy

    CN122417432A