Method and system for predicting gestational diabetes risk in early pregnancy
Through efficient fine-tuning and data enhancement of multi-dimensional physiological rhythm data and pre-trained biomedical large language model parameters, stacked generalized machine learning models are trained to solve the accuracy and personalized problems of early prediction of gestational diabetes, and provide personalized health management suggestions, which improves prediction accuracy and the possibility of early intervention.
Patent Information
- Application Number
- CN202510456198.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-26
AI Technical Summary
The prior art has limited early prediction capabilities, insufficient data utilization, lack of personalization and real-time performance in the diagnosis of gestational diabetes (GDM), making it difficult to achieve accurate individualized risk prediction.
Multi-dimensional physiological rhythm data are used, pre-trained biomedical large language model is used for efficient parameter fine-tuning and data augmentation, stacked generalized machine learning models are trained, and personalized health management suggestions are generated.
Accurate prediction of gestational diabetes risk in early pregnancy is achieved, identification accuracy is improved, personalized health management advice is provided, and early intervention opportunities are strived for.
Smart Images

Figure CN120544863A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gestational diabetes prediction, and in particular to a method and system for predicting the risk of gestational diabetes in early pregnancy. Background Art
[0002] Gestational diabetes mellitus (GDM) is a glucose metabolism disorder that develops or is first diagnosed during pregnancy and is one of the most common complications of pregnancy. The prevalence of GDM in China is approximately 13.0%-20.9% and is on the rise. GDM not only significantly increases the risk of adverse pregnancy outcomes such as cesarean section, postpartum hemorrhage, and macrosomia, but also negatively impacts the long-term health of pregnant women and their offspring, increasing their risk of developing type 2 diabetes and cardiovascular disease in the future.
[0003] Currently, the diagnosis of GDM relies primarily on the oral glucose tolerance test (OGTT), which is typically performed during the second trimester (24-28 weeks). However, this method has significant limitations. Diagnosis is delayed, missing the optimal time for early pregnancy intervention and management. Furthermore, the OGTT itself places a burden on pregnant women and is difficult to implement widely in areas with limited medical resources.
[0004] To achieve early prediction and screening of GDM, existing technologies have explored some approaches, such as assessment models based on traditional risk factors (such as age, pre-pregnancy body mass index (BMI), family history, etc.), or data mining using electronic health records (EHR). However, these methods still have shortcomings:
[0005] 1. Limited early prediction capabilities. Traditional risk factor assessment models are often static and struggle to capture the dynamic physiological changes of early pregnancy, resulting in limited predictive accuracy. EHR-based analyses are also often limited by the quality and quantity of available early pregnancy data.
[0006] 2. Insufficient data utilization. Existing data analysis and prediction technologies cannot effectively utilize the large amount of multi-dimensional data contained in EHR, especially the dynamic data reflecting changes in physiological rhythms, and cannot make accurate individualized risk predictions.
[0007] 3. Data quality and quantity bottlenecks. Obtaining high-quality, large-scale early pregnancy clinical data, especially privacy-sensitive circadian rhythm data, faces ethical, cost, and technical challenges. Problems such as missing data, noise, and class imbalance severely impact the training effectiveness and reliability of machine learning models.
[0008] 4. Lack of personalization and real-time performance. Traditional management systems have single functions and lack dynamic risk prediction capabilities based on real-time data and personalized health management recommendations, making it difficult to meet the needs of clinical precision management.
[0009] Therefore, it is necessary to improve the existing gestational diabetes prediction technology to overcome the defects of the existing technology. Summary of the Invention
[0010] To overcome the problems existing in the related art, the present invention aims to provide a method and system for accurately predicting GDM risk in early pregnancy based on multi-dimensional circadian rhythm data and utilizing data augmentation and machine learning techniques. This overcomes the limited early prediction capabilities of the existing technology.
[0011] A first aspect of the present invention provides a method for predicting the risk of gestational diabetes in early pregnancy, comprising the following steps:
[0012] Acquiring a multi-dimensional physiological indicator data sample of a source pregnant woman in early pregnancy, wherein the multi-dimensional physiological indicator data sample includes at least one of sleep rhythm data, hormone cycle rhythm data, and circadian rhythm data;
[0013] Using a pre-trained large language model in the biomedical field, fine-tuning the large language model in the biomedical field through efficient parameter fine-tuning technology to generate a fine-tuned large language model;
[0014] Using the fine-tuned large language model, performing data augmentation processing on the multi-dimensional physiological indicator data samples to generate enhanced physiological data, wherein the data augmentation processing includes expanding the sample size of the multi-dimensional physiological indicator data samples, filling missing values in the multi-dimensional physiological indicator data samples, and / or balancing the distribution of the multi-dimensional physiological indicator data samples;
[0015] training a machine learning prediction model using a training dataset comprising the enhanced physiological data;
[0016] Receiving physiological indicator data of the target pregnant woman;
[0017] Inputting the physiological indicator data of the target pregnant woman into the machine learning prediction model;
[0018] A prediction result indicating the risk of the target pregnant woman developing gestational diabetes in early pregnancy is output.
[0019] By obtaining specific types of multi-dimensional physiological indicator data in early pregnancy and using a pre-trained large language model for fine-tuning and data enhancement, the problems of scarcity, low quality and uneven distribution of early pregnancy data are overcome, making the trained machine learning prediction model more robust and accurate, thereby significantly improving the accuracy of identifying gestational diabetes risk in early pregnancy and the early warning period.
[0020] Furthermore, the sleep rhythm data includes sleep duration, sleep onset time, sleep efficiency and / or sleep stage data;
[0021] The hormone cycle rhythm data includes change trend data of basal body temperature, estrogen level, progesterone level and / or gonadotropin level;
[0022] The circadian rhythm data includes heart rate variability, cortisol level rhythm, melatonin secretion rhythm and / or activity level change rhythm data;
[0023] The multi-dimensional physiological indicator data sample also includes triglyceride glucose index data.
[0024] By clarifying specific physiological indicators, the dimensionality of the input data and its relevance to the pathophysiological mechanisms of GDM are enhanced, providing a higher-quality, more informative data foundation for subsequent model training and helping to improve the accuracy and interpretability of the prediction model. Incorporating triglyceride-glucose index (TyG) data into the input data significantly enhances the direct representation of core metabolic status within the data dimensionality, enabling the model to more comprehensively and deeply assess GDM risk in pregnant women, particularly with an advantage in identifying early insulin resistance, thereby improving the overall performance and accuracy of the prediction model.
[0025] Furthermore, the pre-trained large language model in the biomedical field is a BioMedLM model.
[0026] The BioMedLM model was selected and the biomedical knowledge gained from its pre-training was utilized to make model fine-tuning more efficient and data augmentation more consistent with physiological reality, thereby improving the professionalism of the entire prediction process and the reliability of the final prediction results.
[0027] Furthermore, the expanding the sample size of the multi-dimensional physiological indicator data sample includes generating synthetic sample data based on the data distribution of the scarce physiological rhythm pattern using the fine-tuned large language model, wherein the synthetic sample data statistically reflects the characteristics of the scarce physiological rhythm pattern, and the synthetic sample data is a new data instance different from the existing samples in the multi-dimensional physiological indicator data sample;
[0028] Filling missing values in the multi-dimensional physiological indicator data sample includes generating a complete simulated sample containing filling values for the physiological indicator data sample with missing attributes;
[0029] Balancing the distribution of the multi-dimensional physiological indicator data samples includes generating corresponding enhanced physiological data that conforms to physiological relevance, in combination with the pregnant woman's age, pre-pregnancy BM or imaging data.
[0030] Through specific data augmentation strategies, common problems such as data scarcity, missing data, and imbalance are specifically addressed, while ensuring the quality and physiological relevance of the augmented data, laying the foundation for training more robust, more generalized, and closer to the real world machine learning prediction models.
[0031] Furthermore, the machine learning prediction model is a stacked generalization model;
[0032] The training of the machine learning prediction model includes:
[0033] Using the training data set to train multiple base classifier models;
[0034] Aggregating the prediction results of the basic classifier model on the training data set to form meta-features;
[0035] A meta-model is trained using the meta-features, and the meta-model constitutes the stacked generalization model.
[0036] By adopting a stacked generalization model architecture, by integrating the predictive capabilities of multiple basic classifiers and optimizing them by a meta-model, the advantages of different algorithms can be fully utilized, thereby obtaining superior prediction accuracy and stronger stability than a single model, ultimately improving the overall performance of GDM risk prediction.
[0037] Furthermore, the plurality of basic classifier models include at least three selected from a K-nearest neighbor model, a logistic regression model, a random forest model, an extreme gradient boosting model, an adaptive boosting model, an optical gradient boosting machine model, and a gradient boosting decision tree model;
[0038] The meta-model is a logistic regression model.
[0039] We selected a variety of models as base classifiers, including K-nearest neighbor (instance-based), logistic regression (linear), random forest (bagging ensemble), XGBoost / AdaBoost / LightGBM / GBDT (boosting ensemble), to ensure model diversity. These models operate based on different principles and can learn data features from different perspectives. The diversity of base classifiers ensures that the meta-features input to the meta-model contain richer and more differentiated information, thus providing the meta-model with a better foundation for learning how to combine them. The logistic regression model was chosen as the meta-model because it is stable and well-interpreted. It excels at handling linearly separable or nearly linearly separable meta-feature spaces, can effectively combine the predictions of the base models, and is not prone to overfitting.
[0040] Furthermore, the prediction result includes a risk level score and / or personalized health management recommendations generated based on the risk level score.
[0041] Providing risk level scores (such as low, medium, high risk, or continuous scores) is more informative than simple "yes / no" predictions and can reflect differences in the degree of risk. Risk grading helps doctors and pregnant women understand the risk status more clearly, so as to develop differentiated monitoring plans and management strategies (for example, routine prenatal checkups for low-risk patients, enhanced monitoring, nutritional guidance, lifestyle interventions, etc. for high-risk patients). Personalized health management recommendations generated based on risk level scores transform abstract risk assessments into specific, executable action guidelines (such as dietary adjustment recommendations, exercise recommendations, monitoring frequency recommendations, etc.), improving the compliance and effectiveness of pregnant women's self-management. Outputting risk level scores and personalized recommendations makes the prediction results not only a notification of risk, but also the starting point for subsequent health management actions, significantly improving the clinical application value of the prediction method, and can directly guide personalized intervention measures to promote maternal and child health.
[0042] A second aspect of the present invention provides a system for predicting the risk of gestational diabetes in early pregnancy, comprising:
[0043] A data acquisition unit, used to obtain multi-dimensional physiological indicator data samples of the source pregnant woman in the early stages of pregnancy;
[0044] A model fine-tuning unit, configured to fine-tune the pre-trained large language model in the biomedical field to generate a fine-tuned large language model;
[0045] a data enhancement unit, configured to perform data enhancement processing on the multi-dimensional physiological indicator data samples using the fine-tuned large language model to generate enhanced physiological data;
[0046] a model training unit, configured to train a machine learning prediction model using a training data set comprising the enhanced physiological data;
[0047] A data receiving unit, configured to receive physiological indicator data of a target pregnant woman;
[0048] The risk prediction unit is used to input the physiological indicator data of the target pregnant woman into the machine learning prediction model and output a prediction result indicating the risk of the target pregnant woman developing gestational diabetes in the early stages of pregnancy.
[0049] Preferably, the system further comprises:
[0050] A backend processing server, which is built using a Spring Boot-based technical framework and is used to host and execute logical operations of the model fine-tuning unit, data enhancement unit, model training unit, and risk prediction unit;
[0051] A data storage component, comprising a MySQL database and a Redis cache database. The MySQL database is a relational database used to persistently store the multi-dimensional physiological indicator data samples, the enhanced physiological data, the trained machine learning prediction model, and basic information of the user and pregnant woman. The Redis cache database is a cache mechanism used to cache hot data to improve system response speed.
[0052] The backend processing server interacts with the relational database through the data persistence layer.
[0053] Hot data refers to data with high access frequency and high value, and typically requires caching to improve performance and user experience. Hot data can be categorized as static or dynamic. Static data can be predicted in advance, while dynamic data requires real-time monitoring and discovery. Hot data can be discovered through two methods: manual prediction and system inference. Manual prediction relies on the experience of operators, while system inference can utilize access logs or the Redis LFU mechanism.
[0054] Preferably, the system further comprises:
[0055] A user management unit for managing user account registration, login, information modification, and authority control; an information management unit for interacting with the data acquisition unit to implement the addition, editing, deletion, and batch import and export of pregnant women's basic information and physiological indicator data samples;
[0056] Medical record management unit, used for storing, structurally displaying and conditionally retrieving electronic medical record information related to the target pregnant woman;
[0057] The interactive interface unit is built based on a preset front-end technology framework and is used to provide a graphical operation interface to receive user input instructions and display the prediction results and management information output by the risk prediction unit.
[0058] The beneficial effects of the present invention are:
[0059] The present invention provides a method for predicting the risk of gestational diabetes mellitus in early pregnancy. The method for predicting the risk of gestational diabetes mellitus in early pregnancy obtains multi-dimensional physiological indicator data of pregnant women in early pregnancy (especially sleep, hormone and circadian rhythm data), uses a biomedical large language model with efficient parameter fine-tuning (PEFT) for data enhancement, and then trains a machine learning model for prediction, thereby achieving accurate prediction of the risk of gestational diabetes mellitus (GDM) in early pregnancy.
[0060] Changes in physiological rhythms (sleep, hormones, day and night) in early pregnancy are early sensitive indicators of endocrine and metabolic status. Capturing these multi-dimensional data provides a basic information source for early risk identification, which is superior to traditional methods that rely on mid- and late-stage blood glucose testing or limited risk factors.
[0061] The biomedical Large Language Model (LLM) is pre-trained and capable of understanding complex biomedical information. Fine-tuning it through PEFT enables it to more expertly understand and process circadian rhythm data. Using the fine-tuned LLM for data augmentation (expansion, padding, and balancing) effectively overcomes the challenges of insufficient sample size, missing data, and class imbalance often faced by clinical data, generating high-quality, physiologically accurate augmented data.
[0062] Using a training set containing enhanced physiological data to train machine learning prediction models, the model's training data is richer, more complete, and more balanced, allowing it to learn more robust and generalized patterns, significantly improving the accuracy and reliability of the prediction model.
[0063] Ultimately, by feeding the trained model with the target pregnant woman's early physiological indicators, it is possible to output GDM risk prediction results in the early stages of pregnancy, before physiological changes are apparent and routine checkups have been performed. This buys valuable time for early intervention and management. This allows for early identification of high-risk women, enabling preventive measures and personalized management. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a system block diagram of a system for predicting the risk of gestational diabetes in early pregnancy provided in this application;
[0065] Figure 2 It is a construction diagram of the Stacking model provided in this application. DETAILED DESCRIPTION
[0066] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although preferred embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present invention more thorough and complete and to fully convey the scope of the present invention to those skilled in the art.
[0067] For ease of understanding, some key technical terms involved in this application are explained as follows:
[0068] Gestational diabetes mellitus (GDM) refers to abnormal glucose tolerance that occurs or is first discovered during pregnancy.
[0069] Multi-dimensional physiological indicator data samples: refers to data records containing multiple variables from different physiological systems that reflect the body's state, especially those that change over time and can reflect physiological rhythms.
[0070] Physiological Rhythm: refers to the phenomenon of periodic changes in life activities, such as sleep-wake cycles, hormone secretion cycles, and diurnal fluctuations in heart rate and body temperature.
[0071] Pre-trained Large Language Model (LLM): refers to a deep learning model that is pre-trained on a large-scale text or data corpus, typically with hundreds of millions to trillions of parameters, and can understand and generate human-like text or other sequence data.
[0072] Large language models in the biomedical field: These refer to LLMs that are pre-trained or optimized on biomedical-related professional text and data (such as medical literature, clinical records, and gene sequences), giving them enhanced capabilities for understanding and processing biomedical knowledge. BioMedLM is an example of such a model.
[0073] Parameter-Efficient Fine-Tuning (PEFT): refers to a set of techniques for updating only a small fraction of the model parameters (for example, by introducing additional adapter layers or low-rank matrices) when fine-tuning a pre-trained model for downstream tasks, thereby significantly reducing computational and storage costs and helping to prevent catastrophic forgetting (i.e., the model forgets pre-trained knowledge when learning a new task).
[0074] Data augmentation: refers to a strategy to increase the amount or diversity of training data by applying various transformations or generation techniques, with the aim of improving the generalization and robustness of machine learning models, especially when the original data is limited or unbalanced.
[0075] Machine learning predictive model: refers to a model that uses algorithms to learn patterns from data and is used to make predictions (such as classification or regression) on new data.
[0076] Ensemble Learning: A machine learning paradigm that builds and combines multiple learners (often called base learners) to complete learning tasks, and its performance is usually better than any single base learner.
[0077] Stacking Generalization (Stacking): A specific ensemble learning method that trains a "meta-model" to learn how to best combine the predictions of multiple different base learners.
[0078] Base Classifier: In the first layer of the Stacking model, a learner that is used to make predictions directly on the original (or augmented) data.
[0079] Meta-features: In the Stacking model, the prediction output of the base classifier on the training data is used as the input feature of the training meta-model.
[0080] Meta-model (Meta-learner): The second-layer learner in the Stacking model, which takes the predictions (meta-features) of the base classifiers as input and produces the final integrated prediction results.
[0081] Example
[0082] like Figure 1 and Figure 2 As shown, this embodiment provides a method and system for predicting the risk of gestational diabetes in early pregnancy. The method for predicting the risk of gestational diabetes in early pregnancy includes the following steps:
[0083] Step 1: Obtain multi-dimensional physiological indicator data samples of the source pregnant woman in the early stages of pregnancy
[0084] The core of this step is to collect data for subsequent model training and fine-tuning. The system obtains multi-dimensional physiological indicator data records of a group of "source pregnant women" (i.e., the historical pregnant population used to build the model) during early pregnancy (usually before 12 or 13 weeks of gestation) through an interface or data import method.
[0085] These data samples are time series data, meaning they contain observations from multiple time points, reflecting the changing trends and rhythmicity of physiological indicators over time. The emphasis on "multidimensionality" is because a single indicator often cannot fully reflect complex physiological states. These data capture key physiological rhythm information by covering at least one or more of the following aspects:
[0086] 1. Sleep rhythm data: For example, daily sleep duration (hours), sleep onset time (specific time), sleep efficiency (total sleep time / percentage of time in bed) collected through wearable devices (such as smart bracelets and watches) or sleep diaries, and more detailed sleep staging data (such as the duration or proportion of deep sleep, light sleep, and REM sleep). These indicators together indicate the regularity and quality of the pregnant woman's sleep-wake cycle.
[0087] 2. Hormone cycle rhythm data: For example, daily morning basal body temperature (BBT) records, which form a temperature change curve; and regular (e.g., weekly or biweekly) blood or urine tests, which measure the changing trends of estrogen, progesterone, and gonadotropin levels (e.g., LH, FSH). These data reflect the dynamic changes in the endocrine environment during early pregnancy.
[0088] 3. Circadian rhythm data: For example, continuous heart rate variability (HRV) data collected by wearable devices reflects the diurnal regulation of the autonomic nervous system; daily rhythms of cortisol levels detected by timed saliva or blood samples (such as morning, noon, evening, and bedtime sampling); melatonin secretion rhythms (or the excretion rhythms of its metabolites such as 6-hydroxymelatonin sulfate) detected by urine or blood; and hourly or more granular activity rhythm data recorded by devices (such as mobile phones and wristbands). These indicators depict the cyclical fluctuations of key bodily functions within a 24-hour period.
[0089] The acquired data needs to be standardized, such as unifying the timestamp format, handling obvious collection errors, associating a unique pregnant woman identifier, etc., to form a structured data sample library.
[0090] Step 2: Fine-tune using a pre-trained large language model in the biomedical field
[0091] The inventors realized that directly using raw physiological data may make it difficult to train a high-performance model due to limited data volume, noise, or missing data. This application introduces an advanced pre-trained large language model in the biomedical field to prepare for data enhancement.
[0092] The BioMedLM model was used as the model. This model was chosen because it has been pre-trained on a large amount of biomedical text and data, contains rich biomedical prior knowledge, and has a strong ability to understand physiological concepts and relationships.
[0093] Since the pre-trained model is general, it needs to be adapted to the specific data and objectives of this task. Therefore, the parameter efficient fine-tuning technique (PEFT) is used to fine-tune the model. The core idea of PEFT is:
[0094] Using the multi-dimensional physiological indicator data samples obtained in step 1, only a small portion of the pre-trained model parameters are adjusted (for example, by adding and training low-rank matrices through the low-rank adaptation LoRA method, or by introducing the Adapter module), rather than training all parameters. Doing so can significantly reduce the computing resources and time required for fine-tuning; effectively avoid "catastrophic forgetting," which is the loss of valuable general knowledge learned in the pre-training phase when learning new tasks; enable the model to effectively capture and learn to generate data patterns that conform to the specific data patterns of this task—that is, the physiological rhythm changes and characteristics unique to early pregnancy, such as changes in sleep structure, fluctuation patterns of specific hormones, and the stability or disorder of circadian rhythms; while maintaining the model's original biomedical knowledge generalization capabilities, it helps generate more reasonable data that is more consistent with physiological logic.
[0095] After PEFT fine-tuning, a fine-tuned large language model optimized for "early pregnancy physiological rhythm data" was obtained.
[0096] Step 3: Perform data augmentation using the fine-tuned large language model
[0097] Using the fine-tuned large language model obtained in the previous step, perform data augmentation on the original multi-dimensional physiological indicator data samples. The purpose of data augmentation is to artificially expand and optimize the training dataset to overcome potential limitations of the original data and ultimately improve the performance of the downstream prediction model. Data augmentation can be implemented in one or more of the following ways:
[0098] 1. Expand the sample size, especially for scarce patterns: If the original data contains very few samples of certain specific physiological rhythm patterns (for example, abnormal rhythm combinations that indicate high risk), or if certain subgroups of pregnant women (such as those in a specific age group or within a specific BMI range) are underrepresented, a fine-tuned large language model can be used to generate new synthetic physiological data samples. The key is that these generated samples should be similar to the real scarce samples in terms of statistical properties and dynamic dependencies of the time series (such as autocorrelation and periodicity), but not identical. This increases the diversity of the data and helps the model learn more robust features.
[0099] 2. Filling missing values: Missing data is inevitable during the raw data collection process. Using a fine-tuned large language model, we infer and generate complete simulated samples with filled-in values based on other available physiological indicator data in the sample as context. For example, the model can reasonably estimate the missing hormone values for a particular day based on a pregnant woman's sleep and activity data over a certain period of time, as well as hormone levels from the days before and after. This filling based on context and learned physiological correlations is more likely to be closer to the actual situation than simple mathematical interpolation or mean filling, thus generating high-quality simulated data.
[0100] 3. Balanced sample distribution / generation of conditional data: In order to enable the model to better learn the combined impact of different factors on GDM risk, conditional data generation can be performed. For example, combined with the age of the pregnant woman, pre-pregnancy BMI (body mass index), or early imaging data (such as ultrasound measurement data, if available and found to be associated with rhythm or GDM risk), the fine-tuned large language model is used to generate enhanced physiological data corresponding to this conditional information and consistent with known physiological associations (such as possible hormone rhythm characteristics under specific age and BMI combinations). This helps the model make more accurate predictions under different conditions and improves the overall distribution balance.
[0101] Through these enhancement methods, an enhanced physiological dataset is finally obtained for model training. Compared with the original dataset, this dataset is usually larger in scale, more complete in information, and more balanced in distribution.
[0102] Step 4: Train the machine learning prediction model
[0103] A machine learning prediction model for GDM risk was trained using a training dataset containing augmented physiological data. Stacking is a powerful ensemble learning method whose core concept is to build a hierarchical model structure, including:
[0104] 1. Training Base Classifiers: Use the augmented physiological dataset to train multiple different base classifier models. Heterogeneous base learners are generally more effective because they capture data characteristics from different perspectives. These base classifiers can be selected from: K-Nearest Neighbor (KNN), Logistic Regression (LR), Random Forest (RF), Extreme Gradient Boosting (XGBoost), Adaptive Boosting (AdaBoost), Light Gradient Boosting Machine (LightGBM), and Gradient Boosted Decision Tree (GBDT). These are all mature and well-performing classification algorithms.
[0105] 2. Generate meta-features: Apply the trained base classifiers to the training dataset (usually using K-fold cross-validation to avoid data leakage) to obtain the prediction results (e.g., class probabilities) given by each base classifier for each sample. These prediction results constitute a new feature set, namely meta-features.
[0106] 3. Meta-model training: Using the generated meta-features as input, a meta-model is trained. The meta-model's task is to learn how to make a final, better prediction based on the predictions of the individual base classifiers. A logistic regression model is preferred for its simplicity, efficiency, and interpretability.
[0107] The meta-model that is finally trained is the entire Stacking prediction model.
[0108] Feature engineering may be necessary before feeding augmented data into the base classifier. Time-domain features (such as mean, variance, kurtosis, and autocorrelation coefficient) and / or frequency-domain features (such as power spectral density and dominant frequency obtained through Fourier transform) can be explicitly extracted from time series data. These features can directly quantify the periodicity, volatility, or multi-rhythm synchronization of physiological rhythms. These extracted features can be used as model input alongside or in place of the original data points.
[0109] The performance evaluation of the model is shown in the following table:
[0110]
[0111]
[0112] In addition, feature selection (such as based on correlation, mutual information, and model built-in importance scores) or dimensionality reduction (such as PCA) techniques can be applied to screen out the most valuable features for prediction and reduce model complexity.
[0113] The model's hyperparameters (such as the number and depth of trees in LightGBM, the K value in KNN, the combination of base learners in Stacking, etc.) need to be carefully adjusted on the validation set through cross-validation combined with grid search or other optimization algorithms to find the optimal configuration.
[0114] Step 5: Receive physiological indicator data of the target pregnant woman
[0115] This step marks the beginning of the model application phase. When a prediction is needed for a new pregnant woman (the "target pregnant woman"), the system receives the target pregnant woman's physiological indicator data from early pregnancy through the user interface, API, or other means. The data type and format should remain consistent with the training phase.
[0116] Step 6: Input the physiological indicator data of the target pregnant woman into the machine learning prediction model
[0117] The received target pregnant data is preprocessed and feature extracted in the same way as the training data, and then input into the machine learning prediction model (i.e., the Stacking model) trained and optimized in step 4.
[0118] Step 7: Output prediction results
[0119] The model calculates based on the input pregnant woman's data and ultimately outputs a prediction result of the target pregnant woman's risk of developing GDM in early pregnancy. The result includes:
[0120] 1. Risk level score: A quantitative score or probability value that intuitively reflects the level of risk. The system can map the score to a specific risk level, such as "low risk," "medium risk," or "high risk," based on preset thresholds (for example, based on the model's performance on a validation set).
[0121] 2. Personalized health management recommendations: Based on the predicted risk level score, the system can automatically generate or prompt the generation of corresponding personalized health management recommendations. For example, high-risk pregnant women may be advised to increase the frequency of blood sugar monitoring, receive nutritional counseling, and adjust their lifestyle; low-risk pregnant women may be advised to maintain regular prenatal checkups. These recommendations are intended to assist doctors and pregnant women in implementing early and effective interventions.
[0122] This embodiment also provides a system for predicting the risk of gestational diabetes in early pregnancy, which is used to perform the above method steps. The system includes the following functional units:
[0123] Data acquisition unit: responsible for interacting with external data sources (such as hospital information system HIS, electronic medical record EHR, wearable device platform API, or manual data entry interface) to obtain and preliminarily process the early pregnancy multi-dimensional physiological indicator data samples of the source and target pregnant women.
[0124] Model fine-tuning unit: Contains the logic for executing the parameter efficient fine-tuning (PEFT) algorithm. It can load a pre-trained large language model in the biomedical field (such as BioMedLM) and fine-tune it using the acquired source pregnant woman's physiological data samples. Finally, it generates and stores a fine-tuned large language model suitable for this task.
[0125] Data enhancement unit: calls the fine-tuned large language model generated by the model fine-tuning unit, processes the original physiological data samples according to the preset data enhancement strategy (such as targeted expansion of scarce samples, filling missing values, generating data based on conditions, etc.), and generates an enhanced physiological data set.
[0126] Model training unit: Implements the training process of a machine learning prediction model (preferably a stacking model), including training a base classifier using augmented physiological datasets, generating meta-features, training a meta-model, and performing necessary feature engineering, feature selection, and hyperparameter optimization. The trained model is persistently stored.
[0127] Data receiving unit: provides a standardized interface (such as RESTful API) or an interactive interface to receive physiological indicator data of target pregnant women who need risk prediction.
[0128] Risk prediction unit: loads the final prediction model trained and stored by the model training unit, receives the target pregnant woman data from the data receiving unit, performs prediction calculations, and outputs structured prediction results (including risk level scores and / or personalized recommendations).
[0129] Backend processing servers: Serving as the system's core processing engine, they are responsible for executing the compute-intensive tasks and business logic of all the aforementioned core units. Backend services are built using a Spring Boot-based framework, facilitating rapid development, deployment, and scalability. The persistence layer interacts with the database through data persistence frameworks such as MyBatis, decoupling business logic from data storage.
[0130] Data storage components: These components provide stable and efficient data storage and access capabilities, including relational databases (such as MySQL) and cache databases (such as Redis). Relational databases serve as the primary persistent storage solution, storing structured data, including but not limited to: user account information, basic information about pregnant women, original and enhanced multi-dimensional physiological indicator data samples, trained machine learning prediction models (or their key parameters / structures), generated prediction result records, and related electronic medical record summaries. The cache database serves as a caching mechanism, storing hot data (e.g., user session information, commonly used configurations, some frequently accessed prediction results, etc.) to improve system response speed and reduce access pressure on the main database.
[0131] User management unit: implements user identity authentication and authorization management, including user registration, login, information modification, password management, role assignment and permission control, to ensure that only authorized users (such as doctors and researchers) can access corresponding functions and data.
[0132] Information Management Unit: This unit manages the core data within the system, specifically the basic information of pregnant women and physiological indicator data samples. It supports adding, editing, deleting, and querying operations, and provides batch import and export functions to facilitate data initialization and maintenance. This unit works closely with the Data Acquisition Unit.
[0133] Medical record management unit: used to store and manage electronic medical record information related to the target pregnant woman (or source pregnant woman) (which can be a structured summary or key information points), supports structured display (such as by timeline, by indicator type) and conditional retrieval (such as by pregnant woman ID, time range, diagnostic keywords), and provides a more comprehensive clinical background reference for risk prediction.
[0134] Interactive interface unit: Serves as the front-end interface for users to interact with the system. It is built based on a preset front-end technology framework, uses Thymeleaf as the server-side template engine to render pages, combines the Bootstrap framework to achieve responsive layout, and uses JavaScript libraries such as JQuery to handle dynamic interactions and asynchronous requests on the browser side. It provides a graphical operation interface (GUI) that allows authorized users to perform various operations, such as logging into the system, entering / managing pregnant women's information and data, triggering model training or prediction tasks, intuitively displaying the prediction results output by the risk prediction unit (such as charts, risk levels, text suggestions), viewing medical records, and performing other information management operations.
[0135] The method and system provided in this embodiment integrate multi-dimensional early circadian rhythm data, the advanced Biomedical Large Language Model (BioMedLM) and Parameter Efficient Fine-Tuning (PEFT) technology for data enhancement, and employ a stacking machine learning model for prediction. Ultimately, combined with a comprehensive system architecture, they achieve highly accurate, personalized, and clinically relevant predictions of gestational diabetes mellitus (GDM) risk in early pregnancy. This system also provides a complete technical implementation solution with significant beneficial effects:
[0136] 1. Improved early prediction and accuracy: By collecting and analyzing multi-dimensional data such as sleep, hormones, and circadian rhythms during early pregnancy, we captured subtle but critical physiological changes that precede the onset of GDM. Data augmentation using BioMedLM, fine-tuned with PEFT, effectively overcomes the common challenges of insufficient sample size, missing values, and class imbalance in clinical data, generating high-quality, physiologically consistent augmented data and laying a solid foundation for subsequent model training. Combined with a powerful stacking ensemble learning model, which integrates the strengths of multiple basic classifiers, this further enhances the accuracy and robustness of the prediction model, enabling more reliable risk assessments at an early stage, when clinical manifestations of the disease are not yet apparent.
[0137] 2. Enhanced clinical utility of prediction results: Prediction results not only provide a quantitative risk score but also generate personalized health management recommendations. This transforms predictions into actionable clinical guidance, moving beyond simply informing about risk. This helps physicians and pregnant women implement targeted preventive measures (such as lifestyle adjustments, nutritional interventions, and enhanced monitoring) based on their specific risk levels, enabling early intervention and improving maternal and infant outcomes.
[0138] 3. Improved feasibility and efficiency of technical solutions: Parameter-efficient fine-tuning (PEFT) technology is used to significantly reduce the computing resources and time required to fine-tune large biomedical language models, making the application of advanced AI technologies more economical and feasible. At the system level, a stable, efficient, and scalable backend support system is built based on the mature Spring Boot backend framework, MySQL database, and Redis cache. At the same time, it provides a complete functional unit including user management, information management, medical record management, and a graphical interactive interface, encapsulating complex prediction processes into an easy-to-use and conveniently managed application system, greatly improving deployment efficiency and user experience in actual clinical or research environments.
[0139] Unless otherwise specifically stated, the relative arrangement, numerical expression and numerical value of the parts and steps set forth in these embodiments do not limit the scope of the application. In all examples shown and discussed here, any specific value should be interpreted as merely exemplary, rather than as a restriction. Therefore, other examples of exemplary embodiments can have different values. It should be noted that: similar reference numerals and letters represent similar items in the accompanying drawings below, and therefore, once a certain item is defined in an accompanying drawing, it does not need to be further discussed in the accompanying drawings subsequently.
[0140] In addition, it should be noted that the use of terms such as "first" and "second" for limitation is only for the convenience of distinction. Unless otherwise stated, the above terms have no special meaning and therefore cannot be understood as limiting the scope of protection of this application.
[0141] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for predicting the risk of gestational diabetes in early pregnancy, characterized in that: The following steps are involved: Acquiring a multi-dimensional physiological indicator data sample of a source pregnant woman in early pregnancy, wherein the multi-dimensional physiological indicator data sample includes at least one of sleep rhythm data, hormone cycle rhythm data, and circadian rhythm data; Using a pre-trained large language model in the biomedical field, fine-tuning the large language model in the biomedical field through efficient parameter fine-tuning technology to generate a fine-tuned large language model; Using the fine-tuned large language model, performing data augmentation processing on the multi-dimensional physiological indicator data samples to generate enhanced physiological data, wherein the data augmentation processing includes expanding the sample size of the multi-dimensional physiological indicator data samples, filling missing values in the multi-dimensional physiological indicator data samples, and / or balancing the distribution of the multi-dimensional physiological indicator data samples; training a machine learning prediction model using a training dataset comprising the enhanced physiological data; Receiving physiological indicator data of the target pregnant woman; Inputting the physiological indicator data of the target pregnant woman into the machine learning prediction model; A prediction result indicating the risk of the target pregnant woman developing gestational diabetes in early pregnancy is output.
2. The method for predicting the risk of gestational diabetes in early pregnancy according to claim 1, characterized in that: The sleep rhythm data includes sleep duration, sleep onset time, sleep efficiency and / or sleep stage data; The hormone cycle rhythm data includes change trend data of basal body temperature, estrogen level, progesterone level and / or gonadotropin level; The circadian rhythm data includes heart rate variability, cortisol level rhythm, melatonin secretion rhythm and / or activity level change rhythm data; The multi-dimensional physiological indicator data sample also includes triglyceride glucose index data.
3. The method for predicting the risk of gestational diabetes in early pregnancy according to claim 1, characterized in that: The pre-trained large language model in the biomedical field is the BioMedLM model.
4. The method for predicting the risk of gestational diabetes in early pregnancy according to claim 1, wherein: Expanding the sample size of the multi-dimensional physiological indicator data samples includes generating synthetic sample data based on the data distribution of the scarce physiological rhythm pattern using the fine-tuned large language model, wherein the synthetic sample data statistically reflects the characteristics of the scarce physiological rhythm pattern, and the synthetic sample data is a new data instance different from existing samples in the multi-dimensional physiological indicator data samples; Filling missing values in the multi-dimensional physiological indicator data sample includes generating a complete simulated sample containing filling values for the physiological indicator data sample with missing attributes; Balancing the distribution of the multi-dimensional physiological indicator data samples includes generating corresponding enhanced physiological data that conforms to physiological relevance, in combination with the pregnant woman's age, pre-pregnancy BM or imaging data.
5. The method for predicting the risk of gestational diabetes mellitus in early pregnancy according to any one of claims 1 to 4, characterized in that: The machine learning prediction model is a stacked generalization model; The training of the machine learning prediction model includes: Using the training data set to train multiple base classifier models; Aggregating the prediction results of the basic classifier model on the training data set to form meta-features; A meta-model is trained using the meta-features, and the meta-model constitutes the stacked generalization model.
6. The method for predicting the risk of gestational diabetes in early pregnancy according to claim 5, characterized in that: The plurality of base classifier models include at least three selected from a K-nearest neighbor model, a logistic regression model, a random forest model, an extreme gradient boosting model, an adaptive boosting model, a light gradient boosting machine model, and a gradient boosting decision tree model; The meta-model is a logistic regression model.
7. The method for predicting the risk of gestational diabetes in early pregnancy according to claim 1, wherein: The prediction result includes a risk level score and / or personalized health management recommendations generated based on the risk level score.
8. A system for predicting the risk of gestational diabetes in early pregnancy, characterized in that: include: A data acquisition unit, used to obtain multi-dimensional physiological indicator data samples of the source pregnant woman in the early stages of pregnancy; A model fine-tuning unit, configured to fine-tune the pre-trained large language model in the biomedical field to generate a fine-tuned large language model; a data enhancement unit, configured to perform data enhancement processing on the multi-dimensional physiological indicator data samples using the fine-tuned large language model to generate enhanced physiological data; a model training unit, configured to train a machine learning prediction model using a training data set comprising the enhanced physiological data; A data receiving unit, configured to receive physiological indicator data of a target pregnant woman; The risk prediction unit is used to input the physiological indicator data of the target pregnant woman into the machine learning prediction model and output a prediction result indicating the risk of the target pregnant woman developing gestational diabetes in the early stages of pregnancy.
9. The system for predicting the risk of gestational diabetes mellitus in early pregnancy according to claim 8, wherein: The system further comprises: A backend processing server, which is built using a Spring Boot-based technical framework and is used to host and execute logical operations of the model fine-tuning unit, data enhancement unit, model training unit, and risk prediction unit; A data storage component, comprising a MySQL database and a Redis cache database. The MySQL database is a relational database used to persistently store the multi-dimensional physiological indicator data samples, the enhanced physiological data, the trained machine learning prediction model, and basic information of the user and pregnant woman. The Redis cache database is a cache mechanism used to cache hot data to improve system response speed. The backend processing server interacts with the relational database through the data persistence layer.
10. The system for predicting the risk of gestational diabetes mellitus in early pregnancy according to claim 8, wherein: The system further comprises: User management unit, used to manage user account registration, login, information modification and permission control; An information management unit interacts with the data acquisition unit to implement the addition, editing, deletion, and batch import and export of pregnant woman information and physiological indicator data samples; Medical record management unit, used for storing, structurally displaying and conditionally retrieving electronic medical record information related to the target pregnant woman; The interactive interface unit is built based on a preset front-end technology framework and is used to provide a graphical operation interface to receive user input instructions and display the prediction results and management information output by the risk prediction unit.
Citation Information
Cited By
Early pregnancy risk prediction method and system based on multi-source data and medium
CN121439240A