A PGT-A outcome prediction method for Gn dose recommendation

By constructing a gradient boosting regression and random forest classification model, the total number of eggs and the probability of OHSS were predicted based on the patient's physiological indicators, which solved the problem of linear assumption in Gn dose recommendation, achieved more accurate egg maturation prediction and OHSS risk control, and improved the success rate and safety of PGT-A.

CN119170094BActive Publication Date: 2025-09-19GUANGDONG UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411338689.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-09-19
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

In the existing technology, the Gn dose recommendation method assumes that the patient response is linear, fails to accurately predict the egg maturation, and does not take into account individual differences and the influence of multiple factors, which increases the treatment cost and the risk of OHSS.

Method used

A gradient boosting regression model and a random forest classification model were constructed, and preprocessing was performed based on the patient's physiological indicators to generate a prediction model to predict the total number of eggs, the number of euploid blastocysts, and the probability of OHSS, providing personalized treatment plans.

Benefits of technology

It improves the accuracy of egg maturity prediction, reduces the risk of OHSS, ensures the success rate of PGT-A and patient safety, and provides accurate treatment cost and success rate expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119170094B_ABST
    Figure CN119170094B_ABST
Patent Text Reader

Abstract

The present invention discloses a PGT-A outcome prediction method for Gn dose recommendation, belonging to the field of prediction technology, comprising: obtaining patient physiological indicators, preprocessing the patient physiological indicators, and generating a preprocessing data set; constructing two gradient boosting regression models and a random forest classification model; training the two gradient boosting regression models and the random forest classification model using the preprocessing data set to obtain a total egg number prediction model, a euploid blastocyst number prediction model, and a severe OHSS probability prediction model; generating prediction results based on the total egg number prediction model, the euploid blastocyst number prediction model, and the severe OHSS probability prediction model and feeding them back to the doctor. The present invention can provide early warning in the treatment stage, reduce the risk of OHSS, and ensure the safety of patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of prediction technology, and in particular relates to a PGT-A outcome prediction method for Gn dosage recommendation. Background Art

[0002] Preimplantation genetic testing for aneuploidy (PGT-A), as a selection aid, has been shown to significantly increase live birth rates per embryo transfer while minimizing the risks of clinical miscarriage, persistent aneuploid pregnancies, and implantation failure. The PGT-A process involves the following steps: first, ovarian stimulation is used to produce and retrieve multiple eggs, followed by in vitro fertilization (IVF). Subsequently, after the embryos reach a certain cell division stage, a small number of embryonic cells are collected for genomic analysis to determine the normality of their chromosome number and structure. This screening eliminates embryos carrying chromosomal abnormalities and retrieves transferable euploid blastocysts, thereby improving implantation success and pregnancy rates. For patients with high-risk factors, such as advanced female age, recurrent miscarriage, and recurrent implantation failure, PGT-A may be a more effective treatment option. The success rate of embryo transfer depends primarily on the quality and quantity of transferable embryos, which are directly influenced by individual patient characteristics and the ovulation stimulation regimen. Collecting a larger number of oocytes increases the chance of obtaining transferable euploid embryos, which typically requires the use of a higher initial dose of gonadotropin (Gn). On the other hand, the use of high-dose Gn may lead to more intense ovarian stimulation, increasing treatment costs and the risk of iatrogenic complications such as moderate to severe ovarian hyperstimulation syndrome (OHSS). Furthermore, patients undergoing PGT-A often have complex and diverse characteristics, which increases the uncertainty of treatment outcomes. Therefore, the development of an ovarian stimulation plan that focuses on individualized, cost-effective initial Gn doses is of great clinical and economic significance.

[0003] However, current methods all assume a linear response to Gn dose. However, in reality, a patient's physiological response can be more complex and potentially influenced by numerous factors, resulting in a nonlinear curve. Therefore, for some patients, a linear model may not accurately predict their actual egg maturation. Furthermore, ovarian stimulation therapy in practice can be influenced by numerous factors, such as treatment cycle, ovulation regimen, height and weight, which were not included in the data set in this study. Furthermore, the study did not investigate the interpretability of the model, further limiting the research. Furthermore, depending on the patient's constitution, the curve should be either linear or convex, with a dose that maximizes the number of eggs. This method does not explicitly model the possibility of avoiding the side effect of OHSS. Although physicians can empirically limit the number of eggs obtained to minimize this possibility, empirical evidence may not be accurate due to individual variability. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a PGT-A outcome prediction method for Gn dose recommendation to solve the problems existing in the above-mentioned prior art.

[0005] To achieve the above objectives, the present invention provides a PGT-A outcome prediction method for Gn dosage recommendation, comprising:

[0006] Acquiring physiological indicators of the patient, preprocessing the physiological indicators of the patient, and generating a preprocessed data set;

[0007] Build two gradient boosting regression models and one random forest classification model;

[0008] The two gradient boosting regression models and one random forest classification model are trained using the preprocessed data set to obtain a total egg number prediction model, a euploid blastocyst number prediction model, and a severe OHSS probability prediction model;

[0009] Based on the total egg number prediction model, euploid blastocyst number prediction model and severe OHSS probability prediction model, prediction results are generated and fed back to the doctor.

[0010] Preferably, the patient's physiological indicators include: female age, number of follicles, basal gonadotropin level, anti-flunoxotic acid level, ovulation induction regimen, height, weight, BMI, initial Gn dose and treatment results.

[0011] Preferably, the process of preprocessing the patient's physiological indicators includes: filling missing values ​​and abnormal values ​​of the patient's physiological indicators by median to generate the preprocessed data set.

[0012] Preferably, the process of obtaining a total egg number prediction model, a euploid blastocyst number prediction model, and a severe OHSS probability prediction model comprises:

[0013] Dividing the preprocessed data set into a training set and a test set based on a five-fold cross validation method;

[0014] The two gradient boosting regression models and one random forest classification model are trained based on the training set, and the models are verified by the test set to generate the total egg number prediction model, euploid blastocyst number prediction model and severe OHSS probability prediction model.

[0015] Preferably, after verifying the model through the test set, the method further includes: gradually removing the input features of the verified model, evaluating the impact of each feature on the model performance, selecting the input feature combination with the best performance, and retraining and evaluating the model.

[0016] Preferably, after verifying the model using the test set, the method further comprises: evaluating the trained gradient boosting regression model using absolute error and mean square error;

[0017] The accuracy rate was used to evaluate the random forest classification model.

[0018] Preferably, the random forest classification model uses a Gini importance calculation method to evaluate the contribution of features to model prediction.

[0019] Preferably, the process of generating prediction results based on the total egg number prediction model, the euploid blastocyst number prediction model and the severe OHSS probability prediction model and feeding back to the doctor includes:

[0020] Predicting the total number of eggs of the patient based on the total egg number prediction model, predicting the number of euploid blastocysts of the patient based on the euploid blastocyst number prediction model, and predicting the probability of OHSS based on the severe OHSS probability prediction model;

[0021] The total number of the patient's eggs, the number of euploid blastocysts, and the probability of OHSS occurrence are fed back to the doctor.

[0022] Compared with the prior art, the present invention has the following advantages and technical effects:

[0023] Traditional linear models assume that patients respond linearly to Gn doses, but in reality, patients' physiological responses may be nonlinear. This invention can capture this nonlinear relationship, improve the accuracy of predicting the number of retrieved eggs, and help develop more accurate ovulation induction plans.

[0024] To ensure the success rate of PGT-A, the number of euploid blastocysts available for transfer is crucial. This invention accurately predicts the number of euploid blastocysts based on the patient's physiological characteristics and Gn dose, helping doctors better select appropriate embryos for transfer. It also provides patients with a reasonable expectation of cost and success rate, ultimately improving the success rate of PGT-A.

[0025] By predicting the probability of OHSS occurrence, the present invention can provide early warning in the early stages of treatment, help doctors adjust the Gn dosage, reduce the risk of OHSS occurrence, and ensure patient safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0027] Figure 1 This is a flowchart of the initial Gn dose prediction model training and prediction according to an embodiment of the present invention;

[0028] Figure 2 A partial dependence diagram for predicting the total number of eggs according to an embodiment of the present invention;

[0029] Figure 3 This is a partial dependence diagram for predicting the number of euploid blastocysts according to an embodiment of the present invention. DETAILED DESCRIPTION

[0030] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0031] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0032] Example 1

[0033] like Figure 1 As shown, this embodiment provides a PGT-A outcome prediction method for Gn dose recommendation, comprising:

[0034] Acquiring physiological indicators of the patient, preprocessing the physiological indicators of the patient, and generating a preprocessed data set;

[0035] Build two gradient boosting regression models and one random forest classification model;

[0036] The two gradient boosting regression models and one random forest classification model are trained using the preprocessed data set to obtain a total egg number prediction model, a euploid blastocyst number prediction model, and a severe OHSS probability prediction model;

[0037] Based on the total egg number prediction model, euploid blastocyst number prediction model and severe OHSS probability prediction model, prediction results are generated and fed back to the doctor.

[0038] The specific process includes:

[0039] (1) Data collection and preprocessing: The doctor collected data including the patient's physiological indicators such as the woman's age, follicle count (AFC), basal gonadotropin level (bFSH), anti-fluoroquinolone level (AMH), ovulation induction regimen, height, weight, body mass index (BMI), initial Gn dose, and treatment outcomes (number of retrieved oocytes, number of euploid blastocysts, and occurrence of OHSS). The data was then cleaned and the median was used to fill in missing values ​​and outliers.

[0040] (2) Model training: The dataset was split into 80% training set and 20% test set. Five-fold cross-validation was used to select the best performing dataset split. The training set data, including the patient's physiological indicators and initial Gn dose, were input. A gradient boosting regression model was trained to predict the total number of eggs, and the model output was the total number of eggs. A gradient boosting regression model was trained to predict the number of euploid blastocysts, and the model output was the number of euploid blastocysts. A random forest classification model was trained to predict the probability of moderate to severe OHSS, and the model output was the probability of moderate to severe OHSS.

[0041] (3) Model evaluation: The performance of the three models was evaluated using the test set data. For the regression model, the evaluation indicators MAE and RMSE were calculated. For the classification model, the accuracy (ACC) was calculated.

[0042] (4) Ablation experiment: Conduct an ablation experiment to evaluate the impact of each feature on model performance by gradually removing input features. Select the input feature combination with the best performance and retrain and evaluate the model.

[0043] (5) Interpretability Assessment: First, we perform feature importance analysis to identify the features that have the greatest impact on the prediction results. Next, we use partial dependency plots to further analyze these key features and understand the specific impact of different values ​​on the prediction results.

[0044] (6) Prediction: For new patients, their physiological indicators are input and the initial Gn dose is traversed. The number of retrieved oocytes and euploid blastocysts is predicted using the trained gradient boosting regression model. The probability of OHSS is predicted using the trained random forest classification model. The prediction results are fed back to the physician to assist in developing a personalized treatment plan.

[0045] This example proposes three machine learning models. The first, a gradient regression boosting model, addresses the problem of predicting the total number of eggs at a given starting Gn dose. The second, a gradient regression boosting model, addresses the key issue: predicting the number of embryos that can be transferred based on the starting Gn dose. Furthermore, this example establishes a random forest classification model for predicting the probability of OHSS. Based on the outputs of these three models, physicians can make better decisions and provide patients with a comprehensive outlook.

[0046] First, the first gradient regression boosting model. This model optimizes model performance by iteratively constructing multiple decision trees, achieving high prediction accuracy in small sample size datasets. It excels at capturing complex nonlinear relationships in data and exhibits a certain degree of robustness to data noise. Furthermore, the model possesses feature selection capabilities, identifying the variables most important to the prediction results. Unlike other models, the first gradient regression boosting model proposed in this example predicts the slope of a linear dose-response function. Instead, it outputs the total number of oocytes. This method allows for a better fit of the drug response function, as this function is often non-monotonic. To achieve a given total number of oocytes, two different starting Gn dose solutions may exist. These solutions may exist on either side of the extreme point, either a low dose or a high dose. Furthermore, the high dose also carries the risk of side effects. By fitting and predicting the response curve, the optimal solution can be determined based on the actual situation. The first gradient regression boosting model simulates the PGT ovulation induction process. Next, the second gradient regression boosting model also uses a gradient boosting model, but unlike the first, it outputs the number of embryos available for transfer.

[0047] To ensure the success rate of PGT-A, a starting dose that produces one or more euploid blastocysts is often chosen. This results in a higher number of eggs, but also a higher probability of side effects. Therefore, a random forest classification model was developed. Unlike the previous two models, this one uses a random forest classification model. Similar to the gradient boosting model, the random forest classifier constructs multiple decision trees and combines their predictions to effectively capture nonlinear relationships in the data and provide interpretability. The random forest classification model takes as input physiological indicators and a starting Gn dose and predicts the probability of OHSS at that dose.

[0048] In this study, this embodiment utilized the model's Gini importance calculation method to assess the contribution of features to model predictions. Since all models are tree models, the feature screening process based on Gini importance is determined by analyzing the reduction in Gini impurity when each feature splits a node in the model. Based on the Gini importance evaluation results and experimental results, this embodiment selected and deleted some less important features, thereby optimizing the model's predictive power and generalization performance. Furthermore, by observing importance, the model has a certain degree of interpretability.

[0049] The dataset contains 2506 data from 2014 to 2024, including 1852 collected from the First Affiliated Hospital of Sun Yat-sen University and 654 collected from Nanning People's Hospital.

[0050] The dataset consists of multiple correlated factors, mediating variables, PGT indications, outcome variables, and PGT indication-related characteristics. Relevant factors include female age, follicle count (AFC), basal gonadotropin level (bFSH), anti-flunomide level (AMH), ovulation induction regimen, height, weight, BMI, basal luteinizing hormone level (bLH), estradiol level (E2), FSH / LH ratio, male age, number of spontaneous abortions, and previous live births. The mediating variable is the initial Gn dose. The outcome variables include the total number of eggs, the number of embryos available for transfer, and whether or not there is severe OHSS. Finally, PGT indication-related characteristics include whether or not there is severe oligoasthenoteratozoospermia, whether or not there is PCOS, and PGT indication. Among them, ovulation induction regimen, whether or not there is severe OHSS, whether or not there is severe oligoasthenoteratozoospermia, whether or not there is PCOS, and PGT indication, and the rest are continuous variables. In addition, BMI and FSH / LH ratio are calculated based on the baseline characteristics.

[0051] Furthermore, this example compares the proposed first gradient regression boosting model with a model for predicting the slope of the response function, both of which use the same feature inputs. Furthermore, given that the total number of oocytes and the number of euploid blastocysts are positively correlated when all other features are the same, this example uses the slope model with the number of transferable embryos as the fitting target and compares its performance with that of the second gradient regression boosting model. This example randomly selects five samples and, by adjusting key features and Gn dosage, obtains the function predicted by the model to examine the model's fitting and learning performance.

[0052] Table 1 shows the model performance on the test set. Based on the experimental results, Model 1 achieved a MAE of 4.19 and an RMSE of 5.67, outperforming the Pova model's 4.43 and 6.27, and the Slope model's 4.48 and 6.35. Model 2 achieved MAE and RMSE of 1.19 and 1.71, respectively. Finally, the classification model, Model 3, achieved an ACC of 0.99.

[0053] Table 1

[0054]

[0055]

[0056] To make the model more interpretable, Table 2 reports the importance of the model input features and the features selected for prediction in this example. Important features are highlighted. Not all features in the dataset were selected as input. This example conducted ablation experiments based on importance to select features that achieved the best results, with increasing or decreasing features impairing model performance.

[0057] Table 2

[0058]

[0059] In order to examine whether the model can learn the important features and the relationship between the initial Gn dose and the output, this example draws the partial dependence diagram of the first gradient regression boosting model, as shown in Figure 2 As shown, not all patients' response curves are linear. There may be an extreme point where the maximum number of eggs can be obtained when using this dose. Increasing or decreasing the dose will reduce the number. This also confirms that for the recommended starting Gn dose, a nonlinear model should be used. In addition, it can be observed that some patients are not sensitive to the drug, and increasing or decreasing the dose has little effect on the output.

[0060] In addition, if Figure 3 As shown, this embodiment also draws a partial dependence graph of the second gradient regression boosting model. It can be observed that the trend of the curve is similar to the partial dependence graph of the first gradient regression boosting model.

[0061] The above content proves that the method of this embodiment is superior to the existing method, can capture the nonlinear characteristics of the drug response curve, and can effectively predict the number of euploid blastocysts in PGT-A.

[0062] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A PGT-A outcome prediction method for Gn dosage recommendation, characterized in that: The following steps are involved: Acquiring physiological indicators of the patient, preprocessing the physiological indicators of the patient, and generating a preprocessed data set; Build two gradient boosting regression models and one random forest classification model; The two gradient boosting regression models and one random forest classification model are trained using the preprocessed data set to obtain a total egg number prediction model, a euploid blastocyst number prediction model, and a severe OHSS probability prediction model; Generate prediction results based on the total egg number prediction model, euploid blastocyst number prediction model and severe OHSS probability prediction model and feed them back to the doctor; The process of obtaining the total egg number prediction model, the euploid blastocyst number prediction model and the severe OHSS probability prediction model includes: Dividing the preprocessed data set into a training set and a test set based on a five-fold cross validation method; The two gradient boosting regression models and one random forest classification model are trained based on the training set, and the models are verified by the test set to generate the total egg number prediction model, the euploid blastocyst number prediction model and the severe OHSS probability prediction model; After verifying the model using the test set, the method further includes: gradually removing input features of the verified model, evaluating the impact of each feature on the model performance, selecting the input feature combination with the best performance, and retraining and evaluating the model; After verifying the model using the test set, the method further includes: evaluating the trained gradient boosting regression model using absolute error and mean square error; Using accuracy to evaluate the random forest classification model; The random forest classification model uses the Gini importance calculation method to evaluate the contribution of features to model predictions; The process of generating prediction results based on the total egg number prediction model, the euploid blastocyst number prediction model, and the severe OHSS probability prediction model and feeding back to the doctor includes: Predicting the total number of eggs of the patient based on the total egg number prediction model, predicting the number of euploid blastocysts of the patient based on the euploid blastocyst number prediction model, and predicting the probability of OHSS based on the severe OHSS probability prediction model; Feedback to the doctor the total number of the patient's eggs, the number of euploid blastocysts, and the probability of OHSS; The patient's physiological indicators include: female age, number of follicles, basal gonadotropin level, anti-fluminoic acid level, ovulation induction regimen, height, weight, BMI, initial Gn dose and treatment results; The process of preprocessing the patient's physiological indicators includes: filling the missing values ​​and abnormal values ​​of the patient's physiological indicators by using the median to generate the preprocessed data set.

Citation Information

Patent Citations

  • Method and device for training OHSS risk screening model

    CN116189886A