Method and system for constructing a model for predicting poor ovarian response and deploying individualized ovarian stimulation strategies
By constructing a machine learning-based model for POR diagnosis and ovarian stimulation strategy, the limitations of existing POR diagnostic standards and the lack of individualization in ovarian stimulation strategies have been addressed. This has enabled more accurate POR risk assessment and individualized treatment strategies, improving the effectiveness and transparency of assisted reproductive technologies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WOMEN S HOSPITAL ZHEJIANG UNIVERSITY SCHOOL OF MEDICINE
- Filing Date
- 2022-06-28
- Publication Date
- 2026-06-02
AI Technical Summary
Existing diagnostic criteria for ovarian regeneration (POR) have limitations for patients undergoing assisted reproductive treatment for the first time. Traditional models fail to fully consider nonlinear relationships and interactions, resulting in poor predictive performance and low transparency in clinical application. Existing ovarian stimulation strategy models lack individualization and universality, leading to biased treatment outcomes.
A POR diagnostic model and an ovarian stimulation strategy deployment model are constructed based on machine learning methods. Real-world data is used to screen key risk factors and stimulation strategy features. Multiple algorithms are applied to optimize the model and the SHAP method is used to interpret the feature contributions, providing individualized ovarian stimulation strategies.
It improves the accuracy of POR diagnosis and the individualized adaptability of ovarian stimulation strategies, enhances the interpretability and universality of the model, optimizes treatment outcomes and economic burden, and improves predictive performance.
Smart Images

Figure CN115331803B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical decision support systems, specifically relating to a method and system for constructing a model based on machine learning to predict the risk of low ovarian response after controlled ovarian stimulation in assisted reproductive technology and to deploy the optimal ovarian stimulation strategy in an individualized manner. Background Technology
[0002] According to WHO predictions, infertility will become the third leading cause of disease in this century, after cancer and cardiovascular diseases. The emergence and development of assisted reproductive technology (ART) has provided new avenues for the diagnosis and treatment of infertility, improving fertility rates, and ensuring the reproductive health of women and the health of their offspring. Poor ovarian response (POR) refers to a pathological state in which the ovaries do not respond well to gonadotropin stimulation, with an incidence rate of 6%-35%, and this rate is increasing annually. POR can lead to fewer follicles developing during ovarian stimulation cycles, low estrogen levels, low number of oocytes retrieved, few transferable embryos, high cycle cancellation rates, low pregnancy and live birth rates, and health problems in offspring. It also places enormous psychological and economic pressure on patients and has been one of the most challenging problems in the field of assisted reproduction for decades.
[0003] Currently, the Bologna criteria (2011) and the Poseidon criteria (2016) are widely used internationally and domestically for diagnosing ovarian reserve (POR). The main criteria for both are: age, AMH (Anti-Mullerian Hormone), basal follicle-stimulating hormone (FSH), basal antral follicle count (AFC), and whether POR has occurred after previous ovarian stimulation. Age, AMH, FSH, and AFC are commonly used clinical indicators for evaluating ovarian reserve. However, the requirement of whether POR has occurred after previous ovarian stimulation requires the patient to have previously undergone ovarian stimulation treatment. Both criteria have certain limitations for patients undergoing assisted reproductive technology for the first time. Currently, most diagnostic and predictive methods for ovarian reserve (POR) rely on classic logistic regression, and most fail to validate or address nonlinear relationships. Some even simply classify continuous features, which has limitations: firstly, it fails to consider clinical mechanisms, such as categorizing age into young (<30 years), middle (30-40 years), and old (>40 years). In reality, ovarian reserve declines sharply in women over 35, and there is a clear nonlinear relationship between age and ovarian function, making such grouping clinically inaccurate; secondly, classification disrupts the original continuity to some extent, resulting in information loss, which is not only unrealistic but may also negatively impact model predictive performance. Furthermore, most models only utilize a few "star" indicators for assessing ovarian reserve (such as age, AFC, AMH, FSH, etc.), and rarely explore and screen numerous patient baseline indicators (such as allergy history, obstetric history, family history, complications, basal hormone levels, blood routine, metabolic indicators, cervical secretions, various clinical diagnoses, etc.) based on real-world electronic medical record data. This ignores other factors that may have a significant impact on POR, and does not consider the interaction of other factors with the "star" indicators. This may lead to amplifying or diminishing the impact of these "star" indicators on POR, resulting in a certain bias in clinicians' understanding of their true role.
[0004] In addition, although it is generally believed that diminished ovarian reserve is the pathological basis of POR, and clinicians generally use ovarian reserve assessment to predict ovarian responsiveness, there are still many cases in clinical practice where low ovarian reserve presents with normal (or high) ovarian responsiveness, or where normal (or high) ovarian reserve eventually leads to POR. That is, ovarian reserve does not completely match ovarian responsiveness. Therefore, compared with indirectly predicting POR through ovarian reserve, direct prediction of POR may have higher value.
[0005] Controlled ovarian stimulation (COS) is a crucial step in assisted reproductive technology (ART) for infertility. The main intervention strategies that clinicians need to pre-determine include: the ovarian stimulation protocol, the initial dose of exogenous FSH, the FSH formulation, and whether to use exogenous luteinizing hormone (LH). There are dozens of commonly used ovarian stimulation protocols, such as minimal stimulation, antagonist, agonist, and high-progesterone ovulation induction protocols, all of which are widely used clinically. Even so, which protocol is most suitable for patients with progesterone-related ovarian dysfunction (POR) remains inconclusive. Besides treatment efficacy, these intervention strategies are closely related to the medical costs incurred by patients. Current research on developing individualized ovarian stimulation strategies, mostly using logistic regression or combined with Norman plots, is limited to the deployment of a single ovarian stimulation strategy (such as the initial FSH dose) and must be based on a specific ovarian stimulation protocol. This limits its clinical applicability and does not consider the impact of other interventions on POR. Summary of the Invention
[0006] Current diagnostic criteria for ovarian reserve (POR) have limitations for patients undergoing controlled ovarian stimulation (CORS) during their first assisted reproductive technology (ART) process. They rely solely on a few key indicators for assessing ovarian reserve (such as age, baseline AFC, baseline FSH, and AMH). Furthermore, in actual clinical practice, assessments of ovarian reserve do not perfectly correlate with ovarian responsiveness. Most existing POR diagnostic prediction models are built using traditional statistical models, which have limited information utilization as the data volume increases. They may also fail to address nonlinear relationships and interactions between variables, or, when using machine learning methods, fail to consider real-world data and adequately account for other risk factors that could affect POR. These shortcomings lead to poor predictive performance, lack of interpretability, and insufficient transparency or reliability in clinical application.
[0007] Currently, clinicians often formulate ovarian stimulation strategies based on their own experience, considering factors such as age, weight, baseline AFC, and sex hormone levels, which introduces a significant degree of subjectivity. Existing POR (Positive Ovary Orientation) ovarian stimulation strategy model studies are often limited to a single intervention (such as the initial FSH dose) or confined to specific scenarios (such as a particular protocol or population), resulting in poor clinical applicability. Furthermore, failing to consider the influence of other factors often leads to either amplifying or diminishing the therapeutic effect of the intervention, thus causing actual deviations in clinical decisions. Existing POR ovarian stimulation strategy model studies also have limitations in terms of methodology and transparency of application.
[0008] In light of the above, this invention, starting from real-world data, constructs two interpretable models based on machine learning methods: a POR diagnostic model and a POR ovarian stimulation strategy deployment model. The former predicts the risk of POR following ovarian stimulation during assisted reproductive technology (ART) for infertility treatment, while the latter uses the predicted risk of POR to individualize the deployment of ovarian stimulation strategies. Both models are interpretable at both the overall and individual levels. This patent provides technical support for early, rapid, and accurate POR screening in ART patients, exploring the pathogenic mechanisms of POR risk factors, assisting physicians in deploying treatment and cost-optimized clinical intervention strategies, and monitoring the effects of various interventions on POR treatment.
[0009] Specifically, the present invention relates to the following technical solutions:
[0010] A method for constructing a model to predict low ovarian response and deploy individualized ovarian stimulation strategies includes a POR diagnostic model and a POR ovarian stimulation strategy deployment model, the specific steps of which are as follows:
[0011] (1) Based on the raw data of patients extracted from the electronic medical record system, preliminary screening of candidate POR risk factors and ovarian stimulation strategy characteristics was conducted.
[0012] (2) Divide the data obtained in step (1) into training set and test set according to the proportion;
[0013] (3) Based on the candidate POR risk factors and ovarian stimulation strategy features in step (1), various different machine learning algorithms are applied to construct candidate POR ovarian stimulation strategy deployment models or candidate POR diagnostic models on the training set described in step (2). The input of the POR diagnostic model is the POR risk factor, and the output is the probability of POR disease risk. The input of the POR ovarian stimulation strategy deployment model is the POR risk factor and the ovarian stimulation strategy features, and the output is the probability of POR disease risk under different combinations of ovarian stimulation strategy features. The candidate models are evaluated on the test set and the best machine learning algorithm is selected.
[0014] (4) The best candidate model selected in step (3) is interpreted by the SHAP (Shapley Additive exPlanations) method. Based on the contribution of the features, 13 key risk factors of POR are obtained.
[0015] (5) Based on the optimal machine learning algorithm selected in step (3), the final POR ovarian stimulation strategy deployment model is constructed using ovarian stimulation strategy features and key POR risk factors selected in step (4); the final POR diagnostic model is constructed using the key POR risk factors selected in step (4). The input to the POR diagnostic model is the key POR risk factors, and the output is whether the patient has the disease and the probability of developing POR. The input to the POR ovarian stimulation strategy deployment model is the POR risk factors and ovarian stimulation strategy features, and the output is the probability of developing POR under the combination of ovarian stimulation strategy features without the disease.
[0016] Further, the preliminary screening of POR candidate risk factors in step (1) includes the following 50 factors: age, education level, height, weight, baseline blood pressure, allergy history, duration of infertility, age at menarche, menstrual cycle, number of days of menstruation, history of dysmenorrhea, premature abortion, adverse fertility history, patient's father's medical condition, patient's mother's medical condition, female partner's (patient's) diagnosis, previous secondary infertility, male partner's (patient's husband's) diagnosis, baseline antral follicle count, baseline endocrine levels (FSH; luteinizing hormone, LH; estradiol, E2; progesterone, P; prolactin, PRL; testosterone, T), AMH, and complete blood count (red blood cells). Cells, white blood cells, hemoglobin, platelets, hematocrit / hematocrit), biochemical indicators (total protein; albumin; alanine aminotransferase, ALT; aspartate aminotransferase, AST; fasting blood glucose, homocysteine, creatinine, blood urea nitrogen, CA125), thyroid hormones (thyroid-stimulating hormone, TSH; thyroid immunoglobulin antibody, A-TG; anti-thyroid peroxidase antibody, A-TPO), coagulation function (activated partial thromboplastin time, APTT; prothrombin time, PT), blood type, chromosome examination, gonococcal DNA, chlamydia DNA, mycoplasma DNA.
[0017] Furthermore, the ovarian stimulation strategy described in step (1) is characterized by: ovarian stimulation protocol, initial FSH dose, FSH formulation (drug name) used during ovulation induction, and whether at least two of LH are used during ovulation induction.
[0018] Further, step (2) specifically involves removing features with more than 15% missing samples from the data described in step (1), then stratifying according to the POR event, then randomly selecting 70% of the samples as the pre-imputation training set and the remaining 30% of the samples as the pre-imputation test set, and then performing multiple imputation separately on the pre-imputation training set and the pre-imputation test set.
[0019] Furthermore, the missing value multiple imputation method is: chain equation multiple imputation method based on random forest, wherein the implementation method is random forest, the prediction mean matching method is used for value selection after prediction, the number of candidate matching values is 5, and the number of iterations is 100.
[0020] Furthermore, the methods for verifying the results after multiple interpolation include: comparing the distribution of the new data after interpolation with the distribution of the original data, and performing iterative convergence diagnosis during the interpolation process.
[0021] Furthermore, the four candidate machine learning algorithms mentioned in step (3) include at least two of the following: LASSO-Logistic regression, support vector machine with RBF as the kernel function, multilayer perceptron, and XGBoost.
[0022] Furthermore, when constructing candidate models based on various machine learning algorithms, the hyperparameter optimization method used is: Bayesian optimization algorithm combined with 5-fold cross-validation, where the simplified approach of the Bayesian optimization algorithm is as follows:
[0023] Suppose a set of hyperparameter combinations is: X = x1, x2, x3, ..., x n ,but,
[0024] Input: D←InitSamples(f,X)
[0025] for i←|D|to T do
[0026] p(y|x,D)←FitModel(M,D)
[0027] x i ←arg max x∈X S(x, p(y|x, D))
[0028] y i ←f(x i )
[0029] D←D∪(x i y i )
[0030] end
[0031] f is the model to be tuned; X is the hyperparameter search space; D is a dataset consisting of several pairs of data, each pair of arrays is represented as (x, y), where x is a set of hyperparameters and y represents the output result of the model under that set of hyperparameters; S is the acquisition function, which is used to select the optimal x; M is the Gaussian model obtained by fitting the dataset D; T is the number of times the parameters are selected in the loop.
[0032] Further, the feature selection process after feature interpretation in step (4) is as follows: all features with non-zero SHAP values in the model constructed by the best algorithm in step (3) are arranged in descending order. Then, a dot-line graph is plotted based on the sequence number and the corresponding SHAP value. Then, based on the degree of SHAP value decrease, 13 key risk factors for POR and 4 inherent ovarian stimulation strategy features are selected from largest to smallest. Previous studies have mostly used the method of selecting features with a certain number of rankings, such as the top 10 or top 15, which is simple and effective, but has a high degree of subjectivity.
[0033] Furthermore, the 13 key risk factors associated with POR were: age, weight, diastolic blood pressure, duration of infertility, AMH level, basal antral follicle count, basal FSH level, basal P level, basal LH level, red blood cell count, white blood cell count, ALT, and a diagnosis containing POI (early-onset ovarian insufficiency) or DOR (decreased ovarian reserve).
[0034] Furthermore, in step (5), the final model needs to be built on the total data after the training set and the test set are combined, and the hyperparameters need to be re-optimized. The specific parameter tuning method is: Bayesian optimization algorithm combined with 5-fold cross-validation.
[0035] Furthermore, the steps for using the model constructed by the method of the present invention are as follows:
[0036] A. Obtain the raw data of the patients to be predicted, specifically the data of the first ovarian stimulation cycle of infertile patients undergoing IVF / ICSI / PGT, including 13 key risk factors for POR and 4 characteristics of ovarian stimulation strategies (the doctor develops multiple possible combinations of ovarian stimulation strategies based on clinical experience (different combinations of 4 ovarian stimulation interventions)).
[0037] B. Input the 13 key risk factors for POR obtained in step A into the POR diagnostic model to obtain the prediction result (whether it is POR) and the risk value of having POR; input the 13 key risk factors for POR and the 4 ovarian stimulation strategy feature data obtained in step A into the POR ovarian stimulation strategy deployment model to obtain the risk value (probability) of having POR under the ovarian stimulation strategy feature combination that predicts that it is not POR. Based on the patient's economic burden, the ovarian stimulation strategy that optimizes the POR risk and the patient's economic burden is then selected.
[0038] Furthermore, the individual-level interpretability diagnostic method for new patients using the POR diagnostic model is as follows: obtain information on 13 key risk factors for POR in new patients and input them into the model to obtain prediction results and risk values for POR, calculate the Shapley values corresponding to the 13 key risk factors for POR and use them to explain the contribution of different risk factors to pathogenic POR.
[0039] Corresponding to the above method, the present invention also provides a machine learning-based system for predicting low ovarian response in assisted reproductive technology and constructing a model for deploying individualized ovarian stimulation strategies. This system includes:
[0040] The screening algorithm module (2) constructs candidate POR ovarian stimulation strategy deployment models or candidate POR diagnostic models on the training set based on various machine learning algorithms. The input of the candidate POR diagnostic model is POR risk factors, and the output is the probability of POR disease risk. The input of the candidate POR ovarian stimulation strategy deployment model is POR risk factors and four ovarian stimulation strategy features, and the output is the probability of POR disease risk under different combinations of ovarian stimulation strategy features. The candidate models are evaluated on the test set based on ROC curves, and the best machine learning algorithm is selected.
[0041] The feature selection and interpretation module (3) is based on the model constructed by the best algorithm described in module (2) and selects key risk factors of POR according to the size of feature contribution.
[0042] The model building module (4) reconstructs the POR ovarian stimulation strategy deployment model based on the best algorithm selected by module (2) and the key risk factors selected by module (3) on the total dataset after merging the training set and the test set; and reconstructs the POR diagnostic model using the key risk factors and ovarian stimulation strategy features selected by module (3).
[0043] Furthermore, the system also includes:
[0044] The model validation and interpretation individual prediction module (5) obtains data from other hospitals or other reproductive centers and performs external data validation on the POR ovarian stimulation strategy deployment model and POR diagnostic model constructed in module (4); obtains new patient data and performs interpretability prediction on the POR ovarian stimulation strategy deployment model and POR diagnostic model constructed in module (4) at the individual level, explaining the contribution of various risk factors and different ovarian stimulation interventions to the POR prediction results.
[0045] The ovarian stimulation strategy deployment module (6) obtains 13 key risk factor values for new patients. Doctors develop multiple possible combinations of ovarian stimulation strategies based on clinical experience (different combinations of 4 ovarian stimulation interventions). After inputting into the POR ovarian stimulation strategy deployment model, the risk prediction of POR corresponding to different combinations is obtained. At the same time, considering the patient's economic burden, the ovarian stimulation strategy with the optimal POR risk and patient economic burden is then screened.
[0046] The present invention has the following beneficial effects:
[0047] (1) Using data from an electronic medical record system that is closer to the real world, compared to the previously determined data, other features that may affect POR can be considered to the greatest extent. At the same time, this invention is applicable to all assisted reproductive patients who receive ovarian stimulation during IVF / ICSI / PGT, and compared to the previous specific population or specific ovarian stimulation protocols, it has a wider range of applications and higher universality.
[0048] (2) This invention is based on machine learning methods. Compared with traditional statistical methods, it can make full use of data information when the amount of data increases, can handle complex feature relationships, and has better prediction results.
[0049] (3) This invention provides both a method for diagnosing POR and a method for deploying a POR ovarian stimulation strategy. The former can assess and diagnose the risk of POR in women who are assisted in conceiving early, quickly and accurately. It can also assess the ovarian reserve function of women who have fertility plans but are unsure when to give birth and make reasonable fertility arrangements. The latter can assist clinicians in developing an individualized ovarian stimulation strategy with the best treatment effect based on the specific POR risk, while also taking into account the economic pressure of patients for balance and optimization.
[0050] (4) The currently widely used diagnostic criteria for POR are the Bologna criteria and the Poseidon criteria. Both of them have specific requirements on the responsiveness of the ovaries in the patient's previous ovarian stimulation cycles. Therefore, when diagnosing POR in patients who have not received any ovarian stimulation treatment, both criteria have limitations. The POR diagnostic method provided by this invention mainly predicts the patient's first ovarian stimulation cycle, thus making up for the inherent defects of the two criteria.
[0051] (5) The POR diagnosis method provided by the present invention has characteristic interpretability and can reveal the specific impact and relative magnitude of key risk factors related to POR on POR risk at the overall level.
[0052] (6) When assessing POR in women of childbearing age, the POR diagnostic method provided by this invention can explain the specific pathogenic effects of each key risk factor at the individual level.
[0053] (7) Compared with the previous methods, it has better predictive effect, with the AUC value of POR diagnosis method being 0.920 and the AUC value of POR ovarian stimulation strategy deployment method being 0.929. Attached Figure Description
[0054] Figure 1 A framework diagram of the POR diagnosis and POR ovarian stimulation strategy deployment method provided by the present invention;
[0055] Figure 2ROC curves (with AUC values and Brier scores) of candidate POR ovarian stimulation strategy deployment models constructed for different machine learning algorithms;
[0056] Figure 3 This is a diagram illustrating the feature selection process based on the SHAP method.
[0057] Figure 4 ROC curves (with AUC values and Brier score) for external validation of the POR diagnostic model and POR ovarian stimulation strategy deployment model;
[0058] Figure 5 This is a diagram illustrating the different impacts of 13 key risk factors for POR at the overall level on POR in the POR diagnostic methodology (SHAP overall screenshot).
[0059] Figure 6 This diagram illustrates the interpretability of diagnostic methods for Prostate Occurrence (POR) at the individual level. (a) shows the prediction for a woman of reproductive age at high risk of POR; (b) shows the prediction for a woman of reproductive age at low risk of POR. The diagram demonstrates that the impact of key risk factors for POR on prediction varies at different individual levels and is clinically significant. Detailed Implementation
[0060] The specific embodiments of the present invention will now be described in more detail with reference to the accompanying drawings and tables. While the embodiments illustrate specific implementation processes of the present invention, the scope of protection of the present invention is not limited by these embodiments. On the contrary, the specific implementation of the present invention is not limited to this one; therefore, without inventive or innovative effort in the art, all other embodiments within the scope of the claims of the present invention are within the protection scope of the present invention.
[0061] The present invention relates to ovarian hyporesponsiveness (POR), a pathological condition characterized by poor ovarian response to gonadotropin stimulation. This can manifest as a low number of follicles developed during ovarian stimulation cycles, low peak estrogen levels, high gonadotropin dosage, high cycle cancellation rate, low oocyte retrieval count, and low clinical pregnancy and live birth rates. This invention references relevant content from the widely used Bologna and Poseidon criteria, using a retrieval count of less than four oocytes after ovarian stimulation as the diagnostic criterion. Specifically, a retrieval count of less than four oocytes indicates ovarian hyporesponsiveness, and the dataset is labeled accordingly.
[0062] like Figure 1 As shown, the specific steps of the method for constructing a machine learning-based model for predicting low ovarian response in assisted reproductive technology and deploying individualized ovarian stimulation strategies provided by this invention are as follows:
[0063] (1) Based on data that is closer to the real world, namely, data of patients undergoing their first ovarian stimulation cycle before IVF / ICSI / PGT are extracted from the electronic medical record system of hospitals or reproductive centers, and 50 candidate POR risk factors and 4 ovarian stimulation strategy intervention features are preliminarily screened.
[0064] The specific subjects selected were 12,012 women undergoing assisted reproduction and their first ovarian stimulation cycle, which served as the total data for training the model. Among them, 1,714 were patients with progesterone-induced ovarian reserve (median age / interquartile range: 34 years / 30-39 years) and 10,298 were non-progesterone-induced ovarian reserve (median age / interquartile range: 30 years / 28-33 years). The obtained data were imported into R software, and the main clinical characteristics are shown in Table 1 of the total training data.
[0065] Table 1. Main clinical characteristics based on POR classification (total training data from 12012 cases and external validation data from 5702 cases).
[0066]
[0067] The table displays the data as median (interquartile range) or quantity (percentage).
[0068] The preliminary screening of POR candidate risk factors included the following 50 factors: age, education level, height, weight, baseline blood pressure, allergy history, duration of infertility, age at menarche, menstrual cycle, number of days of menstruation, history of dysmenorrhea, premature abortion, adverse fertility history, patient's father's medical condition, patient's mother's medical condition, female partner's (patient's) diagnosis, previous secondary infertility, male partner's (patient's husband's) diagnosis, baseline antral follicle count, baseline endocrine levels (FSH, LH, E2, P, PRL, T), AMH, complete blood count (red blood cells, white blood cells, hemoglobin, platelets, hematocrit / hematocrit), biochemical indicators (total protein, albumin, ALT, AST, fasting blood glucose, homocysteine, creatinine, blood urea nitrogen, CA125), thyroid hormones (TSH, A-TG, A-TPO), coagulation function (APTT, PT), blood type, chromosome examination, gonococcal DNA, chlamydia DNA, and mycoplasma DNA.
[0069] The four ovarian stimulation strategies are characterized by: ovarian stimulation protocol, initial FSH dose, FSH formulation (drug name) used during ovulation induction, and whether LH is used during ovulation induction. Ovarian stimulation protocols include: long agonist protocol, ultra-long agonist protocol, short agonist protocol, antagonist protocol, ovulation induction protocol under high progesterone conditions, minimal stimulation or natural cycle protocol, and other protocols. Initial FSH doses include: <100 IU, 150 IU, 200 IU, 225 IU, ≥300 IU, corresponding to non-agonist or non-antagonist protocols. FSH formulations used include: recombinant, urinary, corresponding to non-agonist or non-antagonist protocols. LH use is: no, yes, corresponding to non-agonist or non-antagonist protocols.
[0070] (2) Data preparation: After removing features with more than 15% missing samples from the data described in step (1), in order to ensure that the POR ratio in the training set and the test set is equal to that in the original data, the data described in step (1) is divided into a training set and a test set before imputation in a ratio of 7:3, and multiple imputation of missing values is performed separately on the two datasets before imputation to obtain the training set and the test set.
[0071] The missing value multiple imputation method is a chain equation multiple imputation method based on random forest, where the implementation method is random forest, the prediction value selection method is prediction mean matching, the number of candidate matching values is 5, and the number of iterations is 100. After multiple imputation, the imputation results need to be verified, including comparing the distribution of the new data after imputation with the distribution of the original data, and diagnosing the convergence of the iteration during the imputation process.
[0072] (3) Based on candidate POR risk factors and four ovarian stimulation strategy features, four different machine learning algorithms were applied to construct candidate POR ovarian stimulation strategy deployment models on the training set described in step (2). The candidate models were evaluated on the test set, and the best machine learning algorithm was selected. The input of the candidate POR ovarian stimulation strategy deployment model is the POR risk factors and four ovarian stimulation strategy features, and the output is the probability of POR disease risk under different combinations of ovarian stimulation strategy features.
[0073] The four machine learning algorithms used are: LASSO-Logistic Regression, Support Vector Machine with RBF kernel function, Multilayer Perceptron, and XGBoost. The hyperparameter optimization method used in constructing the candidate model is Bayesian optimization combined with 5-fold cross-validation. The simplified approach of the Bayesian optimization algorithm is as follows:
[0074] Suppose a set of hyperparameter combinations is: X = x1, x2, x3, ..., x n ,but,
[0075] Input: D←InitSamples(f,X)
[0076] for i←|D|_to T do
[0077] p(y|x,D)←FitModel(M,D)
[0078] x i ←arg max x∈X S(x, p(y|x, D))
[0079] y i ←f(x i )
[0080] D←D∪(x i y i )
[0081] end
[0082] f is the model to be tuned; X is the hyperparameter search space; D is a dataset consisting of several pairs of data, each pair of arrays is represented as (x, y), where x is a set of hyperparameters and y represents the output result of the model under that set of hyperparameters; S is the acquisition function, which is used to select the optimal x; M is the Gaussian model obtained by fitting the dataset D; T is the number of times the parameters are selected in the loop.
[0083] Figure 2 The ROC curves, AUC values, and Brier scores of candidate POR ovarian stimulation strategy deployment models constructed using different machine learning algorithms are presented. Among them, the XGBoost algorithm performed best, with an AUC value of 0.930 and a Brier score of 0.064.
[0084] (4) The best candidate model selected in step (3) is interpreted using the SHAP method. Based on the contribution of the features, 13 key risk factors for POR are obtained. That is, as follows: Figure 3 As shown: In the model constructed by the optimal algorithm in step (3), all features with non-zero SHAP values are arranged in descending order. Then, a dot-line graph is plotted based on the sequence number and the corresponding SHAP value. Then, based on the degree of SHAP value decrease, 13 key risk factors for POR and 4 inherent ovarian stimulation strategy features are selected from largest to smallest. The 13 key risk factors for POR are: age, weight, diastolic blood pressure, duration of infertility, AMH level, basal antral follicle count, basal FSH level, basal P level, basal LH level, red blood cell count, white blood cell count, ALT, and whether the diagnosis includes POI or DOR.
[0085] (5) To fully utilize the data, the training set and test set are merged in this invention. Based on the best machine learning algorithm selected in step (3), the hyperparameters are re-optimized. Then, using four ovarian stimulation strategy features and 13 key POR risk factors selected in step (4), the final POR ovarian stimulation strategy deployment model is constructed; using the 13 key POR risk factors selected in step (4), the final POR diagnostic model is constructed. The input of the POR diagnostic model is the key POR risk factors, and the output is whether the disease is present and the probability of POR disease. The input of the POR ovarian stimulation strategy deployment model is the POR risk factors and ovarian stimulation strategy features, and the output is the probability of POR disease under the combination of ovarian stimulation strategy features without disease.
[0086] (6) Perform external data validation on the two models described in step (5). The specific steps are as follows:
[0087] A. External validation data were obtained, specifically from 5702 women undergoing assisted reproduction and their first ovarian stimulation cycle at other hospitals or reproductive centers. Among them, there were 882 patients with progesterone (POR) (median age / interquartile range: 33 years / 30-37 years) and 4820 non-POR patients (median age / interquartile range: 30 years / 28-33 years). The obtained data were imported into R software, and the main clinical characteristics were selected as shown in Table 1 of the external validation data.
[0088] B. Input the external validation data described in step A into the final POR diagnostic model to obtain the prediction result (whether it is POR) and the risk value of having POR; input the 13 key risk factors for POR and the 4 ovarian stimulation strategy feature data obtained in step A into the POR ovarian stimulation strategy deployment model to obtain the risk value (probability) of having POR under the ovarian stimulation strategy feature combination where the prediction result is not POR.
[0089] C. Based on the original external validation data from step A and the prediction results and POR risk value obtained from step B, the model's discriminative ability is evaluated using the ROC curve and AUC (area under the ROC curve), and the model's calibration ability is evaluated using the Brier score (Boolean value). The Brier score (BS) is calculated as follows:
[0090]
[0091] f t Let be the predicted probability for sample t, and o t Let t be the actual POR label for sample t (POR with the disease is 1, POR without the disease is 0), and N be the total number of samples.
[0092] like Figure 4As shown, the POR diagnostic model and the POR ovarian stimulation strategy deployment model still have good predictive performance on external data.
[0093] (7) At the overall level, the SHAP method was applied to interpret the characteristics of the POR diagnostic model, which can demonstrate the different impacts of 13 key POR risk factors on POR, such as Figure 5 As shown, AMH, baseline AFC, diagnosis of POI or DOR, baseline FSH, and age have the top five effects on POR. The lower the AMH and baseline AFC, the higher the risk of POI or DOR diagnosis, and the higher the FSH and age, the higher the risk of POR. In addition, factors such as diastolic blood pressure, ALT, white blood cell count, and red blood cell count are not currently considered to be related to POR in clinical practice or research, and deserve further investigation and related studies.
[0094] At the individual level, such as Figure 6 As shown, the POR diagnostic model described in step (5) was used to predict the interpretability of POR risk for two women of childbearing age (sub-figure a represents women at high risk of POR, and sub-figure b represents women at low risk of POR). The specific method was as follows: information on 13 key risk factors for POR in the new patient was obtained and input into the POR diagnostic model to obtain the prediction results and the risk value of POR. The Shapley values corresponding to the 13 key risk factors for POR were calculated and used to explain the contribution of different risk factors to pathogenic POR. In addition, the POR ovarian stimulation strategy deployment model described in step (5) was used to deploy an individualized optimal ovarian stimulation strategy. The specific method was as follows: values of 13 key risk factors in the new patient were obtained, and then doctors formulated multiple possible combinations of ovarian stimulation strategies (different combinations of 4 ovarian stimulation interventions) based on clinical experience. At the same time, 13+4 features were input into the POR ovarian stimulation strategy deployment model to obtain the risk prediction of POR corresponding to different combinations. Based on the actual economic burden of patients, the ovarian stimulation strategy with the optimal POR risk and patient economic burden was then selected.
Claims
1. A method for constructing a model to predict low ovarian response and deploy individualized ovarian stimulation strategies, characterized in that, This includes a POR diagnostic model and a POR ovarian stimulation strategy deployment model, with the following specific steps: (1) Based on the raw data of patients extracted from the electronic medical record system, preliminary screening of candidate POR risk factors and ovarian stimulation strategy characteristics was conducted. (2) Divide the data obtained in step (1) into training set and test set according to the proportion; (3) Based on the candidate POR risk factors and ovarian stimulation strategy features in step (1), different machine learning algorithms are applied to construct candidate POR ovarian stimulation strategy deployment models or candidate POR diagnostic models on the training set described in step (2), respectively. The input of the candidate POR diagnostic model is the POR risk factor, and the output is the probability of POR disease risk. The input of the candidate POR ovarian stimulation strategy deployment model is the POR risk factor and the ovarian stimulation strategy features, and the output is the probability of POR disease risk under different combinations of ovarian stimulation strategy features. The candidate models are evaluated on the test set and the best machine learning algorithm is selected. (4) The best candidate model selected in step (3) is interpreted by the SHAP method, and the key risk factors of POR are obtained according to the size of the feature contribution. (5) Based on the best machine learning algorithm selected in step (3), the final POR ovarian stimulation strategy deployment model is constructed using the ovarian stimulation strategy features and the key risk factors of POR selected in step (4); the final POR diagnostic model is constructed using the key risk factors of POR selected in step (4).
2. The method according to claim 1, characterized in that, The raw data extracted from the electronic medical record system in step (1) is: data of the first ovarian stimulation cycle of infertile patients undergoing IVF / ICSI / PGT.
3. The method according to claim 1, characterized in that, The preliminary screening of POR candidate risk factors in step (1) includes the following 50 factors: age, education level, height, weight, baseline blood pressure, allergy history, duration of infertility, age at menarche, menstrual cycle, number of days of menstruation, history of dysmenorrhea, premature abortion, adverse fertility history, paternal illness, maternal illness, female diagnosis, primary or secondary infertility, husband's diagnosis, baseline antral follicle count, baseline FSH, baseline LH, baseline E2, baseline P, baseline PRL, baseline T, AMH, red blood cell count, white blood cell count, hemoglobin count, platelet count, hematocrit / hematocrit, total protein, albumin, ALT, AST, fasting blood glucose, homocysteine, creatinine, blood urea nitrogen, CA125, TSH, A-TG, A-TPO, APTT, PT, blood type, chromosome examination, gonococcal DNA, chlamydia DNA, and mycoplasma DNA.
4. The method according to claim 1, characterized in that, The ovarian stimulation strategy described in step (1) is characterized by an ovarian stimulation protocol, an initial FSH dose, an FSH formulation used during ovulation induction, and whether at least two of the following are used during ovulation induction: ovarian stimulation protocol, FSH initial dose, FSH formulation used during ovulation induction, and whether LH is used during ovulation induction.
5. The method according to claim 1, characterized in that, Step (2) specifically involves removing features with more than 15% missing samples from the data described in step (1), then stratifying the data according to the POR event, then randomly selecting 70% of the samples as the pre-imputation training set and the remaining 30% of the samples as the pre-imputation test set, and then performing multiple imputation on the pre-imputation training set and the pre-imputation test set separately to obtain the training set and the test set.
6. The method according to claim 5, characterized in that, The multiple interpolation method is: chain equation multiple interpolation method based on random forest, where the implementation method is random forest, the prediction mean matching method is used for value selection after prediction, the number of candidate matching values is 5, and the number of iterations is 100.
7. The method according to claim 1, characterized in that, The machine learning algorithms mentioned in step (3) include at least two of the following: LASSO-Logistic regression, support vector machine with RBF as kernel function, multilayer perceptron, and XGBoost.
8. The method according to claim 1, characterized in that, The key risk factors for POR mentioned in step (4) are: age, weight, diastolic blood pressure, duration of infertility, AMH level, basal antral follicle count, basal FSH level, basal P level, basal LH level, red blood cell count, white blood cell count, ALT, and whether the diagnosis includes POI or DOR.
9. The method according to claim 1, characterized in that, The methods for verifying the interpolation results include: comparing the distribution of the new data after interpolation with the distribution of the original data, and performing iterative convergence diagnosis during the interpolation process.
10. A system for predicting low ovarian response and constructing a model for deploying individualized ovarian stimulation strategies, characterized in that, The system includes: The data preparation module (1) is used to extract the patient's raw data from the electronic medical record system, preliminarily screen candidate POR risk factors and ovarian stimulation strategy characteristics, and obtain training and test sets after data preprocessing. The screening algorithm module (2) constructs a candidate POR ovarian stimulation strategy deployment model or a candidate POR diagnostic model on the training set based on various machine learning algorithms. The input of the candidate POR diagnostic model is the POR risk factor, and the output is the POR disease risk probability. The input of the candidate POR ovarian stimulation strategy deployment model is the POR risk factor and four ovarian stimulation strategy features, and the output is the POR disease risk probability under different combinations of ovarian stimulation strategy features. The candidate models are evaluated on the test set and the best machine learning algorithm is selected. Feature selection and interpretation module (3) is a model built based on the best machine learning algorithm selected by module (2), and key risk factors of POR are selected according to the size of feature contribution. The model building module (4) reconstructs the POR ovarian stimulation strategy deployment model based on the best machine learning algorithm selected in module (2) and the key risk factors selected in module (3) on the total dataset after merging the training set and the test set; and reconstructs the POR diagnostic model using the key risk factors and ovarian stimulation strategy features selected in module (3). Model validation and interpretation of individual prediction module (5) acquires data from other hospitals or other reproductive centers and performs external data validation on the POR ovarian stimulation strategy deployment model and POR diagnostic model constructed in module (4); acquires new patient data and performs interpretability prediction on the POR ovarian stimulation strategy deployment model and POR diagnostic model constructed in module (4) at the individual level; The ovarian stimulation strategy deployment module (6) obtains 13 key risk factor values for new patients. Doctors develop multiple possible combinations of ovarian stimulation strategies based on clinical experience (different combinations of 4 ovarian stimulation interventions). After inputting into the POR ovarian stimulation strategy deployment model, the risk prediction of POR corresponding to different combinations is obtained. Based on the patient's economic burden, the ovarian stimulation strategy that optimizes POR risk and patient economic burden is then selected.