Method and device for predicting infertility risk based on multi-dimensional factors in combination with XGBoost

The risk prediction model constructed by multi-dimensional data collection and XGBoost algorithm solves the problem of accurate prediction of early infertility risks for couples of childbearing age, provides personalized suggestions, reduces the probability of infertility, and improves the level of reproductive health.

CN120544833APending Publication Date: 2025-08-26INST OF HEALTH & MEDICINE HEFEI COMPREHENSIVE NAT SCI CENT +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510434704.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The prior art cannot accurately predict the risk of infertility in the early stages of pregnancy preparation for couples of childbearing age, and traditional diagnostic methods have problems such as time lag, invasiveness and high cost.

Method used

Based on multi-dimensional factors combined with XGBoost algorithm, by collecting multi-dimensional data of couples of childbearing age, using the Boruta algorithm to screen features, construct a risk prediction model, and calculate SHAP values ​​to achieve accurate prediction of early infertility risk.

Benefits of technology

Provide personalized risk assessments for couples of childbearing age in the early stages of pregnancy preparation, helping them take timely intervention measures, reduce psychological and economic pressure, and improve reproductive health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544833A_ABST
    Figure CN120544833A_ABST
Patent Text Reader

Abstract

The invention provides a method for predicting an infertility risk based on a multi-dimensional factor in combination with XGBoost, and belongs to the field of medical technology and machine learning, and the method comprises the steps: collecting multi-dimensional original data of childbearing age couples and women, and carrying out the data cleaning of the original data, and obtaining the cleaned data; performing feature selection on the cleaned data by using a Boruta algorithm, and screening to obtain a data set containing important feature variables; a risk prediction model is constructed based on an XGBoost algorithm, the risk prediction model is trained with the important feature variables as input and the infertility probability as output, and when a loss function is minimum, the trained risk prediction model is obtained; inputting feature data of a to-be-predicted child-bearing age couple into the trained risk prediction model to obtain a prediction result of an infertility risk, and calculating an SHAP value of each risk feature; the invention further provides a device for predicting the infertility risk. Data are collected from multiple dimensions, variables important for the outcome are screened, and the infertility risk probability is accurately predicted for the childbearing-age couple in the early pregnancy preparation stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of medical technology and machine learning technology, and in particular to a method and device for predicting infertility risk based on multi-dimensional factors combined with XGBoost. Background Art

[0002] Infertility is a global public health issue that places immense psychological and financial pressure on many couples of childbearing age. Accurately predicting the risk of infertility is crucial for couples to proactively implement interventions, develop family planning, and rationally allocate medical resources.

[0003] Currently, clinical infertility diagnosis relies primarily on a series of tests and assessments, such as semen analysis, ovulation monitoring, and fallopian tube patency testing, after couples have tried unsuccessfully for a period of time. However, this traditional diagnostic approach has limitations. First, it is often performed after a couple has tried unsuccessfully for pregnancy, missing the optimal time for early intervention. Second, these tests can be invasive, costly, and physically and psychologically taxing for couples.

[0004] With the continuous development of medical technology and information technology, the application of machine learning algorithms in the medical field is becoming increasingly widespread. For example, the Chinese invention patent application with publication number CN117995410A, "A Method and Apparatus for Predicting Male Infertility Risk and Describing Risk Characteristics," discloses a male infertility risk prediction method, comprising the following steps: Step S1: Obtaining the data to be investigated, which includes the patient's demographic data, occupational and environmental exposure data, nutritional status data, lifestyle data, and family history; Step S2: Inputting the data to be investigated into a pre-trained male infertility risk prediction model; Step S3: The output of the male infertility risk prediction model is the probability of infertility in the subject to be investigated; Step S4: Calculating the SHAP value and analyzing the risk characteristic status of the subject to be investigated. This prediction method greatly improves the accuracy and efficiency of male infertility prediction for patients, but it cannot accurately predict the risk of infertility in couples of childbearing age in the early stages of pregnancy preparation. Summary of the Invention

[0005] The technical problem to be solved by the present invention is how to accurately predict the risk of infertility in couples of childbearing age in the early stages of pregnancy preparation.

[0006] The present invention solves the above technical problems through the following technical solutions: a method for predicting infertility risk based on multi-dimensional factors combined with XGBoost, the method comprising:

[0007] Collect multi-dimensional raw data from couples of childbearing age and clean the raw data to obtain cleaned data;

[0008] Use the Boruta algorithm to perform feature selection on the cleaned data to obtain a data set containing important feature variables;

[0009] A risk prediction model was constructed based on the XGBoost algorithm. The risk prediction model was trained with important feature variables as input and the probability of infertility as output. When the loss function was minimized, the trained risk prediction model was obtained.

[0010] The characteristic data of the childbearing-age couples to be predicted are input into the trained risk prediction model to obtain the prediction results of infertility risk, and the SHAP value of each risk feature is calculated.

[0011] Beneficial effects: The present invention collects data samples of couples of childbearing age from multiple dimensions, uses the Boruta algorithm to screen out variables that are important to the outcome, utilizes the powerful predictive ability of the XGBoost algorithm to construct a risk prediction model, and trains the model with important variables. The trained risk prediction model can accurately predict the probability of infertility risk for each couple of childbearing age in the early stages of pregnancy preparation through analysis of multidimensional behavioral factors. Couples of childbearing age can take timely intervention measures or medical decisions based on the risk probability, adjust their lifestyles, and help couples of childbearing age plan the time and method of childbearing more scientifically, reduce the psychological and economic pressures caused by infertility, thereby reducing the probability of infertility and improving reproductive health.

[0012] Preferably, the multi-dimensional original data includes: demographic and socioeconomic characteristic data, lifestyle and behavioral factor data, dietary factor data, life event data, physical characteristic data, and reproductive history data.

[0013] Beneficial effects: Based on the analysis of multidimensional behavioral factors, the present invention can provide each couple with personalized risk assessment and targeted recommendations, thereby improving the accuracy and efficiency of medical services.

[0014] Preferably, data cleaning of the original data includes sequentially processing outliers and missing values, recoding, and resampling the original data, wherein the resampling is first performing SMOTE oversampling and then performing Tomek Links undersampling.

[0015] Beneficial effect: The present invention first performs SMOTE oversampling and then performs Tomek Links undersampling to achieve the purpose of data balance.

[0016] Preferably, the process of using the Boruta algorithm to perform feature selection on the cleaned data includes:

[0017] For each original feature in the cleaned data, generate a corresponding shadow feature;

[0018] The original features and shadow features are combined as input features, and the random forest model is trained until all the original features are labeled to obtain the trained random forest model.

[0019] Input the original features and shadow features into the trained random forest model to obtain the importance scores of the original features and shadow features;

[0020] The original features whose importance scores are significantly higher than the importance scores of their corresponding shadow features are identified as important feature variables.

[0021] Beneficial effects: The present invention uses the Boruta algorithm to perform feature selection on the cleaned data and screens out 20 important feature variables from 128 candidate variables.

[0022] Preferably, important characteristic variables include: male age, female age, whether contraceptive measures have been used before, female hip circumference, female waist circumference, whether there is pregnancy preparation behavior, female education level, male education level, male waist circumference, male hip circumference, female BMI, male BMI, male smoking years, female smoking years, female nighttime sleep time, female age of first sexual intercourse, female stress at work or study, female tense interpersonal relationships, female being misunderstood or talked about, and the time when men prepare to go to bed on weekends.

[0023] Preferably, the predicted value Y of the risk prediction model i for:

[0024]

[0025] The loss function L is:

[0026]

[0027] Among them, X i is the feature vector of the i-th sample, f k (X i ) is the predicted value of the k-th tree for the i-th sample, K is the number of decision trees, n is the total number of samples, y i The probability of infertility.

[0028] Preferably, the loss function L uses the approximate value L after Taylor expansion t for:

[0029]

[0030] Among them, f t (x i ) is the predicted value of the model for the i-th sample in the t-th iteration, g i is the loss function with respect to the predicted value f t (xi ), h i is the loss function with respect to the predicted value f t (x i )’s second-order derivative;

[0031] The optimal weight w of the jth leaf node j for:

[0032]

[0033] Among them, I j is the set of the jth leaf node, and λ is the regularization parameter.

[0034] Preferably, when training a risk prediction model with important characteristic variables as input and infertility probability as output, the important characteristic variables are divided into a training set and a validation set. The training set is used for model learning and parameter adjustment, and the validation set is used to evaluate the performance of the model under different hyperparameter combinations. The trained risk prediction model is further evaluated and verified by an external validation set, which is sample data collected based on important characteristic variables.

[0035] Preferably, the SHAP value of each risk feature i for:

[0036]

[0037] Among them, S is any subset of the feature set N, indicates that subset S does not contain risk feature i, |S| represents the number of risk features in subset S, v(S) represents the contribution of subset S to the model output, that is, the predicted value of the model when only the risk features in subset S are considered, and v(S∪{i}) represents the contribution of subset S to the model output after adding risk feature i.

[0038] The present invention also provides a device for predicting infertility risk based on multi-dimensional factors combined with XGBoost, the device comprising:

[0039] The data processing module is used to collect multi-dimensional raw data of couples of childbearing age and clean the raw data to obtain cleaned data;

[0040] The feature selection module is used to perform feature selection on the cleaned data using the Boruta algorithm to screen out a data set containing important feature variables;

[0041] The model training module is used to build a risk prediction model based on the XGBoost algorithm. The risk prediction model is trained with important feature variables as input and the probability of infertility as output. When the loss function is minimized, the trained risk prediction model is obtained.

[0042] The risk prediction module is used to input the characteristic data of the childbearing-age couples to be predicted into the trained risk prediction model, obtain the prediction results of infertility risk, and calculate the SHAP value of each risk feature.

[0043] The advantages provided by the present invention also include: the prediction results of the present invention on the infertility risk of couples of childbearing age can provide strong support and reference basis for doctors' diagnosis and treatment decisions, thereby improving the accuracy of medical decisions; and early risk prediction and intervention can reduce unnecessary medical examinations and treatments, thereby saving limited medical resources and allocating them more effectively to patients who really need them. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A flowchart of a method for predicting infertility risk based on multi-dimensional factors combined with XGBoost provided in an embodiment of the present invention;

[0045] Figure 2 A schematic diagram of a method for predicting infertility risk based on multi-dimensional factors combined with XGBoost, provided in an embodiment of the present invention;

[0046] Figure 3 A schematic diagram of important characteristic variables screened out in the method for predicting infertility risk based on multi-dimensional factors combined with XGBoost provided in an embodiment of the present invention;

[0047] Figure 4 The area under the curve of the internal validation set subjects in the method for predicting infertility risk based on multidimensional factors combined with XGBoost provided in an embodiment of the present invention;

[0048] Figure 5 A plot of the area under the curve for subjects in the external validation set of the method for predicting infertility risk based on multidimensional factors combined with XGBoost provided in an embodiment of the present invention;

[0049] Figure 6 The SHAP values ​​of each risk feature in the method for predicting infertility risk based on multi-dimensional factors combined with XGBoost provided in an embodiment of the present invention;

[0050] Figure 7 A schematic diagram of a device for predicting infertility risk based on multi-dimensional factors combined with XGBoost provided by an embodiment of the present invention;

[0051] Figures 8-10 Schematic diagram of the prediction results of the method for predicting infertility risk based on multi-dimensional factors combined with XGBoost provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the following describes the technical solutions of the present invention clearly and completely with reference to specific embodiments and the accompanying drawings. It is obvious that the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0053] like Figure 1 and Figure 2 As shown, this embodiment provides a method for predicting infertility risk based on multi-dimensional factors combined with XGBoost, including the following steps:

[0054] Step 1: Collect multi-dimensional original data from both parties of childbearing age, including: demographic and socioeconomic characteristics data, lifestyle and behavioral factors data, dietary factors data, life event data, physical characteristics data, and reproductive history data, and clean the original data. The cleaning includes processing outliers and missing values, recoding, and resampling in sequence to obtain cleaned data.

[0055] Among them, when collecting demographic and socioeconomic characteristic data, the ordered categorical variables are defined as 1, 2, 3..., and the unordered binary categorical variables are defined as 0, 1 among the obtained age, education level, personal annual income, and occupation.

[0056] When collecting data on lifestyle and behavioral factors, ordered categorical variables were defined as 1, 2, 3..., and unordered binary variables were defined as 0 / 1, among the factors obtained, such as smoking frequency, drinking frequency, outdoor activity time, nighttime sleep duration, sleep chronotype, screen use frequency, sleep quality, snoring frequency, sleep apnea, lunch break length, sleepiness scale (ESS) score, depression scale (PHQ9) score, sedentary time, physical activity time, and shoulder, neck, and back pain.

[0057] When collecting data on dietary factors, ordered categorical variables were defined as 1, 2, 3..., and unordered binary variables were defined as 0, 1 for factors such as fruits, vegetables, grains, eggs, beans and their products, nuts, red meat, poultry, aquatic products, smoked, barbecued, and fried foods, sugary soft drinks, milk and dairy products, takeout, and the frequency of use of disposable tableware.

[0058] When collecting life event data, binary unordered variables were converted to 0 / 1 for factors such as whether parents were not in harmony, whether the family had housing shortages, whether the family had moved, whether a family member had been subject to criminal or civil penalties, whether a family member was seriously ill or injured, whether a family member had died, whether the family was unemployed or unemployed, whether the family member had high work or study pressure, whether the family member had tense relationships with family members, colleagues, or friends, whether the family member had been misunderstood, falsely accused, or gossiped about, whether the family member had property damage or investment errors, and whether the family member had been frightened or had an accident or natural disaster.

[0059] When collecting physical characteristic data, the physical characteristics obtained included height, weight, BMI, waist circumference, and hip circumference.

[0060] When collecting reproductive history data, among the factors obtained, such as the age of female menarche, the age of male first nocturnal emission, whether contraceptive tools have been used, fertility intention, pregnancy history, newborn birth history, stillbirth history, spontaneous abortion history, induced abortion history, and age of first sexual intercourse, ordered categorical variables were defined as 1, 2, 3..., and unordered binary variables were defined as 0, 1.

[0061] The fertility status of the couples was obtained through follow-up and divided into normal fertility and infertility, and the fertility status outcome was defined as 0 / 1.

[0062] The methods used to process outliers and missing values ​​in the original data are as follows: outliers are eliminated and missing values ​​are multiple-interpolated using the MICE method. Recoding refers to correcting the variable names to understandable names. Resampling involves first performing SMOTE oversampling and then Tomek Links undersampling.

[0063] The Synthetic Minority Over-sampling Technique (SMOTE) is a method used to address class imbalance in datasets. In many real-world applications, the number of samples in different classes often varies significantly. This imbalance can lead to poor learning of the minority class during model training, which in turn affects the model's overall performance. SMOTE generates new minority class samples to increase the number of minority class samples, thereby making the class distribution in the dataset more balanced.

[0064] The SMOTE oversampling process includes:

[0065] (1) Determine minority class samples: First, identify a smaller number of class samples from the original data. These samples are the objects that need to be oversampled.

[0066] (2) Select neighborhood: For each minority class sample x, use the Euclidean distance to select k neighbor samples from its k nearest neighbors (k is a pre-set parameter);

[0067] (3) Generate new samples: Randomly select a neighbor sample xneighbor from the selected k neighbors, and then randomly select a point on the line between xneighbor and xneighbor to generate a new minority class sample.

[0068] The formula is: new =x+rand(0,1)×(x 邻近 -x), where x new represents a final synthesized sample, x represents an input minority class sample, xnearby represents a neighboring sample of the selected x, and rand(0,1) is a random number between 0 and 1.

[0069] Tomek Links undersampling is a method for dealing with class imbalance in datasets. Based on the concept of Tomek Links, this method adjusts the ratio of the number of samples of different categories by removing specific pairs of samples in the dataset, making the dataset more balanced and thus improving the performance of the machine learning model. Tomek Links ratio makes the dataset more balanced and thus improves the performance of the machine learning model. Tomek Links is a pair of samples in the dataset that belong to different categories and the distance between them is the shortest among all pairs of samples. If there are two samples X i and X j , satisfying: X i and X j Belong to different categories, there is no other sample pair ((X m , X n )), so that d(X m , X n ) <d(X i , X j ), d represents some distance metric (Euclidean distance).

[0070] The majority class examples in Tomek Links are often located near class boundaries. These examples can interfere with the classifier's ability to learn the correct decision boundary, resulting in a decrease in the model's ability to identify minority class examples. Therefore, this undersampling method identifies all Tomek Links in the dataset and removes the majority class examples. Tomek Links generates new minority class examples, increasing the number of minority class examples and achieving a more balanced class distribution in the dataset.

[0071] The present invention uses a combination of these two methods, that is, first performing SMOTE oversampling and then performing Tomek Links undersampling to achieve the purpose of data balancing.

[0072] Step 2: Use the Boruta algorithm to perform feature selection on the cleaned data to obtain a data set containing important feature variables.

[0073] The Boruta algorithm evaluates the importance of each original feature by creating shadow features. Shadow features are randomly shuffled versions of the original features, and theoretically have no true correlation with the target variable. By comparing the importance of the original and shadow features in the random forest model, the importance of the original features can be determined.

[0074] The process of using the Boruta algorithm to select features from cleaned data includes:

[0075] Step 2.1: Create shadow features: For each original feature, generate a corresponding shadow feature. Shadow features are obtained by randomly scrambling the values ​​of the original feature. Their purpose is to serve as a benchmark for comparing the importance of the original feature.

[0076] Step 2.2: Combine the original features and shadow features as input features and train a random forest model using these features until all original features are labeled, resulting in a trained random forest model. Each tree in the random forest randomly selects and splits features during construction, allowing us to calculate the importance score for each feature. The feature importance metric used is based on the Gini impurity metric.

[0077] Step 2.3: Input the original features and shadow features into the trained random forest model to obtain the importance scores of the original features and shadow features;

[0078] Step 2.4: Compare the importance scores of the original feature and the shadow feature. If the importance score of an original feature is significantly higher than that of its corresponding shadow feature, then the original feature is considered important; conversely, if the importance score of the original feature is similar to or lower than that of the shadow feature, then the original feature may be unimportant.

[0079] Original features with significantly higher importance scores than shadow features are marked as "confirmed." These features are considered important features related to the target variable. Original features with significantly lower importance scores than shadow features are marked as "rejected." These features are considered unimportant features unrelated to the target variable and can be deleted. Original features with importance scores similar to those of shadow features are marked as "tentative." The importance of these features cannot be determined yet and requires further processing.

[0080] This paper uses the Boruta algorithm to perform feature selection on the cleaned data, see Figure 3 Among the 128 candidate variables, we screened out the variables that are important to the outcome, such as male age, female age, whether contraceptive measures have been used before, female hip circumference, female waist circumference, whether there is preparation for pregnancy, female education level, male education level, male waist circumference, male hip circumference, female BMI, male BMI, male smoking years, female smoking years, female nighttime sleep time, age of first sexual intercourse, female stress at work or study, female tense interpersonal relationships, female being misunderstood or talked about, and the time men prepare to go to bed on weekends.

[0081] Step 3: Build a risk prediction model based on the XGBoost algorithm, use important feature variables as input and the probability of infertility as output to train the risk prediction model. When the loss function is minimized, the trained risk prediction model is obtained.

[0082] The goal of XGBoost is to use a set of decision trees to predict a target variable. The predicted value Y of the risk prediction model is i is the weighted sum of each tree:

[0083]

[0084] Among them, X i is the feature vector of the i-th sample, f k (X i ) is the predicted value of the k-th tree for the i-th sample, K is the number of decision trees, and n is the total number of samples.

[0085] The loss function L uses the mean square error, and the calculation formula is:

[0086]

[0087] Among them, y i The probability of infertility.

[0088] XGBoost trains each tree by iteratively optimizing the loss function. In each round, a new tree f is added t , hoping to minimize losses.

[0089] The loss function L uses the approximation L after Taylor expansion t for:

[0090]

[0091] Among them, f t (x i ) is the predicted value of the model for the i-th sample in the t-th iteration, g i is the loss function with respect to the predicted value f t (xi ), h i is the loss function with respect to the predicted value f t (x i ). By minimizing this approximate loss function, the optimal solution can be found more efficiently.

[0092] For each leaf node, we can calculate an optimal weight to reduce the error. The optimal weight w of the jth leaf node is j for:

[0093]

[0094] Among them, I j is the set of the jth leaf node, and λ is the regularization parameter used to control the complexity of the model and prevent overfitting.

[0095] The present invention inputs a dataset containing important characteristic variables into the risk prediction model for training. The dataset is usually divided into a training set and a validation set in an 8:2 ratio. The training set is used for model learning and parameter adjustment. The validation set is used to evaluate the performance of the model under different hyperparameter combinations to select the optimal hyperparameters. To further evaluate the trained model, the present invention also collects an external validation set based on the selected important characteristic variables to ultimately objectively evaluate the performance and generalization ability of the model. The indicators used when evaluating the trained model include accuracy, sensitivity, specificity, recall rate, F1 value and area under the ROC curve. The internal validation set refers to a portion of data separated from the same study population and used to verify the performance of the model. These data come from the same population as the model training data and have similar feature distributions. The external validation set refers to a dataset from a different population or different study than the model training data and is used to verify the generalization ability of the model in independent samples. These data differ from the training data in terms of feature distribution, sample source, etc. Figure 4 This is the area under the curve of the internal validation set of subjects using XGBoost to predict infertility in couples of childbearing age. Figure 5 The area under the curve for the external validation set of subjects using XGBoost to predict infertility in couples of reproductive age.

[0096] Step 4: Input the characteristic data of the childbearing-age couples to be predicted into the trained risk prediction model to obtain the prediction results of infertility risk and calculate the SHAP value of each risk feature.

[0097] SHAP value of each risk feature i for:

[0098]

[0099] Among them, S is any subset of the feature set N, indicates that subset S does not contain risk feature i, |S| represents the number of risk features in subset S, v(S) represents the contribution of subset S to the model output, that is, the prediction value of the model when only considering the risk features in subset S, v(S∪{i}) represents the contribution of subset S plus risk feature i to the model output, that is, the prediction value of the model when considering subset S and risk feature i, v(S∪{i})-v(S) represents the marginal contribution of risk feature i after adding it to subset S, that is, the incremental contribution of feature i to the model output.

[0100] The above formula is summed over all possible subsets S, and the contribution of each subset S is weighted by the number of combinations N! |S|! (N-|S|-1)!. This number of combinations actually calculates the number of different positions that risk feature i can be inserted into subset S, ensuring that the contribution of each feature is calculated fairly. Figure 6 As shown, from the y-axis, the importance of variables is arranged from high to low, and from the x-axis, the closer the color is to yellow, the higher the risk of infertility.

[0101] The present invention collects data samples of couples of childbearing age from multiple dimensions, uses the Boruta algorithm to screen out variables that are important to the outcome, utilizes the powerful predictive ability of the XGBoost algorithm to construct a risk prediction model, and trains the model with important variables. The trained risk prediction model can accurately predict the infertility risk probability for each couple of childbearing age in the early stages of pregnancy preparation through analysis of multidimensional behavioral factors. Couples of childbearing age can take timely intervention measures or medical decisions based on the risk probability, adjust their lifestyles, and help couples of childbearing age plan the time and method of childbearing more scientifically, reduce the psychological and economic pressures caused by infertility, thereby reducing the probability of infertility and improving reproductive health.

[0102] The prediction results of the present invention on the infertility risk of couples of childbearing age can provide strong support and reference basis for doctors' diagnosis and treatment decisions, improve the accuracy of medical decisions, and early risk prediction and intervention can reduce unnecessary medical examinations and treatments, thereby saving limited medical resources and allocating them more effectively to patients who really need them.

[0103] See also Figure 7 The present invention also provides a device for predicting infertility risk based on multi-dimensional factors combined with XGBoost, which corresponds to the above method. For details not disclosed in the device of the present invention, please refer to the method of the present invention, which will not be repeated here. The device includes:

[0104] The data processing module is used to collect multi-dimensional raw data from couples of childbearing age and clean the raw data to obtain cleaned data; the data processing module has the ability to access multi-source data and can receive influencing parameters from multiple channels such as medical institutions, user self-filling, third-party testing, etc., and supports data input in different formats and types.

[0105] The feature selection module is used to perform feature selection on the cleaned data using the Boruta algorithm to screen out data sets containing important feature variables.

[0106] The model training module is used to build a risk prediction model based on the XGBoost algorithm. The risk prediction model is trained with important feature variables as input and the probability of infertility as output. When the loss function is minimized, the trained risk prediction model is obtained.

[0107] The risk prediction module is used to input the characteristic data of the childbearing-age couples to be predicted into the trained risk prediction model, obtain the prediction results of infertility risk, and calculate the SHAP value of each risk feature.

[0108] The result output module is used to output and display the prediction results of infertility risk.

[0109] See also Figures 8-10 The infertility risk prediction results are presented to users in an intuitive and easy-to-understand manner, such as providing risk levels (high, medium, and low) and risk probability values. A detailed description of the risk profile is also provided, along with explanations and recommendations, such as lifestyle adjustments and further testing recommendations. The user-friendly interface and simple operation make it easy for medical staff and couples of childbearing age to use, without the need for complex training or specialized knowledge.

[0110] The entire device boasts excellent scalability and compatibility, enabling continuous updating and optimization of prediction models and parameters as new research findings and data accumulate. The device incorporates strict data security and privacy protection mechanisms to ensure the security and confidentiality of input personal information and influencing parameters of reproductive-age couples during processing and storage.

[0111] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for predicting infertility risk based on multi-dimensional factors combined with XGBoost, characterized by: Methods include: Collect multi-dimensional raw data from couples of childbearing age and clean the raw data to obtain cleaned data; Use the Boruta algorithm to perform feature selection on the cleaned data to obtain a data set containing important feature variables; A risk prediction model was constructed based on the XGBoost algorithm. The risk prediction model was trained with important feature variables as input and the probability of infertility as output. When the loss function was minimized, the trained risk prediction model was obtained. The characteristic data of the childbearing-age couples to be predicted are input into the trained risk prediction model to obtain the prediction results of infertility risk, and the SHAP value of each risk feature is calculated.

2. The method for predicting infertility risk based on multi-dimensional factors combined with XGBoost according to claim 1, characterized in that: The multi-dimensional original data include: demographic and socioeconomic characteristics data, lifestyle and behavioral factors data, dietary factors data, life event data, physical characteristics data, and reproductive history data.

3. The method for predicting infertility risk based on multi-dimensional factors combined with XGBoost according to claim 1, characterized in that: Data cleaning of the original data includes processing outliers and missing values, recoding, and resampling of the original data in sequence. Among them, the resampling is first SMOTE oversampling and then Tomek Links undersampling.

4. The method for predicting infertility risk based on multi-dimensional factors combined with XGBoost according to claim 1, characterized in that: The process of using the Boruta algorithm to select features from cleaned data includes: For each original feature in the cleaned data, generate a corresponding shadow feature; The original features and shadow features are combined as input features, and the random forest model is trained until all the original features are labeled to obtain the trained random forest model. Input the original features and shadow features into the trained random forest model to obtain the importance scores of the original features and shadow features; The original features whose importance scores are significantly higher than the importance scores of their corresponding shadow features are identified as important feature variables.

5. The method for predicting infertility risk based on multi-dimensional factors combined with XGBoost according to claim 1, characterized in that: Important characteristic variables include: male age, female age, whether contraceptive measures have been used before, female hip circumference, female waist circumference, whether there is any preparation for pregnancy, female education level, male education level, male waist circumference, male hip circumference, female BMI, male BMI, male smoking years, female smoking years, female nighttime sleep time, female age of first sexual intercourse, female stress at work or study, female tense interpersonal relationships, female being misunderstood or talked about, and the time men prepare to go to bed on their days off.

6. The method for predicting infertility risk based on multi-dimensional factors combined with XGBoost according to claim 1, characterized in that: Predicted value Y of the risk prediction model i for: The loss function L is: Among them, X i is the feature vector of the i-th sample, f k (X i ) is the predicted value of the k-th tree for the i-th sample, K is the number of decision trees, n is the total number of samples, y i The probability of infertility.

7. The method for predicting infertility risk based on multi-dimensional factors combined with XGBoost according to claim 6, characterized in that: The loss function L uses the approximation L after Taylor expansion t for: Among them, f t (x i ) is the predicted value of the model for the i-th sample in the t-th iteration, g i is the loss function with respect to the predicted value f t (x i ), h i is the loss function with respect to the predicted value f t (x i )’s second-order derivative; The optimal weight w of the jth leaf node j for: Among them, I j is the set of the jth leaf node, and λ is the regularization parameter.

8. The method for predicting infertility risk based on multi-dimensional factors combined with XGBoost according to claim 1, characterized in that: When training a risk prediction model with important characteristic variables as input and infertility probability as output, the important characteristic variables are divided into a training set and a validation set. The training set is used for model learning and parameter adjustment, and the validation set is used to evaluate the performance of the model under different hyperparameter combinations. The trained risk prediction model is further evaluated and verified through an external validation set, which is sample data collected based on important characteristic variables.

9. The method for predicting infertility risk based on multi-dimensional factors combined with XGBoost according to claim 1, characterized in that: SHAP value of each risk feature i for: Among them, S is any subset of the feature set N, indicates that subset S does not contain risk feature i, |S| represents the number of risk features in subset S, v(S) represents the contribution of subset S to the model output, that is, the predicted value of the model when only the risk features in subset S are considered, and v(S∪{i}) represents the contribution of subset S to the model output after adding risk feature i.

10. A device for predicting infertility risk based on multi-dimensional factors combined with XGBoost, characterized by: The device includes: The data processing module is used to collect multi-dimensional raw data of couples of childbearing age and clean the raw data to obtain cleaned data; The feature selection module is used to perform feature selection on the cleaned data using the Boruta algorithm to screen out a data set containing important feature variables; The model training module is used to build a risk prediction model based on the XGBoost algorithm. The risk prediction model is trained with important feature variables as input and the probability of infertility as output. When the loss function is minimized, the trained risk prediction model is obtained. The risk prediction module is used to input the characteristic data of the childbearing-age couples to be predicted into the trained risk prediction model, obtain the prediction results of infertility risk, and calculate the SHAP value of each risk feature.

Citation Information

Patent Citations

  • Male infertility risk prediction and risk feature description method and device

    CN117995410A