Ovarian reactivity prediction method and device based on machine learning technology

Through the ovarian responsiveness prediction method based on machine learning, the nonlinear regression model and DeLong test were used to determine the minimum number of tests and screen key indicators, which solved the problem of prediction accuracy of high-dimensional continuous data and achieved efficient and accurate ovarian responsiveness prediction.

CN120597012APending Publication Date: 2025-09-05CHANGSHA CENT HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411737218.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing ovarian responsiveness assessment methods have poor prediction accuracy when processing high-dimensional continuous data, and traditional statistical methods are difficult to effectively handle nonlinear patterns, resulting in insufficient accuracy and generalization ability of ovarian responsiveness prediction.

Method used

An ovarian responsiveness prediction method based on machine learning technology is adopted. By obtaining ovarian data from multiple consecutive tests, data preprocessing and statistical transformation are performed to determine key indicators. A nonlinear regression machine learning model is used for prediction. Combined with the DeLong test, the minimum test number threshold is determined, key indicators are screened, and an efficient prediction model is constructed.

Benefits of technology

It improves the accuracy and generalization ability of ovarian responsiveness prediction, reduces the number of unnecessary tests, reduces costs and improves testing efficiency, and provides guidance for clinical treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597012A_ABST
    Figure CN120597012A_ABST
Patent Text Reader

Abstract

The invention discloses an ovarian reactivity prediction method and device based on a machine learning technology, and relates to the field of medical data analysis and artificial intelligence, and the method comprises the steps: obtaining to-be-predicted ovarian reactivity detection data of a patient, and carrying out the data preprocessing and statistical transformation, and obtaining the ovarian detection data after statistical transformation; and determining data corresponding to the key indexes in the ovarian detection data after statistical transformation, and inputting the data corresponding to the key indexes into the ovarian reactivity prediction model to obtain an ovarian reactivity prediction result. According to the method, the to-be-predicted ovarian reactivity detection data with high-dimensional continuity is subjected to statistical transformation to obtain the feature data set with statistical significance, so that the generalization ability of processing high-dimensional data for continuous multiple days can be improved when the ovarian reactivity prediction model performs ovarian reactivity prediction, and the accuracy of ovarian reactivity prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of medical data analysis and artificial intelligence, and in particular to a method and device for predicting ovarian responsiveness based on machine learning technology. Background Art

[0002] In existing medical research, the assessment of female ovarian responsiveness is a crucial yet complex component of assisted reproductive technology. Ovarian responsiveness refers to the ovarian response to exogenous hormones and other stimuli, which directly impacts treatment success and complication rates. Currently, clinical assessment of ovarian responsiveness relies primarily on indicators such as patient age, hormone levels, and ovarian reserve. However, ovarian responsiveness often exhibits dynamic changes, and these traditional assessment methods are mostly based on a single or a few indicators acquired at a single time, resulting in limitations in data comprehensiveness and predictive accuracy. Traditional statistical methods often struggle with this type of high-dimensional, continuous monitoring data, particularly when dealing with nonlinear patterns, which limits their predictive power. With the rapid accumulation of medical data and the development of artificial intelligence, new opportunities are emerging in the field of medical data analysis. Support vector machines (SVMs) are a commonly used approach for predicting ovarian responsiveness. They utilize kernel techniques to effectively process high-dimensional data, maintaining good performance even when the original feature space is very high dimensional. Furthermore, due to the maximum margin principle, they are robust to noise and outliers in the data. However, they perform poorly for continuous, high-dimensional data, resulting in inaccurate predictions. Summary of the Invention

[0003] The purpose of this application is to provide an ovarian responsiveness prediction method and device based on machine learning technology, which can improve the generalization ability of the ovarian responsiveness prediction model in processing high-dimensional data for multiple consecutive days when making ovarian responsiveness predictions, thereby improving the accuracy of ovarian responsiveness predictions.

[0004] To achieve the above objectives, this application provides the following solutions:

[0005] In a first aspect, the present application provides a method for predicting ovarian responsiveness based on machine learning technology, comprising:

[0006] Obtaining the patient's ovarian response test data to be predicted; the ovarian response test data to be predicted includes ovarian data of multiple consecutive tests; the number of tests in the ovarian response test data to be predicted is greater than or equal to a minimum test number threshold;

[0007] performing data preprocessing and statistical transformation on the ovarian response detection data to be predicted to obtain statistically transformed ovarian detection data; the statistically transformed ovarian detection data includes the preprocessed ovarian detection data and the statistically transformed data of the ovarian response detection data to be predicted;

[0008] Determining data corresponding to key indicators in the ovarian detection data after the statistical transformation;

[0009] The data corresponding to the key indicators in the ovarian detection data after the statistical transformation are input into the ovarian response prediction model to obtain the ovarian response prediction result; the ovarian response prediction model is a nonlinear regression machine learning model.

[0010] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned ovarian responsiveness prediction method based on machine learning technology.

[0011] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned ovarian responsiveness prediction method based on machine learning technology.

[0012] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0013] The present application provides an ovarian responsiveness prediction method and device based on machine learning technology. The method includes obtaining the patient's ovarian responsiveness test data to be predicted and performing data preprocessing and statistical transformation to obtain the ovarian test data after statistical transformation; determining the data corresponding to the key indicators in the ovarian test data after statistical transformation, and inputting the data corresponding to the key indicators into the ovarian responsiveness prediction model to obtain the ovarian responsiveness prediction result. In the present application, on the one hand, time series data is used to predict ovarian responsiveness, which has high prediction accuracy compared to the existing prediction based on a single indicator or multiple indicators at a single time point; on the other hand, the high-dimensional continuous ovarian responsiveness test data to be predicted is statistically transformed to obtain a feature data set with statistical significance, so as to better fit the actual clinical situation and improve the prediction accuracy. It can also improve the generalization ability of the ovarian responsiveness prediction model to process high-dimensional data for multiple consecutive days when predicting ovarian responsiveness, thereby improving the accuracy of ovarian responsiveness prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0015] Figure 1This is a diagram of an application environment of an ovarian responsiveness prediction method based on machine learning technology in one embodiment of the present application;

[0016] Figure 2 A flowchart of a method for predicting ovarian responsiveness based on machine learning technology provided in one embodiment of the present application;

[0017] Figure 3 A schematic diagram of the technical concept of a method for predicting ovarian responsiveness based on machine learning technology provided in one embodiment of the present application;

[0018] Figure 4 A schematic diagram of a set of key indicators for predicting ovarian responsiveness provided in one embodiment of the present application;

[0019] Figure 5 Schematic diagram of ROC curves of training models corresponding to different data subsets provided in one embodiment of the present application;

[0020] Figure 6 A schematic diagram of the statistical difference P-value between the models corresponding to each subset provided in an embodiment of the present application;

[0021] Figure 7 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0024] The ovarian response prediction method based on machine learning technology provided in the embodiment of the present application is specifically an ovarian response prediction method based on high-dimensional time series data and machine learning technology, which can be applied to Figure 1In the application environment shown, the terminal communicates with the server via a network. The data storage system can store data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the ovarian response test data to be predicted to the server. After the server receives the ovarian response test data to be predicted, the server performs data preprocessing and statistical transformation on the ovarian response test data to be predicted, and obtains the ovarian test data after the statistical transformation; determines the data corresponding to the key indicators in the ovarian test data after the statistical transformation; inputs the data corresponding to the key indicators in the ovarian test data after the statistical transformation into the ovarian response prediction model to obtain the ovarian response prediction result. The server can feedback the obtained ovarian response prediction result to the terminal. In addition, in some embodiments, the ovarian response prediction method based on machine learning technology can also be implemented separately by the server or the terminal. For example, the terminal can directly perform ovarian response prediction processing on the ovarian response test data to be predicted, or the server can obtain the ovarian response test data to be predicted from the data storage system and perform ovarian response prediction processing.

[0025] Terminals include, but are not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices include smart watches, smart bracelets, and head-mounted devices. Servers can be implemented as standalone servers or server clusters consisting of multiple servers, or even cloud servers.

[0026] In an exemplary embodiment, Figure 2 As shown, a method for predicting ovarian response based on machine learning technology is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The server in is used as an example to illustrate, including the following steps 101 to 104.

[0027] in:

[0028] Step 101: Obtain the patient's ovarian response test data to be predicted. The ovarian response test data to be predicted includes ovarian data from multiple consecutive tests. The number of tests in the ovarian response test data to be predicted is greater than or equal to a minimum test number threshold. The ovarian test data includes ovarian hormone indicators, ovarian ultrasound data, and physical index data.

[0029] Step 102: preprocess and statistically transform the ovarian response detection data to be predicted to obtain statistically transformed ovarian detection data. The statistically transformed ovarian detection data includes the preprocessed ovarian detection data and the statistically transformed data of the ovarian response detection data to be predicted.

[0030] Step 103: Determine data corresponding to key indicators in the ovarian detection data after the statistical transformation.

[0031] Step 104: inputting the data corresponding to the key indicators in the statistically transformed ovarian detection data into an ovarian response prediction model to obtain an ovarian response prediction result; the ovarian response prediction model is a nonlinear regression machine learning model.

[0032] By implementing steps 101 to 104 above, the present application uses a statistical transformation method to process multi-day continuous high-dimensional ovarian test data. By statistically transforming the high-dimensional continuous data, a statistically significant feature data set is obtained, providing high-quality data for model prediction. This overcomes the shortcomings of previous ovarian responsiveness modeling methods for processing multi-day continuous high-dimensional data, which have poor generalization ability and low classification accuracy. In addition, when predicting ovarian responsiveness, the data corresponding to the key indicators in the statistically transformed ovarian test data are input into the ovarian responsiveness prediction model. The data corresponding to the key indicators are input into the model, and both the ovarian responsiveness test data to be predicted and the statistically transformed ovarian test data are input into the model, which can more accurately predict ovarian responsiveness.

[0033] In this application, in the existing research on ovarian response prediction models, no effective solution is provided to determine the minimum number of tests required for ovarian response prediction. Therefore, in order to ensure the accuracy of ovarian response prediction and to ensure that patients undergo ovarian data testing for the minimum number of tests, this application provides a method for determining the minimum number of tests threshold by splitting the complete historical ovarian test data set into different data subsets, such as Figure 3 As shown, one or more models are trained on each data subset, and then the DeLong test is used to compare whether there is a significant difference in the area under the curve (AUC) between the models trained on different data subsets. The performance difference of the models trained on different data subsets under different monitoring times is evaluated by the DeLong test, thereby providing a basis for determining the minimum detection number threshold. Therefore, when obtaining the patient's ovarian response test data to be predicted, the minimum detection number threshold is first determined, and then the ovarian response test data to be predicted with a detection number greater than or equal to the minimum detection number threshold is obtained. Therefore, as an optional embodiment, a model is used to determine the minimum detection number threshold. Step 101, obtaining the patient's ovarian response test data to be predicted, specifically includes:

[0034] (a1) Acquire the patient's first ovarian history test data; the first ovarian history test data includes multiple consecutive ovarian test data.

[0035] (a2) performing data preprocessing and statistical transformation on the first ovarian historical detection data to obtain the ovarian historical detection data after the first statistical transformation.

[0036] (a3) Dividing the ovarian historical test data after the first statistical transformation into multiple data subsets according to the number of tests; one data subset corresponds to one number of tests; one data subset includes the previous m ovarian historical test data; m is the number of tests corresponding to the data subset; the data subset corresponding to the maximum number of tests is recorded as the full data set.

[0037] (a4) Using each data subset and the entire data set, a nonlinear regression machine learning model is trained.

[0038] (a5) Determine the first AUC evaluation metric of the trained nonlinear regression machine learning model corresponding to each data subset.

[0039] (a6) Determine a second AUC evaluation metric of the trained nonlinear regression machine learning model corresponding to the entire data set.

[0040] (a7) DeLong's test is used to evaluate the difference between each first AUC evaluation index and the second AUC evaluation index. The AUC evaluation index refers to the area under the ROU curve of the model.

[0041] (a8) The minimum number of detections corresponding to the minimum difference is used as the minimum detection threshold.

[0042] Taking the XGBoost model as an example, each data subset corresponds to a trained XGBoost model, and each trained XGBoost model has a first AUC evaluation indicator. Using the DeLong test method, the difference between each first AUC evaluation indicator and the second AUC evaluation indicator of the XGBoost model trained on the entire data set is evaluated. In the DeLong test evaluation results, a p-value less than 0.05 is considered to have a large difference, and a p-value greater than 0.05 is considered to have a small difference. Therefore, the p-value greater than 0.05 and the minimum number of detections are selected as the minimum detection threshold. Therefore, the data subset with the smallest difference between the first AUC evaluation indicator and the second AUC evaluation indicator corresponding to the entire data set and the smallest number of detections can be considered to have the minimum detection threshold.

[0043] (a9) Obtaining ovarian test data of the patient with at least a minimum test number threshold, and obtaining ovarian response test data to be predicted.

[0044] As another optional embodiment, a plurality of models are used to determine the minimum detection number threshold. By constructing XGBoost, Catboost, LightFBM, ANN and SVM model pairs on the complete data set, the first four times (a data subset), the first three times (a data subset), the first two times (a data subset) and the univariate data set (a data subset), respectively, the statistical differences of the models trained on different subsets are calculated by DeLong test to determine the minimum detection number threshold. Therefore, step 101, obtaining the patient's ovarian response test data to be predicted, specifically includes:

[0045] (b1) Acquire the patient's first ovarian history test data; the first ovarian history test data includes multiple consecutive ovarian test data.

[0046] (b2) performing data preprocessing and statistical transformation on the first ovarian historical detection data to obtain the ovarian historical detection data after the first statistical transformation.

[0047] (b3) Divide the ovarian historical test data after the first statistical transformation into multiple data subsets according to the number of tests; one data subset corresponds to one number of tests; one data subset includes the previous m ovarian historical test data; m is the number of tests corresponding to the data subset; the data subset corresponding to the maximum number of tests is recorded as the full data set.

[0048] (b4) For each data subset, use the data subset to train multiple nonlinear regression machine learning models, such as XGBoost, CatBoost, LightGBM, artificial neural network (ANN), and support vector machine (SVM).

[0049] (b5) Use the entire data set to train multiple nonlinear regression machine learning models.

[0050] (b6) Determine the first AUC evaluation indicator of each trained nonlinear regression machine learning model corresponding to each data subset.

[0051] (b7) Determine a second AUC evaluation indicator of the trained nonlinear regression machine learning model corresponding to the entire data set.

[0052] (b8) The DeLong test method is used to evaluate the difference between each first AUC evaluation indicator and the second AUC evaluation indicator corresponding to the same nonlinear regression machine learning model.

[0053] For the same nonlinear regression machine learning model, taking the XGBoost model as an example, each data subset corresponds to a trained XGBoost model, and each trained XGBoost model has a first AUC evaluation indicator. The DeLong test method is used to evaluate the difference between each first AUC evaluation indicator and the second AUC evaluation indicator of the XGBoost model trained with the entire data set. In the DeLong test evaluation results, a p-value less than 0.05 is considered to have a large difference, and a p-value greater than 0.05 is considered to have a small difference. Therefore, the p-value greater than 0.05 and the minimum number of detections are selected as the minimum detection number threshold.

[0054] (b9) For each nonlinear regression machine learning model, the minimum number of detections corresponding to the minimum difference is used as the reference detection threshold.

[0055] (b10) According to the reference detection number threshold corresponding to each nonlinear regression machine learning model, the minimum detection number threshold is obtained.

[0056] For example, each model in XGBoost, CatBoost, LightGBM, artificial neural network (ANN) and support vector machine (SVM) corresponds to a reference detection number threshold. The minimum detection number threshold is obtained by comprehensively considering the reference detection number thresholds of each model.

[0057] (b11) Obtaining ovarian test data from the patient that meets a minimum threshold number of tests, and obtaining the ovarian response test data to be predicted. Once the minimum threshold number of tests is determined, collecting ovarian test data that is greater than or equal to the minimum threshold number of tests ensures the accuracy of the predicted data. Furthermore, this allows for early prediction and the avoidance of complications.

[0058] In another exemplary embodiment of the present application, for data preprocessing and statistical transformation: One-Hot Encoding is performed on the text type attributes in the data, the result attributes (ovarian response results) are discretized, missing values ​​are filled, duplicate values ​​and outliers are removed, and then the mean, standard deviation, maximum value, minimum value and variation of each indicator in the high-dimensional continuous data set monitored for multiple consecutive days are statistically transformed. In the original data, each patient has various hormone recording data of different times. The variation calculated in the statistical transformation here is the last test data of the current indicator minus the first test data. The original high-dimensional continuous data is converted into a set of statistically significant features. Finally, the preprocessing.scale method in sklearn is used for standardization processing, and the statistical data of each indicator is processed into data that obeys a normal distribution with a mean of 0 and a variance of 1. Standardization processes data after data cleaning and statistical transformation. The data processing process for tabular data is: data cleaning of the original data set (including One-Hot Encoding of text type attributes, discretization of result attributes, filling default values, and removing duplicate values ​​and outliers) and statistical transformation (including statistical transformation of the mean, standard deviation, maximum, minimum, and variation of each indicator). Only after these two major steps (data cleaning and statistical transformation) is the data clean and complete, and standardization can be performed on the cleaned and transformed complete data. Standardization is performed after data cleaning and any form of transformation.

[0059] Therefore, in step 102, data preprocessing and statistical transformation are performed on the ovarian response detection data to be predicted to obtain the ovarian detection data after statistical transformation, which specifically includes:

[0060] (c1) performing missing value filling and outlier processing on the ovarian response detection data to be predicted to obtain preprocessed ovarian detection data.

[0061] (c2) For each indicator in the preprocessed ovarian detection data, the mean, standard deviation, maximum value, minimum value and variation of the indicator are calculated to obtain statistical transformation data of each indicator.

[0062] (c3) The statistically transformed data of each indicator and the pre-processed ovarian detection data are combined to form statistically transformed ovarian detection data.

[0063] In another exemplary embodiment of the present application, in addition, existing prediction models based on machine learning mostly focus on single measurements or periodic change data, and the training feature indicators they rely on are also different. Therefore, the present application selects key indicator data by using recursive feature elimination (RFE) to perform feature screening and select N most predictive indicators from the influencing indicators. Therefore, in step 103, the data corresponding to the key indicators in the ovarian response test data to be predicted and the ovarian test data after statistical transformation are determined, specifically including:

[0064] (d1) Based on the ovarian detection data after the statistical transformation, a recursive feature elimination method is used to screen indicators to select key indicators for ovarian responsiveness prediction.

[0065] (d2) Determining data corresponding to key indicators in the ovarian detection data after the statistical transformation.

[0066] In another exemplary embodiment of the present application, the recursive feature elimination method is as follows: let the original indexes x1, x2, ..., x p , initialize the indicator set S={x1,x2,…,x p The base model is set as support vector machine (SVM) and scene basis function (RBF) is selected as kernel function.

[0067] For the current indicator set S, use the SVM training model with RBF kernel function and calculate the importance of each feature.

[0068] For a given training data (X, y), the optimization goal of the SVM with RBF kernel function is:

[0069]

[0070] Among them, ξ i is the slack variable, C is the penalty parameter, and the RBF kernel function is:

[0071] K(x i ,x j )=exp(-γ||x i -x j || 2 )

[0072] Among them, γ is the kernel function, which controls the width of the Gaussian function.

[0073] After training, the decision function of the SVM model is:

[0074]

[0075] Among them, α i is the Lagrange multiplier, y iis the data point x i labels. b is the bias term in the SVM model. In the SVM decision boundary equation, b is used to adjust the position of the decision boundary so that it better separates data points of different categories. N represents the number of support vectors. Support vectors are data points in the training data that are crucial for constructing the decision boundary. They are the points closest to the decision boundary and determine the width of the maximum margin. Both b and n are automatically learned and determined during the SVM model training process.

[0076] Use the trained SVM model to calculate the weight of each indicator. The importance of the indicator is measured by the impact of the indicator on the decision function. Calculate the importance of each indicator:

[0077]

[0078] For a dataset of n samples, each indicator x i The mean importance value over the entire dataset is expressed as:

[0079]

[0080] Finally, select the indicator x with the lowest indicator importance min .

[0081] Remove x from the index set S min Update the indicator set S and repeat the indicator importance calculation and screening until the predetermined number of indicators is reached.

[0082] Based on the above, in step (d1), the recursive feature elimination method is used to screen the key indicators for ovarian responsiveness prediction based on the ovarian test data after the statistical transformation, specifically including:

[0083] (e1) Inputting data corresponding to each indicator in the ovarian detection data after the statistical transformation into the SVM model to obtain a weight value for each indicator.

[0084] (e2) Deleting the indicator with the lowest weight from the ovarian detection data after the statistical transformation.

[0085] (e3) Inputting the data corresponding to each remaining indicator in the ovarian detection data after the statistical transformation into the SVM model to obtain the weight value of each remaining indicator.

[0086] (e4) Return to step “delete the indicator with the lowest weight from the ovarian detection data after the statistical transformation” until the number of remaining indicators reaches the preset number of indicators.

[0087] (e5) The remaining preset number of indicators are considered as key indicators for predicting ovarian responsiveness.

[0088] In another exemplary embodiment of the present application, before predicting ovarian responsiveness, it is necessary to use the historical ovarian test data to determine the optimal prediction model, and then input the data corresponding to the key indicators in the statistically transformed ovarian test data into the optimal ovarian responsiveness prediction model. Using the N key indicators screened out, five different nonlinear regression machine learning models are constructed, including XGBoost, CatBoost, LightGBM, artificial neural network (ANN) and support vector machine (SVM), and the optimal prediction model is selected from them, such as Figure 3 Therefore, in step 104, the data corresponding to the key indicators in the ovarian detection data after the statistical transformation are input into the ovarian response prediction model to obtain the ovarian response prediction result, which specifically includes:

[0089] (f1) Obtain the patient's second ovary historical detection data.

[0090] (f2) performing data preprocessing and statistical transformation on the second ovarian historical detection data to obtain the ovarian historical detection data after the second statistical transformation.

[0091] (f3) constructing the second ovarian detection data and the second statistically transformed ovarian historical detection data into a training and validation data set.

[0092] (f4) training multiple nonlinear regression machine learning models using the key features in the training and validation dataset as input and the corresponding ovarian response results as labels, and determining the nonlinear regression machine learning model with the best prediction effect as the ovarian response prediction model. The multiple nonlinear regression machine learning models include XGBoost, CatBoost, LightGBM, artificial neural network (ANN), and support vector machine (SVM).

[0093] (f5) Inputting data corresponding to key indicators in the ovarian detection data after the statistical transformation into an ovarian responsiveness prediction model to obtain an ovarian responsiveness prediction result.

[0094] In another exemplary embodiment of the present application, in order to make the model as universal as possible, the K-fold cross validation technique is used to repeatedly split the data set into a training set and a validation set. As an example, a 5-fold cross validation can be used, and the complete data set is evenly divided into 5 subsets of equal size, each of which can be regarded as an independent validation set. Five independent training and validation cycles are performed. In each cycle: one of the subsets is randomly selected as the validation set. The remaining 4 subsets are combined as the training set. Therefore, in step (f4), the key features in the training validation data set are used as input, and the corresponding ovarian responsiveness results are used as labels to train multiple nonlinear regression machine learning models, and the nonlinear regression machine learning model with the best prediction effect is determined as the ovarian responsiveness prediction model, specifically including:

[0095] (g1) For each nonlinear regression machine learning model, the training and validation sample set is divided into K parts of data.

[0096] (g2) Each of the K pieces of data is used as a validation sample set, and the remaining K-1 pieces are used as training sample sets.

[0097] (g3) For each training sample set, a nonlinear regression machine learning model is trained using the key features in the training sample set as input and the corresponding ovarian response results as labels to obtain a trained nonlinear regression machine learning model.

[0098] (g4) Using K validation sample sets to respectively validate the corresponding trained nonlinear regression machine learning model, and obtaining the model evaluation index of the K-times validated trained nonlinear regression machine learning model.

[0099] (g5) The model evaluation indicators of the K-times validation are averaged to obtain the final model evaluation indicators of the trained nonlinear regression machine learning model.

[0100] (g6) Compare the final model evaluation indicators of each trained nonlinear regression machine learning model to determine the nonlinear regression machine learning model with the best prediction effect.

[0101] (g7) The nonlinear regression machine learning model with the best prediction effect is used as the ovarian response prediction model. After comparison, it was found that the XGBoost model had the best prediction effect among XGBoost, CatBoost, LightGBM, artificial neural network (ANN), and support vector machine (SVM). The test data is substituted into the model, and the model will predict the test data based on the learned rules. The predicted value is converted into a classification label. If the model output is close to 0, it is judged as a low response; if it is close to 1, it is judged as a high response.

[0102] The optimization goal of the model is to minimize the objective function:

[0103]

[0104] Where l is the logarithmic loss function, which is used to measure the error between the predicted value and the true value. Ω is the regularization term, which is usually in the form of:

[0105]

[0106] Among them, T is the number of leaf nodes in the tree, w i is the weight of the leaf node. Regularization is used to control the complexity of the model and prevent overfitting.

[0107] When training the model, the number of leaf nodes T and regularization parameters γ and λ of each tree t are initialized. For each node, the structural score is calculated:

[0108]

[0109] Among them, G j and H j The sum of the gradient and second-order derivative of each node respectively. The structural score is used to measure the splitting quality of the current node. The higher the score, the better the splitting effect. Then the optimal weight of each node is calculated:

[0110]

[0111] The optimal weight is used to optimize the predicted value of the leaf node so that the objective function is minimized, thereby improving the prediction accuracy of the model. Finally, the node split gain is calculated:

[0112]

[0113] Among them, G L , G R and H L 、H R The node splitting gain is used to determine whether to split the current node. The larger the gain, the better the effect of the split, thus selecting the optimal split point.

[0114] Finally, a model that conforms to the above function is obtained, and the weight values ​​of each key indicator value are obtained to determine the importance of each key indicator.

[0115] In this application, the defects of poor generalization ability and low classification accuracy in the previous modeling of ovarian responsiveness for processing high-dimensional data for multiple consecutive days are overcome. By statistically transforming the high-dimensional continuous data, a statistically significant feature data set is obtained. At the same time, this application analyzes the model effects of different indicator groups selected and finds that different models rely on different numbers of key indicators and the importance ranking of key indicators when predicting the degree of ovarian responsiveness. After in-depth comparison and analysis, a set of optimal key indicators is determined from the key indicators. These optimal key indicators show high predictive efficiency in multiple models, providing important guidance and benchmarks for subsequent clinical treatment, hormone and ultrasound indicator testing. In addition, this application also uses models trained with different data subsets to determine the threshold of the minimum number of tests required for ovarian responsiveness prediction through the DeLong test. This discovery helps to optimize the clinical monitoring process, reduce unnecessary examinations, and ensure the accuracy of predictions, further promoting the development of precision medicine. In the solution of the present application, by performing statistical transformation on high-dimensional data to ensure the prediction effect, the minimum detection number threshold and the combination of key indicators for efficient prediction are determined to solve the limitations of the existing model's poor performance when facing high-dimensional continuous data and the subjective number of tests required for ovarian response prediction.

[0116] The following is an application example of this application: 902 groups of data from 205 patients are used as training sets. Python coding is used to process high-dimensional continuous data sets into data sets with statistical characteristics according to statistical transformation rules. The indicators are arranged according to the importance calculated by the recursive feature screening method, and key indicators are screened out. Then, different models are trained based on the key indicators, and the effects of different models are compared. Figure 4 As shown in the figure, the top 25 most important indicators were selected, and the training set was reorganized with these 25 indicators. XGBoost, CatBoost, LightGBM, ANN and SVM were used to model them respectively, and they were divided into two results: high response and low response. 307 groups of data from 71 patients were used as test sets, and they were divided into two results using the same method. The accuracy of the model was verified by the five-fold cross-validation method. The original data set was divided into the first four times, the first three times, the first two times and the single indicator data set. The same data preprocessing method and indicator statistical transformation method were used to obtain the benchmark data set of the subset, and the training models corresponding to different subsets were trained. The training model corresponding to the complete data set and the training model corresponding to each subset were optimized, and the area under the ROC curve of the model, AUC, was calculated, as shown in the figure. Figure 5 As shown, the Delong test was used to calculate the statistical difference P-value between each subset model, as shown in Figure 6 As shown. Figure 6The results show that the statistical differences calculated for the five models are all greater than a p-value of 0.05, which means that statistically, there is insufficient evidence to reject the null hypothesis. Based on these p-values, it can be concluded that the performance of these models on the full dataset and the first two datasets is similar, with no significant statistical differences. This shows that the model can maintain good performance even on smaller datasets. The number of tests can be reduced from 5 or more to 2, and the model can still produce the same prediction results. The minimum number of tests that guarantees good prediction results has been found. In practical applications, reducing the number of tests can reduce costs, improve efficiency, and reduce resource requirements.

[0117] The present application also provides an application scenario, which applies the above-mentioned ovarian responsiveness prediction method based on machine learning technology. Specifically: the ovarian responsiveness prediction method based on machine learning technology provided in this embodiment can be applied in a prediction scenario for female ovarian responsiveness. The scenario includes a data collection link, an ovarian responsiveness prediction link and a prediction result display link; the data collection link is used to collect the patient's ovarian test data for multiple consecutive times; the ovarian responsiveness prediction link is used to predict ovarian responsiveness based on the collected ovarian test data; the prediction result display link is used to display the ovarian responsiveness prediction results. The ovarian responsiveness prediction method based on machine learning technology provided in this embodiment belongs to the ovarian responsiveness prediction link.

[0118] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store ovarian responsiveness prediction results. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an ovarian responsiveness prediction method based on machine learning technology is implemented.

[0119] Those skilled in the art will understand that Figure 7The structure shown in the figure is merely a block diagram of a portion of the structure related to the present application solution and does not constitute a limitation on the computer device to which the present application solution is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned method embodiments are implemented.

[0120] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0121] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0122] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0123] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0124] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0125] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for predicting ovarian responsiveness based on machine learning technology, characterized in that: The ovarian responsiveness prediction method based on machine learning technology includes: Obtaining the patient's ovarian response test data to be predicted; the ovarian response test data to be predicted includes ovarian data of multiple consecutive tests; the number of tests in the ovarian response test data to be predicted is greater than or equal to a minimum test number threshold; performing data preprocessing and statistical transformation on the ovarian response detection data to be predicted to obtain statistically transformed ovarian detection data; the statistically transformed ovarian detection data includes the preprocessed ovarian detection data and the statistically transformed data of the ovarian response detection data to be predicted; Determining data corresponding to key indicators in the ovarian detection data after the statistical transformation; The data corresponding to the key indicators in the ovarian detection data after the statistical transformation are input into the ovarian response prediction model to obtain the ovarian response prediction result; the ovarian response prediction model is a nonlinear regression machine learning model.

2. The ovarian responsiveness prediction method based on machine learning technology according to claim 1, characterized in that: Obtain the patient's ovarian response test data to be predicted, including: Acquiring first historical ovarian test data of the patient; the first historical ovarian test data includes multiple consecutive ovarian test data; performing data preprocessing and statistical transformation on the first ovarian historical detection data to obtain the ovarian historical detection data after first statistical transformation; According to the number of tests, the ovarian historical test data after the first statistical transformation is divided into multiple data subsets; one data subset corresponds to one number of tests; one data subset includes the previous m ovarian historical test data; m is the number of tests corresponding to the data subset; the data subset corresponding to the maximum number of tests is recorded as the full data set; Use each data subset and the entire data set to train a nonlinear regression machine learning model; Determine the first AUC evaluation metric of the trained nonlinear regression machine learning model corresponding to each data subset; Determine a second AUC evaluation metric of the trained nonlinear regression machine learning model corresponding to the entire data set; The DeLong test method was used to evaluate the difference between each first AUC evaluation index and the second AUC evaluation index; The minimum number of detections corresponding to the minimum difference is used as the minimum detection number threshold; Obtain the ovarian test data of the patient that meets at least a minimum test number threshold, and obtain the ovarian response test data to be predicted.

3. The ovarian responsiveness prediction method based on machine learning technology according to claim 1, characterized in that: Obtain the patient's ovarian response test data to be predicted, including: Acquiring first historical ovarian test data of the patient; the first historical ovarian test data includes multiple consecutive ovarian test data; performing data preprocessing and statistical transformation on the first ovarian historical detection data to obtain the ovarian historical detection data after first statistical transformation; According to the number of tests, the ovarian historical test data after the first statistical transformation is divided into multiple data subsets; one data subset corresponds to one number of tests; one data subset includes the previous m ovarian historical test data; m is the number of tests corresponding to the data subset; the data subset corresponding to the maximum number of tests is recorded as the full data set; For each data subset, multiple nonlinear regression machine learning models are trained using the data subsets; Use the entire data set to train multiple nonlinear regression machine learning models; Determine the first AUC evaluation metric for each trained nonlinear regression machine learning model corresponding to each data subset; Determine a second AUC evaluation metric of the trained nonlinear regression machine learning model corresponding to the entire data set; The DeLong test method was used to evaluate the difference between each first AUC evaluation index and the second AUC evaluation index corresponding to the same nonlinear regression machine learning model; For each nonlinear regression machine learning model, the minimum number of detections corresponding to the minimum difference is used as the reference detection number threshold; According to the reference detection number threshold corresponding to each nonlinear regression machine learning model, the minimum detection number threshold is obtained; Obtain the ovarian test data of the patient that meets at least a minimum test number threshold, and obtain the ovarian response test data to be predicted.

4. The ovarian responsiveness prediction method based on machine learning technology according to claim 1, characterized in that: Performing data preprocessing and statistical transformation on the ovarian response detection data to be predicted to obtain statistically transformed ovarian detection data specifically includes: Filling missing values ​​and processing outliers on the ovarian response test data to be predicted to obtain preprocessed ovarian test data; For each indicator in the pre-processed ovarian detection data, the mean, standard deviation, maximum value, minimum value and variation of the indicator are calculated to obtain statistical transformation data of each indicator; The statistically transformed data of each indicator and the pre-processed ovarian detection data constitute statistically transformed ovarian detection data.

5. The ovarian responsiveness prediction method based on machine learning technology according to claim 1, characterized in that: Determining data corresponding to key indicators in the ovarian detection data after the statistical transformation specifically includes: Based on the ovarian test data after the statistical transformation, a recursive feature elimination method is used to screen indicators to screen out key indicators for ovarian responsiveness prediction; Determine data corresponding to key indicators in the ovarian detection data after the statistical transformation.

6. The ovarian responsiveness prediction method based on machine learning technology according to claim 5, characterized in that: Based on the ovarian test data after the statistical transformation, the recursive feature elimination method is used to screen the key indicators for predicting ovarian responsiveness, including: Inputting the data corresponding to each indicator in the ovarian detection data after the statistical transformation into the SVM model to obtain the weight value of each indicator; Deleting the indicator with the lowest weight from the ovarian detection data after the statistical transformation; Inputting the data corresponding to each remaining indicator in the ovarian detection data after the statistical transformation into the SVM model to obtain a weight value for each remaining indicator; Return to step "deleting the indicator with the lowest weight from the ovarian detection data after the statistical transformation" until the number of remaining indicators reaches the preset number of indicators; The final remaining indicators are considered as key indicators for predicting ovarian responsiveness.

7. The ovarian responsiveness prediction method based on machine learning technology according to claim 1, characterized in that: Inputting the data corresponding to the key indicators in the statistically transformed ovarian test data into the ovarian response prediction model to obtain the ovarian response prediction results, specifically including: Obtain the patient's historical second ovary testing data; performing data preprocessing and statistical transformation on the second ovarian historical detection data to obtain the ovarian historical detection data after second statistical transformation; constructing the second ovarian detection data and the second statistically transformed ovarian historical detection data into a training and verification data set; Using the key features in the training and validation dataset as input and the corresponding ovarian response results as labels, multiple nonlinear regression machine learning models are trained, and the nonlinear regression machine learning model with the best prediction effect is determined as the ovarian response prediction model; The data corresponding to the key indicators in the ovarian detection data after the statistical transformation are input into the ovarian response prediction model to obtain the ovarian response prediction result.

8. The ovarian responsiveness prediction method based on machine learning technology according to claim 7, characterized in that: Using the key features in the training and validation dataset as input and the corresponding ovarian response results as labels, multiple nonlinear regression machine learning models are trained, and the nonlinear regression machine learning model with the best prediction effect is determined as the ovarian response prediction model, specifically including: For each nonlinear regression machine learning model, the training and validation sample sets are divided into K parts of data; Each of the K pieces of data is used as a validation sample set, and the remaining K-1 pieces are used as training sample sets; For each training sample set, a nonlinear regression machine learning model is trained using the key features in the training sample set as input and the corresponding ovarian response results as labels to obtain a trained nonlinear regression machine learning model; Use K validation sample sets to respectively validate the corresponding trained nonlinear regression machine learning model, and obtain the model evaluation index of the K-times validation of the trained nonlinear regression machine learning model; The model evaluation indicators of the K-times validation are averaged to obtain the final model evaluation indicators of the trained nonlinear regression machine learning model; Compare the final model evaluation indicators of each trained nonlinear regression machine learning model to determine the nonlinear regression machine learning model with the best prediction effect; The nonlinear regression machine learning model with the best prediction effect is used as the ovarian responsiveness prediction model.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the ovarian responsiveness prediction method based on machine learning technology according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting ovarian responsiveness based on machine learning technology according to any one of claims 1 to 8 is implemented.