Method and device for establishing a prognostic outcome classification prediction model for cesarean scar pregnancy

By classifying data of cesarean scar pregnancy patients and optimizing machine learning models, an accurate classification model of prognostic outcomes is solved, and the problem of difficult-to-predict bleeding risks caused by differences in treatment methods in the existing technology is improved, and the accuracy and safety of treatment are improved.

CN119560166BActive Publication Date: 2025-07-25PEKING UNION MEDICAL COLLEGE HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411602159.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-07-25
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The prior art lacks fine and accurate classification standards in the classification of prognostic outcomes for cesarean scar pregnancy, which makes it difficult to predict and control the bleeding risk caused by differences in treatment methods.

Method used

By obtaining the patient database, classifying patient data according to treatment methods and prognosis, and using ultrasound and clinical characteristics combined with machine learning models, a GDBT model is established for optimization, identifying patients with high prognosis risks to select appropriate treatment methods and avoiding overtreatment.

Benefits of technology

A more accurate classification of prognosis outcomes for cesarean scar pregnancy is achieved, helping clinicians choose appropriate treatment methods, identify high-risk patients and avoid unnecessary losses, and improve the accuracy and safety of treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119560166B_ABST
    Figure CN119560166B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy. The method for establishing the classification prediction model for the prognosis outcome of cesarean scar pregnancy includes: obtaining a patient database, where the patient database includes at least one patient data; classifying the patient data in the patient database according to the treatment method and prognosis, so as to obtain the first type of patient data and the second type of patient data; establishing a model for the first type of patient data based on the first type of patient data; establishing a model for the second type of patient data based on the first type of patient data, the second type of patient data, and the third type of patient data. The present application utilizes ultrasound and clinical features, combined with a machine learning model, to establish a more accurate classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of model establishment, and particularly relates to a method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy and a device for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy. Background Art

[0002] Cesarean scar pregnancy (CSP) is a special type of pregnancy in which the gestational sac implants in the cesarean scar at the lower part of the anterior uterine wall. In recent years, with the increase in the cesarean section rate and the progress of imaging diagnosis technology, the incidence of cesarean scar pregnancy (CSP) has been increasing continuously. The pregnancy outcomes and prognoses of CSP cases vary, and some cases are accompanied by serious risks, such as life-threatening massive hemorrhage, uterine rupture, and even death. Transvaginal ultrasound is the preferred evaluation method, and CSP can be divided into different types according to US results. For example, by observing the relationship between the gestational sac and the uterine cavity line and the serosa line, the Delphi consensus divides it into 3 types. According to the remaining myometrial thickness and whether the gestational sac protrudes into the serosa layer, the Chinese consensus divides it into 3 types. At present, some studies have tried to establish a CSP clinical classification model to identify patients with a large amount of intraoperative bleeding, so as to provide reference for treatment decisions. For example, based on the data of patients with cesarean scar pregnancy in Qilu Hospital of Shandong Province, univariate analysis and multivariate logistic regression analysis were used to explore the independent risk factors for bleeding (≥300 ml) during cesarean scar pregnancy surgery. The receiver operating characteristic curve method was used to determine the optimal threshold of the identified risk factors, and a clinical classification of five cesarean scar pregnancies was established using two variables, the anterior scar myometrial thickness and the average diameter of the gestational sac.

[0003] However, due to the continuous changes in the understanding and treatment of CSP, the treatment of CSP in different periods, regions, and different classifications varies, including surgical treatments such as uterine curettage, lesion resection, and even hysterectomy, simple drug treatment, or drug combined with surgical treatment. In order to better avoid the bleeding risk, auxiliary means such as uterine artery embolization and balloon compression hemostasis are also included. Therefore, it is not fine and accurate enough to use only the intraoperative blood loss as the classification standard without considering the differences in treatment methods.

[0004] Therefore, it is hoped that there is a technical solution to solve or at least alleviate the above deficiencies of the prior art. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy to at least solve one of the above technical problems.

[0006] The present invention provides the following solutions:

[0007] According to one aspect of the present invention, there is provided a method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy, and the method for establishing the classification prediction model for the prognosis outcome of cesarean scar pregnancy includes:

[0008] Step 1: Obtain a patient database, where the patient database includes at least one patient's data;

[0009] Step 2: Classify the patient data in the patient database according to the treatment method and prognosis, so as to obtain the first type of patient data, the second type of patient data, and the third type of patient data;

[0010] Step 3: Establish a model for the first type of patient data and a model for the second type of patient data respectively according to the first type of patient data, the second type of patient data, and the third type of patient data.

[0011] Optionally, the Step 3: Establish a model for the first type of patient data respectively according to the first type of patient data, the second type of patient data, and the third type of patient data includes:

[0012] Step 31: Obtain the first-round multimodal data corresponding to the first type of patient data as positive samples, and obtain the first-round multimodal data of the second type of patient data and the third type of patient data as negative samples;

[0013] Step 32: Obtain a trained large model;

[0014] Step 33: Input the first-round multimodal data into the trained large model to obtain the first-round feature selection information;

[0015] Step 34: Establish a first GDBT model according to the first-round feature selection information;

[0016] Step 35: Optimize the first GDBT model to obtain a model for the first type of patient data.

[0017] Optionally, the Step 35: Optimize the first GDBT model to obtain a model for the first type of patient data includes:

[0018] Step 351: Obtain feature selection information for optimization based on the first GDBT model;

[0019] Step 352: Obtain an optimized GDBT model according to the feature selection information for optimization;

[0020] Step 353: Evaluate the model performance of the first GDBT model and the optimized GDBT model respectively, and retain the one with better performance as the model for the first type of patient data.

[0021] Optionally, step 35: optimizing the first GDBT model to obtain a model for the first type of patient data further includes:

[0022] Step 354: Obtain optimization feature selection information according to the optimization GDBT model obtained in step 352;

[0023] Step 355: Generate a new optimization GDBT model according to the optimization feature selection information;

[0024] Step 356: Evaluate the performance of the new optimization GDBT model and compare the performance with the previously retained model, so as to retain the model with better performance as the model for the first type of patient data.

[0025] Optionally, step 35: optimizing the first GDBT model to obtain a model for the first type of patient data further includes:

[0026] Step 357: Repeat generating a new optimization GDBT model and comparing the performance with the previously retained model until the model prediction performance converges or decreases to a preset value.

[0027] Optionally, the first-round multimodal data includes patient disease-related clinical knowledge and diagnosis and treatment data, and a set of clinical features available for the model.

[0028] Optionally, step 351: obtaining optimization feature selection information based on the first GDBT model includes:

[0029] Step 3511: Obtain the first-round multimodal data corresponding to the first type of patient data, the model data obtained through the established first GDBT model, the feature set selected by the current-round machine learning model, the performance information obtained by the current-round machine learning model during training and testing, and the model feature selection and performance iteration history information;

[0030] Step 3512: Obtain a trained large model;

[0031] Step 3513: Input the first-round multimodal data, the model data obtained through the established first GDBT model, the feature set selected by the current-round machine learning model, the performance information obtained by the current-round machine learning model during training and testing, and the model feature selection and performance iteration history information into the trained large model, so as to obtain optimization feature selection information.

[0032] Optionally, establishing a model for the second type of patient data according to the first type of patient data, the second type of patient data, and the third type of patient data includes:

[0033] Obtain the set of univariate significantly correlated variables for the first type of patient data and the set of univariate significantly correlated variables for the third type of patient data;

[0034] Plot the ROC curve of the set of univariate significantly correlated variables for the first type of patient data against the first type of patient data and the ROC curve of the set of univariate significantly correlated variables for the third type of patient data against the third type of patient data;

[0035] Find the threshold with the largest Youden index on each ROC curve, and convert the variable into a 0-1 categorical variable based on this threshold;

[0036] For each variable converted into a 0-1 categorical variable, assign four weight scores of 0, 1, 2, and 3, and find the variable weight combination with the largest AUC value based on a computer algorithm.

[0037] This application also provides a device for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy. The device for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy includes:

[0038] A patient database acquisition module, which is used to acquire a patient database, and the patient database includes at least one patient data;

[0039] A classification module, which is used to classify the patient data in the patient database according to the treatment method and prognosis, so as to obtain the first type of patient data, the second type of patient data, and the third type of patient data;

[0040] A first model establishment module, which is used to establish a model for the first type of patient data according to the first type of patient data, the second type of patient data, and the third type of patient data;

[0041] A second model establishment module, which is used to establish a model for the second type of patient data according to the first type of patient data, the second type of patient data, and the third type of patient data.

[0042] The method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy in this application uses ultrasound and clinical features, combines with a machine learning model, and establishes a more accurate classification model to help clinicians select appropriate treatment methods, identify patients with higher prognosis risks so as to give sufficient attention and treatment, and identify patients with lower prognosis risks to avoid unnecessary losses caused by over-treatment. Description of the Drawings

[0043] Figure 1 is a schematic flowchart of the method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy in an embodiment of this application.

[0044] Figure 2 It is a schematic structural diagram of an electronic device in an embodiment of the present application.

[0045] Figure 3 It is a schematic diagram of a large model selecting features in an embodiment of the present application.

[0046] Figure 4 It is a schematic diagram of the feature distributions of the first type of patient data, the second type of patient data, and the third type of patient data in an embodiment of the present application.

[0047] Figure 5 It is a schematic diagram of the importance scores of ultrasound and clinical features in an embodiment of the present application.

[0048] Figure 6 It is a schematic diagram of an ROC curve in an embodiment of the present application.

[0049] Figure 7 It is a schematic diagram of the ROC curve of the machine learning model for the first type of patient data in an embodiment of the present application. Detailed implementation manners

[0050] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] As Figure 1 shown, the method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy includes:

[0052] Step 1: Obtain a patient database, where the patient database includes at least one patient data;

[0053] Step 2: Classify the patient data in the patient database according to the treatment method and prognosis, so as to obtain the first type of patient data, the second type of patient data, and the third type of patient data;

[0054] Step 3: Establish a model for the first type of patient data and a model for the second type of patient data respectively according to the first type of patient data, the second type of patient data, and the third type of patient data.

[0055] The method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy in the present application utilizes ultrasound and clinical features, combines a machine learning model, and establishes a more accurate classification model to help clinicians select appropriate treatment means, identify patients with higher prognosis risks so as to give sufficient attention and treatment, and identify patients with lower prognosis risks to avoid unnecessary losses caused by over-treatment.

[0056] In this embodiment, step 3: Establishing a model for the first type of patient data based on the first type of patient data includes:

[0057] Step 31: Obtain the first-round multimodal data corresponding to the first type of patient data;

[0058] Step 32: Obtain a trained large model;

[0059] Step 33: Input the first-round multimodal data into the trained large model to obtain the first-round feature selection information;

[0060] Step 34: Establish a first GDBT model based on the first-round feature selection information;

[0061] Step 35: Optimize the first GDBT model to obtain a model for the first type of patient data.

[0062] In this embodiment, step 35: Optimizing the first GDBT model to obtain a model for the first type of patient data includes:

[0063] Step 351: Obtain feature selection information for optimization based on the first GDBT model;

[0064] Step 352: Obtain an optimized GDBT model according to the feature selection information for optimization;

[0065] Step 353: Evaluate the model performance of the first GDBT model and the optimized GDBT model respectively, and retain the one with better performance as the model for the first type of patient data.

[0066] In this embodiment, step 35: Optimizing the first GDBT model to obtain a model for the first type of patient data further includes:

[0067] Step 354: Obtain feature selection information for optimization according to the optimized GDBT model obtained in step 352;

[0068] Step 355: Generate a new optimized GDBT model according to the feature selection information for optimization;

[0069] Step 356: Evaluate the model performance of the new optimized GDBT model and compare it with the previously retained model, and retain the one with better performance as the model for the first type of patient data.

[0070] In this embodiment, step 35: Optimizing the first GDBT model to obtain a model for the first type of patient data further includes:

[0071] Step 357: Repeatedly generate a new optimized GDBT model and compare its performance with the previously retained model until the model prediction performance converges or drops to a preset value.

[0072] In this embodiment, the first-round multimodal data includes patient disease-related clinical knowledge and diagnosis and treatment data, and a set of clinical features available for the model.

[0073] For example, the GDBT model of the present application needs to obtain the optimal GDBT model (including the model parameters of the GDBT model and the selected features) through iterative training. For example, in the large model obtained in step 32, at this time, since it is the first-round training, there are only patient disease-related clinical knowledge and diagnosis and treatment data, and a set of clinical features available for the model. The features output by this large model can be considered as the first-round features. However, the first-round features may not be the best. Because after having the first-round features, a first GDBT model can be established and the performance of the first GDBT model can be tested. After the test, it may not meet the requirements or may not be the best. Therefore, a new round of feature selection will be carried out again, that is, the first-round multimodal data corresponding to the first type of patient data, the model data obtained through the established first GDBT model, the set of features selected by this round of machine learning model, the performance information obtained by this round of machine learning model during the training and testing set, and the model feature selection and performance iteration history information are re-input into the trained large model, so as to obtain the optimized feature selection information for establishing the GDBT model in the second round.

[0074] Then, a new GDBT model (temporarily called the second GDBT model) is established according to the optimized feature selection information.

[0075] It can be understood that the second GDBT model may not be the best model either. At this time, the model data obtained through the established second GDBT model, the set of features selected by this round of machine learning model, the performance information obtained by this round of machine learning model during the training and testing set, and the model feature selection and performance iteration history information are re-input into the trained large model, so as to obtain the optimized feature selection information for establishing the GDBT model in the third round.

[0076] A new GDBT model (temporarily called the third GDBT model) is established according to the optimized feature selection information for establishing the GDBT model in the third round.

[0077] It can be understood that the third GDBT model may not be the best model either. In this case, the model data obtained through the established third GDBT model, the feature set selected by this round of machine learning model, the performance information obtained by this round of machine learning model during training and testing, and the model feature selection and performance iteration history information are re-input into the trained large model, so as to obtain the optimized feature selection information for establishing the GDBT model in the fourth round.

[0078] Repeat the above iterative rotation until the preset requirements are met.

[0079] In this embodiment, the step 351: obtaining the optimized feature selection information based on the first GDBT model includes:

[0080] Step 3511: Obtain the first-round multimodal data corresponding to the first type of patient data, the model data obtained through the established first GDBT model, the feature set selected by this round of machine learning model, the performance information obtained by this round of machine learning model during training and testing, and the model feature selection and performance iteration history information;

[0081] Step 3512: Obtain the trained large model;

[0082] Step 3513: Input the first-round multimodal data, the model data obtained through the established first GDBT model, the feature set selected by this round of machine learning model, the performance information obtained by this round of machine learning model during training and testing, and the model feature selection and performance iteration history information into the trained large model, so as to obtain the optimized feature selection information.

[0083] In this embodiment, establishing the model for the second type of patient data based on the first type of patient data, the second type of patient data, and the third type of patient data includes:

[0084] Obtain the set of single-factor significantly correlated variables of the first type of patient data and the set of single-factor significantly correlated variables of the second type of patient data; In this embodiment, the initial variables are derived by doctors from the electronic medical records or test records of patients. For example, the indicators recorded in the electronic medical record information system. After manually extracting and summarizing the initial continuous variables, subsequent processing will be carried out.

[0085] In this embodiment, single-factor significantly correlated variables are selected from each initial variable through single-factor analysis.

[0086] The single-factor significant variables are evaluated through univariate logistic regression. If the P-value < 0.05, the univariate variable is considered significantly correlated. This single-factor significant variable is obtained through statistical analysis for Group A and Group C respectively. That is, there will be a set of single-factor significantly correlated variables for Group A and another set for Group C, and the variable sets of the two groups are not necessarily the same.

[0087] Plot the ROC curve of the set of single-factor significantly correlated variables of the first type of patient data and the ROC curve of the set of single-factor significantly correlated variables of the third type of patient data;

[0088] Find the threshold with the largest Youden index on each ROC curve, and convert the variable into a 0-to-1 categorical variable based on this threshold. In this embodiment, different points on the ROC curve correspond to the sensitivity and specificity of the model under different classification thresholds. The Youden index = sensitivity + specificity - 1. We select a point with the largest Youden index, and use the classification threshold corresponding to this point to determine whether the patient should belong to Group A or Group C. For example, the threshold corresponding to the point with the largest Youden index on the ROC curve is a probability value between 0 and 1. For example, the threshold is 0.6. Then, based on the patient characteristics, if the probability predicted by the model is 0.8, which is greater than 0.6, the patient's label is considered 1, and this 1 can represent Group A or Group C. This rule applies to the two models constructed for Group A and Group C. If it is a model constructed for Group A, then 1 represents belonging to Group A, and 0 represents not belonging to Group A (B + C).

[0089] For each variable converted into a 0-to-1 categorical variable, assign four weight scores of 0, 1, 2, and 3, and find the variable weight combination with the largest AUC value based on a computer algorithm. For example, the python hyperopt package can be used to achieve this.

[0090] The following further elaborates on this application by way of example. It can be understood that this example does not constitute any limitation to this application.

[0091] The data used in this application is a dataset constructed based on inpatients with cesarean scar pregnancy (CSP) in a certain hospital in Beijing. The inclusion criteria are: being diagnosed with cesarean scar pregnancy (CSP) in a certain hospital in Beijing from January 2006 to October 2023; the final preoperative ultrasound examination was performed by a senior sonographer with more than 15 years of experience. The exclusion criteria include patients in the second trimester of pregnancy, or patients diagnosed with inevitable abortion, lower uterine segment pregnancy, or cervical pregnancy. Patients with incomplete ultrasound or clinical data are also excluded.

[0092] The diagnosis of cesarean scar pregnancy is based on transvaginal ultrasound examination and is carried out according to the following criteria: (1) The uterine cavity and cervical canal are empty; (2) The gestational sac is implanted at the cesarean scar on the lower anterior wall of the uterus; (3) The myometrium at the cesarean scar on the lower anterior wall of the uterus is thinned or absent; (4) Color Doppler ultrasound examination shows trophoblastic blood flow at the implantation site of the gestational sac. This study was approved by the Ethics Committee of Peking Union Medical College Hospital (Ethical review number: K3071).

[0093] Eligible patients were retrieved according to the diagnosis code of cesarean scar pregnancy (ICD code: O00.807) in the hospital's electronic medical record system. Through the hospital's electronic medical record system, three eligible doctors collected patient identity information and risk factors related to the prognosis of cesarean scar pregnancy, which were reviewed by an obstetric and gynecological ultrasound expert with more than 15 years of experience.

[0094] The variables included the long diameter of the gestational sac or mass (GSCL), the short diameter of the gestational sac or mass (GSCH), the width of the gestational sac or mass (GSCW), the maximum diameter of the gestational sac or mass (GSDmax), the average outer diameter of the gestational sac or mass (GSDavg), the long diameter of the implantation part (IMPL), the short diameter of the implantation part (IMPW), the remaining myometrial thickness (RMT), the adjacent myometrial thickness (AMT), the fetal bud (Fetus), the length of the fetal bud (FL), the fetal heart beat (FHB), the protrusion of the serosa layer of the lower uterine segment (Protrusion), the blood flow distribution around the gestational sac (blood sup), the blood flow grade of the gestational sac (CDFI), the peak blood flow velocity (PSV), the resistance index (RI), the cheese sign (Lacunae), the height of the gestational sac protruding from the UCL (GSUCL), the height of the gestational sac protruding from the SL (GSSL), the anteroposterior diameter of the gestational sac at the diverticulum level (GSSH), Unamed29, GS-UA, the number of days of amenorrhea (D), the number of pregnancies (P), the number of deliveries (G), the number of cesarean sections (CS), the interval time from the last cesarean section (CST), the blood β-HCG value (HCG), abdominal pain, and bleeding.

[0095] All patients were divided into 3 groups according to different treatment methods and prognosis (data of the first type of patients, data of the second type of patients, and data of the third type of patients). Patients with better prognosis were recorded as group A (data of the first type of patients): those who only received simple drug treatment and had a rapid remission of the condition, or those who underwent simple uterine curettage, did not undergo uterine artery embolization, had a smooth operation, and the blood loss was <200 ml.

[0096] Patients with severe illness and poor prognosis are recorded as Group C (data of the third type of patients): Those who need hysterectomy due to illness, or the surgical blood loss during uterine curettage or lesion resection is ≥200 ml, or massive hemorrhage occurs after surgery and surgical intervention is required, or the originally scheduled surgery is not carried out smoothly and the surgical method is changed midway are recorded as Group C [the originally scheduled surgery is not carried out smoothly and the surgical method needs to be changed midway, the intraoperative blood loss is ≥200 ml, hysterectomy is performed, or massive hemorrhage and other severe complications occur after surgery].

[0097] Other patients with medium prognosis are recorded as Group B (data of the second type of patients): including drug + uterine curettage, for example: surgical treatment after treatment with MTX, KCL, etc.; UAE + uterine curettage; or MTX supplementary treatment is carried out after uterine curettage due to residue (the decline of HCG is not satisfactory or ultrasound suggests residue).

[0098] The reason why this application analyzes the data of three types of patients and uses two models is that, based on the variables with significant single factors, through principal component analysis (PCA), the patient populations of Groups A, B, and C are visualized to analyze the distribution rules of clinical variables of the three types of populations. It is found that when the people in Groups A, B, and C are displayed in three-dimensional space, each person is a point on the graph, so that the similarity of the people in Groups A, B, and C in three-dimensional space can be evaluated. The farther the distance, the easier it is to distinguish. Through PCA dimensionality reduction and visualization into three-dimensional space, it can be evaluated and found that Group A is not easy to distinguish and a more complex model needs to be used, so the GBDT model is adopted; the population of Group C is relatively far from Groups A and B and is easy to distinguish, so a linear weighted model is adopted.

[0099] In this embodiment, considering that there is a high correlation between the included variables, for the variables with significant single-factor correlation, the machine learning model (Gradient Boosting Decision Tree, GDBT) will be used to score and rank the importance of the variables. If the variable contributes more to the prediction of patients in Group A or Group C by the machine learning model, the higher the importance score of the variable. Subsequent modeling will further screen based on the importance of the variable to the prediction of the artificial intelligence model, and gradually remove the variables with lower weights to streamline the model.

[0100] In this embodiment, the construction of the prognostic prediction artificial intelligence model using Gradient Boosting Decision Tree (GDBT) includes: randomly extracting 25% of the data from the patient data of groups A, B, and C as the model test set, and the remaining 75% as the model training set. First, all variables are included for model construction, and then variables with lower weights are deleted according to the importance of the variables, and the model is rebuilt. If the model prediction performance improves after deleting this variable, the deletion operation is retained; otherwise, variables are reselected for deletion; the above process is repeated multiple times until the model prediction performance converges or decreases. The model performance will be evaluated by calculating the model sensitivity, specificity, and the area under the ROC curve (AUC).

[0101] In this embodiment, the model for establishing the second patient data is specifically as follows:

[0102] For variables significantly correlated in a single factor, they will be further converted into categorical variables. The conversion method is to plot the ROC curve corresponding to the prognostic prediction of this variable and group A or group C, and find the threshold with the largest Youden index on the ROC curve. Based on this threshold, this variable is converted into a 0 / 1 categorical variable. Then, for each categorical variable, four weight scores of 0, 1, 2, and 3 are assigned, and the optimal variable weight combination with the largest AUC value is found based on the computer algorithm. After obtaining the optimal variable weights, the optimal scoring model division threshold will be selected by analyzing the ROC curve. When the score exceeds this threshold, the patient is considered to belong to group A or group C.

[0103] Continuous variables are expressed as the median and quartiles, and categorical variables are expressed as non-zero counts and proportions. Variables with a P-value less than 0.05 are statistically significant. Statistical analysis is performed using IBM SPSS and Python 3.8.8. The machine learning model is constructed through the scikit-learn toolkit. The hyperparameters max_depth and n_estimators of the GBDT model are set to 2 and 50 respectively. The optimal combination of the overall weights of the variables is selected through the "Hyperopt" Python package. The SHAP values of the variables are calculated through the "shap" Python package.

[0104] Based on 20 univariate significant variables (GSCL, GSCH, GSCW, GSDmax, GSDavg, IMPL, RMT, AMT, FL, GSUCL, GSSH, Protrusion, GSTP, Blood sup, CDFI, PSV, RI, Lacunae, gestational age, abdominal pain) in group A or group C, after removing variables with more than 50% missing values (IMPW, GSSL, IMPA, and GSUA), unsupervised clustering analysis was performed on CSP patients. The PCA method was used to reduce the dimension of each patient's feature vector, and Figure 1The visualization results are shown. It can be observed that the patients in group C (marked with red dots) are mainly distributed in the upper right quadrant of the figure, showing obvious distribution differences compared with group A (blue dots) and group B (green dots). From the proximity of the green and blue dots in the figure, it can be seen that the patients in group B and group A show relatively similar characteristic distributions( Figure 4 ).

[0105] Figure 4 . Visualization of the distribution of CSP patients. Each point in the figure represents a patient, and groups A, B, and C are marked with blue, green, and red respectively. The distance between different points in the figure can represent the similarity degree of the clinical variables of two patients.

[0106] Based on the results of univariate analysis, each feature predicting the prognostic importance of patients in group C and group A was further evaluated through machine learning modeling and interpretive analysis. Before modeling, univariate significant variables were filtered according to the proportion of missing values and clinical feasibility. Subsequently, GBDT models were established for group A and group C respectively, and the SHAP values of each variable were calculated to obtain the importance scores of ultrasound and clinical features( Figure 5 ). The average absolute value of the SHAP value of each feature reflects its importance in predicting the prognosis. The variables in the figure are sorted from high to low according to their weights. It can be seen that for the patients in group C, IMPL, GSUCL, and GSSH made the greatest contributions to prognosis prediction. For group A, these variables are IMPL, RMT, and PSV respectively. The importance scores of these three variables, IMPL, GSUCL, and RMT, always ranked among the top five in group A and group C.

[0107] Figure 5 . Variable importance for predicting the prognosis of group C and group A. The variables are sorted in descending order according to their average contribution to the model prediction results. Larger values of the variables mean they are more important for the model.

[0108] For the patients in group C, a machine learning model and a traditional linear scoring model were constructed respectively, and the ROC curves of the models are shown in Figure 6 . After variable selection, the final GBDT model contains three features, and the corresponding feature importance values are as follows: IMPL is 0.762, GSUCL is 0.550, and GSSH is 0.376. The GBDT model achieved a sensitivity of 0.875 (0.473, 0.997), a specificity of 0.857 (0.728, 0.941), and an AUC value of 0.927 (0.856, 0.999) (classification threshold is 0.244) on the test dataset (Table 7).

[0109] Figure 6. ROC curves of two prognostic models in Group C. The ROC curve of the GBDT model is on the left, and the ROC curve of the traditional linear scoring model is on the right. The blue line is the ROC curve on the test dataset, and the orange line is the ROC curve on the training dataset.

[0110] Based on feature importance and clinical feasibility, a simplified linear scoring model was derived, as listed in Table 3. In this model, IMPL, GSUCL, and RMT were transformed into categorical variables and assigned integer scores (IMPL ≥ 2.43 cm was 3 points, GSUCL ≥ 1.4 cm was 2 points, and missing RMT was 1 point). The model achieved a sensitivity of 0.857 (0.421, 0.996), a specificity of 0.840 (0.709, 0.928), and an AUC value of 0.939 (0.872, 1.000) on the test set (Table 4). If the weighted score ≥ 3 points, the patient was classified as high-risk, that is, a value of IMPL ≥ 2.43 cm, or GSUCL ≥ 1.5 cm and RMT was missing.

[0111] Table 3. Scoring model for Group C.

[0112]

[0113] Note: A weighted score ≥ 3 points was considered high-risk and stratified into Group C, while a score below 3 points was considered low-risk and stratified into Group A or B.

[0114] For patients in Group A, a GBDT machine learning prognostic model containing 13 significant variables was established by combining a large model with traditional machine learning for automated feature selection and model optimization, and the weights of each feature variable are shown in Figure 5 ROC curves were plotted in Figure 7 The sensitivity and specificity of the model were 0.867 (0.595, 0.983) and 0.881 (0.744, 0.960) respectively (Table 4). The AUC value of the model on the test dataset was 0.917 (0.842, 0.993) (classification threshold was 0.233).

[0115] Figure 7 . ROC curve of the machine learning model in Group A. The blue line is the ROC curve on the test dataset, and the orange line is the ROC curve on the training dataset.

[0116] Table 4. Prediction validity of two groups of machine learning and linear scoring models.

[0117]

[0118] AUC, area under the curve; PPV, positive predictive value; NPV, negative predictive value; PLHR, positive likelihood ratio; NLHR, negative likelihood ratio.

[0119] The method of this application has the following advantages:

[0120] 1. Based on the clustering results of scar pregnancy patients, evaluate the classification prediction difficulty of groups A / B / C. Since the distribution distance of group C is far from that of groups A / B, a traditional linear scoring model is selected for group C for differentiation; a machine learning model is constructed for group A for identification. In this application, model selection based on the results of population clustering can more accurately select a more suitable model and method according to the population distribution and modeling complexity; for the convenience of clinical application, when group C patients are relatively easy to distinguish, there is no need to use a more complex model for modeling prediction, but a more concise model with good interpretability should be preferentially selected for construction.

[0121] 2. Identify single-factor and multi-factor significant variables based on traditional logistic regression, and combine computer algorithms to find the optimal variable combination. The scoring model for group C includes three 0 / 1 variables: IMPL≥2.43 cm, GSUCL≥1.5 cm, and RMT missing, with the scores of each variable being 3 points, 2 points, and 1 point respectively. For any patient, weighted summation is performed according to the three variables. A weighted score ≥3 points is regarded as high risk and stratified into group C, while a score lower than 3 points is regarded as low risk and stratified into group A or B. In this application, converting continuous variables into categorical variables and assigning scores can have better interpretability, making it easy for doctors to remember and apply clinically.

[0122] 3. For non-group C patients, combine large models with traditional machine learning algorithms for automatic selection of feature variables and model construction, and select the optimal variable combination and model prediction results. In this embodiment, combining large models and machine learning algorithms can automatically perform feature selection and model optimization, and based on the powerful text understanding ability of large models, it can better combine clinical expertise and select features from a clinical perspective like a doctor. On the one hand, it avoids the time-consuming and laborious problem of traditional manual feature selection, and on the other hand, it has better interpretability compared to the strategy of selecting features purely based on statistical methods.

[0123] Finally, 13 variables are selected from 20 variables: IMPL, RMT, PSV, GSUCL, GSSH, GSCL, FL, GSDmax, RI, GSSL, GSCH, GSDavg, CDFI. The feature weights of the 13 variables in GBDT are respectively:

[0124]

[0125]

[0126] In this embodiment, the large model receives five aspects of input: a. Medical knowledge and clinical diagnosis and treatment data related to the population of patients with scar pregnancy; b. Features selected by the current machine learning model; c. The current available set of clinical features; d. The performance metrics of the current machine learning model on the training and test sets; e. Historical information on model feature and performance iteration. Then, the large model outputs a new set of features for training the machine learning model. After multiple rounds of iteration, until the performance of the machine learning model does not improve significantly, the training stops, and the currently selected clinical feature weights and the trained model are output.

[0127] This application also provides a device for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy. The device for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy includes a patient database acquisition module, a classification module, a first model establishment module, and a second model establishment module, wherein,

[0128] The patient database acquisition module is used to acquire a patient database, and the patient database includes at least one patient data;

[0129] The classification module is used to classify the patient data in the patient database according to the treatment method and prognosis, so as to obtain first-class patient data, second-class patient data, and third-class patient data;

[0130] The first model establishment module is used to establish a model for first-class patient data according to the first-class patient data, the second-class patient data, and the third-class patient data;

[0131] The second model establishment module is used to establish a model for second-class patient data according to the first-class patient data, the second-class patient data, and the third-class patient data.

[0132] Figure 2 It is a block diagram of the structure of an electronic device provided by one or more embodiments of the present invention.

[0133] Such as Figure 2As shown, the present application also discloses an electronic device, including: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; a computer program is stored in the memory. When the computer program is executed by the processor, the processor is caused to execute the steps of the method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy.

[0134] The present application also provides a computer-readable storage medium, which stores a computer program executable by an electronic device. When the computer program runs on the electronic device, it can implement the steps of the method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy.

[0135] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0136] The electronic device includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system. The hardware layer includes hardware such as a Central Processing Unit (CPU), a Memory Management Unit (MMU), and a memory. The operating system can be any one or more computer operating systems that implement the control of the electronic device through a process. For example, the Linux operating system, the Unix operating system, the Android operating system, the iOS operating system, or the windows operating system, etc. And in the embodiments of the present invention, the electronic device can be a handheld device such as a smart phone or a tablet computer, or an electronic device such as a desktop computer or a portable computer. It is not particularly limited in the embodiments of the present invention.

[0137] The execution subject of the control of the electronic device in the embodiments of the present invention can be the electronic device, or a functional module in the electronic device that can call and execute the program. The electronic device can obtain the firmware corresponding to the storage medium. The firmware corresponding to the storage medium is provided by the supplier. The firmware corresponding to different storage media can be the same or different, and this is not limited here. After the electronic device obtains the firmware corresponding to the storage medium, it can write the firmware corresponding to the storage medium into the storage medium. Specifically, it burns the firmware corresponding to the storage medium into the storage medium. The process of burning the firmware into the storage medium can be implemented by the prior art and will not be elaborated in the embodiments of the present invention.

[0138] The electronic device can also obtain the reset command corresponding to the storage medium, and the reset command corresponding to the storage medium is provided by the supplier. The reset commands corresponding to different storage media can be the same or different, which is not limited herein.

[0139] At this time, the storage medium of the electronic device is the storage medium written with the corresponding firmware. The electronic device can respond to the reset command corresponding to the storage medium in the storage medium written with the corresponding firmware, so that the electronic device resets the storage medium written with the corresponding firmware according to the reset command corresponding to the storage medium. The process of resetting the storage medium according to the reset command can be implemented by the prior art and will not be elaborated in the embodiments of the present invention.

[0140] For the convenience of description, when describing the above device, various units and modules are described separately according to their functions. Of course, when implementing the present application, the functions of each unit and module can be implemented in the same or multiple software and / or hardware.

[0141] Those skilled in the art of this technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined.

[0142] For the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0143] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy, characterized in that, The method for establishing the classification prediction model for the prognosis outcome of cesarean scar pregnancy includes: Step 1: Obtain a patient database, where the patient database includes at least one patient's data; Step 2: Classify the patient data in the patient database according to the treatment method and prognosis, so as to obtain the first-class patient data, the second-class patient data, and the third-class patient data; the first-class patient data is the data of patients with better prognosis; the second-class patient data is the data of other patients with medium prognosis; the third-class patient data is the data of patients with severe illness and poor prognosis; Step 3: Establish a model for the first-class patient data and a model for the second-class patient data respectively according to the first-class patient data, the second-class patient data, and the third-class patient data; The model for the first-class patient data is a GBDT model; the model for the second-class patient data is a linear weighted model; Through principal component analysis, visualize the first-class patient data, the second-class patient data, and the third-class patient data. Through PCA dimensionality reduction and visualization into a three-dimensional space, it is evaluated and found that the first-class patient data is not easy to distinguish, and a GBDT model is used for modeling; the third-class patient data is far from the first-class patient data and the second-class patient data, and a linear weighted model is used; The construction of the model for the first-class patient data includes: Step 31: Obtain the first-round multimodal data corresponding to the first-class patient data as positive samples, and obtain the first-round multimodal data of the second-class patient data and the third-class patient data as negative samples; Step 32: Obtain a trained large model; Step 33: Input the first-round multimodal data into the trained large model to obtain the first-round feature selection information; Step 34: Establish a first GDBT model according to the first-round feature selection information; Step 35: Optimize the first GDBT model to obtain the model for the first-class patient data; The construction of the model for the second-class patient data includes: Obtain the set of single-factor significantly correlated variables of the first-class patient data and the set of single-factor significantly correlated variables of the third-class patient data; Draw the ROC curve of the set of single-factor significantly correlated variables of the first-class patient data and the first-class patient data, and the ROC curve of the set of single-factor significantly correlated variables of the third-class patient data and the third-class patient data; Find the threshold with the largest Youden index on each ROC curve, and convert the variable set into a 0 / 1 classification variable based on this threshold; For each variable converted into a 0 / 1 classification variable, assign four weight scores of 0, 1, 2, and 3, and find the variable weight combination with the largest AUC value based on a computer algorithm.

2. The method for establishing the classification prediction model for the prognosis outcome of cesarean scar pregnancy according to claim 1, wherein The step 35: Optimize the first GDBT model to obtain the model for the first-class patient data includes: Step 351: Obtain the feature selection information for optimization based on the first GDBT model; Step 352: Obtain the optimized GDBT model according to the feature selection information for optimization; Step 353: Evaluate the performance of the first GDBT model and the optimized GDBT model respectively, and retain the one with better performance as the model for the first type of patient data.

3. The method for establishing a prediction model for the classification of the prognosis outcome of cesarean scar pregnancy according to claim 2, characterized in that, The step 35 of optimizing the first GDBT model to obtain the model for the first type of patient data further includes: Step 354: Obtain the optimized feature selection information according to the optimized GDBT model obtained in step 352; Step 355: Generate a new optimized GDBT model according to the optimized feature selection information; Step 356: Evaluate the performance of the new optimized GDBT model and compare its performance with the previously retained model, so as to retain the one with better performance as the model for the first type of patient data.

4. The method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy according to claim 3, characterized in that, The step 35 of optimizing the first GDBT model to obtain the model for the first type of patient data further includes: Step 357: Repeat generating a new optimized GDBT model and comparing its performance with the previously retained model until the model prediction performance converges or drops to a preset value.

5. The method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy according to claim 4, wherein The first-round multimodal data includes patient disease-related clinical knowledge and diagnosis and treatment data, and the set of clinical features available for the model.

6. The method for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy according to claim 5, wherein, The step 351 of obtaining the optimized feature selection information based on the first GDBT model includes: Step 3511: Obtain the first-round multimodal data corresponding to the first type of patient data, the model data obtained through the established first GDBT model, the set of features selected by the current-round machine learning model, the performance information obtained by the current-round machine learning model during training and testing, and the model feature selection and performance iteration history information; Step 3512: Obtain the trained large model; Step 3513: Input the first-round multimodal data, the model data obtained through the established first GDBT model, the set of features selected by the current-round machine learning model, the performance information obtained by the current-round machine learning model during training and testing, and the model feature selection and performance iteration history information into the trained large model, so as to obtain the optimized feature selection information.

7. An apparatus for establishing a classification prediction model for the prognosis outcome of cesarean scar pregnancy, characterized in that, The device for establishing the cesarean scar pregnancy prognosis outcome classification prediction model includes: A patient database acquisition module, which is used to acquire a patient database, and the patient database includes at least one patient data; A classification module, which is used to classify the patient data in the patient database according to the treatment method and prognosis, so as to obtain the first type of patient data, the second type of patient data, and the third type of patient data; the first type of patient data is the data of patients with better prognosis; the second type of patient data is the data of other patients with medium prognosis; the third type of patient data is the data of patients with severe illness and poor prognosis; A first model establishment module, which is used to establish a model for the first type of patient data according to the first type of patient data, the second type of patient data, and the third type of patient data; A second model establishment module, which is used to establish a model for the second type of patient data according to the first type of patient data, the second type of patient data, and the third type of patient data; The model for the first type of patient data is the GBDT model; the model for the second type of patient data is the linear weighted model; Through principal component analysis, the first type of patient data, the second type of patient data, and the third type of patient data are visualized. Through PCA dimensionality reduction and visualization into three-dimensional space, it is evaluated that the first type of patient data is not easy to distinguish, and the GBDT model is used for modeling; the third type of patient data is far from the first type of patient data and the second type of patient data, and the linear weighted model is used; The first model building module includes: Obtain the first-round multimodal data corresponding to the first type of patient data as positive samples, and obtain the first-round multimodal data of the second type of patient data and the third type of patient data as negative samples; Obtain a trained large model; Input the first-round multimodal data into the trained large model to obtain the first-round feature selection information; Establish a first GDBT model according to the first-round feature selection information; Optimize the first GDBT model to obtain the model for the first type of patient data; The second model building module includes: Obtain the set of single-factor significantly correlated variables of the first type of patient data and the set of single-factor significantly correlated variables of the third type of patient data; Plot the ROC curve of the set of single-factor significantly correlated variables of the first type of patient data and the first type of patient data, and the ROC curve of the set of single-factor significantly correlated variables of the third type of patient data and the third type of patient data; Find the threshold with the largest Youden index on each ROC curve, and convert the variable set into 0 / 1 classification variables based on this threshold; For each variable converted into 0 / 1 classification variables, assign four weight scores of 0, 1, 2, and 3, and find the variable weight combination with the largest AUC value based on a computer algorithm.

Citation Information

Patent Citations

  • Clinical predictors based on multiple machine learning models

    CN115699204A