A novel PCOS subphenotype precision diagnosis and treatment platform

By constructing a novel PCOS subphenotype precision diagnosis and treatment platform, and using the XGBoost model to classify PCOS patients into 5 subtypes, the platform addresses the shortcomings of existing diagnostic criteria, achieves accurate subtype prediction and treatment guidance, and significantly improves patients' health indicators.

CN121054237BActive Publication Date: 2026-04-28RENJI HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RENJI HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
Filing Date
2025-11-04
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The existing diagnostic criteria for PCOS lack a deep understanding of the pathophysiological and prognostic differences between different phenotypes, making it difficult to support precision medicine and personalized treatment plans.

Method used

A novel PCOS subphenotype precision diagnosis and treatment platform was constructed. Through case sample inclusion and cohort construction, data collection and processing, key feature screening, unsupervised learning and XGBoost model, PCOS patients were divided into 5 subtypes, and an XGBoost model was constructed for prediction.

Benefits of technology

It enables precise differentiation of PCOS subtypes and treatment guidance, improves model prediction accuracy to 92%, significantly improves patients' metabolic indicators, androgenemia, and ovarian volume, and supports individualized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121054237B_ABST
    Figure CN121054237B_ABST
Patent Text Reader

Abstract

The application relates to the field of clinical molecular typing, and discloses a novel PCOS subphenotype precision diagnosis and treatment platform, a case sample inclusion and cohort construction module, a data acquisition and processing module, a key feature screening module, 12-dimensional feature data of a patient used for subphenotype clustering, an unsupervised learning module, a K-Means clustering algorithm used for unsupervised clustering of 12-dimensional feature data of the case, a PCOS patient group divided into five new subtypes, an XGBoost model construction module, the five new subtypes as classification labels, and an XGBoost model constructed, and a typing prediction module, the XGBoost model constructed used for prediction of PCOS subphenotypes. The application provides a new fine-grained typing, a prediction model and a precision treatment scheme based on subtypes, and is suitable for subtype differentiation and treatment guidance of PCOS patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of clinical molecular subtyping, and in particular to a novel PCOS subphenotype precision diagnosis and treatment platform. Background Technology

[0002] Polycystic ovary syndrome (PCOS), a common endocrine disorder affecting women of reproductive age, has a global prevalence that fluctuates due to differences in diagnostic criteria, and is estimated to affect up to 15% of women in this age group. The syndrome presents with diverse clinical manifestations, primarily encompassing menstrual cycle irregularities (such as infrequent or anovulatory ovulation), clinical or biochemical signs of hyperandrogenemia, and often associated metabolic abnormalities. PCOS is closely associated with a range of health problems, including insulin resistance, central obesity, potential cardiovascular risks, decreased fertility, and even endometrial cancer, and remains one of the leading endocrine disorders in women of reproductive age. These complex reproductive and metabolic dysfunctions persist throughout a woman's lifespan and are key factors contributing to ovulatory infertility and type 2 diabetes (T2D). Therefore, PCOS poses a significant challenge and burden on both individual patients' quality of life and the healthcare system.

[0003] The patent application number CN202310250696.2 from Shandong University uses nine final clinical variables to give four subphenotypes, but the predictive effect is poor, with an AUC of only 0.81, and lacks clinical application results for the subtypes. Its clinical application value is still unclear.

[0004] Therefore, developing targeted treatment strategies for different clinical characteristics and building a comprehensive management model that meets individualized needs are key issues that urgently need to be addressed in the clinical diagnosis and treatment of PCOS. Summary of the Invention

[0005] The main objective of this invention is to address the technical problem in existing technologies where clinical efficacy exhibits significant individual differences, and patients' long-term prognoses also vary. A novel PCOS subphenotype precision diagnosis and treatment platform includes:

[0006] The case sample inclusion and cohort construction module builds multiple clinical cohorts for model development and validation, including a main cohort for subphenotype clustering, a test cohort for external model testing, and a treatment analysis cohort for validation of specific treatment regimens.

[0007] The data acquisition and processing module collects multidimensional clinical variables from the research cohort and performs rigorous preprocessing on the raw data, including outlier and error value handling, high missing value feature removal, and missing value imputation.

[0008] The key feature screening module obtains 12-dimensional feature data of patients for subphenotype clustering. The key features are: BMI, T, SHBG, DHEAS, A2, FT, FAI, O'PG, 120'INS, HOMA-IR, MI, and DI.

[0009] The unsupervised learning module uses the K-Means clustering algorithm to perform unsupervised clustering on the 12-dimensional feature data of cases, dividing the PCOS patient population into 5 new subtypes;

[0010] The XGBoost model building module uses the five new subtypes as classification labels to build the XGBoost model.

[0011] The genotyping prediction module uses a pre-built XGBoost model to predict PCOS subphenotypes.

[0012] The main cohort included 1550 PCOS patients; the external testing cohort included 200 PCOS patients; and the treatment analysis cohort included 599 PCOS patients with both initial diagnosis and follow-up data. All patients were Chinese women aged 18-45 years diagnosed according to the 2003 Rotterdam criteria, with a BMI between 18.5 and 50.0 kg / m². 2 .

[0013] In the preprocessing, features with a missing rate greater than 30% are removed, and the K nearest neighbor algorithm is used to fill in the missing values.

[0014] The platform also includes a verification module, which uses bootstrapping sampling of the main queue data; external verification performs isoparametric clustering on an independent queue containing 200 patients; and the reliability of the results is verified by comparing combinations of various clustering algorithms and data imputation methods.

[0015] The platform also includes a model evaluation module, which is used to verify the overall performance of the constructed XGBoost model and ultimately select the XGBoost model with the best overall performance.

[0016] The overall performance evaluation metrics include accuracy, F1 score, confusion matrix, area under the receiver operating characteristic curve, and clinical decision curve.

[0017] The present invention has the following beneficial effects:

[0018] This invention provides a novel fine-grained subtyping + predictive model + subtype-based precision treatment plan, applicable to PCOS patients for subtype differentiation and treatment guidance. It directly links subtyping results to clinical decision-making, and patients who meet the subtyping treatment plan show better metabolic, androgen, and ovarian volume outcomes. This represents a leap from "diagnostic subtyping" to "treatment decision-making," and has significant clinical implications.

[0019] In this invention, the model's prediction accuracy is 92%, and the ROC evaluation is 0.99. The average Jaccard index is 0.71, the area under the ROC curve is >0.99, and the optimized accuracy reaches 92%. Attached Figure Description

[0020] Figure 1 This is a technical roadmap of the present invention;

[0021] Figure 2 This is a scatter plot of UMAP dimensionality reduction clustering in this invention;

[0022] Figure 3-5 A scatter plot showing the internal verification of other methods in this invention;

[0023] Figure 6 A pie chart showing the distribution ratio of cluster subtypes in this invention;

[0024] Figure 7-9 This is a radar diagram illustrating the physiological characteristics between different subtypes in this invention;

[0025] Figure 10 This is a heatmap of the clustering features in this invention;

[0026] Figure 11 The confusion matrix diagram of the prediction model constructed for this invention;

[0027] Figure 12 ROC curve of the prediction model constructed in this invention;

[0028] Figure 13-17 The SHAP diagram of the prediction model constructed in this invention;

[0029] Figure 18-22 This is a DCA plot used in this invention to measure the prediction model curve;

[0030] Figure 23 This is a forest plot used in this invention to measure the improvement effect of specific treatments;

[0031] Figure 24 This is a waterfall chart used in this invention to measure the effectiveness of specific treatments. Detailed Implementation

[0032] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” or “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] As a complex syndrome involving reproductive disorders and metabolic disturbances, PCOS has a core clinical management objective of early identification and prevention of serious secondary organ complications such as ovulatory dysfunction-related infertility and type 2 diabetes. However, the currently widely used PCOS classification based on the Rotterdam criteria (typically divided into four classic phenotypes), and several revisions and updates to international PCOS diagnosis and treatment guidelines in recent years, have failed to fully reveal the deep-seated pathophysiological and prognostic differences between different phenotypes in practice, thus limiting the translation towards precision medicine and personalized treatment plans. Although the classification and diagnostic criteria for PCOS are continuously evolving, the core parameters and modeling foundations upon which current classifications rely still largely focus on the perspectives of obstetrics and gynecology and reproductive medicine, with relatively insufficient attention paid to the long-term prognosis and mechanisms of metabolic disturbances in patients. This limitation inevitably leads to limited contributions of current classifications to assessing and improving metabolic outcomes, making it difficult to support precise clinical risk stratification for individual patients and posing challenges to the development of effective personalized interventions. Based on our current understanding of the known pathophysiological mechanisms of PCOS related to specific organs, this study explores and constructs a novel subtype classification system that can more precisely reflect the inherent heterogeneity of PCOS, and evaluates the effectiveness of individualized intervention strategies based on this new subtype.

[0034] In this invention, the 12 key features selected for clustering are: body mass index (BMI), testosterone (T), sex hormone-binding globulin (SHBG), dehydroepiandrosterone sulfate (DHEAS), androstenedione (A2), free testosterone (FT), free androgen index (FAI), 0-minute blood glucose (0'PG), 120-minute insulin (120'INS), insulin resistance homeostasis model assessment index (HOMA-IR), Matsuda index (MI), and disposition index (DI).

[0035] This invention identified five PCOS subtypes, each with unique pathophysiological characteristics:

[0036] Metabolic healthy obesity phenotype, adrenal-derived phenotype, ovarian-derived phenotype, pancreatic-derived phenotype, liver-derived phenotype.

[0037] In this invention, the model used for the above subtype classification maintains stable accuracy in both internal and external validation (average Jaccard index: 0.71, area under the ROC curve > 0.99, and accuracy of 92% after optimization).

[0038] The enhancement model showed that patients recommended treatment regimens by subtype had significant improvements in metabolic parameters, hyperandrogenemia, and ovarian volume (mean enhancement score = 0.18).

[0039] Innovation of different classification and treatment models based on clinical experience and pathological mechanisms.

[0040] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 , Figure 1 As shown in the technical roadmap of this invention, the first embodiment of the novel PCOS subphenotype precision diagnosis and treatment platform in this invention includes:

[0041] The case sample inclusion and cohort construction module builds multiple clinical cohorts for model development and validation, including a main cohort for subphenotype clustering, a test cohort for external model testing, and a treatment analysis cohort for validation of specific treatment regimens.

[0042] The data acquisition and processing module collects multidimensional clinical variables from the research cohort and performs rigorous preprocessing on the raw data, including outlier and error value handling, high missing value feature removal, and missing value imputation.

[0043] The key feature screening module obtains 12-dimensional feature data of patients for subphenotype clustering. The key features are: BMI, T, SHBG, DHEAS, A2, FT, FAI, O'PG, 120'INS, HOMA-IR, MI, and DI.

[0044] The unsupervised learning module uses the K-Means clustering algorithm to perform unsupervised clustering on the 12-dimensional feature data of cases, dividing the PCOS patient population into 5 new subtypes;

[0045] The XGBoost model building module uses the five new subtypes as classification labels to build the XGBoost model.

[0046] The genotyping prediction module uses a pre-built XGBoost model to predict PCOS subphenotypes.

[0047] The novel PCOS subtype precision diagnosis and treatment platform provided by this invention is constructed through the following steps:

[0048] 1. Case Sample Inclusion and Cohort Construction: To conduct subtype cluster analysis, predictive model construction, and validation of specific treatment methods, this embodiment constructed three clinical research cohorts. The main cohort included 1550 PCOS patients from Renji Hospital affiliated with Shanghai Jiao Tong University School of Medicine for the discovery of novel subtype clusters. The external testing cohort included 200 PCOS patients from Shanghai Tenth People's Hospital to validate the generalization of the clustering results. The treatment analysis cohort included 599 patients from Renji Hospital who had both complete initial diagnosis data and long-term treatment follow-up data to validate the effectiveness of specific treatment regimens.

[0049] All enrolled patients met the following criteria: (1) Chinese women diagnosed with PCOS according to the 2003 Rotterdam criteria; (2) aged between 18 and 45 years; and (3) body mass index (BMI) between 18.5 kg / m². 2 Up to 50.0 kg / m 2 All participants signed written informed consent forms, and this study has been approved by the relevant ethics committee.

[0050] 2. Collection and Preprocessing of Clinical Variables: By consulting the electronic medical record system, 47 clinical variables covering metabolism, reproduction, and other aspects were initially obtained. Rigorous preprocessing was performed on the original dataset: First, outliers and erroneous values ​​were identified using appropriate statistical methods and treated as missing values; second, to ensure data quality, data features with an overall missing value rate exceeding 30% were excluded; finally, the K-nearest neighbor algorithm was used to impute the remaining missing values ​​to preserve the original distribution characteristics of the data to the greatest extent possible.

[0051] 3. Key feature screening and novel subphenotype clustering:

[0052] Based on variable contribution analysis and clinical expert evaluation, 12 key features were selected from the preprocessed data. Feature importance was determined using a combination of statistical and machine learning methods, including variance-based selection and variable contribution calculation using the XGBoost model. In the clustering phase, the K-Means algorithm was used for unsupervised learning. To determine the optimal number of clusters K, the elbow method and clinical analysis were combined, ultimately determining K=5. The model iteration count was set to 1000, and the initial centroids were generated using the K-Means++ algorithm to improve convergence stability.

[0053] After data preprocessing, and considering feature importance and expert clinical experience, 12 key features most representative of PCOS heterogeneity were selected for further analysis. These 12 features are: body mass index (BMI), testosterone (T), sex hormone-binding globulin (SHBG), dehydroepiandrosterone sulfate (DHEAS), androstenedione (A2), free testosterone (FT), free androgen index (FAI), 0-minute blood glucose (0'PG), 120-minute insulin (120'INS), insulin resistance homeostasis model assessment index (HOMA-IR), Matsuda index (MI), and disposition index (DI). The K-Means clustering algorithm was used to perform unsupervised clustering of the 12-dimensional feature data of the main cohort (1550 patients), ultimately dividing the PCOS patient population into 5 new subtypes.

[0054] 4. Verification of the stability and reproducibility of clustering results:

[0055] To verify the robustness of the clustering results, 100 bootstrapping samples were performed on the main cohort, and the clustering stability index was calculated. Simultaneously, the same parameters were applied to an independent external cohort containing 300 patients for clustering, to externally evaluate the model's reproducibility. The consistency and reliability of subphenotype segmentation were further verified by comparing results with other methods such as hierarchical clustering. This confirms that the subtyping scheme is not accidental and possesses high stability and reproducibility across datasets.

[0056] 5. Construction and evaluation of subphenotypic prediction models:

[0057] Based on clustering results as labeled data, three supervised learning models—logistic regression, random forest, and XGBoost—were constructed. Hyperparameters were optimized using five-fold cross-validation, with a learning rate of 0.08 and a maximum tree depth of 3. The XGBoost model, which had the highest AUC, was ultimately selected as the optimal prediction model. Model performance was comprehensively evaluated using accuracy, F1 score, confusion matrix, ROC curve, and DCA curve. Furthermore, SHAP values ​​were used to analyze the model's feature contribution, enabling model interpretability visualization.

[0058] 6. Proposal and efficacy verification of subphenotype-specific treatment plans:

[0059] For each subphenotype identified in the aforementioned steps, and considering its clinical and metabolic characteristics, corresponding specific treatment interventions are proposed. To scientifically evaluate the therapeutic advantages guided by subtyping, this invention utilizes follow-up data from a treatment analysis cohort to construct an uplift model based on causal inference for efficacy verification. The specific process is as follows: the baseline physiological characteristics and clinical indicators of the follow-up patients are used as the input variable set X; whether or not specific treatment is received is used as a binary treatment variable T (where T=1 indicates receiving specific treatment, and T=0 indicates receiving conventional treatment); and the magnitude of change in key physiological indicators before and after treatment is used as the outcome variable Y. By training the model, the uplift score under different feature combinations is calculated to quantify the additional therapeutic effect brought about by subtyping intervention.

[0060] 7. Platform Deployment and Clinical Application:

[0061] This embodiment encapsulates the final selected optimal XGBoost prediction model and deploys it as a publicly accessible web platform. The method for using this platform is as follows:

[0062] 1. Data input: On the platform's prediction interface, input 12 key clinical characteristics of a PCOS patient (i.e., BMI, T, SHBG, DHEAS, A2, FT, FAI, 0'PG, 120'INS, HOMA-IR, MI, DI).

[0063] Get the results: Click the "Predict" button, and the system backend will call the deployed XGBoost model to perform real-time calculations and return the subtype classification and specific treatment suggestions on the front-end page in an instant.

[0064] The experimental results from the aforementioned steps are interpreted as follows:

[0065] Figure 2 This is a scatter plot of UMAP dimensionality reduction clustering in this invention. This invention first uses the KMeans clustering algorithm to divide the sample data into subtypes, and then uses the UMAP algorithm to reduce the dimensionality of the high-dimensional data to visualize the distribution characteristics between different subtypes. The scatter points of different colors in the figure represent each cluster subtype.

[0066] Figure 3-5 The scatter plots used in this invention to internally validate other methods involved repeated clustering analysis of the sample data using different feature engineering techniques and various clustering algorithms. By comparing the distribution patterns and subtype classifications of the clustering results obtained by different methods, it can be observed that each group of samples exhibits a relatively similar clustering structure under different algorithms, indicating that the clustering model of this invention has high stability.

[0067] Figure 6This is a pie chart showing the distribution ratio of cluster subtypes in this invention, representing the proportion of people distributed among the various subtypes of the cluster.

[0068] Figure 7-9 This invention provides a radar chart describing the physiological characteristics between different subtypes, illustrating the distribution differences of key physiological indicators among the various subtypes in a cluster.

[0069] Figure 10 This is a heatmap of clustering features in this invention. The color intensity of each cell in the graph represents the relative level of the corresponding feature value, and different color gradients reflect the degree of difference between different features.

[0070] Figure 11 The confusion matrix diagram of the prediction model constructed in this invention reflects the recognition accuracy and classification error of each category by comparing the correspondence between the model's prediction results and the true labels. The values ​​on the diagonal of the diagram represent the number of samples correctly classified by the model, while the off-diagonal portion represents the distribution of misclassified samples. By observing the color intensity and numerical distribution of the confusion matrix, the model's ability to distinguish between different subtypes and its overall classification performance can be intuitively evaluated.

[0071] Figure 12 The image shows the ROC curve of the prediction model constructed in this invention. The horizontal axis represents the false positive rate (FPR), the vertical axis represents the true positive rate (TPR), and the area under the curve (AUC) is used to quantify the overall classification ability of the model. Each colored curve corresponds to the prediction result of a different subtype. The closer the AUC value is to 1, the higher the accuracy of the model in identifying that subtype.

[0072] Figure 13-17 This is a SHAP plot of the prediction model constructed in this invention. The horizontal axis represents the SHAP value of each feature, i.e., the direction and intensity of the feature's influence on the model's output; the vertical axis represents different input variables; and the color gradient reflects the level of feature values. This plot visually displays the degree of influence of various physiological and metabolic indicators on the prediction results of specific subtypes, thereby revealing the key features upon which the model bases its judgments on different subtypes.

[0073] Figure 18-22 This is a DCA plot measuring the predictive model curve of this invention. The horizontal axis represents the probability of the prediction threshold, and the vertical axis represents the net benefit value. The curve "Model" indicates the net benefit of the model established by this invention. "All" and "None" represent two extreme strategies: assuming intervention on all samples or not intervening on any samples, respectively. This plot reflects the clinical decision-making value and application advantages of the model of this invention within various threshold ranges, providing a basis for the use of the model in actual diagnosis and treatment scenarios.

[0074] Figure 23This forest plot, used in this invention to measure the effectiveness of specific treatments, shows the odds ratio (OR) and its 95% confidence interval (CI) on the horizontal axis, and different stratification variables on the vertical axis. Black squares represent the odds ratios of each subgroup, with the square size reflecting the sample size, and the horizontal lines representing the confidence interval range. The dashed line is the reference line for OR=1, used to distinguish the difference in efficacy between subtype-stratified treatment and control treatment.

[0075] Figure 24 This invention uses a waterfall plot to measure the effectiveness of specific treatment. The waterfall plot illustrates the changes in multiple physiological indicators across different subtypes under two intervention methods: specific treatment (T=1) and non-specific treatment (T=0). The horizontal axis represents individual samples of different subtypes, and the vertical axis represents the percentage change relative to baseline. Each bar corresponds to the direction and magnitude of indicator change for one sample.

[0076] The above experiments demonstrate that this invention employs the K-Means clustering algorithm to perform unsupervised clustering of the 12-dimensional feature data of cases, dividing the PCOS patient population into five new subtypes. These five new subtypes are then used as classification labels to construct an XGBoost model. The average Jaccard index is 0.71, the area under the ROC curve is >0.99, and the model's prediction accuracy reaches 92%, with an ROC evaluation of 0.99. This significantly improves the model's prediction accuracy. The uplift model shows that patients recommended treatment plans according to subtype exhibit significant improvements in metabolic indicators, hyperandrogenemia, and ovarian volume, with an average improvement score of 0.18.

[0077] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A novel PCOS subphenotype precision diagnosis and treatment platform, characterized in that, The platform includes: The case sample inclusion and cohort construction module constructs multiple clinical cohorts for model development and validation, including a main cohort for subphenotype clustering, a test cohort for external model testing, and a treatment analysis cohort for specific treatment validation. The main cohort includes 1550 PCOS patients; the test cohort for external model testing includes 200 PCOS patients; and the treatment analysis cohort includes 599 PCOS patients with both initial diagnosis and follow-up data. All patients are Chinese women aged 18-45 years diagnosed according to the 2003 Rotterdam criteria, with a BMI between 18.5 and 50.0 kg / m². The data acquisition and processing module collects multidimensional clinical variables from the research cohort and performs rigorous preprocessing on the raw data, including outlier and error value handling, high missing value feature removal, and missing value imputation. The key feature screening module obtains 12-dimensional feature data of patients for subphenotype clustering. The key features are: BMI, T, SHBG, DHEAS, A2, FT, FAI, O'PG, 120'INS, HOMA-IR, MI, and DI. The unsupervised learning module uses the K-Means clustering algorithm to perform unsupervised clustering on the 12-dimensional feature data of the cases, dividing the PCOS patient population into 5 new subtypes; the 5 new subtypes are metabolically healthy obesity phenotype, adrenal-derived phenotype, ovarian-derived phenotype, pancreatic-derived phenotype, and liver-derived phenotype. The XGBoost model building module uses the five new subtypes as classification labels to build the XGBoost model. The genotyping prediction module uses a pre-built XGBoost model for predicting PCOS subphenotypes. The specific treatment validation module utilizes follow-up data from the treatment analysis cohort to construct an uplift model based on causal inference for efficacy validation. The specific process is as follows: the baseline physiological characteristics and clinical indicators of the follow-up patients are used as the input variable set X; whether or not specific treatment is received is used as a binary treatment variable T, where T=1 indicates receiving specific treatment and T=0 indicates receiving conventional treatment; the change in key physiological indicators before and after treatment is used as the outcome variable Y; by training the model, the improvement score under different feature combinations is calculated to quantify the additional therapeutic effect brought about by the subtype intervention.

2. The novel PCOS subphenotype precision diagnosis and treatment platform according to claim 1, characterized in that, In the preprocessing, features with a missing rate greater than 30% are removed, and the K-nearest neighbor algorithm is used to fill in the missing values.

3. The novel PCOS subphenotype precision diagnosis and treatment platform according to claim 1, characterized in that, The platform also includes a verification module, which uses bootstrapping sampling of the main queue data; external verification performs isoparametric clustering on an independent queue containing 200 patients; and the reliability of the results is verified by comparing combinations of various clustering algorithms and data imputation methods.

4. A novel PCOS subphenotype precision diagnosis and treatment platform according to claim 1, characterized in that, The platform also includes a model evaluation module, which is used to verify the overall performance of the constructed XGBoost model and ultimately select the XGBoost model with the best overall performance.

5. A novel PCOS subphenotype precision diagnosis and treatment platform according to claim 4, characterized in that, The overall performance evaluation metrics include accuracy, F1 score, confusion matrix, area under the receiver operating characteristic curve, and clinical decision curve.

Citation Information

Patent Citations

  • Psychological pre-judgment method and system based on K-means clustering and XGBoost algorithm

    CN112530546A

  • Typing system for polycystic ovarian syndrome

    CN118629558A

  • Typing model based on susceptibility genes and clinical characteristics of polycystic ovarian syndrome

    CN118866112A