A biomarker combination for diagnosing chronic endometritis and use thereof
The diagnostic model constructed using a combination of biomarkers and the XGBoost algorithm solves the problem of invasive diagnosis of chronic endometritis, achieving non-invasive and accurate early diagnosis, improving the specificity and sensitivity of the diagnosis, and providing support for personalized treatment.
Patent Information
- Application Number
- CN202510557152.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Existing diagnostic methods for chronic endometritis have risks of invasive procedures, insufficient sensitivity and specificity, lack of standardized thresholds and non-invasive detection methods, resulting in a high rate of missed diagnoses and making it difficult to achieve early and accurate diagnosis.
A diagnostic model is constructed using a combination of biomarkers such as lymphocyte count, CD3+ T cell percentage, CD8+ T cell percentage, CD19+ B cell count, NK cell count, IL10, IgG, and IgG4, combined with the XGBoost algorithm, to achieve non-invasive detection and rapid diagnosis.
It enables non-invasive and non-surgical detection of chronic endometritis, improves the accuracy of early diagnosis, achieves a specificity of 88.9%, is suitable for routine screening of asymptomatic high-risk groups, reduces the risk of endometrial damage and infection, and provides theoretical guidance for personalized treatment.
Smart Images

Figure CN120404536B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biotechnology, in particular to a biomarker combination for diagnosing chronic endometritis and application thereof. BACKGROUND
[0002] Chronic endometritis (CE) is a chronic inflammatory disease characterized by endometrial stromal plasma cell infiltration, which is closely related to infertility, repeated implantation failure and recurrent pregnancy loss. Epidemiological data shows that about 10%-30% of infertile women have CE, and the misdiagnosis rate is as high as 30%. The pathogenesis of CE is complex, involving bacterial infection, uterine cavity microecological imbalance, immune microenvironment disorder and abnormal immune cell infiltration. Since CE lacks typical clinical symptoms, its early diagnosis is crucial to improve pregnancy outcome.
[0003] Currently, the clinical diagnosis of chronic endometritis (CE) mainly relies on the following two techniques: histopathological examination and hysteroscopy. Histopathological examination is the gold standard, which detects CD138-positive plasma cells through endometrial biopsy to confirm the status of plasma cell infiltration. Hysteroscopy can directly observe the changes of endometrium, such as "strawberry sign" (large area of mucosal congestion with a white center point), focal congestion and endometrial micro-polyps (small lumps with a diameter of <1mm, with obvious connective vascular axis) and other characteristics.
[0004] Currently, the diagnosis of CE has the following key defects and deficiencies: 1. Risk of invasive operation: The existing diagnostic methods mainly rely on hysteroscopy and histopathological examination. These operations not only have the risk of causing damage to the endometrium, but also can cause bleeding or infection. For patients with infertility, invasive examination methods can also have an adverse effect on subsequent reproductive treatment. 2. Insufficient sensitivity and specificity: Although endometrial biopsy to detect CD138 is considered the gold standard for diagnosing CE, a small number of CD138+ plasma cells may appear in the endometrium of healthy women of childbearing age, resulting in a false positive rate of 20%-30%. Moreover, hysteroscopy examination to some extent depends on the subjective experience of the operator, and there is a certain degree of uncertainty. 3. Lack of standardized threshold: Currently, the diagnosis of CE relies on the counting of CD138+ plasma cells, but the threshold values of different studies are not completely consistent. Some studies use a threshold of 1-5 / HPF, while some studies use a standard of 1 / 10 HPF. In addition, inflammation can be distributed in foci, and single biopsy has the risk of missing detection; the influence of the menstrual cycle and the sampling volume also has a certain influence on the diagnostic results. In summary, there is no uniform and standardized diagnostic standard. 4. Lack of non-invasive detection means: The existing technology has not been able to achieve non-invasive screening of CE through peripheral blood biomarkers, making it difficult for high-risk groups (such as patients receiving in vitro fertilization embryo transfer) to receive effective intervention at an early stage. At the same time, the sensitivity and specificity of conventional single blood tests are not sufficient to replace pathological biopsy, thereby limiting the application of CE in early detection and monitoring. SUMMARY
[0005] The purpose of the present application is to provide a biomarker combination for diagnosing chronic endometritis and its application, in order to solve the problems existing in the prior art. The present application detects 8 biomarkers related to chronic endometritis, which can realize the non-invasive detection and rapid diagnosis of CE, avoid the damage caused by invasive surgical operation and long diagnosis time. The area under the curve of the receiver operating characteristic curve of the diagnostic model based on the above markers is 0.929, which has high specificity (88.9%), significantly improves the early diagnosis accuracy of CE, provides a new method for rapid diagnosis and early diagnosis of CE, provides theoretical guidance for personalized treatment of CE, and also provides technical support for in-depth study of the treatment mechanism of CE.
[0006] To achieve the above purpose, the present application provides the following solutions:
[0007] The present application provides a biomarker combination for diagnosing chronic endometritis, which comprises lymphocyte count, CD3+ T cell percentage, CD8+ T cell percentage, CD19+ B cell count, NK cell count, IL10, IgG and IgG4.
[0008] The application also provides application of the biomarker combination in preparation of a product for early diagnosis of chronic endometritis.
[0009] Optionally, the product comprises a reagent or a kit.
[0010] Further, the product diagnoses chronic endometritis by detecting the content of the biomarker combination in peripheral blood.
[0011] The application also provides a kit for early diagnosis of chronic endometritis, wherein the kit comprises reagents for detecting the content of lymphocyte count, CD3+ T cell percentage, CD8+ T cell percentage, CD19+ B cell count, NK cell count, IL10, IgG and IgG4 in peripheral blood.
[0012] The application also provides application of the biomarker combination in construction of a diagnostic model for chronic endometritis.
[0013] The application also provides a construction method of a diagnostic model for chronic endometritis, wherein the diagnostic model for chronic endometritis is constructed by using an XGBoost algorithm.
[0014] Further, the construction method comprises the following steps:
[0015] acquiring data of lymphocyte count, CD3+ T cell percentage, CD8+ T cell percentage, CD19+ B cell count, NK cell count, IL10 content, IgG content and IgG4 content in peripheral blood of a patient as a data set;
[0016] randomly dividing the data set into a training set, a validation set and a test set;
[0017] constructing a diagnostic model based on an XGBoost algorithm by using the training set, performing cross-validation on the diagnostic model by using the validation set, and evaluating clinical applicability and diagnostic performance of the diagnostic model by using the test set.
[0018] The application also provides a diagnostic model for chronic endometritis constructed by the construction method.
[0019] The application discloses the following technical effects:
[0020] The application adopts LASSO regression analysis to screen out 8 biomarkers related to chronic endometritis (CE), and realizes non-invasive and non-invasive detection of CE by detecting 8 biomarkers (lymphocyte count, IL10, CD3+ percentage, CD8+ percentage, CD19+ B cell count, NK cell count, IgG, IgG4) in peripheral blood, which can avoid the risk of endometrial damage, bleeding and infection caused by hysteroscopy or endometrial biopsy, and reduce the discomfort of patients. The detection of the 8 biomarkers can realize early detection and diagnosis of CE, and is suitable for routine screening of asymptomatic high-risk groups (such as IVF patients), the receiver operating characteristic curve AUC = 0.929 has high specificity (88.9%), and can significantly improve the early diagnosis accuracy of CE.
[0021] The application combines multi-dimensional immune indicators (cell subgroups + inflammatory factors), uses XGBoost machine learning algorithm to integrate data, optimizes classification performance, outputs CE probability based on biomarker data, and uses SHAP explanation model to predict logic, which can standardize and visualize the diagnostic results and has interpretability, providing a new method for rapid diagnosis and early diagnosis of CE, providing theoretical guidance for personalized treatment of CE, and providing technical support for in-depth research on the treatment mechanism of CE. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0023] Figure 1 LASSO coefficient path diagram;
[0024] Figure 2 LASSO regression analysis cross-validation curve;
[0025] Figure 3 Variance inflation factor (VIF) for screening biomarkers;
[0026] Figure 4 ROC curve of the training set of various machine learning classification models;
[0027] Figure 5 Calibration curve of the training set of various machine learning classification models;
[0028] Figure 6 Precision recall curve of the training set of various machine learning classification models;
[0029] Figure 7 Validation set ROC curve for the diagnostic model; where different colored lines represent 10 cross-validations;
[0030] Figure 8 Test set ROC curve for the diagnostic model;
[0031] Figure 9 Test set DCA curve for the diagnostic model;
[0032] Figure 10 Test set calibration curve for the diagnostic model;
[0033] Figure 11 SHAP ranking plot for biomarker importance;
[0034] Figure 12 SHAP dependence plot;
[0035] Figure 13 SHAP waterfall plot for personalized interpretation of diagnostic results for patient A. DETAILED DESCRIPTION
[0036] A number of exemplary embodiments of the present application are now described in detail below. The following description provides specific details in order to provide a thorough understanding of the present application. However, the description itself and the specific examples described therein are not intended to limit the present application in any way.
[0037] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. In addition, for a range of values of or between the stated values and intervening values, each intermediate value of the stated range is also specifically disclosed. Each smaller range between any stated value or intervening value in the stated range and any other stated or intervening value in that stated range is also specifically disclosed. The upper and lower limits of these smaller ranges can independently be included or excluded in the range, and the endpoints are included in the smaller ranges.
[0038] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. Although preferred methods and materials are described herein, any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present application. All documents mentioned herein are incorporated by reference to disclose and describe the methods and / or materials in connection with which the documents are cited. In the case of conflict between the present specification and any document incorporated by reference, the present specification will control.
[0039] Many modifications and variations to the illustrative embodiments described herein will be apparent to those of ordinary skill in the art from the foregoing description. Such variations may not depart from the scope or spirit of the present disclosure. Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the disclosure disclosed herein. The specification and examples given are exemplary only.
[0040] As used herein, the terms "comprises", "comprising", "includes", "including", "has", "having", and the like are open-ended terms that are intended to permit but not limit the inclusion of elements or the number of elements, as the case can be.
[0041] The technical idea of the present disclosure is as follows:
[0042] ① Collecting patient clinical diagnosis results and laboratory test data; the diagnosis result is whether chronic endometritis occurs.
[0043] The laboratory test data includes 77 test parameters, specifically including:
[0044] Complete blood count (20 items): white blood cell count (WBC), red blood cell count (RBC), hemoglobin (Hb), mean corpuscular volume (MCV), mean corpuscular hemoglobin (MCH), mean corpuscular hemoglobin concentration (MCHC), platelet count (PLT), red blood cell distribution width coefficient of variation (RDW-CV), mean platelet volume (MPV), platelet distribution width (PDW), percentage and absolute count of white blood cell classification (lymphocytes, neutrophils, eosinophils, basophils, monocytes);
[0045] Serum infection markers and biochemical analysis (27 items): troponin T, ferritin, procalcitonin, C-reactive protein, alanine aminotransferase, total protein, albumin, total bilirubin, alkaline phosphatase, aspartate aminotransferase, gamma-glutamyl transferase, cholinesterase, lipase, amylase, glucose, uric acid, creatine kinase isoenzyme MB, creatine kinase, lactate dehydrogenase, urea, creatinine, potassium, sodium, chloride, calcium, magnesium, total carbon dioxide;
[0046] Immunological analysis (17 items): immunoglobulin IgG, IgA, IgM and 4 IgG subclasses, lymphocyte subsets: percentage and absolute count of CD3+ T cells, CD4+ T helper cells, CD8+ T cells, CD19+ B cells, natural killer (NK) cells;
[0047] Cytokine profile (12 items) interleukin (IL-1β, IL-2, IL-4, IL-5, IL-6, IL-8, IL-10, IL-12P70, IL-17), interferon (IFN-α, IFN-γ) and tumor necrosis factor-α (TNF-α);
[0048] Coagulation index D-dimer.
[0049] ②Standardized data processing methods were used to normalize the test results.
[0050] ③Through LASSO regression, key biomarkers including lymphocyte count, CD3+ T cell percentage, CD8+ T cell percentage, CD19+ B cell count, NK cell count, IL-10, IgG, and IgG4 were selected from blood routine index, blood biochemical index, coagulation index, lymphocyte subset index, and cytokine index.
[0051] ④XGBoost algorithm was used to construct a prediction model. According to the model score, the probability of patients suffering from CE was determined.
[0052] Example 1
[0053] 1. Collection of clinical patient samples
[0054] Patients participating in the screening were infertile patients who received in vitro fertilization (IVF) treatment for the first time at Peking University Third Hospital from 2018 to 2022. All patients were screened according to strict inclusion and exclusion criteria.
[0055] Inclusion criteria include: (1) infertile patients who underwent hysteroscopy and endometrial biopsy; (2) endometrial biopsy specimens were stained with immunohistochemistry (IHC) for standard histopathological examination of CD138 and CD38; (3) complete laboratory records, including complete blood count, biochemical parameters, coagulation function, immunoglobulin and IgG subclass, lymphocyte subsets, and cytokine profile.
[0056] Exclusion criteria include: (1) history of recurrent miscarriage or severe endocrine disease (e.g., thyroid dysfunction); (2) endometriosis; (3) uterine malformation, intrauterine adhesions, submucosal myoma; (4) chromosomal abnormalities; (5) abnormal preoperative vaginal secretion examination; (6) polycystic ovary syndrome (PCOS); (7) abnormal uterine bleeding (AUB); (8) patients with missing data for any of the above information or key laboratory indicators.
[0057] After screening, 542 participants were finally included, and the recruited patients were diagnosed with chronic endometritis (CE) according to the following criteria: (1) histopathological examination: clear plasma cell infiltration; (2) clinical symptoms: including infertility and recurrent miscarriage; (3) microbial detection: used as an auxiliary tool, including PCR and culture method. Finally, 290 patients with chronic endometritis (CE) and 252 patients without CE were determined.
[0058] 2. Laboratory test data
[0059] Complete blood cell count (20 items): white blood cell count (WBC), red blood cell count (RBC), hemoglobin (Hb), mean corpuscular volume (MCV), mean corpuscular hemoglobin (MCH), mean corpuscular hemoglobin concentration (MCHC), platelet count (PLT), coefficient of variation of red blood cell distribution width (RDW-CV), mean platelet volume (MPV), platelet distribution width (PDW), percentage and absolute count of white blood cell differential (lymphocytes, neutrophils, eosinophils, basophils, monocytes);
[0060] Serum infection markers and biochemical analysis (27 items): Troponin T, ferritin, procalcitonin, C-reactive protein, alanine aminotransferase, total protein, albumin, total bilirubin, alkaline phosphatase, aspartate aminotransferase, gamma-glutamyl transferase, cholinesterase, lipase, amylase, glucose, uric acid, creatine kinase isoenzyme MB, creatine kinase, lactate dehydrogenase, urea, creatinine, potassium, sodium, chloride, calcium, magnesium, and total carbon dioxide;
[0061] Immunological analysis (17 items): Immunoglobulins IgG, IgA, IgM and their four IgG subclasses; Lymphocyte subsets: percentage and absolute count of CD3+ T cells, CD4+ T helper cells, CD8+ T cells, CD19+ B cells, and natural killer (NK) cells;
[0062] Cytokine profile (12 items): interleukins (IL-1β, IL-2, IL-4, IL-5, IL-6, IL-8, IL-10, IL-12P70, IL-17), interferons (IFN-α, IFN-γ) and tumor necrosis factor-α (TNF-α);
[0063] Coagulation marker D-dimer.
[0064] 3. LASSO regression feature selection:
[0065] Parameter settings: LASSO regression analysis was used for variable selection (λ = 0.03042).
[0066] Filtering results: such as Figure 1 and Figure 2 As shown, eight key biomarkers were selected from candidate variables (laboratory test data), including immune cell parameters: lymphocyte count, percentage of CD3+ T cells, percentage of CD8+ T cells, CD19+ B cell count, and NK cell count; and cytokines: IL10, IgG, and IgG4.
[0067] The results of eight biomarkers in clinical CE patients and non-CE patients are shown in Table 1. All biomarkers were higher in CE patients than in non-CE patients.
[0068] Table 1. Detection results of eight biomarkers in CE patients and non-CE patients.
[0069]
[0070] 4. Multicollinearity of key biomarkers
[0071] To assess potential multicollinearity among the selected parameters, the variance inflation factor (VIF) was calculated for the eight key biomarkers selected. The results are as follows: Figure 3 As shown, all VIF values are less than 5, indicating a low degree of multicollinearity. This suggests that the variables are not highly correlated and can be used individually for further analysis.
[0072] 5. Machine Learning Model Building
[0073] 5.1 Selecting the best model from multiple models
[0074] 5.1.1 Data partitioning: The data from 542 patients were randomly divided into a training set (n=325), a validation set (n=108), and a test set (n=109) in a ratio of 6:2:2.
[0075] 5.1.2 Multiple machine learning (ML) classification models were applied for comprehensive analysis. These included Extreme Gradient Boosting (XGBoost), Logistic Regression, Light Gradient Boosting Machine (LightGBM), Random Forest, Gaussian Naive Bayes (GNB), Multilayer Perceptron (MLP), Support Vector Machine (SVN), and k-Nearest Neighbors (KNN).
[0076] The ML model was trained using the training set, and its performance was then evaluated using the validation set. To assess predictive power, various performance metrics were calculated, including AUC, accuracy, sensitivity, specificity, positive predictive value, negative predictive value, F1 score, and Kappa. Furthermore, calibration curves, decision curve analysis (DCA), and precision-recall (PR) curves were evaluated to further validate the model's performance.
[0077] The results are as follows Figures 4-6 As shown in Tables 2-3, the evaluation results of multiple machine learning models reveal different performance trends. XGBoost has the highest area under the AUC curve on both the training set (0.998) and the validation set (0.934), making it the best choice for further development and validation of the prediction model.
[0078] Table 2 Performance metrics of each machine learning classification model based on the training set
[0079]
[0080]
[0081] Table 3 Performance metrics of each machine learning classification model based on the validation set.
[0082]
[0083] 5.2 Construction and Validation of Diagnostic Model
[0084] A diagnostic model for chronic endometritis (CE) was established using the Extreme Gradient Boosting (XGBoost) algorithm. The optimal parameters were determined by grid search: learning rate 0.1, maximum depth 6, and number of trees 200.
[0085] After performing 10 cross-validations on the validation set, the results are as follows: Figure 7 As shown, the model exhibits robust performance metrics, with an AUC of 0.994 (0.989-0.999) on the validation set.
[0086] The test set had an AUC of 0.929 (95% CI: 0.885-0.974), sensitivity of 0.762, specificity of 0.889, and positive predictive value of 0.906. Figure 8 Decision curve analysis (DCA) further validated the clinical applicability of the model, demonstrating stable net clinical benefits over a wide threshold probability range (10%–90%). Within the intermediate threshold range (30%–70%), the net benefit rate significantly outperformed both the "full treatment" and "no treatment" strategies, reinforcing its role in optimizing clinical decision-making. Figure 9 The Brier score on the test set calibration curve was 0.12, indicating a high degree of consistency between the predicted and actual probabilities. Figure 10 ).
[0087] The results demonstrate that the model exhibits excellent performance in distinguishing between patients and non-patients. It possesses strong diagnostic validity, with a sensitivity of 0.762 and a specificity of 0.889, reflecting its dual ability to accurately identify true positive cases while minimizing false positive misclassification. The high positive predictive value (0.906) further underscores its clinical application in high-risk screening contexts, where reliable identification of positive cases is crucial.
[0088] 6. Interpretability of the diagnostic model
[0089] The XGBoost diagnostic model constructed in step 5 can be analyzed for interpretability using the SHAP (Shapley Additive Explanations) method. The specific steps are as follows:
[0090] 1. SHAP value calculation: Based on the training set data, the SHAP algorithm is used to quantify the contribution value of each biomarker to the model output, and the feature contribution matrix of individual samples is generated.
[0091] 2. Feature importance ranking: The SHAP values of all samples are summarized and ranked in descending order of absolute contribution to determine the priority of key biomarkers. As shown in Figure 11 , the top-ranked features such as interleukin 10 (IL-10), CD3+ T cells, IgG4, and NK count have the most significant impact on the model output.
[0092] 3. Feature influence direction analysis: The SHAP dependence plot ( Figure 12 ) shows the correlation direction of each feature value and the prediction result, with red marks indicating positive contribution, such as increased CD3+ T cell ratio increasing CE risk, and blue marks indicating negative contribution.
[0093] 4. Individualized interpretation application: For a single sample (e.g., patient A), the SHAP waterfall plot ( Figure 13 ) is generated to visually display the cumulative impact of each feature on the prediction result, supporting clinical decision-making.
[0094] By using the SHAP method to make the model transparent, not only can the clinical credibility be improved, but also individualized treatment can be guided according to the feature abnormal type, which is significantly better than traditional machine learning models that only rely on probability output.
[0095] The above-described embodiments are only preferred modes of the present application and do not limit the scope of the present application. Without departing from the design spirit of the present application, various modifications and improvements to the technical solutions of the present application made by those skilled in the art shall fall within the protection scope determined by the claims of the present application.
Claims
1. Use of a biomarker combination in the manufacture of a product for the early diagnosis of chronic endometritis; The biomarker combination is lymphocyte count, CD3+ T cell percentage, CD8+ T cell percentage, CD19+ B cell count, NK cell count, IL10, IgG and IgG4.
2. Use according to claim 1, characterized in that, The product comprises a reagent or a kit.
3. Use according to claim 1, characterized in that, The product diagnoses chronic endometritis by detecting the content of the biomarker combination in peripheral blood.
4. Use of a biomarker combination in the construction of a diagnostic model for chronic endometritis for non-diagnostic purposes; The biomarker combination is lymphocyte count, CD3+ T cell percentage, CD8+ T cell percentage, CD19+ B cell count, NK cell count, IL10, IgG and IgG4.
5. A method for constructing a non-diagnostic purpose of a chronic endometritis diagnostic model, characterized by, The diagnostic model for chronic endometritis is constructed based on the content data of the biomarker combination using the XGBoost algorithm; The biomarker combination is lymphocyte count, CD3+ T cell percentage, CD8+ T cell percentage, CD19+ B cell count, NK cell count, IL10, IgG and IgG4.
6. The construction method of claim 5, wherein, The method comprises the following steps: Obtaining the data of peripheral blood lymphocyte count, CD3+ T cell percentage, CD8+ T cell percentage, CD19+ B cell count, NK cell count, IL10 content, IgG content and IgG4 content of a patient as a data set; Randomly dividing the data set into a training set, a validation set and a test set; Using the training set, constructing a diagnostic model based on the XGBoost algorithm; using the validation set, cross-validating the diagnostic model; and using the test set, evaluating the clinical applicability and diagnostic performance of the diagnostic model.
Citation Information
Patent Citations
Reagents, methods and kits for diagnosing primary immunodeficiencies
CN107209101A
Peripheral blood biomarker composition for diagnosing endometriosis and application thereof
CN118501023A
Patient immune response monitoring method and system for cell drug clinical research
CN119361070A