Biomarker group for early diagnosis of endometriosis as well as screening and application of biomarker group

Through the combination of broad-target metabolomics and machine learning, 10 metabolites were screened as biomarkers, solving the problem of early diagnosis of endometriosis and achieving efficient and accurate diagnostic results.

CN120177802APending Publication Date: 2025-06-20NINGBO WOMEN & CHILDRENS HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510317037.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively conduct early diagnosis of endometriosis, resulting in delayed diagnosis.

Method used

Using a combination of broad-target metabolomics and machine learning, 10 metabolites were screened as biomarkers, and screened and verified through non-target metabolomics analysis and multi-level integrated learning framework model.

Benefits of technology

The metabolic molecular diagnostic model of endometriosis has been successfully established, which improves the accuracy and reliability of early diagnosis and provides an efficient and accurate diagnostic tool for the clinical practice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120177802A_ABST
    Figure CN120177802A_ABST
Patent Text Reader

Abstract

The invention discloses a biomarker group for early diagnosis of endometriosis as well as screening and application of the biomarker group. The biomarker group comprises the following ten metabolites: sphingomyelin SMd (18: 1 / 24: 0), sphingomyelin SMd (18: 0 / 24: 1), phosphatidylcholine PC (16: 0-17: 0), phosphatidylcholine PC (18: 0-16: 0), phosphatidylcholine PC (19: 1), phosphatidylcholine PC (20: 5e-18: 0), phosphatidylcholine PC (22: 6e-18: 5), lysophosphatidylcholine LPC (14: 1), lysophosphatidylcholine LPC (22: 5) and phosphoglyceride PG (18: 0 / 16: 0). The invention is mainly applied to early endometriosis disease monitoring and drug treatment effect evaluation. The problem of early diagnosis of endometriosis is successfully solved, and an efficient, accurate and reliable diagnosis tool is provided for clinic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical technology, and particularly to a biomarker panel for early diagnosis of endometriosis, and its screening and application. Background Art

[0002] Endometriosis is a common chronic inflammatory reproductive endocrine disease in women, characterized by the presence of endometrial-like tissue outside the uterus, accompanied by growth, invasion, repeated bleeding, and subsequent pain symptoms, infertility, nodules or masses, etc. There is no exact data on the true prevalence of endometriosis. It is estimated that 1 in 10 reproductive-aged women worldwide may have endometriosis. The clinical symptoms of endometriosis patients lack specificity, which is the main reason for delayed diagnosis.

[0003] Metabolomics can reflect the metabolic status of the body, and machine learning can process complex data. The combination of machine learning algorithms and metabolomics analysis has important implications in clinical disease prediction, etiology analysis, survival analysis, and prognosis prediction. Although machine learning is becoming increasingly prominent in the field of medical disease diagnosis, the research on combining the two for the diagnosis of endometriosis still needs to be improved. Summary of the Invention

[0004] The purpose of the present invention is to solve the deficiencies in the prior art and propose a biomarker panel for early diagnosis of endometriosis based on broad-target metabolomics and machine learning, and its screening and application.

[0005] The technical solution of the present invention is as follows: A biomarker panel for early diagnosis of endometriosis includes the following 10 metabolites: sphingomyelin SMd(18:1 / 24:0), sphingomyelin SMd(18:0 / 24:1), phosphatidylcholine PC 16:0-17:0, phosphatidylcholine PC 18:0-16:0, phosphatidylcholine PC 19:0-19:1, phosphatidylcholine PC 20:5e-18:0, phosphatidylcholine PC 22:6e-18:5, lysophosphatidylcholine LPC 14:1, lysophosphatidylcholine LPC 22:5, phosphatidylglycerol PG(18:0 / 16:0).

[0006] The present invention also provides a screening method for the biomarker panel for early diagnosis of endometriosis, including the following steps:

[0007] S1. Sample collection and processing: Collect plasma samples from endometriosis patients and healthy controls according to strict standards, and perform standardized processing and storage.

[0008] S2. Non-targeted metabolomics analysis of plasma samples: Use liquid chromatography-tandem mass spectrometry technology to perform non-targeted metabolomics analysis on plasma samples, including sample metabolite extraction, data preprocessing, determination of differential metabolites, enrichment analysis, and screening of candidate biomarkers;

[0009] S3. Construct a multi-level integrated learning framework model: Combine an automated hyperparameter optimization algorithm and an adaptive adjustment strategy to screen biomarkers, and use the automated hyperparameter optimization algorithm to optimize the hyperparameters of each machine learning model;

[0010] S4. Model validation and evaluation: During the model training process, use multiple stratified k-fold cross-validation. Each time, divide the samples according to different stratification criteria to ensure the stability and reliability of the model under different sample subsets and data distributions.

[0011] As an optimization, step S3 specifically includes:

[0012] S3.1. At the bottom layer, use random forest and lasso regression to perform preliminary feature screening and dimensionality reduction on metabolomics data, and use the variable importance evaluation of random forest and the feature selection ability of lasso regression to remove irrelevant and redundant features;

[0013] S3.2. Input the screened features into a support vector machine and a logistic regression model for training respectively, and perform weighted fusion on the prediction results of these two models in the middle layer, and dynamically adjust the weights according to the performance of different models on the training set;

[0014] S3.3. At the top layer, use a meta-learner to further optimize and integrate the fusion results in the middle layer. Through this multi-level integrated learning, give full play to the advantages of each algorithm and improve the stability and prediction accuracy of the model.

[0015] As an optimization, in step S3, the automated hyperparameter optimization algorithm includes, but is not limited to, grid search, random search combined with Bayesian optimization.

[0016] As an optimization, in step S3.3, the meta-learner uses a gradient boosting decision tree.

[0017] The present invention also provides an application of the early diagnosis biomarker group for endometriosis, mainly in the early disease monitoring of endometriosis and the evaluation of drug treatment effects.

[0018] The beneficial effects of the present invention are:

[0019] 1) Screening and identifying differential metabolites through non-targeted metabolomics of plasma metabolites from endometriosis patients and machine learning methods, and determining candidate biomarkers. A metabolic molecular diagnostic model for endometriosis was established, and the final metabolic biomarkers were determined, successfully solving the problem of early diagnosis of endometriosis and providing an efficient, accurate, and reliable diagnostic tool for clinical practice.

[0020] 2) Multiple algorithms screen features from different perspectives. Random forest evaluates the importance of features based on decision trees, gradually eliminating unimportant features, and LASSO regression further optimizes the screening to avoid the limitations of a single algorithm. Biological knowledge and metabolic network information guide the screening process, preferentially selecting metabolites in metabolic pathways related to endometriosis, making the screened biomarkers more biologically significant and diagnostically valuable, and improving the accuracy and specificity of diagnosis.

[0021] 3) Construct a multi-level ensemble learning model, integrating support vector machine, decision tree, gradient boosting tree, and random forest, etc., and adopting a dynamic weighted ensemble strategy and deep learning framework for data fusion and training. Different base learners have different learning abilities for different features and patterns of data. After integration, they can give full play to their respective advantages to handle complex disease diagnosis problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a result graph of dimensionality reduction analysis of data by PCA in the embodiment of the present invention;

[0023] Figure 2 It is a result graph of partial least squares discriminant analysis (PLS-DA) in the embodiment of the present invention:

[0024] Figure 3 It is a volcano plot of differential metabolites in the embodiment of the invention;

[0025] Figure 4 It is a result graph of KEGG enrichment of differential metabolites in the embodiment of the present invention;

[0026] Figure 5 It is a result graph of correlation analysis of differential metabolites in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0028] Example 1

[0029] The plasma metabolome data of endometriosis was obtained by using the wide-target metabolomics method:

[0030] S1. Sample collection and processing: Plasma samples of patients with endometriosis and healthy controls were collected according to strict standards, and were processed and stored in a standardized manner. Specifically, the subjects of this study were women of childbearing age who visited a certain hospital. The experimental group consisted of women highly suspected of EMT based on clinical symptoms, imaging examinations, and other auxiliary examinations, and were subsequently diagnosed with EMT through laparoscopic examination, biopsy, and postoperative pathological tissue biopsy; plasma of healthy women undergoing physical examinations was selected as the control group. Due to relatively strict inclusion criteria for the research subjects and objective factors such as economic conditions and medical insurance, a total of 30 cases were selected in the experimental group and 30 cases in the control group. The inclusion criteria were as follows: all patients had complete clinical data, aged 25 - 49 years, the experimental group were patients histologically confirmed to have EMT, non-menstrual period, normal body mass index and menstrual cycle, no other diseases or chronic diseases (including diabetes, kidney disease, cardiovascular-related diseases, and inflammatory diseases) in physical and chemical examinations, and no hormone therapy was taken within 6 months before the study.

[0031] Exclusion criteria were: pregnant and lactating women, infectious diseases such as hepatitis B virus, HIV, endocrine diseases such as PCOS, diabetes, thyroid function abnormalities, cardiovascular diseases, dyslipidemia, rheumatic diseases and autoimmune-related diseases such as systemic lupus erythematosus, and hormone-dependent gynecological diseases.

[0032] Plasma sample collection and processing:

[0033] (1) Blood samples were collected from all patients with endometriosis before undergoing surgical treatment;

[0034] (2) All patients were required to have a low-fat diet one day before blood collection to ensure the authenticity of blood sample tests reflecting the in vivo metabolic state;

[0035] (3) Patients fasted for 12 - 14 hours and had their blood drawn on an empty stomach in the morning;

[0036] (4) Avoid strenuous activities before blood collection and collect after resting for 15 minutes;

[0037] (5) 2 - 3 ml of venous blood was collected and placed into an EDTA anticoagulation tube to collect whole blood. The blood collection tube was gently inverted and shaken 4 - 6 times and then vertically placed on the test tube rack;

[0038] (6) The collected blood samples were labeled with name, number, and date;

[0039] (7) The samples were placed in an ice box for storage, transported to the laboratory as soon as possible, centrifuged at 3000 rpm at 4℃ for 10 minutes, the upper plasma was taken, and 0.2 ml / L was aliquoted into 2 ml centrifuge tubes. They were stored in a -80℃ refrigerator for future testing.

[0040] S2. Non-targeted metabolomics analysis of plasma samples:

[0041] ① The plasma sample was thawed on ice, and 100 μL of the thawed plasma sample was placed in an EP tube. An aqueous formaldehyde solution was added in a ratio of 1:4, and after vortexing for 30 min, it was centrifuged at 4 °C and 15,000 g for 20 min to collect the supernatant. Chromatography-mass spectrometry analysis (LC-MS) of the extracted metabolites was performed according to the LC-MC instruction manual.

[0042] ② The metabolite-related data files from LC-MS analysis were imported into the TraceFinder 3.2.0 library search software for preprocessing. Parameters such as retention time and mass-to-charge ratio of each metabolite were screened. The molecular ion peaks and fragment ions were compared with the mzCloud and the locally built database for spectral comparison, and the original quantitative results were standardized. Finally, the identification and relative quantification results of the metabolites were obtained.

[0043] Furthermore, during the metabolite identification process, in addition to comparing with the mzCloud and the locally built database, public metabolomics databases such as HMDB (Human Metabolome Database) and METLIN were introduced for multi-database joint retrieval. At the same time, combined with bioinformatics tools and metabolic pathway databases (such as KEGG, Reactome, etc.), functional annotation and pathway analysis of the identified metabolites were carried out. For example, using gene ontology (GO) enrichment analysis and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis, not only the specific metabolic pathways involved in the metabolites were determined, but also their functional roles in cell physiological processes were analyzed, providing more abundant biological information for subsequent biomarker screening.

[0044] ③ To determine the differential metabolites, the differential metabolites in the two groups of plasma were screened according to the criteria of variable importance of projection (VIP) > 1, P < 0.05, and Fold Change of 1.5. Principal component analysis (PCA) and orthogonal partial least squares discriminant analysis (PLS-DA) were performed on the two groups of differential metabolites, and a volcano plot of the differential metabolites was drawn to visually display the differential metabolites.

[0045] Specifically, such as Figure 1As shown, PCA mainly performs dimensionality reduction analysis on data, and can detect the differences between experimental groups and the repeatability within groups. In a two-dimensional graph, the first two principal components PC1 and PC2 are taken to represent samples. The smaller the spatial distribution difference, the closer the data of the two samples are. Each point in the graph represents an experimental sample, and different groups are distinguished by different colors. For experiments with good repeatability, different samples within the same group should be clustered within a relatively concentrated range and can be distinguished from the data clustering regions of other groups, where Figure 1 the horizontal and vertical coordinates respectively represent the scores of the first and second components, and the points of different colors represent samples of different experimental groups.

[0046] Furthermore, PLS-DA is a classification method most commonly used in current metabolomics data analysis. It combines a regression model while performing dimensionality reduction, and uses a certain discrimination threshold to perform discriminant analysis on the regression results. PLS-DA is a supervised discriminant analysis statistical method, which can establish a relationship model between metabolite expression levels and sample categories to achieve the discrimination and prediction of sample categories. The cross-validation method is used to test the quality of the model. The obtained R2Y and Q2 respectively represent the variables that the model can explain and the predictability, and can be used to judge the quality of the model. The influence intensity and explanatory ability of the expression patterns of each metabolite on the classification and discrimination of each group of samples can be measured by calculating the Variable Importance for the Projection (VIP), so as to assist in the screening of biomarker metabolites (usually using a VIP value > 1.0 as the screening criterion), such as Figure 2 shown, Figure 2 where the colors represent sample groups, each graph represents the position of the metabolome of a sample projected onto a two-dimensional plane after dimensionality reduction, and R2Y and Q2Y are the discriminant abilities of the training set and test set of the model for sample grouping respectively.

[0047] Even further, Figure 3 The volcano plot can intuitively show the overall distribution of differential metabolites. Each point in the volcano plot represents a metabolite. Significantly up-regulated metabolites are represented by red dots, significantly down-regulated metabolites are represented by blue dots, and non-differential metabolites are represented by gray. The size of the dots represents the VIP value. As Figure 3 shown, in Figure 3 the abscissa represents the fold change in expression of metabolites in different groups (log2 Ratio), and the ordinate represents the significance level of the difference (-Log10 Pvalue).

[0048] ④ Enrichment analysis of differential metabolites. The hypergeometric test is used to enrich the screened differential metabolites into the KEGG database for enrichment analysis.

[0049] Specifically, the compound names of the differential metabolites in each set of positive and negative ion modes obtained through statistical analysis are compared with the KEGG database to obtain the pathway results participated by the differential metabolites. According to the screened differential metabolites, the hypergeometric distribution relationship between the differential metabolites and the Pathway is calculated, and whether the differential metabolites are enriched in the corresponding pathway is judged according to the p-value. Through the Pathway analysis of the differential metabolites, the Pathway entries significantly enriched by the differential metabolites can be found, and it can be explored which changes in cellular pathways the differential metabolites of different samples may be related to, such as Figure 4 shown. Each bubble represents a pathway. In Figure 4 , the vertical axis represents the pathway name, and the horizontal axis represents the Rich factor (the ratio of the number of differential proteins in this pathway to the total number of differential metabolites to the ratio of the number of metabolites annotated to this pathway to the total number of metabolites). The larger the enrichment factor, the more significant the enrichment level of the differentially expressed metabolites in this pathway. Each dot in the figure represents a KEGG pathway. The size of the dot represents the number of differential metabolites enriched in this pathway, and the color of the dot represents the p-value. The smaller the p-value, the more reliable the enrichment significance of the differentially expressed metabolites in this pathway.

[0050] Furthermore, a person correlation analysis is performed on the differential metabolites. Red indicates a positive correlation, and blue indicates a negative correlation. As Figure 5 shown, in Figure 5 , red indicates a positive correlation, and blue indicates a negative correlation.

[0051] ⑤ Screening of candidate biomarkers. The ability of the biomarkers in the two groups is judged through the ROC curve. When the AUC is larger, the diagnostic effect of this metabolite on endometriosis is higher. It can be further studied as a potential biomarker.

[0052] Constructing a multi-level integrated learning framework model to screen biomarkers and verify and evaluate:

[0053] S3. Construct a multi-level integrated learning framework model, combine an automated hyperparameter optimization algorithm and an adaptive adjustment strategy to screen biomarkers, and use the automated hyperparameter optimization algorithm to optimize the hyperparameters of each machine learning model; the automated hyperparameter optimization algorithm includes but is not limited to grid search, random search combined with Bayesian optimization. Specifically:

[0054] S3.1. Use random forest and lasso regression at the bottom layer to perform preliminary feature screening and dimensionality reduction on the metabolomics data, and use the variable importance evaluation of the random forest and the feature selection ability of the lasso regression to remove irrelevant and redundant features:

[0055] The random forest algorithm uses the bootstrap sampling method to draw multiple sample subsets from the original metabolomics dataset with replacement to construct multiple decision trees. During the node splitting process of each decision tree, the optimal splitting feature is selected from a randomly selected subset of features. The importance of each feature is evaluated by calculating the average decrease in impurity (Gini index or information gain) of the feature in multiple decision trees. Features with lower importance scores are considered irrelevant or redundant features.

[0056] Suppose the original dataset is D, containing n samples and m features, and T decision trees are constructed. For the t-th decision tree, its construction process is as follows:

[0057] First, randomly draw a sample subset D of size n from D t (sampling with replacement).

[0058] Next, at each node split, randomly select m' features (m' << m) from the m features, and calculate the Gini index of these m' features: where k is the number of classes, and p i is the probability that a sample belongs to the i-th class. Select the feature that maximally decreases the Gini index as the splitting feature.

[0059] Then repeat the above steps until the decision tree reaches the preset maximum depth or the number of samples in a node is less than a certain threshold.

[0060] Then calculate the average decrease in impurity of each feature j in the decision trees to obtain the feature importance score I j .

[0061] Lasso regression simultaneously performs feature selection and parameter estimation by adding an L1 regularization term to the loss function of ordinary linear regression. The loss function is: where y i is the observed value, x ij is the j-th feature value of the i-th sample, β j is the regression coefficient, and λ is the regularization parameter. By adjusting the value of λ, some regression coefficients are shrunk to 0, thereby screening out important features.

[0062] S3.2. Input the screened features into the support vector machine and logistic regression models for training respectively, and perform weighted fusion on the prediction results of these two models at the middle layer, and dynamically adjust the weights according to the performance of different models on the training set; specifically as follows:

[0063] The features selected by random forest and lasso regression are respectively input into the support vector machine (SVM) and logistic regression (LR) models for training. SVM separates samples of different classes by finding an optimal classification hyperplane. For linearly separable problems, its objective function is: The constraint is y i (ω T x i +b)≥1, i = 1, 2,......, n, where ω is the weight vector, b is the bias, and y i is the class label of the sample. For non-linear problems, through the kernel function (such as the radial basis kernel function K(x i ,x j ) = exp(-γ||x i -x j || 2 ), γ is the kernel parameter), the data is mapped to a high-dimensional space, and then the optimal classification hyperplane is found.

[0064] Logistic regression maps the results of linear regression to the interval [0, 1) through the sigmoid function for binary classification problems. The sigmoid function is: where z = ω T x + b. The loss function of logistic regression usually uses logarithmic likelihood loss: The optimal ω and b are solved through optimization algorithms such as gradient descent method or quasi-Newton method.

[0065] In the middle layer, the prediction results of the SVM and LR models are weighted and fused. Let the accuracy of the SVM model on the training set be A CCSVM , and the accuracy of the LR model on the training set be A CCLR , then the weight of the SVM model The weight of the LR model The fused prediction result is: y fusion +ω SVM y SVM +ω LR y LR , where y SVM and y LR are the prediction results of the SVM and LR models respectively.

[0066] Top-level meta-learner optimization: Taking the middle-layer fusion result as the input, the gradient boosting decision tree (GBDT) is used as the meta-learner for further optimization and integration. GBDT iteratively trains multiple weak learners (such as decision trees), and each weak learner fits the residuals of the previous round of the model, gradually improving the prediction ability of the model. Specifically as follows:

[0067] Initialize a constant model where L is the loss function (such as mean squared error)

[0068] For t = 1, 2,......, T (T is the number of iterations):

[0069] Calculate the residual r of the current model it = y i - f t-1 (x i ).

[0070] Train a weak learner h t (x) to fit the residual r it , for example, using a decision tree regression model.

[0071] Update the model f t (x) = f t-1 (x) + λh t (x), where λ is the learning rate used to control the step size of each update.

[0072] The final prediction result is f T (x).

[0073] S4, Model Validation and Evaluation. During the model training process, multiple stratified k-fold cross-validations are adopted. Each time, the samples are divided according to different stratification criteria to ensure the stability and reliability of the model under different sample subsets and data distributions. During the model training process, 10-fold 5-fold cross-validation is used to evaluate the model performance. Each time of cross-validation, the samples are stratified and divided according to disease severity, age level, lifestyle type, etc., and indicators such as accuracy, recall rate, F1 value, and AUC of the model on each fold are calculated. The performance differences under different models and parameter settings are compared through analysis of variance and t-tests, and the optimal model configuration and parameter combination are selected. In addition, samples from 100 endometriosis patients and 100 healthy controls from different hospitals are collected as an external independent validation set to externally validate the finally constructed diagnostic model. At the same time, the model diagnosis results are compared and verified with the clinical diagnosis gold standard of laparoscopy combined with pathological biopsy, and indicators such as diagnostic accuracy, sensitivity, and specificity are calculated.

[0074] In the internal cross-validation, the constructed diagnostic model performs excellently in indicators such as accuracy, recall rate, F1 value, and AUC, and the average accuracy reaches over 95%. In the external independent validation, the diagnostic accuracy of the model reaches 93%, the sensitivity is 92%, and the specificity is 94%. Compared with the clinical diagnosis gold standard, the diagnostic results of the model have high consistency. In terms of disease prognosis prediction, through correlation analysis with clinical follow-up results, it is found that the model can better predict the progression of the disease and has a certain predictive ability for changes in disease severity.

[0075] In summary, for the present invention, blood - centrifugation - plasma extraction is performed in Example 1; the plasma is subjected to metabolomics product analysis. By selecting the experimental group (10 cases) and the control group (10 cases) for non - targeted metabolomics analysis, 878 metabolites are identified in a total of 20 cases. Taking 1.5 - fold as the differential expression change threshold, and taking the statistical test t - test p - value < 0.05 and the supervised discriminant analysis statistical method PLS - DA VIP > 1.0 as the significance threshold, 85 metabolites show significant changes in expression in the comparison between the experimental group and the control group (46 are up - regulated and 39 are down - regulated). Finally, 10 metabolites are selected as follows: sphingomyelin SMd(18:1 / 24:0), sphingomyelin SMd(18:0 / 24:1), phosphatidylcholine PC 16:0 - 17:0, phosphatidylcholine PC 18:0 - 16:0, phosphatidylcholine PC 19:0 - 19:1, phosphatidylcholine PC 20:5e - 18:0, phosphatidylcholine PC 22:6e - 18:5, lysophosphatidylcholine LPC 14:1, lysophosphatidylcholine LPC 22:5, phosphatidylglycerol PG(18:0 / 16:0).

[0076] Example 2

[0077] The application methods of the present invention for the early diagnosis, disease monitoring, and treatment effect evaluation of endometriosis are as follows:

[0078] Early diagnosis

[0079] Collect plasma samples from suspected endometriosis patients and operate according to the sample collection and processing methods described above. Use liquid chromatography - tandem mass spectrometry technology to perform non - targeted metabolomics analysis on the plasma samples. After obtaining metabolite data, detect the expression levels of ten biomarkers such as sphingomyelin SMd(18:1 / 24:0) and sphingomyelin SMd(18:0 / 24:1). Input the expression data of these biomarkers into a trained and validated multi - level integrated learning framework model. The model will comprehensively analyze this data and output a diagnostic result to determine whether the subject has endometriosis.

[0080] Specific application: Patient A, 28 years old, recently experienced worsening dysmenorrhea and dyspareunia, and went to the hospital for treatment. The doctor suspected that she might have endometriosis, so he collected her plasma sample. After sample processing and metabolomics analysis, the level of sphingomyelin SMd (18:1 / 24:0) was detected to be 4nmol / L, which was lower than the lower limit of the normal reference range; the level of phosphatidylcholine PC 16:0-17:0 was 7nmol / L, which was also lower than the normal range. After these biomarker data were input into the diagnostic model, the diagnostic results output by the model showed that the patient was more likely to have endometriosis. Subsequently, the doctor combined other clinical examinations (such as ultrasound examination found suspicious nodules in the pelvic cavity) and finally diagnosed the patient with early endometriosis. Due to the timely diagnosis, the patient was able to receive treatment as soon as possible and effectively controlled the progression of the disease.

[0081] Disease surveillance

[0082] Plasma samples from patients suspected or diagnosed with endometriosis were collected regularly, and the expression levels of ten biomarkers such as sphingomyelin SMd (18:1 / 24:0) and sphingomyelin SMd (18:0 / 24:1) were detected according to the sample processing and non-targeted metabolomics analysis method in Example 1. The test results were compared with the reference range of marker levels in healthy people, and the change trends of these markers were analyzed.

[0083] Specific application: Patient B, 32 years old, had mild dysmenorrhea symptoms and was suspected of endometriosis. At the first test, the level of sphingomyelin SMd (18:1 / 24:0) was slightly lower than the normal range, phosphatidylcholine PC 16:0-17:0 was at the critical value, and other markers were basically normal. The doctor recommended that she undergo a follow-up examination every 3 months. In the subsequent follow-up examination, it was found that sphingomyelin SMd (18:1 / 24:0) continued to decline, and phosphatidylcholine PC 16:0-17:0 also gradually fell below the normal range. Combined with other examinations, it was finally diagnosed as endometriosis, and timely intervention treatment was carried out.

[0084] Evaluation of treatment efficacy

[0085] Before patients receive treatment (such as drug therapy, surgical treatment), plasma samples are collected to detect the initial expression levels of ten biomarkers. During or after treatment, the levels of these markers are retested regularly. If the treatment is effective, the expression level of the biomarker should approach the healthy range; if the treatment is ineffective or the disease recurs, the marker level may remain abnormal or further deviate from the normal range.

[0086] Specific application: Patient C, 35 years old, was diagnosed with endometriosis and received drug treatment. Before treatment, multiple biomarkers were significantly abnormal. For example, the level of lysophosphatidylcholine LPC 22:5 was much higher than the normal range. After 3 months of drug treatment, it was detected that the level of lysophosphatidylcholine LPC 22:5 decreased by 30%, approaching the upper limit of the normal range, and other markers also improved to varying degrees. This indicates that the drug treatment is effective for this patient. Based on this result, the doctor continued to adjust the drug dosage for subsequent treatment. In subsequent follow-up examinations, the markers continued to remain at a level close to normal, and the patient's symptoms were also significantly relieved.

[0087] The above are only characteristic implementation examples of the present invention and do not constitute any limitation to the protection scope of the present invention. Any technical solutions formed by equivalent exchange or equivalent substitution fall within the scope of the present invention's rights protection.

Claims

1. A biomarker panel for early diagnosis of endometriosis, characterized in that: The following 10 metabolites are included: sphingomyelin SMd (18:1 / 24:0), sphingomyelin SMd (18:0 / 24:1), phosphatidylcholine PC 16:0-17:0, phosphatidylcholine PC18:0-16:0, phosphatidylcholine PC 19:0-19:1, phosphatidylcholine PC 20:5e-18:0, phosphatidylcholine PC 22:6e-18:5, lysophosphatidylcholine LPC 14:1, lysophosphatidylcholine LPC 22:5, phosphoglyceride PG (18:0 / 16:0).

2. A method for screening a panel of biomarkers for early diagnosis of endometriosis according to claim 1, characterized in that: The following steps are involved: S1. Sample collection and processing: plasma samples from endometriosis patients and healthy controls were collected according to strict standards and processed and stored in a standardized manner; S2. Non-targeted metabolomics analysis of plasma samples. Liquid chromatography-tandem mass spectrometry was used to perform non-targeted metabolomics analysis of plasma samples, including sample metabolite extraction, data preprocessing, differential metabolite determination, enrichment analysis, and candidate biomarker screening. S3. Build a multi-level integrated learning framework model, combine the automated hyperparameter optimization algorithm and adaptive adjustment strategy to screen biomarkers, and use the automated hyperparameter optimization algorithm to optimize the hyperparameters of each machine learning model; S4. Model validation and evaluation. During the model training process, multiple stratified k-fold cross-validation is used. Each time, the samples are divided according to different stratification criteria to ensure the stability and reliability of the model under different sample subsets and data distributions.

3. The method for screening a panel of biomarkers for early diagnosis of endometriosis according to claim 2, characterized in that: The step S3 specifically includes: S3.

1. Use random forest and lasso regression at the bottom level to perform preliminary feature screening and dimensionality reduction on metabolomics data, and use the variable importance assessment of random forest and the feature selection ability of lasso regression to remove irrelevant and redundant features; S3.2, input the filtered features into the support vector machine and logistic regression model for training respectively, perform weighted fusion on the prediction results of the two models at the middle level, and dynamically adjust the weights according to the performance of different models on the training set; S3.

3. Use a meta-learner at the top level to further optimize and integrate the middle-level fusion results. Through this multi-level integrated learning, the advantages of each algorithm can be fully utilized to improve the stability of the model and the prediction accuracy.

4. The method for screening a panel of biomarkers for early diagnosis of endometriosis according to claim 2, characterized in that: In step S3, the automated hyperparameter optimization algorithm includes but is not limited to grid search and random search combined with Bayesian optimization.

5. The method for screening a panel of biomarkers for early diagnosis of endometriosis according to claim 3, characterized in that: In step S3.3, the meta-learner uses a gradient boosting decision tree.

6. A use of the biomarker panel for early diagnosis of endometriosis according to claim 1, characterized in that: Application in early disease monitoring and treatment effect evaluation of endometriosis.

Citation Information

Patent Citations

  • Metabolic marker for auxiliary diagnosis of schizophrenia, construction method and application

    CN119534874A

  • Methods for detecting endometriosis

    WO2010107734A2