A depression disease management system based on targeted metabolomics and machine learning models
Through a depression disease management system based on targeted metabolomics and machine learning models, blood samples are collected to detect a combination of metabolic molecular markers, and the optimal model is screened to identify and evaluate patients with depression, providing accurate diagnosis and treatment recommendations.
Patent Information
- Application Number
- CN202411781535.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-12-05
Smart Images

Figure CN119724545B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of disease management systems, and in particular, to a depression disease management system based on targeted metabolomics and machine learning models. Background Art
[0002] Depression seriously affects patients' lives and work, and brings a heavy burden to their families and society.
[0003] A prominent problem in the diagnosis and treatment of depression is the low recognition rate of depression in the medical system. This is mainly because the clinical diagnosis of depression is mainly based on medical history, clinical symptoms, and course of the disease.
[0004] In addition, another prominent problem in the diagnosis and treatment of depression is the current lack of a complete diagnosis and treatment system for depression. Summary of the Invention
[0005] In response to the above-mentioned prominent problems in the diagnosis and treatment of depression in the current medical system, the present invention aims to provide a combination of metabolic molecular markers involved in the full course management of depression and a depression disease management system.
[0006] The above technical objectives of the present invention are achieved through the following technical solutions: a depression disease management system based on targeted metabolomics and machine learning models, comprising: an intelligent diagnosis model module, a differential diagnosis model module, a hierarchical diagnosis model module, a companion diagnosis model module, a treatment endpoint outcome prediction model module, a relapse prediction model module, and a comprehensive judgment module;
[0007] Among them, the intelligent diagnosis model module is used to identify patients with depression; the differential diagnosis model module is used to distinguish bipolar and unipolar depression; the graded diagnosis model module is used to distinguish mild depression from moderate to severe depression; the companion diagnosis model module is used to evaluate whether the treatment of depression patients is effective; the treatment endpoint outcome prediction model module is used to make intelligent predictions of the patient's clinical treatment outcomes and provide a basis for optimizing treatment plans; the recurrence prediction model module is used to predict disease recurrence in patients in a stable period and take timely intervention measures; the comprehensive judgment module is used to integrate the results of the diagnosis, identification and prediction models of the remaining modules to construct a depression disease management system.
[0008] Furthermore, the method for using the depression disease management system includes the following steps:
[0009] S1. Build an intelligent diagnostic model through the intelligent diagnostic model module to identify patients with depression;
[0010] S2. Construct a differential diagnosis model using the differential diagnosis model module to differentiate between bipolar and unipolar depression;
[0011] S3. Construct a hierarchical diagnostic model using the hierarchical diagnostic model module to differentiate between mild depression and moderate to severe depression;
[0012] S4. Build a companion diagnostic model using the companion diagnostic model module to assess whether treatment for depression is effective.
[0013] S5. Construct a treatment endpoint outcome prediction model through the treatment endpoint outcome prediction model module to intelligently predict the patient's clinical treatment outcome and provide a basis for optimizing treatment plans;
[0014] S6. Build a recurrence prediction model through the recurrence prediction model module to predict disease recurrence in patients in the stable phase and take timely intervention measures;
[0015] S7. Build a depression disease management system by integrating various diagnosis, identification and prediction models through a comprehensive judgment module.
[0016] Furthermore, the specific operation of constructing the intelligent diagnostic model through the intelligent diagnostic model module in S1 is: collecting blood samples and detecting the following metabolic molecular marker combinations: proline, betaine, alanine, tryptophan, kynurenine, 5-hydroxytryptamine, creatine, succinate, taurine, and 2-hydroxybutyrate, and constructing different machine learning models to screen the metabolic marker combination under the best model for identifying patients with depression.
[0017] Furthermore, the specific operation of constructing the differential diagnosis model through the differential diagnosis model module in S2 is: collecting blood samples and detecting the following combination of metabolic molecular markers: proline, ornithine, kynurenine, 5-hydroxytryptamine, succinic acid, 2-hydroxybutyric acid, xanthine, bilirubin, γ-aminobutyric acid, allantoin, deoxyribose, pyruvic acid, linoleic acid and uric acid, constructing different machine learning models, and screening the optimal model for distinguishing patients with unipolar and bipolar depression.
[0018] Furthermore, the specific operation of constructing a graded diagnostic model through the graded diagnostic model module in S3 is: collecting blood samples and detecting the following combination of metabolic molecular markers: proline, valine, ornithine, tryptophan, 5-hydroxytryptamine, creatine, glutamine, serine, methionine, guanine, hypoxanthine, androstenedione, cytosine, bilirubin, γ-aminobutyric acid, uric acid, adipic acid, and pseudouridine, and constructing different machine learning models to screen the optimal model for distinguishing mild depression from moderate to severe depression.
[0019] Furthermore, the specific operation of constructing the companion diagnostic model through the companion diagnostic model module in S4 is: collecting blood samples and detecting the following metabolic molecular marker combination: ornithine, betaine, alanine, kynurenine, 5-hydroxytryptamine, glutamate, lysine, methionine, phenylalanine, guanine, hypoxanthine, γ-aminobutyric acid, allantoin and uric acid, constructing different machine learning models, screening the optimal model, and using it to evaluate whether the treatment of depression patients is effective, and helping to screen treatment drug regimens for patients with different disease courses.
[0020] Furthermore, the specific operation of constructing the treatment endpoint outcome prediction model through the treatment endpoint outcome prediction model module in S5 is: tracking the treatment status and recovery status of the depressed patients undergoing treatment, and detecting the following metabolic molecular marker combinations at different stages: valine, betaine, tryptophan, kynurenine, creatine, taurine, phenylalanine, tyrosine, histidine, aspartic acid, threonine, guanine, γ-aminobutyric acid, allantoin and pseudouridine nucleoside, constructing different machine learning models, screening the best model, and using it to make intelligent predictions of the patient's clinical treatment outcomes and provide a basis for optimizing the treatment plan.
[0021] Furthermore, the specific operation of constructing the relapse prediction model through the relapse prediction model module in S6 is: tracking the condition of patients with depression, and detecting the following metabolic molecular marker combinations at different stages: proline, ornithine, alanine, tryptophan, kynurenine, creatine, 2-hydroxybutyric acid, androstenedione, cytosine, xanthine, bilirubin, γ-aminobutyric acid, allantoin, deoxyribose, pyruvic acid and linoleic acid. By constructing different machine learning models, the best model is screened for predicting disease relapse in patients in the stable period and taking timely intervention measures.
[0022] Furthermore, the specific operations of integrating various diagnostic, differential and predictive models through the comprehensive judgment module in S7 are: by tracking the progression of depression patients' illness, collecting peripheral blood metabolite concentration data at different stages, and using machine learning algorithms, developing a model system for depression diagnosis, graded diagnosis, differential diagnosis, concomitant diagnosis, treatment outcome prediction and stable period relapse prediction based on multiple metabolic molecular markers, wherein the disease management system provides an interactive interface that is easy for users to operate based on peripheral blood metabolite expression concentration data, and outputs corresponding results and disease management decision recommendations according to different purposes.
[0023] Furthermore, the machine learning models include: logistic regression, lasso regression, decision tree, neural network, support vector machine, extreme gradient boosting, random forest, principal component analysis, Bayesian network, and linear regression.
[0024] Furthermore, the sample for detecting the metabolic molecular marker combination is human serum, plasma, or dried blood spots.
[0025] In summary, the marker combination and management system provided by the present invention can achieve early screening prediction of depression, prediction of disease treatment outcomes, and prediction of disease recurrence during the stable period. The technical solution of the present invention is used to diagnose depression and provide medication guidance with high specificity, sensitivity, and accuracy. The present invention can make the diagnosis of depression no longer rely entirely on the subjective judgment of clinicians' experience and questionnaire scales, greatly improving the diagnostic accuracy. The present invention has the following beneficial effects:
[0026] (1) The depression disease management system developed by the present invention based on targeted metabolomics and machine learning models can solve the current clinical problems in diagnosis, treatment outcome and recurrence, achieve accurate diagnosis, timely treatment, and improve the mental health level of the whole population.
[0027] (2) The depression disease management system based on targeted metabolomics and machine learning models can elevate the diagnosis and treatment of depression to a new level.
[0028] (3) Widely apply big data information analysis technologies such as artificial intelligence, develop real-time updated and iterative artificial intelligence platforms, and use depression disease cohorts under multi-group samples and multi-time point diagnosis and treatment data to deeply reveal the etiology and pathological mechanisms of depression, thereby making pioneering contributions to the establishment of an objective diagnosis, precise treatment and evaluation system for depression, and developing new depression intervention technologies and means. This field is currently at the forefront of research at home and abroad.
[0029] (4) From the perspective of adolescents, the study aims to clarify the factors and mechanisms that influence the changes in the incidence of depression through gene-environment interaction, provide real-time data support for the country to formulate depression prevention and treatment policies, and meet the urgent needs of the country for the prevention and treatment of adolescent depression. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a general flow chart of an embodiment of the present invention.
[0031] Figure 2 Schematic diagram of the intelligent model construction in an embodiment of the present invention.
[0032] Figure 3 Receiver operating characteristic (ROC) curve for distinguishing healthy subjects from patients with depression using a generalized linear model (GLM).
[0033] Figure 4 Receiver operating characteristic (ROC) curve for logistic regression in differentiating patients with unipolar and bipolar depression.
[0034] Figure 5ROC curve for Naive Bayes to distinguish between mild and moderate to severe depression patients.
[0035] Figure 6 The ROC curve of random forest (RF) for distinguishing patients with depression who have responded to treatment and those who have not seen significant effects.
[0036] Figure 7 Receiver operating characteristic (ROC) curve for the artificial neural network (NNET) in distinguishing patients with depression who had treatment outcome from those who did not.
[0037] Figure 8 ROC curve for support vector machine (SVM) to distinguish between patients with stable relapse and those without relapse of depression. DETAILED DESCRIPTION
[0038] The following is combined with Figure 1-8 The present invention is described in further detail.
[0039] Example 1 Metabolic marker screening and model establishment for depression diagnosis
[0040] 1.1 Experimental subjects
[0041] According to the following inclusion and exclusion criteria, depression samples and healthy control samples were constructed respectively.
[0042] Inclusion criteria for depression samples: (1) All included depression patients met the ICD-10 diagnostic criteria; (2) Hamilton Depression Rating Scale (HDRS) 17-item score greater than or equal to 18 points; (3) Patients with first-episode depression who were not taking any antidepressants; (4) Aged 18 years and younger than 60 years. Exclusion criteria: (1) Patients with other mental illnesses or a history of other mental illnesses; (2) Patients with organic brain diseases and a history of severe brain trauma, or with heart, liver, or kidney diseases, diabetes, or other serious physical diseases; (3) Patients with abnormalities in routine laboratory tests (blood routine, liver function, and urine routine); (4) Female research subjects who were pregnant, breastfeeding, or menstruating; (5) Patients with a history of drug and substance abuse.
[0043] The inclusion criteria for healthy controls were: no history of neuropsychiatric disease, no history of drug abuse or dependence, no systemic physical disease; and no obvious abnormalities in routine laboratory tests.
[0044] 1.2 Dried blood spot preparation and biomarker extraction
[0045] (1) Blood samples were collected from the included depression samples and healthy control samples, and dried blood spot samples were prepared. The specific process was as follows:
[0046] Prepare dried blood spot cards: Cards should be clean, free of contamination and mold, and sealed before use. Each card should have unique identifying information (name, age, gender, and code). The card has a blood collection area. The dried blood spot card is divided into two parts: the first is the blood collection area, which has an inner circle with a radius of 3 mm and an outer circle with a radius of 5 mm. This area is prohibited from contact. The second is the holding area, which requires wearing clean medical gloves. Dried blood spot cards contain antioxidants (VC) and the enzyme inactivator 1-aminobenzotriazole (ABT).
[0047] Blood sampling using capillary blood and venous blood: Collect blood samples from the fingertip capillaries, clean the fingertips in advance, and do not use non-volatile disinfectants such as iodine tincture. Then drop a drop of blood in each inner circle of the dry blood spot card to ensure that the entire circle is covered and the filter paper is completely penetrated, while avoiding blood from penetrating the outer circle. Avoid strong light during the collection process to avoid affecting photosensitive substances. After collecting the blood spot, avoid prolonged exposure to the air, and use nitrogen blowing or vacuum instruments to dry it quickly. The collected dry blood spots are quickly cooled in liquid nitrogen to inactivate the enzyme, and then dried. The cards cannot be stacked. After being divided into separate sealed bags, they are placed in a -80°C freezer.
[0048] (2) Use a punch to cut a 10 mm diameter disc along the outer circle line in the blood collection area, place it in a 1.5 ml centrifuge tube, and then add 500 ml of eluent / precipitant containing internal standard. Vortex for about 5 minutes. Centrifuge at 15000g for 10 minutes, take the supernatant, perform a second centrifugation (or filter with a filter membrane), and take the supernatant for sampling. Then add 500 μL of the methanol-water solution (metabolite extraction solvent) in the depression detection kit and shake at 20°C and 1450 rpm for 45 minutes;
[0049] (3) Take 10 μL of the extract and add 90 μL of water. Then add 200 μL of the internal standard working solution, 50 μL of 1 M NaHCO3 solution, and 100 μL of a 1 mg / mL derivatization reagent in sequence. Oscillate at 30°C and 1450 rpm for 30 minutes.
[0050] 1.3 Detection of biomarkers
[0051] The targeted metabolomics analysis method based on liquid chromatography-tandem mass spectrometry (LC-MS / MS) detects metabolites including: methionine, phenylalanine, guanine, hypoxanthine, γ-aminobutyric acid, allantoin, uric acid, histidine, aspartic acid, threonine, androstenedione, cytosine, xanthine, bilirubin, pyruvate, linoleic acid, glutamine, serine, adipic acid, pseudouridine, tyrosine, proline, valine, ornithine, betaine, alanine, tryptophan, leucine, kynurenine, 5-hydroxytryptamine, creatine, succinate, taurine, 2-hydroxybutyrate, glutamate, glucose, lysine, and arginine. The specific steps are:
[0052] (1) Sample preparation: Take the sample to be analyzed and ensure that the sample is within the required concentration range. Dilute the sample as needed (usually with mobile phase). Filter the sample to remove possible particulate matter, usually using a 0.45 μm or 0.2 μm filter membrane.
[0053] (2) Select mobile phase: inject the metabolic extract into the Extend C18 column (ZORBAXRR Extend-C18, 80A, 4.6×150mm, 3.5μm, USA) was used to separate metabolites in blood samples. The specific liquid chromatography conditions were: injection volume: 5μL; flow rate: 0.6mL / min; column temperature: 40°C; autosampler temperature: 4°C; mobile phase A: 0.1% (volume ratio) formic acid in water; mobile phase B: methanol; elution program was 0-1min, linear change from 10% B phase to 40% B; 1-1.5min, linear change from 40% B phase to 50% B phase; 1.5min-2.0min, linear change from 50% B phase to 60% B phase; 2.0min-3.2min, equilibrium at 80% B phase; 3.2min-3.6min, from 80% B phase to 10% B phase; 3.7-4.5min, equilibrium at 10% B phase.
[0054] (3) Equipment setup: Check all parts of the HPLC system, including the pump, injector, column, and detector. Ensure that they are all in normal working order. Introduce the chromatographically separated metabolites into the triple quadrupole mass spectrometer and scan and detect the metabolites using the multiple reaction detection mode. Use the electrospray ionization source (ESI) with instantaneous positive and negative ion switching mode, and the scan time is 0 min to 4.5 min. CUR: 20; CAD: Medium; IS: 5000; TEM: 500; GS1: 30; GS2: 50; Interface heater (ihe): on; DP: ±36; CE: ±26; EP: ±15; CXP: ±10. Dwell time: 20 ms; Resolution Q1: Unit; Resolution Q3: Unit; Pause between mass: 5.007 ms.
[0055] (4) Sample injection: Use the injector to introduce the prepared sample into the system. The typical injection volume is 10 to 100 μL. Ensure that the sample liquid flows smoothly and record the injection time.
[0056] (5) After data collection, the metabolic marker of each diagnostic marker is quantified using a specific ion pair, the peak area of the collected signal is compared with the working curve of the corresponding standard, and the internal standard solution is added for correction to obtain the concentration value of the metabolic marker and the content of each metabolic marker in the original blood.
[0057] (6) Quantitative determination of metabolites in the dried blood spot paper extract.
[0058] 1.4 Stability test
[0059] After drying, the dried blood spot discs were stored at -80°C, -20°C, and room temperature. Liquid chromatography-mass spectrometry was used to determine the content of substances in the dried blood spots. Samples stored at -80°C were measured monthly, samples stored at -20°C were measured biweekly, and samples stored at room temperature were measured every other day. The relative standard deviation (RSD) was calculated, and the recovery rate was calculated using the initial mean value as the initial value. The results showed that the precision was within 10% and the recovery rate was between 80% and 110%. This indicates that the dried blood spot discs can be stored stably for at least one year at -80°C, at least three months at -20°C, and for 15 days at room temperature.
[0060] 1.5 Test results
[0061] 150 normal human dried blood spot samples were collected, and the metabolites in the dried blood spot samples were quantified to determine the normal reference range of the markers. The chromatogram of the metabolic markers is shown in Figure 2. Figure 1During the analysis of dried blood spot samples, a QC sample was interspersed with every 10 study samples to assess the variability of the derivatization process and instrumental analysis. QC samples are crucial to ensure the reproducibility, reliability, accuracy, and robustness of metabolite quantification.
[0062] At the same time, samples from 150 healthy people and 60 patients with depression were collected and tested. The sample test results showed that the differences in the levels of the following metabolites can well distinguish normal people from patients with depression: proline, betaine, alanine, tryptophan, kynurenine, 5-hydroxytryptamine, creatine, succinate, taurine, and 2-hydroxybutyric acid (Table 1). Principal component analysis (PCA) was performed based on the measured plasma metabolites. The results showed that the QC samples were closely clustered relative to other samples of dried blood spots, indicating that the method of this embodiment has good repeatability. The QC sample here is BSA (bovine serum albumin solution). Orthogonal partial least squares-discriminant analysis (OPLS-DA) analysis was performed, and the results showed that normal and depression samples could be well distinguished, indicating that there are significant differences in metabolism between patients with depression and normal people.
[0063] Table 1. Differentially expressed metabolites in depression group compared with healthy group
[0064] metabolites People with depression Healthy people FC value VIP value P-value Proline (μmol / L) 81.48±15.57 98.71±16.94 0.83 1.56 <0.001 5-HT (ng / mL) 43.00±6.71 100.94±9.10 0.43 5.61 <0.001 Creatine (mg / dL) 0.95±0.23 0.81±0.21 1.17 3.15 <0.001 Betaine (ng / mL) 5674±1840 5017±1679 1.13 2.42 0.0175 Alanine (μmol / L) 328.37±105.55 375.79±88.67 0.87 3.71 0.0024 Tryptophan (μmol / L) 56.11±6.87 89.48±9.41 0.63 4.11 <0.001 Kynurenine (μmol / L) 2.75±0.82 2.23±0.45 1.23 1.87 <0.001 Succinic acid (ng / mL) 334.02±100.8 519.31±141.87 0.64 4.30 <0.001 Taurine (ng / mL) 14451±3199 11814±4429 1.22 1.13 <0.001 2-Hydroxybutyric acid (ng / mL) 6490±2624 4946±1883 1.31 1.91 <0.001
[0065] 1.6 Machine Learning and Feature Selection
[0066] (1) Model construction method
[0067] The inventors aimed to use machine learning techniques to identify metabolite signatures and predict physiological states based on metabolomic data, thereby better differentiating patients with depression. They used machine learning methods including generalized linear models, linear discriminant analysis, K-nearest neighbor, logistic regression, lasso regression, decision trees, artificial neural networks, support vector machines, extreme gradient boosting, random forests, principal component analysis, Bayesian networks, and linear regression to identify metabolite biomarkers for differentiating patients with depression. A DBB strategy was employed to compensate for missing values. Receiver operating characteristic (ROC) and balanced accuracy were calculated to evaluate model performance. ROC curves were constructed using the pROC software package, and area under the curve (AUC) values were calculated. The ROC probability statistic was calculated based on the prediction score formula derived from PLS analysis. ROC curves for the entire metabolite profile allowed comparison of different models based on their corresponding AUC values. Furthermore, variable importance in projection (VIP) values, which reflect the relative importance of each metabolite in the projected model, were analyzed. A total of 10 metabolites were identified through a combination of machine learning (ML) and bioinformatics. To verify whether the selected metabolites were good classifiers capable of distinguishing different groups, the ROC curve was quantified to evaluate their diagnostic performance.
[0068] (2) Analysis of experimental results
[0069] Machine learning was used to identify plasma biomarkers to distinguish patients with depression, to screen out potential specific biomarkers to better distinguish patients with depression, and to find that the symptoms of these diseases were correlated to some extent with objective clinical scales. The inventors used machine learning to identify promising metabolite markers and predict physiological states based on proteomic data. First, different ML methods for classification were constructed (Table 2), and the performance of different models was evaluated by analyzing the confusion matrix and ROC curve. The results showed that among the 13 machine learning methods, different algorithms were further used to construct models. The best model was the generalized linear model (GLM), and the receiver operating characteristic curve (ROC) of the significantly changed metabolites was used. The area under the curve (AUC) value was 1, the accuracy was 99.30%, the sensitivity was 100%, and the specificity was 98.96%. Figure 3 The model showed good diagnostic value. The combination of metabolites proline, betaine, alanine, tryptophan, kynurenine, 5-hydroxytryptamine, creatine, succinate, taurine, and 2-hydroxybutyrate showed excellent specificity and sensitivity.
[0070] Table 2. Comparative analysis of the results of different models for the diagnosis of depression
[0071] Model AUC value Accuracy (%) Sensitivity (%) Specificity (%) Generalized Linear Models 1.000 99.30% 100.00% 98.96% Linear Discriminant Analysis 0.9982 98.60% 97.92% 100.00% K-nearest neighbor algorithm 0.9576 94.41% 97.92% 87.23% Logistic regression 0.9976 97.90% 97.92% 97.87% Lasso regression 0.9981 97.20% 98.96% 93.62% Decision Tree 0.9983 97.20% 96.87% 97.87% Artificial Neural Networks 0.9749 93.71% 93.75% 93.62% Support Vector Machine 0.9978 97.20% 97.92% 95.74% Extreme Gradient Boosting 0.9989 99.30% 98.96% 100.00% Random Forest 0.9996 97.90% 96.87% 100.00% Principal component analysis 0.9980 97.90% 97.92% 97.87% Bayesian Network 0.9844 97.90% 96.87% 100.00% Linear regression 0.9953 99.30% 98.96% 100.00%
[0072] Experimental Example 1 Evaluation of the Diagnostic Effect of the Metabolic Marker Kit and Diagnostic Model Established in Example 1 on Depression
[0073] Dried blood spot samples from outpatients or inpatients of a hospital, as well as dried blood spot samples from recruited healthy controls, were randomly collected (a total of 300 samples; all samples had clear diagnostic data collected by professional physicians according to ICD-10 standards, and informed consent was signed). The ion abundance ratios of biomarkers (proline, betaine, alanine, tryptophan, kynurenine, 5-hydroxytryptamine, creatine, succinate, taurine, and 2-hydroxybutyrate) and internal standards were directly detected according to the method of Example 1. The data were input into the diagnostic model established in Example 1 to obtain diagnostic results. The diagnostic results were compared with the results of depression diagnosis performed by professional physicians on the 300 samples using ICD-10. The results are shown in Table 3.
[0074] Table 3. Evaluation of the effectiveness of metabolic marker kits in diagnosing depression
[0075]
[0076] As can be seen from Table 3, the marker combination of Example 1 and the diagnostic model established therefrom have an accuracy of 100%, a sensitivity of 100%, and a specificity of 100% for diagnosing depression.
[0077] Example 2 Metabolic marker screening and model establishment for differential diagnosis of unipolar depression and bipolar depression
[0078] The same method as in Example 1 was used to identify patients with unipolar and bipolar depression. The test results were as follows:
[0079] Samples from 150 patients with unipolar depression and 60 patients with bipolar depression were collected and tested. The results of the sample tests showed that the differences in the levels of the following metabolites can well distinguish patients with unipolar depression from patients with bipolar depression: proline, ornithine, kynurenine, 5-hydroxytryptamine, succinic acid, 2-hydroxybutyric acid, xanthine, bilirubin, γ-aminobutyric acid, allantoin, deoxyribose, pyruvic acid, linoleic acid and uric acid (Table 4). Different algorithms were further used to construct models (Table 5). The best model was based on logistic regression. The significantly changed metabolites were further used to make a receiver operating characteristic curve (ROC). The area under the curve (AUC) value was 1, the accuracy was 100%, the sensitivity was 100%, and the specificity was 100%. Figure 4 The results showed that the model has good differential diagnostic value. The combination of metabolites proline, ornithine, kynurenine, 5-hydroxytryptamine, succinate, 2-hydroxybutyrate, xanthine, bilirubin, γ-aminobutyric acid, allantoin, deoxyribose, pyruvate, linoleic acid, and uric acid showed excellent specificity and sensitivity.
[0080] Table 4. Differentially expressed metabolites in bipolar depression group compared with unipolar depression group
[0081]
[0082]
[0083] Table 5. Comparative analysis of the results of different models for the differential diagnosis of unipolar depression and bipolar depression
[0084] Model AUC value Accuracy (%) Sensitivity (%) Specificity (%) Generalized Linear Models 0.9896 96.50% 97.92% 93.62% Linear Discriminant Analysis 0.9991 99.30% 98.96% 100.00% K-nearest neighbor algorithm 0.9996 99.30% 98.96% 100.00% Logistic regression 1.000 100.00% 100.00% 100.00% Lasso regression 0.9856 96.50% 97.92% 93.62% Decision Tree 0.9531 94.40% 98.96% 85.11% Artificial Neural Networks 0.9894 96.50% 97.92% 93.62% Support Vector Machine 0.9802 95.80% 100.00% 87.23% Extreme Gradient Boosting 0.9880 95.10% 97.92% 89.36% Random Forest 0.9807 95.10% 97.92% 89.36% Principal component analysis 0.9957 95.80% 98.96% 89.36% Bayesian Network 0.9914 98.60% 100.00% 95.74% Linear regression 0.9938 95.80% 98.96% 89.36%
[0085] Experimental Example 2 Evaluation of the Differential Diagnostic Effect of the Metabolic Marker Combination and Diagnostic Model Established in Example 2 on Unipolar Depression and Bipolar Depression
[0086] Dried blood spot samples were randomly collected from outpatients and inpatients at a hospital, including 220 patients initially diagnosed with monophasic or biphasic disease. All samples were analyzed by a professional physician based on their medical history and informed consent was obtained. The ion abundance ratios of the biomarker panel (proline, ornithine, kynurenine, 5-hydroxytryptamine, succinate, 2-hydroxybutyrate, xanthine, bilirubin, γ-aminobutyric acid, allantoin, deoxyribose, pyruvate, linoleic acid, and uric acid) to the internal standard were directly measured using the method of Example 2. This data was input into the diagnostic model established in Example 2 to obtain diagnostic results, which were then compared with the differential diagnosis results of the professional physicians. The results are shown in Table 6.
[0087] Table 6. Evaluation of the effectiveness of metabolic marker kits in differential diagnosis of unipolar depression and bipolar depression
[0088]
[0089] As can be seen from Table 6, the accuracy of differentiating unipolar depression from bipolar depression using the marker combination of Example 2 and the diagnostic model established therefrom is 97.27%, the sensitivity is 97.10%, and the specificity is 97.56%.
[0090] Example 3 Metabolic marker screening and model establishment for graded diagnosis of mild and moderate to severe depression
[0091] The same method as in Example 1 was used to identify patients with mild, moderate or severe depression. The test results were as follows:
[0092] Samples from 120 patients with mild depression and 50 patients with moderate to severe depression were collected and tested. The results of the sample tests showed that the differences in the levels of the following metabolites can well distinguish patients with mild from those with moderate to severe depression: proline, valine, ornithine, tryptophan, 5-hydroxytryptamine, creatine, glutamine, serine, methionine, guanine, hypoxanthine, androstenedione, cytosine, bilirubin, γ-aminobutyric acid, uric acid, adipic acid, and pseudouridine (Table 7). Different algorithms were further used to construct models (Table 8). The best model was naive Bayes. The receiver operating characteristic curve (ROC) was further constructed using the significantly changed metabolites. The area under the curve (AUC) value was 0.98, with an accuracy of 98.60%, a sensitivity of 97.87%, and a specificity of 98.96%. See Figure 5 The model showed good hierarchical diagnostic value. The combination of metabolites proline, valine, ornithine, tryptophan, 5-hydroxytryptamine, creatine, glutamine, serine, methionine, guanine, hypoxanthine, androstenedione, cytosine, bilirubin, γ-aminobutyric acid, uric acid, adipic acid, and pseudouridine showed excellent specificity and sensitivity.
[0093] Table 7. Differentially expressed metabolites between mild and moderate to severe depression groups
[0094]
[0095]
[0096] Table 8. Comparative analysis of the results of different models for the graded diagnosis of mild and moderate to severe depression
[0097] Model AUC value Accuracy (%) Sensitivity (%) Specificity (%) Generalized Linear Models 0.9395 88.11% 92.71% 78.72% Linear Discriminant Analysis 0.8867 83.22% 83.33% 82.98% K-nearest neighbor algorithm 0.9357 86.71% 92.71% 74.47% Logistic regression 0.9153 90.91% 94.79% 82.98% Lasso regression 0.8991 81.82% 86.46% 72.34% Decision Tree 0.9339 88.11% 90.62% 82.98% Artificial Neural Networks 0.9388 90.21% 92.71% 85.11% Support Vector Machine 0.9457 89.51% 90.62% 87.23% Extreme Gradient Boosting 0.9641 90.21% 90.62% 89.36% Random Forest 0.9386 86.71% 92.71% 74.49% Principal component analysis 0.8204 80.42% 84.37% 72.34% Bayesian Network 0.9800 98.60% 97.87% 98.96% Linear regression 0.9058 82.52% 88.54% 70.21%
[0098] Experimental Example 3 Evaluation of the diagnostic effect of the metabolic marker combination and diagnostic model established in Example 3 on the graded diagnosis of mild and moderate to severe depression
[0099] Dried blood spot samples from outpatients or inpatients of a hospital were randomly collected, including 501 samples. All samples were collected with clear graded diagnosis data of professional doctors based on the ICD-10 standard, and informed consent was signed. The ion abundance ratios of biomarkers (proline, valine, ornithine, tryptophan, 5-hydroxytryptamine, creatine, glutamine, serine, methionine, guanine, hypoxanthine, androstenedione, cytosine, bilirubin, γ-aminobutyric acid, uric acid, adipic acid, and pseudouridine) and internal standards were directly detected according to the method of Example 3. The data were input into the diagnostic model established in Example 3 to obtain graded diagnostic results. The diagnostic results were compared with the graded diagnostic results of professional doctors based on the ICD-10 standard. The results are shown in Table 9.
[0100] Table 9 Evaluation of the effectiveness of metabolic marker kits in grading the diagnosis of mild and moderate to severe depression
[0101]
[0102] As can be seen from Table 9, the marker combination of Example 3 and the hierarchical diagnostic model established therefrom have an accuracy of 100%, a sensitivity of 100%, and a specificity of 100% in distinguishing mild depression from moderate to severe depression.
[0103] Example 4 Metabolic marker screening and model establishment for evaluating whether treatment of depression patients is effective
[0104] The same method as in Example 1 was used to evaluate whether the treatment of depression patients had a significant effect. The test results were as follows:
[0105] Samples from 140 patients who responded to treatment and 60 patients who did not show significant effects were collected and tested. The results of the sample tests showed that the differences in the levels of the following metabolites can well distinguish patients who responded to treatment from those who did not show significant effects: ornithine, betaine, alanine, kynurenine, 5-hydroxytryptamine, glutamic acid, lysine, methionine, phenylalanine, guanine, hypoxanthine, γ-aminobutyric acid, allantoin and uric acid (Table 10). Different algorithms were further used to construct models (Table 11). The best model was random forest (RF). The receiver operating characteristic curve (ROC) was further drawn using the metabolites with significant changes. The area under the curve (AUC) value was 1, the accuracy was 100%, the sensitivity was 100%, and the specificity was 100%. Figure 6 The model showed good companion diagnostic value. The combination of metabolites ornithine, betaine, alanine, kynurenine, 5-hydroxytryptamine, glutamate, lysine, methionine, phenylalanine, guanine, hypoxanthine, γ-aminobutyric acid, allantoin, and uric acid showed excellent specificity and sensitivity.
[0106] Table 10. Differentially expressed metabolites between the effective and non-effective depression groups
[0107]
[0108]
[0109] Table 11. Comparative analysis of the results of different models for evaluating the effectiveness of depression treatment
[0110] Model AUC value Accuracy (%) Sensitivity (%) Specificity (%) Generalized Linear Models 0.9969 97.90% 98.96% 95.74% Linear Discriminant Analysis 0.9576 94.40% 97.92% 87.23% K-nearest neighbor algorithm 0.99649 96.50% 98.96% 91.49% Logistic regression 0.99809 97.20% 98.96% 93.62% Lasso regression 0.9895 96.50% 98.96% 91.49% Decision Tree 0.9976 96.50% 98.96% 91.49% Artificial Neural Networks 0.9946 95.10% 96.87% 91.49% Support Vector Machine 0.9978 97.90% 98.96% 95.74% Extreme Gradient Boosting 0.9980 97.90% 98.96% 95.74% Random Forest 1.000 100.00% 100.00% 100.00% Principal component analysis 0.9969 97.90% 98.96% 95.74% Bayesian Network 0.9078 83.20% 89.58% 70.21% Linear regression 0.9302 83.22% 95.83% 57.45%
[0111] Experimental Example 4: Evaluation of the Concomitant Diagnostic Effect of the Diagnostic Model Established in Example 4 on the Effectiveness of Treatment for Depression Patients
[0112] Dried blood spot samples from outpatients or inpatients of a hospital were randomly collected. All samples were evaluated using the scale method and signed informed consent. The ion abundance ratios of biomarkers (ornithine, betaine, alanine, kynurenine, 5-hydroxytryptamine, glutamic acid, lysine, methionine, phenylalanine, guanine, hypoxanthine, γ-aminobutyric acid, allantoin, and uric acid) and internal standards were directly detected according to the method of Example 4. The data were input into the prediction model established in Example 4 to obtain companion diagnostic results. The diagnostic results were compared and analyzed with the companion diagnostic results of professional doctors. The results are shown in Table 12.
[0113] Table 12. Metabolic marker kit evaluation of whether treatment for depression is effective
[0114]
[0115] As can be seen from Table 12, the marker combination of Example 4 and the companion diagnostic model established therefrom have an accuracy of 95.14%, a sensitivity of 96.40%, and a specificity of 92.00% in assessing whether depression treatment is effective.
[0116] Example 5 Metabolic marker screening and model establishment for predicting treatment outcomes in patients with depression
[0117] The same method as in Example 1 was used to predict the treatment outcome of depression patients. The test results were as follows:
[0118] Samples from 130 patients with depression who had treatment outcomes and 50 patients with depression who had no treatment outcomes were collected and tested. The results of the sample tests showed that the differences in the levels of the following metabolites can well distinguish between patients with treatment outcomes and those without treatment outcomes: valine, betaine, tryptophan, kynurenine, creatine, taurine, phenylalanine, tyrosine, histidine, aspartic acid, threonine, guanine, γ-aminobutyric acid, allantoin and pseudouridine (Table 13). Different algorithms were further used to construct models (Table 14). The best model was an artificial neural network (NNET). The significantly changed metabolites were further used to make a receiver operating characteristic curve (ROC). The area under the curve (AUC) value was 1, the accuracy was 100%, the sensitivity was 100%, and the specificity was 100%. Figure 7 The model showed good predictive value. The combination of metabolites valine, betaine, tryptophan, kynurenine, creatine, taurine, phenylalanine, tyrosine, histidine, aspartic acid, threonine, guanine, γ-aminobutyric acid, allantoin, and pseudouridine showed excellent specificity and sensitivity.
[0119] Table 13. Differentially expressed metabolites between the depression-prone and non-prone groups
[0120]
[0121]
[0122] Table 14. Comparative analysis of different models for predicting the outcome of depression patients after treatment
[0123] Model AUC value Accuracy (%) Sensitivity (%) Specificity (%) Generalized Linear Models 0.9113 86.01% 93.75% 70.21% Linear Discriminant Analysis 0.8547 84.61% 89.58% 74.49% K-nearest neighbor algorithm 0.9071 83.92% 93.75% 63.83% Logistic regression 0.9281 85.31% 96.87% 61.70% Lasso regression 0.8970 84.61% 89.58% 74.47% Decision Tree 0.9191 82.52% 93.75% 59.57% Artificial Neural Networks 1.000 100.00% 100.00% 100.00% Support Vector Machine 0.9477 85.31% 94.79% 65.96% Extreme Gradient Boosting 0.9410 84.61% 88.54% 76.6% Random Forest 0.9160 84.61% 92.71% 68.08% Principal component analysis 0.9253 85.31% 95.83% 63.83% Bayesian Network 0.9188 86.71% 92.71% 74.47% Linear regression 0.9375 85.31% 90.62% 74.47%
[0124] Experimental Example 5 Evaluation of the Prediction Effect of the Prediction Model Established in Example 5 on the Treatment Outcome of Depression Patients
[0125] Dried blood spot samples from outpatient clinics or inpatients of a hospital were randomly collected. All samples had complete scale method evaluation data collected and signed informed consent. The ion abundance ratios of biomarkers (valine, betaine, tryptophan, kynurenine, creatine, taurine, phenylalanine, tyrosine, histidine, aspartic acid, threonine, guanine, γ-aminobutyric acid, allantoin, and pseudouridine) and internal standards were directly detected according to the method of Example 5. The data were input into the prediction model established in Example 5 to obtain prediction results, which were compared with the scale method evaluation results. The results are shown in Table 15.
[0126] Table 15. Evaluation of the effectiveness of metabolic marker kits in predicting treatment outcomes in patients with depression
[0127]
[0128]
[0129] As can be seen from Table 15, the accuracy of predicting the treatment outcome of depression patients using the marker combination of Example 5 and the prediction model established therefrom is 93.91%, the sensitivity is 93.12%, and the specificity is 95.71%.
[0130] Example 6 Metabolic marker screening and model establishment for predicting whether depression patients will relapse during stable period
[0131] The same method as in Example 1 was used to predict whether a depression patient would relapse during the stable period. The test results were as follows:
[0132] Samples from 130 patients with stable depression who did not relapse and 58 patients with stable depression who relapsed were collected and tested. The results of the sample tests showed that the differences in the levels of the following metabolites can well distinguish between patients with stable depression who relapsed and those who did not relapse: proline, ornithine, alanine, tryptophan, kynurenine, creatine, 2-hydroxybutyric acid, androstenedione, cytosine, xanthine, bilirubin, γ-aminobutyric acid, allantoin, deoxyribose, pyruvic acid and linoleic acid (Table 16). Different algorithms were further used to construct models (Table 17). The best model was the support vector machine (SVM). The significantly changed metabolites were further used to make a receiver operating characteristic curve (ROC). The area under the curve (AUC) value was 1, the accuracy was 99.30%, the sensitivity was 100%, and the specificity was 98.96%. Figure 8 The model showed good predictive value. The combination of metabolites proline, ornithine, alanine, tryptophan, kynurenine, creatine, 2-hydroxybutyrate, androstenedione, cytosine, xanthine, bilirubin, γ-aminobutyric acid, allantoin, deoxyribose, pyruvate, and linoleic acid showed excellent specificity and sensitivity.
[0133] Table 16. Differentially expressed metabolites between relapsed and non-relapsed depression groups
[0134]
[0135]
[0136] Table 17. Comparative analysis of different models for predicting relapse in patients with depression during stable period
[0137] Model AUC value Accuracy (%) Sensitivity (%) Specificity (%) Generalized Linear Models 0.9869 96.50% 97.92% 93.62% Linear Discriminant Analysis 0.9511 94.40% 97.92% 87.23% K-nearest neighbor algorithm 0.9867 95.80% 97.92% 91.49% Logistic regression 0.9836 95.80% 98.96% 89.36% Lasso regression 0.9815 94.40% 98.96% 85.11% Decision Tree 0.9776 94.40% 98.96% 85.11% Artificial Neural Networks 0.9742 94.40% 98.96% 85.11% Support Vector Machine 1.000 99.30% 100.00% 98.96% Extreme Gradient Boosting 0.9894 95.80% 96.87% 93.62% Random Forest 0.9874 97.90% 98.96% 95.74% Principal component analysis 0.9865 95.80% 97.92% 91.49% Bayesian Network 0.9869 96.50% 97.92% 93.62% Linear regression 0.9362 86.01% 90.62% 76.59%
[0138] Experimental Example 6 Evaluation of the prediction effect of the prediction model established in Example 6 on the recurrence of depression patients in the stable period
[0139] Dried blood spot samples were randomly collected from outpatients or inpatients of a hospital. All samples were diagnosed by professional physicians according to the ICD-10 standard and signed informed consent. The ion abundance ratios of biomarkers (proline, ornithine, alanine, tryptophan, kynurenine, creatine, 2-hydroxybutyric acid, androstenedione, cytosine, xanthine, bilirubin, γ-aminobutyric acid, allantoin, deoxyribose, pyruvic acid, and linoleic acid) and internal standards were directly detected according to the method of Example 6. The data were input into the prediction model established in Example 6 to obtain prediction results, which were compared with the diagnosis results of professional physicians. The results are shown in Table 18.
[0140] Table 18. Evaluation of the effectiveness of metabolic marker kits in predicting relapse in patients with stable depression
[0141]
[0142]
[0143] As can be seen from Table 18, the marker combination of Example 6 and the prediction model established therefrom have an accuracy of 96.63%, a sensitivity of 90.00%, and a specificity of 97.34% in predicting the relapse of patients with depression during the stable phase.
[0144] The embodiments of the present invention are described in detail above, but the present invention is not limited to the described embodiments. It is apparent to those skilled in the art that various changes, modifications, substitutions, and variations of these embodiments may be made without departing from the principles and spirit of the present invention, and the changes still fall within the scope of protection of the present invention.
Claims
1. A depression disease management system based on targeted metabolomics and machine learning models, characterized in that: include: Intelligent diagnosis model module, differential diagnosis model module, hierarchical diagnosis model module, companion diagnosis model module, treatment endpoint outcome prediction model module, recurrence prediction model module, and comprehensive judgment module; Among them, the intelligent diagnosis model module is used to identify patients with depression; the differential diagnosis model module is used to distinguish between bipolar and unipolar depression; the graded diagnosis model module is used to distinguish between mild depression and moderate to severe depression; the companion diagnosis model module is used to evaluate whether the treatment of depression patients is effective; the treatment endpoint outcome prediction model module is used to make intelligent predictions of the patient's clinical treatment outcomes and provide a basis for optimizing treatment plans; the recurrence prediction model module is used to predict disease recurrence in patients in the stable period and take timely intervention measures; the comprehensive judgment module is used to integrate the results of the diagnosis, identification and prediction models of the remaining modules to build a depression disease management system; The method for using the depression disease management system includes the following steps: S1. Build an intelligent diagnostic model through the intelligent diagnostic model module to identify patients with depression; S2. Construct a differential diagnosis model using the differential diagnosis model module to differentiate between bipolar and unipolar depression; S3. Construct a hierarchical diagnostic model using the hierarchical diagnostic model module to differentiate between mild depression and moderate to severe depression; S4. Build a companion diagnostic model using the companion diagnostic model module to assess whether treatment for depression is effective. S5. Construct a treatment endpoint outcome prediction model using the treatment endpoint outcome prediction model module to intelligently predict the patient's clinical treatment outcome and provide a basis for optimizing treatment plans; S6. Build a recurrence prediction model using the recurrence prediction model module to predict disease recurrence in stable patients and take timely intervention measures; S7. Build a depression disease management system by integrating various diagnostic, identification, and prediction models through a comprehensive judgment module; The specific operations of constructing the intelligent diagnostic model through the intelligent diagnostic model module in S1 are as follows: collecting blood samples and detecting the following metabolic molecular marker combinations: proline, betaine, alanine, tryptophan, kynurenine, 5-hydroxytryptamine, creatine, succinate, taurine, and 2-hydroxybutyrate, and constructing different machine learning models to screen the metabolic marker combination under the best model for identifying patients with depression; The specific operation of constructing the differential diagnosis model through the differential diagnosis model module in S2 is: collecting a blood sample and detecting the following metabolic molecular marker combination: proline, ornithine, kynurenine, 5-hydroxytryptamine, succinic acid, 2-hydroxybutyric acid, xanthine, bilirubin, γ-aminobutyric acid, allantoin, deoxyribose, pyruvic acid, linoleic acid and uric acid, constructing different machine learning models, and screening the optimal model for distinguishing patients with unipolar and bipolar depression; The specific operation of constructing the hierarchical diagnosis model through the hierarchical diagnosis model module in S3 is as follows: collecting blood samples and testing the following metabolic molecular marker combination: proline, valine, ornithine, tryptophan, 5-hydroxytryptamine, creatine, glutamine, serine, methionine, guanine, hypoxanthine, androstenedione, cytosine, bilirubin, γ-aminobutyric acid, uric acid, adipic acid, and pseudouridine, constructing different machine learning models, screening the optimal model, and using it to distinguish mild depression from moderate to severe depression; The specific operation of constructing the companion diagnostic model through the companion diagnostic model module in S4 is as follows: collecting blood samples and testing the following metabolic molecular marker combination: ornithine, betaine, alanine, kynurenine, serotonin, glutamate, lysine, methionine, phenylalanine, guanine, hypoxanthine, γ-aminobutyric acid, allantoin and uric acid, constructing different machine learning models, screening the optimal model, and using it to evaluate whether the treatment of depression patients is effective, and helping to screen treatment drug regimens for patients with different disease courses; The specific operation of constructing the treatment endpoint outcome prediction model through the treatment endpoint outcome prediction model module in S5 is: tracking the treatment status and recovery status of the depressed patients undergoing treatment, and detecting the following metabolic molecular marker combinations at different stages: valine, betaine, tryptophan, kynurenine, creatine, taurine, phenylalanine, tyrosine, histidine, aspartic acid, threonine, guanine, γ-aminobutyric acid, allantoin and pseudouridine nucleoside, constructing different machine learning models, screening the best model, and using it to intelligently predict the patient's clinical treatment outcome and provide a basis for optimizing the treatment plan; The specific operation of constructing the relapse prediction model through the relapse prediction model module in S6 is: tracking the condition of patients with depression and detecting the following metabolic molecular marker combinations at different stages: proline, ornithine, alanine, tryptophan, kynurenine, creatine, 2-hydroxybutyric acid, androstenedione, cytosine, xanthine, bilirubin, γ-aminobutyric acid, allantoin, deoxyribose, pyruvic acid and linoleic acid, constructing different machine learning models, screening the best model, and using it to predict disease relapse in patients in the stable stage and take timely intervention measures; The specific operation of integrating the various diagnostic, differential, and predictive models through the comprehensive judgment module in S7 is as follows: by tracking the progression of depression patients' illness, collecting peripheral blood metabolite concentration data at different stages, and using a machine learning algorithm, developing a model system for depression diagnosis, graded diagnosis, differential diagnosis, concomitant diagnosis, treatment outcome prediction, and stable period relapse prediction based on multiple metabolic molecular markers, wherein the disease management system provides an interactive interface that is easy for users to operate based on the peripheral blood metabolite expression concentration data, and outputs corresponding results and disease management decision-making recommendations according to different purposes; Among them, the machine learning models include: generalized linear model, linear discriminant analysis, K-nearest neighbor algorithm, logistic regression, lasso regression, decision tree, artificial neural network, support vector machine, extreme gradient boosting, random forest, principal component analysis, Bayesian network, and linear regression.
2. The depression disease management system according to claim 1, characterized in that: The sample for detecting the metabolic molecular marker combination is human serum, plasma, or dried blood spots.
Citation Information
Patent Citations
Biomarker and detection reagent for diagnosing depression and predicting curative effect of targeted visual cortex repeated transcranial magnetic stimulation treatment
CN115612726A
Combined marker for severe depression of teenagers and construction method of marker model
CN117612701A
Method for intelligently preventing and controlling juvenile depression based on multi-dimensional information
CN118230957A