A biomarker and artificial intelligence-based depression tendency identification method
By combining LC-MS technology with various machine learning algorithms, a classification and scoring model for depressive tendencies was constructed, which solved the problems of subjectivity and data processing limitations of existing diagnostic methods, and realized the objective quantitative identification and assessment of early-stage depression.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA PHARM UNIV
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-29
Smart Images

Figure CN122117259A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of biomedical engineering and artificial intelligence, specifically, but not limited to, a method for identifying depressive tendencies based on biomarkers and artificial intelligence. Background Technology
[0002] Depression is a common and serious mental illness that significantly impacts a patient's social functioning. Its core symptoms include depressed mood and loss of interest. With its high prevalence and severity, depression has become a pressing global public health issue. Mild depression, in particular, represents an early stage of depression, and early detection and intervention are crucial to preventing its progression. However, because early symptoms of mild depression are often subtle and the rate of seeking medical attention is low, traditional diagnosis relies primarily on patient self-reporting and clinical interviews, leading to frequent missed or misdiagnoses. Existing assessment scales (such as the PHQ-9 and HAMD) are subjective, easily influenced by individual differences in expression and the assessor's experience, and lack objective, quantifiable biological evidence.
[0003] In recent years, with the development of biodetection technology, LC-MS has been widely used in metabolomics research for high-throughput quantitative analysis of small molecules (such as neurotransmitters, amino acids, bile acids, etc.) in body fluid samples, revealing a variety of biomarkers related to mood regulation.
[0004] However, a single biomarker is insufficient to fully characterize the complex biological background of mood disorders, and traditional multivariate statistical methods have limitations when processing high-dimensional, nonlinear biological data. Therefore, there is an urgent need for an artificial intelligence-based data modeling method that can integrate multidimensional LC-MS biomarker information to achieve objective quantification of depression-related states and identification of potential neuro-metabolic perturbation patterns.
[0005] In view of this, a new method is needed to solve at least some of the above problems. Summary of the Invention
[0006] To address one or more problems in existing technologies, this invention proposes a method for identifying depressive tendencies based on biomarkers and artificial intelligence. It utilizes LC-MS technology to quantitatively analyze multidimensional biomarkers in test samples and combines various machine learning algorithms to generate continuous scores and status labels reflecting the degree of an individual's depressive tendency. This provides data support for the objective assessment of mental health status and the development of intervention strategies, offering an objective and efficient quantitative identification method in the initial screening of depressive tendencies. It can assist clinicians in screening, risk stratification, or referencing intervention plans.
[0007] The technical solution to achieve the purpose of this invention is as follows:
[0008] A method for identifying depressive tendencies based on biomarkers and artificial intelligence, characterized by comprising:
[0009] S1. Obtain biological samples, and use liquid chromatography-mass spectrometry (LC-MS) technology to perform high-throughput quantitative detection of several neuro-metabolic biomarkers in the biological samples to obtain biomarker concentration data, and perform data preprocessing on the biomarker concentration data;
[0010] S2. Obtain the mental health assessment scale score corresponding to the biological sample. Based on the score and the preprocessed biomarker concentration data, construct a standardized metabolomics feature dataset and establish a multi-dimensional feature mapping rule based on the biomarker reference range of emotion regulation-related neural-metabolic pathways. The multi-dimensional feature mapping rule maps the standardized concentration values of each biomarker in the standardized metabolomics feature dataset to the biomarker reference range of emotion regulation-related neural-metabolic pathways to obtain pathway-dimensional abnormal / perturbation features. These pathway-dimensional abnormal / perturbation features are used for mechanism-related state label determination and interpretability output, including key biomarker abnormality pattern analysis and pathway abnormality dimension description.
[0011] S3. Employing a multi-algorithm comparison and ensemble strategy, a depressive tendency state classification model and a depressive tendency score regression model are established based on the metabolomics feature dataset. A classification + regression joint model is trained, and an interpretability module is constructed. The depressive tendency state classification model is used to classify the biological samples into mechanism-related states or normal states and output corresponding depressive tendency state labels. The depressive tendency score regression model is used to calculate the continuous depressive tendency score corresponding to the biological samples. The interpretability module is used to analyze abnormal patterns of key biomarkers.
[0012] S4. Based on the classification + regression joint model, the new biological samples are processed to obtain the depression tendency status label, continuous depression tendency score, and abnormal pattern analysis of key biomarkers of the new biological samples.
[0013] Compared with the prior art, the present invention, employing the above technical solution, has the following technical effects:
[0014] 1. This invention uses LC-MS technology to perform quantitative analysis of multidimensional biomarkers, providing a repeatable and quantifiable data foundation, effectively reducing individual bias caused by traditional subjective scale assessments, and the LC-MS platform can simultaneously detect more than 27 small molecule biomarkers, greatly improving analytical throughput and efficiency.
[0015] 2. This invention utilizes artificial intelligence modeling methods to efficiently process high-dimensional, nonlinear metabolomics data, significantly improving the accuracy and stability of identifying states related to the neuro-metabolic mechanisms of depressive tendencies.
[0016] 3. This invention provides a data-driven hierarchical perspective on the heterogeneity of mental health states through clustering of mechanism-related states and association analysis of key biomarkers, and assists in the formulation of differentiated intervention strategies. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the invention and, together with the description, are used to explain embodiments of the invention, but do not constitute a limitation thereof. In the drawings:
[0018] Figure 1 The performance comparison chart of the eight depressive tendency classification models of the present invention is shown.
[0019] Figure 2 The confusion matrix of the AdaBooSt classification model of the present invention is shown.
[0020] Figure 3 The performance comparison chart of the six depression tendency score regression models of the present invention is shown.
[0021] Figure 4 The actual score-predicted score scatter plot of the random forest regression model of the present invention is shown.
[0022] Figure 5 A schematic diagram showing the top 10 biomarkers and their importance in mechanism prediction and score prediction is presented.
[0023] Figure 6 A schematic diagram of biomarkers showing significant statistical differences is shown.
[0024] Figure 7 The histogram of subtype distribution of 4000 virtual generated samples of the present invention is shown.
[0025] Figure 8 The scoring distribution of 4000 virtual generated samples of the present invention is shown.
[0026] Figure 9 A pie chart showing the subtype percentages of the top 10 samples by serial number out of 4000 virtual generated samples of this invention is presented.
[0027] Figure 10 The histogram showing the score distribution of the top 10 samples by serial number out of 4000 virtual generated samples of the present invention is shown.
[0028] Figure 11 A flowchart of the depressive tendency identification method based on biomarkers and artificial intelligence of the present invention is shown. Detailed Implementation
[0029] To further understand the present invention, preferred embodiments of the present invention are described below in conjunction with examples. However, it should be understood that these descriptions are only for further illustrating the features and advantages of the present invention, and not for limiting the scope of the claims of the present invention.
[0030] The description in this section pertains only to typical embodiments, and the present invention is not limited to the scope of the embodiments described. Combinations of different embodiments, substitution of some technical features in different embodiments, and substitution of similar or identical prior art with some technical features in the embodiments are also within the scope of the description and protection of the present invention.
[0031] According to one aspect of the present invention, a method for identifying depressive tendencies based on biomarkers and artificial intelligence, such as... Figure 11 As shown, it includes:
[0032] (1) Quantitative analysis by liquid chromatography-mass spectrometry (LC-MS).
[0033] LC-MS / MS analysis was performed using an Agilent 1290 Infinity LC system and a 6495 triple quadrupole mass spectrometer to provide high-quality, reproducible biomarker concentration data input.
[0034] The chromatographic conditions are as follows:
[0035] Column: Agilent ZORBAX Eclipse Plus C18 (2.1×100 mm, 1.8 μm);
[0036] Mobile phase: A (0.1% formic acid-water), B (0.1% formic acid-acetonitrile);
[0037] Gradient elution: 0-2 min, 5% B; 2-15 min, 5-95% B; 15-18 min, 95% B; 18-20 min, 95-5% B; 20-25 min, 5% B;
[0038] Flow rate: 0.3 mL / min;
[0039] Column temperature: 40℃.
[0040] The mass spectrometry conditions are as follows:
[0041] Electrospray ionization (ESI) source, positive ion mode;
[0042] Multiple reaction monitoring (MRM) mode optimizes collision energy and ion pairs;
[0043] Sheath gas temperature: 350℃, sheath gas flow rate: 11 L / min;
[0044] Capillary voltage: 4000V.
[0045] High-sensitivity and high-specificity quantitative detection of 27 neuro-metabolic-related small molecule biomarkers in subject biological samples was achieved using liquid chromatography-mass spectrometry (LC-MS / MS). These biomarkers cover multiple biological pathways closely related to mood regulation. The 27 biomarkers include: glutamic acid (Glu), glutamine (Gln), aspartic acid (Asp), asparagine (Asn), glycine (Gly), γ-aminobutyric acid (GABA), tryptophan (Trp), valine (Val), alanine (Ala), arginine (Arg), leucine (Leu), isoleucine (Ile), histidine (His), phenylalanine (Phe), tyrosine (Tyr), homocysteine (Hcy), deoxycholic acid (DCA), chenodeoxycholic acid (CDCA), kynurenine (Kyn), taurine (Tau), melatonin (MT), cortisol (CS), adrenaline (E), norepinephrine (NE), 5-hydroxyindoleacetic acid (5-HIAA), serotonin (5-HT), and dopamine (DA).
[0046] All tests were conducted in compliance with ethical guidelines and human genetic resource management requirements. Internal standard method was used to correct systematic errors, and multiple response monitoring (MRM) mode was used to ensure quantitative accuracy and repeatability. The obtained biomarker concentration data served as the core input features for subsequent artificial intelligence modeling.
[0047] (2) Data preprocessing obtained from metabolomics.
[0048] The concentration data of 27 biomarkers obtained from LC-MS quantitative analysis were preprocessed using Agilent MassHunter software to improve the accuracy and stability of the model. This included baseline correction, peak identification, and peak area integration. Specifically, missing values were filled using the mean of each biomarker characteristic; outliers were identified using box plots, and removed or replaced as needed. The internal standard method was used to quantify the biomarker concentrations, followed by normalization.
[0049] (3) Data preprocessing for machine learning modeling.
[0050] The preprocessed concentration data of 27 biomarkers were Z-score standardized using StandardScaler (to make the mean 0 and the standard deviation 1), and the training set and test set were randomly divided in an 8:2 ratio. At the same time, a stratified sampling strategy was adopted to ensure that the distribution of the state labels related to each mechanism was consistent in the two groups.
[0051] (4) Mechanism state classification modeling based on multi-algorithm comparison and integration.
[0052] Eight mainstream machine learning classification models were built using Python's scikit-learn library. These models were then used to identify seven categories of neuro-metabolic mechanism-related states and normal states. All models used the same concentration data of 27 biomarkers as input and output seven categories of mechanism-related state labels and normal state labels.
[0053] The state classification models used for comparison and screening specifically include:
[0054] 1) Random Forest Classification Model: ŷ=mode{f1(X),f2(X),...,f_K(X)}.
[0055] In the formula: ŷ: the predicted depression subtype (AG subtype or N normal type); K: the number of decision trees, K=100 (n_estimators=100); f_K(X): the prediction result of the Kth decision tree; X: the input feature vector, X=[x1,x2,...,x 27 ] contains concentration data for 27 metabolic biomarkers; mode{·}: mode function, selects the category that appears most frequently from all decision tree prediction results.
[0056] The model building process is as follows:
[0057] Parameter settings: n_estimators=100, random_state=42;
[0058] Feature importance assessment: The contribution of each metabolite to subtype classification was assessed using Gini importance.
[0059] Subtype definition: Based on the literature, depression is divided into 7 subtypes (AG) and normal type (N);
[0060] 2) Support Vector Machine (SVC) classification model: min(1 / 2)‖w‖²+C∑ξ_i.
[0061] The constraints are as follows: y_i(w*φ(x_i)+b)≥1-ξ_i,ξ_i≥0;
[0062] In the formula: w: weight vector, w∈R^d, d is the feature dimension; b: bias term, b∈R; φ(x_i): kernel function that maps the input features to a high-dimensional space, using the RBF kernel: K(x_i,x_j)=exp(-γ‖x_i-x_j‖²); C: penalty coefficient, C=1.0, controlling the balance between model complexity and misclassification; ξ_i: slack variable, ξ_i≥0, the degree to which the i-th sample is allowed to be misclassified; y_i: true label of the i-th sample, y_i∈{-1,+1} (for binary classification) or multi-class expansion; x_i: feature vector of the i-th sample.
[0063] 3) Logistic Regression Classification Model: P(y=c|X)=exp(w_c*X+b_c) / ∑_{j=1}^Cexp(w_j·X+b_j).
[0064] In the formula: P(y=c|X): the probability that sample X belongs to category c; w_c: the weight vector of category c, w_c∈R^d; b_c: the bias term of category c, b_c∈R; C: the total number of categories, C=8 (7 subtypes + normal type); X: the input feature vector; the linear output is converted into a probability distribution through the softmax function.
[0065] 4) K-Nearest Neighbors Classifier: ŷ=argmax_{c}∑_{i∈N_k(x)}I(y_i=c).
[0066] In the formula: ŷ: predicted class label; N_k(x): set of k nearest neighbors of sample x, k=5; I(·): indicator function, returns 1 when the condition is true, otherwise returns 0; y_i: true label of the i-th neighbor; c: class label; distance metric: Euclidean distance d(x_i,x_j)=√∑{m=1} 27 (x{im}-x_{jm}) 2 .
[0067] 5) Decision Tree Classifier: Gini(D)=1-∑_{k=1}^Cp_k², ΔGini=Gini(D)-(|D_left| / |D|)Gini(D_left)-(|D_right| / |D|)Gini(D_right).
[0068] In the formula: Gini(D): Gini impurity of dataset D; p_k: proportion of class k samples in the dataset; C: total number of classes; D_left, D_right: left and right child node datasets after splitting; ΔGini: reduction in Gini impurity, the feature and threshold that maximize ΔGini are selected for splitting.
[0069] 6) Gaussian Naive Bayes model: ŷ=argmax_{c}P(c)∏_{j=1}^dP(x_j|c), P(x_j|c)=(1 / √(2πσ_{c,j}²))exp(-(x_j-μ_{c,j})² / (2σ_{c,j}²)).
[0070] In the formula: P(c): prior probability of category c; P(x_j|c): likelihood probability of feature x_j given category c; μ_{c,j}: mean of feature x_j in category c; σ_{c,j}: standard deviation of feature x_j in category c; d: feature dimension, d=27.
[0071] 7) AdaBoost ensemble classification model (AdaBoostClassifier): F(X)=sign(∑_{m=1}^Mα_mh_m(X)), α_m=(1 / 2)ln((1-err_m) / err_m).
[0072] In the formula: F(X): final ensemble classifier; h_m(X): prediction result of the m-th weak classifier; α_m: weight of the m-th weak classifier; err_m: weighted error rate of the m-th weak classifier; M: number of weak classifiers, M=50; sign(·): sign function.
[0073] 8) Gradient Boosting Classifier: F_m(X)=F_{m-1}(X)+ν*h_m(X), h_m(X)≈-[∂L / ∂F_{m-1}).
[0074] In the formula: F_m(X): the ensemble model after the m-th iteration; F_{m-1}(X): the model of the previous iteration; h_m(X): the m-th base learner, which fits the negative gradient of the current model; ν: learning rate, ν=0.1; L: loss function, usually log loss; M: number of iterations, M=100.
[0075] Under a unified data partitioning (training set / test set = 8:2) and preprocessing workflow, each state classification model was trained and optimized using five-fold cross-validation. Systematic performance comparisons were conducted based on metrics such as accuracy, F1 score, and cross-validation F1 score (CVF1 score). The optimal set of state classification models and their corresponding weights, exhibiting suitable performance in classification consistency, generalization ability, and computational efficiency, were selected according to preset evaluation metrics to construct an ensemble state classification model. The optimal model set and its corresponding weights can be updated as the training dataset is updated. This classification model construction method reduces the risk of overfitting or bias in a single model, providing reliable state priors for subsequent depression tendency scoring and mechanism analysis.
[0076] In one embodiment, the comparison results of the accuracy, F1 score, and cross-validation F1 score of the eight machine learning classification models are as follows: Figure 1 As shown, according to Figure 1 Based on the comparison results, the AdaBoost classification model was selected as the classification model for depressive tendencies. The confusion matrix of the AdaBoost classification model is as follows: Figure 2 As shown, the AdaBoost classification model has high accuracy in identifying mechanisms of depressive tendencies. Figure 2 In the coordinate system, A represents a deficiency of monoamine neurotransmitters; B represents an abnormality of the glutamatergic system; C represents inhibition of the GABAergic system; D represents HPA axis dysfunction; E represents gut microbiota-bile acid metabolism disorder; F represents neuroinflammation and immune metabolism abnormalities; G represents circadian rhythm disorder; and N represents normal. Figure 2 This demonstrates the matching degree between the actual mechanism and the predicted mechanism, and can evaluate the AdaBooSt classification model's ability to identify various mechanisms. Figure 2 It can be seen that the identification of type B and type D is more accurate.
[0077] (5) Regression modeling of depression tendency scores based on multi-algorithm comparison.
[0078] A quantitative prediction method for depressive tendency scores was constructed, employing six mainstream regression algorithms to continuously quantify individuals' mental health status. All models were trained using the same concentration data of 27 biomarkers as input and scores from mental health assessment scales as reference labels, outputting a continuous depressive tendency score.
[0079] The regression models used for comparison and screening specifically include:
[0080] 1) Random Forest Regressor: ŷ=(1 / K)Σ_{k=1}^Kf_k(X).
[0081] In the formula: ŷ: predicted depression severity score, ŷ∈R; K: number of decision trees, K=100; f_k(X): prediction result of the k-th regression tree; X: input feature vector; each regression tree is constructed by minimizing the mean squared error.
[0082] 2) Support Vector Regression (SVR) model: min(1 / 2)∣∣w∣∣²+C∑(ξ_i+ξ_i^*).
[0083] The constraints are as follows:
[0084] y_i-w·φ(x_i)-b≤ε+ξ_i;
[0085] w·φ(x_i)+b-y_i≤ε+ξ_i^*;
[0086] ξ_i,ξ_i^*≥0;
[0087] In the formula: w: weight vector; b: bias term; φ(x_i): feature mapping function, using RBF kernel; C: penalty parameter, C=1.0; ε: tolerance error, ε=0.1, defining the insensitive region; ξ_i,ξ_i^*: slack variables, the degree to which samples are allowed to exceed the ε pipeline; y_i: the true score of the i-th sample.
[0088] 3) Linear Regression Model: ŷ=w·X+b,min‖y-Xw‖².
[0089] In the formula: ŷ: predicted depression score; w: weight coefficient vector, w∈R^d; b: bias term, b∈R; X: design matrix, X∈R^{n×d}; y: true score vector, y∈R^n; solved by least squares: w=(X^TX) -1 X^Ty.
[0090] 4) K-NeighborsRegressor: ŷ=(1 / k)Σ_{i∈N_k(x)}y_i.
[0091] In the formula: ŷ: predicted depression score; k: number of nearest neighbors, k=5; N_k(x): set of k nearest neighbors of sample x; y_i: true score of the i-th neighbor; prediction is based on the principle of local averaging.
[0092] 5) Decision Tree Regressor: Var(D) = (1 / |D|)Σ_{i∈D}(y_i-ȳ) 2, ΔVar=Var(D)-(|D_left| / |D|)Var(D_left)-(|D_right| / |D|)Var(D_right).
[0093] In the formula: Var(D): variance of dataset D; ȳ: mean of target values in dataset D; D_left, D_right: left and right child node datasets after splitting; ΔVar: variance reduction amount, the feature and threshold that maximize ΔVar are selected for splitting.
[0094] 6) Gradient Boosting Regressor: F_m(X)=F_{m-1}(X)+v*h_m(X), h_m(X)≈-[∂L(y,F_{m-1}(X)) / ∂F_{m-1}(X)].
[0095] In the formula: F_m(X): the ensemble model after the m-th iteration; F_{m-1}(X): the model of the previous iteration; h_m(X): the m-th base learner, which fits the negative gradient of the current model; ν: learning rate, ν=0.1; L: loss function, usually the mean squared error loss; M: number of iterations, M=100.
[0096] Based on unified data preprocessing and an 8:2 training / test split, the performance of each rating regression model was evaluated using five-fold cross-validation. Root mean square error (RMSE), coefficient of determination (R²), and explanatory power (CV R²) were used as core metrics for cross-sectional comparison. Based on the validation results, a regression model with suitable generalization ability and predictive stability was selected from the various rating regression models as the final predictor of depression tendency scores.
[0097] In one embodiment, the comparison results of the error (RMSE), coefficient of determination (R²), and explanatory power (CV R²) of six machine learning regression models are as follows: Figure 3 As shown, this illustrates the performance of each regression model in the "score prediction" scenario. According to... Figure 3 Based on the comparison results, the Random Forest regression model was selected as the regression model for the depression tendency score. The scatter plot of the actual score versus predicted score from the Random Forest regression model is shown below. Figure 4 As shown, the dashed line represents the ideal fit line, used to observe the accuracy of the regression model in predicting the scores. From Figure 4 The high accuracy of the Random Forest regression model in scoring depression tendencies can be observed.
[0098] (6) Evaluation and validation of the classification + regression joint model.
[0099] Establish multi-level model evaluation and statistical validation to ensure the reliability and interpretability of the established classification + regression joint model.
[0100] 1) Hierarchical five-fold cross-validation is employed, and model selection and classification performance are validated using accuracy, F1 score, and cross-validation F1 score. The evaluation metrics include:
[0101] Accuracy: Accuracy = (TP+TN) / (TP+TN+FP+FN). Where TP: true positive, TN: true negative, FP: false positive, FN: false negative.
[0102] F1 score: F1 = 2×(Precision×Recall) / (Precision+Recall).
[0103] Cross-validation: 5-fold stratified cross-validation.
[0104] 2) Hierarchical five-fold cross-validation was employed to evaluate the performance of the systematic regression based on the root mean square error (RMSE), coefficient of determination (R²), and explanatory power (CV R²). The evaluation metrics included:
[0105] Mean squared error: MSE = (1 / n)∑(y_i - ŷ_i) 2 Where y_i: the true value, ŷ_i: the predicted value, and n: the number of samples.
[0106] Root mean square error: RMSE = √MSE.
[0107] Coefficient of determination: R² = 1 - ∑(y_i - ŷ_i)² / ∑(y_i - ȳ)². Where, ȳ: the mean of the true values.
[0108] Cross-validation: 5-fold stratified cross-validation.
[0109] 3) The bootstrap method is used to resample and calculate the 95% confidence interval of key performance indicators to quantify the uncertainty of prediction. The bootstrap confidence interval refers to the 95% confidence interval of the model performance calculated through 1000 resampling.
[0110] Effect size calculation: Cohen'sd effect size analysis is introduced to assess the practical significance of differences between groups and avoid misjudgment caused by relying solely on statistical significance (p value).
[0111] The random forest algorithm combined with recursive feature elimination technique was used to analyze the obtained biomarker data, evaluate and screen the top 10 important metabolic features, such as... Figure 5As shown, Tau has the highest weight, followed by 5-HT, illustrating the core role of these markers in "subtype classification + score prediction".
[0112] The Kruskal-Wallis test was used to assess the significance of differences in metabolites among different subtypes. The significance characteristics of the Kruskal-Wallis test (p<0.05) are as follows: Figure 6 As shown, the metabolites MT, Tau, and Kyn all exhibit statistically significant differences, which is key evidence that "subtypes have a molecular basis".
[0113] The Kruskal-Wallis test and feature importance analysis together revealed the discriminative value of key biomarkers and provided a molecular basis and interpretability for seven types of neuro-metabolic mechanism-related states. Thus, from the three dimensions of model performance, biological rationality, and output credibility, this identification method fully demonstrates its innovation and practicality in objectively quantifying mental health status.
[0114] This evaluation and validation system not only verified the superiority of the classification + regression joint model in terms of classification consistency and regression accuracy, but also provided solid statistical support for the discriminative value of biomarkers and the credibility of the system output.
[0115] The following analysis and validation of the classification + regression joint model is performed using 4000 virtual samples. The subtype distribution of the 4000 virtual samples is as follows: Figure 7 As shown, the E type had the most cases (1052) and the N type had the fewest cases (33), reflecting the macroscopic distribution pattern of subtypes in a large sample.
[0116] The distribution of predicted scores for 4000 virtual generated samples is as follows: Figure 8 As shown, the mean is 5.82, and the 95% confidence interval is [0.66, 0.87], demonstrating the overall characteristics of the large sample scores.
[0117] The pie chart showing the subtype percentages of the top 10 samples by serial number out of 4000 virtual generated samples is as follows: Figure 9 As shown, type B accounts for 50% (the most), type E 30%, and types D and F each 10%, reflecting the distribution characteristics of subtypes in a small sample.
[0118] The ID, subtype, and score of the top 10 samples by serial number out of 4000 virtual generated samples are shown in the table below, which is a supplement to the quantitative results.
[0119] Sample Name Predicted subtype Predicted score Sample 0001 B 5.25 Sample 0002 D 7.16 Sample 0003 E 7.97 Sample 0004 B 4.97 Sample 0005 B 5.87 Sample 0006 B 6.05 Sample 0007 F 8.30 Sample 0008 B 7.67 Sample 0009 E 7.82 Sample 0010 E 6.13
[0120] The histogram of the score distribution of the top 10 samples by serial number out of 4000 virtual generated samples is shown below. Figure 10 As shown, the mean score is 6.72, indicating a central tendency in the scores.
[0121] It can be seen that the classification + regression joint model can accurately distinguish 4,000 virtual large samples and accurately predict and quantify depression scores.
[0122] The description and application of the present invention herein are illustrative and not intended to limit the scope of the invention to the embodiments described above. The effects or advantages described in the specification may not be apparent in actual experimental cases due to uncertainties in specific conditions or other factors, and such descriptions are not intended to limit the scope of the invention. Variations and modifications to the embodiments disclosed herein are possible, and various substitutions and equivalents of the components in the embodiments are well known to those skilled in the art. It should be understood by those skilled in the art that the invention can be implemented in other forms, structures, arrangements, proportions, and with other components, materials, and parts without departing from the spirit or essential characteristics of the invention. Other variations and modifications can be made to the embodiments disclosed herein without departing from the scope and spirit of the invention.
Claims
1. A method for identifying depressive tendencies based on biomarkers and artificial intelligence, characterized in that, include: S1. Obtain biological samples, and use liquid chromatography-mass spectrometry to perform high-throughput quantitative detection of several neuro-metabolic biomarkers in the biological samples to obtain biomarker concentration data, and perform data preprocessing on the biomarker concentration data. S2. Obtain the mental health assessment scale score corresponding to the biological sample. Based on the score and the preprocessed biomarker concentration data, construct a standardized metabolomics feature dataset and establish a multi-dimensional feature mapping rule based on the reference range of biomarkers for emotion regulation-related neural-metabolic pathways. S3. Employing a multi-algorithm comparison and ensemble strategy, a depressive tendency state classification model and a depressive tendency score regression model are established based on the metabolomics feature dataset. A classification + regression joint model is trained, and an interpretability module is constructed. The depressive tendency state classification model is used to classify the biological samples into mechanism-related states or normal states and output corresponding depressive tendency state labels. The depressive tendency score regression model is used to calculate the continuous depressive tendency score corresponding to the biological samples. The interpretability module is used to analyze abnormal patterns of key biomarkers. S4. Based on the classification + regression joint model, process the new biological samples and output the depression tendency status label, continuous depression tendency score and abnormal pattern analysis of key biomarkers of the new biological samples.
2. The method for identifying depressive tendencies based on biomarkers and artificial intelligence according to claim 1, characterized in that, The neuro-metabolic biomarkers described in S1 include: glutamate (Glu), glutamine (Gln), aspartic acid (Asp), asparagine (Asn), glycine (Gly), γ-aminobutyric acid (GABA), tryptophan (Trp), valine (Val), alanine (Ala), arginine (Arg), leucine (Leu), isoleucine (Ile), histidine (His), phenylalanine (Phe), tyrosine (Tyr), homocysteine (Hcy), deoxycholic acid (DCA), chenodeoxycholic acid (CDCA), kynurenine (Kyn), taurine (Tau), melatonin (MT), cortisol (CS), adrenaline (E), norepinephrine (NE), 5-hydroxyindoleacetic acid (5-HIAA), serotonin (5-HT), and dopamine (DA).
3. The method for identifying depressive tendencies based on biomarkers and artificial intelligence according to claim 1, characterized in that, S3 specifically includes: S31. Using an ensemble learning strategy, with the metabolomics feature dataset as input and mechanism-related biomarkers as feature labels, multiple machine learning classifiers are trained simultaneously to obtain multiple state classification models. S32. The various state classification models are trained and optimized using five-fold cross-validation, and their performance is compared based on metrics such as accuracy, F1 score, and cross-validation F1 score. Based on preset evaluation metrics, the optimal set of state classification models and their corresponding weights are determined from the various state classification models to construct an integrated state classification model. The integrated state classification model is used to predict the biological sample and output the corresponding depressive tendency state label, which includes mechanism-related state labels and normal state labels. S33. A regression modeling strategy with multiple algorithms is adopted. The metabolomics feature dataset is used as input, and the score of the mental health assessment scale is used as a continuous reference label. Multiple machine learning regression models are trained simultaneously to obtain multiple scoring regression models. S34. Five-fold cross-validation is used to train and optimize the various rating regression models. The root mean square error, coefficient of determination and cross-validation R² are used as core indicators for horizontal comparison. The optimal rating regression model is selected from the various rating regression models. The optimal rating regression model is used to calculate and output the continuous score of depression tendency corresponding to the biological sample. S35. Based on the integrated state classification model and the optimal rating regression model, construct a joint model of depression tendency classification and regression, and verify its classification performance and regression performance. S36. The generalization ability of the classification + regression joint model is evaluated by using the hierarchical five-fold cross-validation method, and simulation sample data is generated based on the statistical distribution characteristics of the metabolomics feature dataset to conduct a system robustness stress test on the classification + regression joint model. S37. Using a random forest algorithm combined with recursive feature elimination technology, a subset of key biomarkers that contribute most to predicting depressive tendency state labels and calculating continuous depressive tendency scores are selected from the metabolomics feature dataset. An association matrix is established between the subset of key biomarkers and the emotion regulation-related neural-metabolic pathways. An interpretability module is constructed based on the association matrix.
4. The method for identifying depressive tendencies based on biomarkers and artificial intelligence according to claim 3, characterized in that, The optimal state classification model set described in S32 includes at least one of the following: random forest classification model, support vector machine classification model, logistic regression classification model, K-nearest neighbor classification model, decision tree classification model, Gaussian Naive Bayes classification model, AdaBooSt classification model, and gradient boosting tree classification model.
5. The method for identifying depressive tendencies based on biomarkers and artificial intelligence according to claim 3, characterized in that, The optimal score regression model described in S34 includes one of the following: random forest regression model, support vector machine regression model, linear regression model, K-nearest neighbor regression model, decision tree regression model, and gradient boosting regression model.
6. The method for identifying depressive tendencies based on biomarkers and artificial intelligence according to claim 1 or 3, characterized in that, The mechanisms-related states include monoamine neurotransmitter deficiency, glutamatergic system abnormalities, GABAergic system inhibition, HPA axis dysfunction, gut microbiota-bile acid metabolism disorders, neuroinflammation and immune metabolism abnormalities, and circadian rhythm disorders.
7. The method for identifying depressive tendencies based on biomarkers and artificial intelligence according to claim 1 or 6, characterized in that, The combination of neuro-metabolic biomarkers corresponding to the mechanism-related state is as follows: deficiency of monoamine neurotransmitters: histidine (His), phenylalanine (Phe), tyrosine (Tyr), norepinephrine (NE), 5-hydroxyindoleacetic acid (5-HIAA), 5-hydroxytryptamine (5-HT), and dopamine (DA). Abnormalities in the glutamatergic system: glutamate (Glu), glutamine (Gln), aspartate (Asp), asparagine (Asn), glycine (Gly); GABAergic system inhibition: glutamate (Glu), aspartate (Asp), γ-aminobutyric acid (GABA), homocysteine (Hcy); HPA axis dysfunction: Valine (Val), Alanine (Ala), Arginine (Arg), Taurine (Tau), Cortisol (CS); Gut microbiota-bile acid metabolism disorder: deoxycholic acid (DCA), chenodeoxycholic acid (CDCA); Neuroinflammation and abnormal immune metabolism: Tryptophan (Trp), Leucine (Leu), Kynuronic acid (Kyn); Circadian rhythm disorder: melatonin MT.
8. The method for identifying depressive tendencies based on biomarkers and artificial intelligence according to claim 1, characterized in that, The multi-dimensional feature mapping rule described in S2 maps the standardized concentration values of each biomarker in the standardized metabolomics feature dataset to the reference range of biomarkers in emotion regulation-related neuro-metabolic pathways, thereby obtaining abnormal / perturbation features of the pathway dimension. These abnormal / perturbation features of the pathway dimension are used for mechanism-related state label determination and interpretability output, including analysis of abnormal patterns of key biomarkers and description of abnormal pathway dimensions.
9. A depressive tendency identification system based on biomarkers and artificial intelligence, characterized in that, It includes a classification + regression joint model, a standardized parameter module, and an interpretability module. The classification + regression joint model models the metabolomics characteristic data of biomarkers and outputs a classification label for depressive tendencies and a continuous score for depressive tendencies. The standardized parameter module completes internal standard correction, missing / outlier handling, and normalization, and maps the concentration of each biomarker to abnormal features at the pathway dimension. The interpretability module outputs an interpretation analysis of abnormal patterns and pathway perturbations based on key biomarkers and their pathway relationships.
10. A detection reagent, characterized in that, It includes several neuro-metabolic biomarkers as described in claims 1-2.