Biomarker combination related to depression, product and application
The depression diagnosis model constructed through biomarker combination and machine learning algorithms solves the problem of lack of efficient and accurate diagnosis in the prior art, realizes accurate prediction and diagnosis of depression, and provides kits and computer program products.
Patent Information
- Application Number
- CN202510655680.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-02
AI Technical Summary
The prior art lacks biomarker combinations and related prediction models that can be used to diagnose depression, resulting in relatively subjective diagnosis and efficacy prediction, and lack of efficient and accurate diagnostic methods.
A biomarker combination is provided, including 4-bromophenacetic acid, L-threonine-L-leucine, carnitine C20:2, carnitine C18:0 and L-proline-L-phenylalanine, combined with machine learning algorithms to build a diagnostic model, use gradient enhancement tree and LASSO regression to screen out key metabolites, and build an accurate diagnostic model for depression.
It achieves efficient and accurate prediction and diagnosis of depression, provides kits and computer program products based on biomarker combinations, improves the sensitivity and specificity of diagnosis and provides better diagnostic and treatment indicators.
Smart Images

Figure CN120577533A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of the combination of biomedicine and machine learning, and specifically relates to a biomarker combination, product and application related to depression. Background Art
[0002] Depression is a common mental illness characterized by low mood and loss of interest, and is associated with high rates of disability and relapse. According to the World Health Organization, over 300 million people worldwide suffer from depression. Over 95 million people in China suffer from depression, and the "2022 National Depression Blue Book" survey shows that 50% of those with depression are students. Currently, there are no objective diagnostic markers for this condition, making diagnosis and treatment efficacy prediction relatively subjective, relying primarily on physicians' clinical experience and patient descriptions.
[0003] Metabolites, as intermediate or end products of various biochemical reactions in the body, are closely associated with the development and progression of diseases. Therefore, they serve as potential biomarkers for disease diagnosis, prognosis, and treatment monitoring. To date, hundreds of studies have explored metabolomic alterations associated with depression. These studies have employed a variety of detection techniques, primarily liquid chromatography-mass spectrometry, gas chromatography-mass spectrometry, and NMR methods, analyzing biological samples from various fluid types, including plasma, serum, urine, and cerebrospinal fluid. Studies have found that multiple metabolites are significantly associated with depression. These metabolites are involved in multiple biological processes, including cell signaling, cell membrane composition, neurotransmitter metabolism, inflammatory and immune regulation, hormone activation and its precursors, and sleep regulation. However, existing research data have not yet identified specific metabolite changes in depression that can be directly applied clinically. Furthermore, single biomarkers seem unlikely to meet the needs of diagnosis and treatment. Integrating multiple molecules to generate biomarker panels can provide better indicators for disease diagnosis and treatment decisions, with improved sensitivity and specificity.
[0004] Therefore, there is currently a lack of a biomarker combination and related prediction model that can be used to diagnose depression to achieve efficient and accurate prediction and diagnosis. Summary of the Invention
[0005] In response to the above technical problems, the present application provides a biomarker combination, product and application related to depression.
[0006] The technical solutions provided in this application are as follows: In a first aspect, a biomarker combination is provided, comprising: 4-bromophenylacetic acid, L-threonine-L-leucine, carnitine C20:2, carnitine C18:0, and L-proline-L-phenylalanine.
[0007] In a second aspect, a reagent for detecting a biomarker combination is provided for use in preparing a product for diagnosing depression, wherein the biomarker combination includes 4-bromophenylacetic acid, L-threonine-L-leucine, carnitine C20:2, carnitine C18:0, and L-proline-L-phenylalanine.
[0008] In a third aspect, a kit is provided, comprising a detection reagent for detecting the biomarker combination described in the first aspect.
[0009] In a fourth aspect, a method for preparing a product for detecting depression is provided, wherein the kit described in the third aspect is used for preparing a product for detecting depression.
[0010] In a fifth aspect, there is provided use of the detection reagent in the kit described in the third aspect in preparing a kit for diagnosing depression.
[0011] In a sixth aspect, a program product related to depression is provided, wherein the computer program product is used to diagnose the risk of a subject suffering from depression, comprising the following steps: Obtaining the level of each biomarker in the plasma of the test subject; the biomarkers include 4-bromophenylacetic acid, L-threonine-L-leucine, carnitine C20:2, carnitine C18:0, and L-proline-L-phenylalanine; Substituting the content of each biomarker into the model composed of the gradient boosting tree to calculate the probability y that the subject to be tested has depression; Based on the comparison result of the probability y and the reference value, diagnose or predict whether the subject suffers from depression or has the risk of depression.
[0012] In one possible implementation, the method for obtaining the model composed of the gradient boosting tree includes: Obtain the levels of biomarkers in healthy individuals and patients with depression; The model constructed by gradient boosting tree is trained with the content of biomarkers as input and whether the disease is present as output.
[0013] In a seventh aspect, a method for screening the biomarker according to the first aspect is provided, comprising: Obtaining data on the levels of biomarkers in the plasma of healthy individuals and patients with depression; Compare the metabolite expression levels of depression patients and healthy subjects and perform differential metabolite detection; Draw principal component analysis graphs and expression level heat maps of differential metabolites; Use machine learning algorithms to build an initial depression diagnosis model, and use cross-validation to evaluate the model performance and obtain evaluation results; The top-ranked biomarkers in the initial diagnostic model were screened based on the SHAP values of the evaluation results; Biomarkers were further screened based on LASSO regression, and the final diagnostic model was established using machine learning algorithms; The expression levels of metabolite groups of the screened biomarker combinations were compared and correlation analysis was performed with the depression severity scale to verify disease specificity.
[0014] In one possible implementation, the machine learning algorithm includes a gradient boosting tree algorithm.
[0015] In one possible implementation, the cross-validation method is a K-fold cross-validation method.
[0016] The beneficial effects of this application are as follows: 1. This application provides a biomarker combination and application, which can be used to prepare products for diagnosing depression.
[0017] 2. This application provides related products based on this biomarker combination, including a kit and a computer program product. This computer program product, combined with artificial intelligence technology, constructs a diagnostic model with excellent performance, enabling accurate prediction and diagnosis of depression, and providing new insights into the precise diagnosis and treatment of depression.
[0018] 3. This application provides a method for screening the biomarkers, based on which biomarkers related to depression can be efficiently screened out. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 shows the differential expression of metabolites between depression and the control group. (A) Volcano plot of differentially expressed metabolites, with each dot representing a metabolite, red representing upregulated metabolites and green representing downregulated metabolites. The critical values in the figure are |log2FC| ≥ 0.58 (vertical dashed line) and P < 0.05 (horizontal dashed line); (B) Sample PCA plot of differential metabolic characterization, with each dot representing a sample, green representing HC and red representing MDD samples; (C) Quantitative heat map of differential metabolites between samples; rows represent different differential metabolites, columns represent different samples, and colors indicate the expression levels of differential metabolites.
[0020] Figure 2 shows the diagnostic performance of machine learning. (A) ROC curve for the diagnostic model constructed using 107 differentially expressed metabolites; (B) the top 20 important features output by the initial model; (C) ROC curve for the diagnostic model constructed using the top 20 important metabolites; (D) SHAP value ranking of the top 20 important features.
[0021] Figure 3 shows LASSO regression for screening metabolites and constructing a diagnostic model. (A) Coefficient curves are generated based on the logarithmic (lambda) sequence, with the optimal lambda yielding nonzero coefficients; (B) The optimal parameter (lambda) of the LASSO model is obtained through 10-fold cross-validation using the minimum standard deviation plus one standard error (right vertical line); (C) Diagnostic model constructed based on the important metabolites screened by LASSO regression; (D) SHAP value ranking of important metabolites.
[0022] Fig. 4 Differences in the expression of depression-related metabolites among groups. DETAILED DESCRIPTION
[0023] The content of this application is further described below with reference to specific embodiments, but the content of this application is not limited thereto.
[0024] Currently, there is a lack of a biomarker combination and related prediction model that can be used to diagnose depression to achieve efficient and accurate prediction and diagnosis.
[0025] In view of this, this embodiment provides a biomarker combination, product and application related to depression.
[0026] For the convenience of description, relevant professional terms involved in the embodiments are explained in a unified manner.
[0027] Biomarkers: 4-Bromophenylacetic acid, the structural formula is .
[0028] L-Threonine-L-Leucine (Thr-Leu), the structural formula is .
[0029] Carnitine C20:2, the structural formula is .
[0030] Carnitine C18:0, the structural formula is .
[0031] L-Proline-L-phenylalanine (Pro-Phe), the structural formula is .
[0032] Example 1 The steps for biomarker screening and model construction are as follows: Step 1: Data collection and detection A total of 118 participants were enrolled, including 81 patients with depression and 37 healthy controls. Venous blood samples were collected from all participants, centrifuged to obtain serum and plasma, which were immediately stored in a -80°C ultra-low temperature freezer. In November 2024, plasma samples were analyzed for targeted metabolomics using liquid chromatography-tandem mass spectrometry. A sociodemographic questionnaire was designed to collect basic information about participants, including gender, age, education level, and BMI. Depressive symptoms were assessed using the 9-item Patient Health Questionnaire-9 (PHQ-9) and the 17-item Hamilton Depression Rating Scale-17 (HAMD-17); anxiety symptoms were assessed using the Generalized Anxiety Disorder Scale-7 (GAD-7) and the 14-item Hamilton Anxiety Scale-14 (HAMA-14); and somatic symptoms were assessed using the Patient Health Questionnaire-15 (PHQ-15). The PHQ-9, GAD-7, and PHQ-15 are self-assessments completed independently by participants, while the HAMD-17 and HAMA-14 are peer-assessed assessments administered by trained psychiatrists through clinical observation and semi-structured interviews.
[0033] Step 2: Targeted metabolomics analysis and data preprocessing The experiment used liquid chromatography-tandem mass spectrometry to achieve accurate qualitative and quantitative analysis of metabolites.
[0034] The following process was used for data processing: first, missing values were filled with 20% of the minimum value of the metabolites; then the coefficient of variation of the quality control samples was calculated, and metabolite data with a coefficient of variation lower than 0.3 were screened for subsequent analysis.
[0035] Step 3: Detection of differential metabolites Metabolic profiling based on a broadly targeted metabolomics approach was performed on the 118 samples included, detecting 811 metabolites. Intergroup variability analysis revealed that 107 metabolites were significantly differentially expressed in the depression group compared with the control group (|log2FC| ≥ 0.58, P < 0.05), of which 82 were upregulated and 25 were downregulated (Figure 1A).
[0036] Step 4: Principal Component Analysis A principal component analysis plot was created based on the 107 differentially expressed metabolites identified, with different points representing different samples. The results showed that patients with depression and healthy controls could be clearly distinguished by the differentially expressed metabolites (Figure 1B). Heat maps generated using the quantitative information of the differentially expressed metabolites revealed differences in metabolite expression levels between the two groups (Figure 1C).
[0037] Step 5: Use the machine learning algorithm gradient boosting tree to build a diagnostic model and screen important features In this example, a model was constructed using 107 differentially expressed metabolites using the Catboost (Categorical Boosting) algorithm, and the top 20 most important metabolites were output, as shown in Figure 2. The average area under the curve (AUC) for this model in the test set was 0.982 (Figure 2A). Further diagnostic model construction using only the top 20 most important metabolites in the model achieved an average AUC of 0.989 in the test set (Figure 2C), demonstrating no significant degradation in model performance.
[0038] Step 6: Simplify the model based on LASSO regression The LASSO regression model was used to further screen out 5 metabolites (4-Bromophenylacetic acid, Thr-Leu, Carnitine C20:2, Carnitine C18:0, Pro-Phe) from the 20 important metabolites in step 5 to simplify the diagnostic model, and the Catboost (Categorical Boosting) algorithm was used to construct the diagnostic model.
[0039] As shown in Figure 3, the diagnostic model constructed using these five metabolites achieved an AUC of 0.990, an accuracy of 1.000, a sensitivity of 0.938, a specificity of 1.000, an ACC of 0.957, and an F1 value of 0.966 (Figure 3C). Table 1 summarizes the performance metrics of the different diagnostic models.
[0040] Table 1 Performance evaluation of different diagnostic models
[0041] Note: Marker 107 represents the differential metabolites screened out in step 3, marker 20 represents the top 20 important markers output by the machine learning model in step 5, and marker 5 represents the metabolites screened out in step 6.
[0042] Step 7. Comparison of expression levels between metabolite groups Figure 4 shows five depression-associated metabolites identified by LASSO regression. Only Pro-Phe was downregulated in the depression group; the others were significantly upregulated. Furthermore, the relationship between these five metabolites and the Depression Severity Scale was further analyzed. The results showed that the expression of these metabolites was not correlated with the Depression Scale, Anxiety Scale, or Somatic Symptom Scale, with P-values > 0.05. This further suggests that the identified metabolites may be disease-specific rather than correlated with depressive symptom severity.
[0043] Example 2 This embodiment provides a computer program product for diagnosing the risk of a subject suffering from depression, comprising the following steps: 2.1. Substituting the content of each biomarker into the model constructed by the gradient boosting tree, and calculating the probability y that the subject to be tested has depression; 2.2. Based on the comparison result of the probability y and the reference value, diagnose or predict whether the subject suffers from depression or has a risk of depression.
[0044] In practice, when the y value is greater than 0.5, it indicates a high probability that the subject has depression; when the y value is less than 0.5, it indicates a low probability. When the y value is 0.5, the subject may be healthy or depressed, and further testing is necessary using other methods, such as clinical assessment scales. Furthermore, the closer the y value is to 0.5, the more necessary it is to use other testing methods.
[0045] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent replacements and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention should be included in the scope of protection of the invention.
Claims
1. A biomarker combination, characterized in that: include: 4-Bromophenylacetic acid, L-threonine-L-leucine, carnitine C20:2, carnitine C18:0, and L-proline-L-phenylalanine.
2. Use of a reagent for detecting a biomarker combination in the preparation of a product for diagnosing depression, wherein the biomarker combination includes 4-bromophenylacetic acid, L-threonine-L-leucine, carnitine C20:2, carnitine C18:0, and L-proline-L-phenylalanine.
3. A kit, characterized in that The kit comprises a detection reagent for detecting the biomarker combination according to claim 1.
4. Use of the kit according to claim 3 in preparing a product for detecting depression.
5. Use of the detection reagent in the kit according to claim 3 in preparing a kit for diagnosing depression.
6. A program product related to depression, characterized in that: The computer program product is used to diagnose the risk of a subject suffering from depression, comprising the following steps: Obtaining the level of each biomarker in the plasma of the test subject; the biomarkers include 4-bromophenylacetic acid, L-threonine-L-leucine, carnitine C20:2, carnitine C18:0, and L-proline-L-phenylalanine; Substituting the content of each biomarker into the model composed of the gradient boosting tree to calculate the probability y that the subject to be tested has depression; Based on the comparison result of the probability y and the reference value, diagnose or predict whether the subject suffers from depression or has the risk of depression.
7. The program product according to claim 6, wherein The method for obtaining the model composed of the gradient boosting tree comprises: Obtain the levels of biomarkers in healthy individuals and patients with depression; The model constructed by gradient boosting tree is trained with the content of biomarkers as input and whether the disease is present as output.
8. The biomarker screening method according to claim 1, characterized in that: include: Obtaining data on the levels of biomarkers in the plasma of healthy individuals and patients with depression; Compare the metabolite expression levels of depression patients and healthy subjects and perform differential metabolite detection; Draw principal component analysis graphs and expression level heat maps of differential metabolites; Use machine learning algorithms to build an initial depression diagnosis model, and use cross-validation to evaluate the model performance and obtain evaluation results; The top-ranked biomarkers in the initial diagnostic model were screened based on the SHAP values of the evaluation results; Biomarkers were further screened based on LASSO regression, and the final diagnostic model was established using machine learning algorithms; The expression levels of metabolite groups of the screened biomarker combinations were compared and correlation analysis was performed with the depression severity scale to verify disease specificity.
9. The biomarker screening method according to claim 8, characterized in that: The machine learning algorithm includes a gradient boosting tree algorithm.
10. The biomarker screening method according to claim 8, characterized in that: The cross validation method is a K-fold cross validation method.