Biomarker combination, model for evaluating depression risk and application

By constructing a biomarker combination containing specific metabolites and microorganisms, combined with machine learning models, the problem of insufficient accuracy of depression risk assessment in the prior art is solved, and a more efficient depression risk assessment is achieved.

CN120446459APending Publication Date: 2025-08-08HUMAN METABOLOMICS INST INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510313065.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-06
Filing Date
2025-03-17
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Inadequate accuracy and stability of models for evaluating the risk of depression in the prior art lead to difficulty in diagnosis of depression.

Method used

A biomarker combination, including specific metabolites and microorganisms, assess the risk of depression through machine learning models, and use serum concentrations of metabolites and the abundance of microorganisms in feces for risk assessment.

Benefits of technology

It improves the accuracy and specificity of risk assessment of depression and provides a more robust assessment tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120446459A_ABST
    Figure CN120446459A_ABST
Patent Text Reader

Abstract

The invention discloses a biomarker combination, a model for evaluating depression risk and application. The biomarker has relatively high correlation with the depression, and the model for evaluating the depression risk has relatively high accuracy, sensitivity and specificity in the aspect of evaluating the depression risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of clinical risk assessment, and specifically relates to a biomarker combination, a method for constructing a model for assessing depression risk, a model for assessing depression risk, a system for assessing depression risk, and applications thereof. Background Art

[0002] Depression, also known as depressive disorder, is characterized by a pronounced and persistent low mood and is the primary type of mood disorder. Clinically, the depressed mood is disproportionate to the individual's situation, ranging from melancholy to profound grief, low self-esteem, depression, and even pessimism and world-weariness. Suicidal attempts or behaviors may occur, and even catatonia may occur. Some cases present with significant anxiety and motor agitation. In severe cases, psychotic symptoms such as hallucinations and delusions may develop. Each episode lasts for at least two weeks, and in some cases, even for several years. Most cases tend to relapse, and while most episodes resolve, some may have residual symptoms or become chronic.

[0003] There is no definitive theory for the origin and pathogenesis of depression, as both environmental and physiological factors can have an impact. It is now generally accepted that depression may be caused by a deficiency or dysfunction of neurotransmitters such as serotonin and adrenaline in the brain. Therefore, the metabolic levels of neurotransmitters in an organism, such as dopamine, adrenaline, serotonin, and gamma-aminobutyric acid (GABA), can reflect to a certain extent whether a person suffers from depression. These compounds can be used as diagnostic biomarkers to assist in the early diagnosis of depression based on the metabolic levels of specific compounds in the body, so as to detect mild patients or people with depression tendencies at an early stage, conduct early intervention, and reduce the harm caused by depression. International and domestic depression diagnostic methods based on ICD-10 rely heavily on the subjective judgment of doctors, making the diagnosis of depression very difficult and making the treatment of depression difficult.

[0004] The literature (PhD dissertation of Chongqing Medical University, author: Zheng Peng, Screening of diagnostic markers for depression in plasma and urine based on metabolomics strategy, 2013 and Zheng P, Wang Y, Chen L, et al. Identification and validation of urinary metabolite biomarkers for major depressive disorder. Mol Cell Proteomics. 2013; 12(1): 207-214) disclosed that 126 urine samples of depression patients and 134 normal controls were measured by nuclear magnetic resonance, and a total of 23 differential metabolites were identified. Logistic regression analysis found that the biomarker group composed of five metabolites, namely malonate, formate, N-methylnicotinamide, M-hydroxyphenylacetate and alanine, has potential clinical diagnostic value. The biomarker panel could distinguish 82 depression patients from 82 normal controls in the training set with an AUC value of 0.812 (95% confidence interval: 0.829–0.961); and distinguish 44 depression patients from 52 normal controls in the test set with an AUC value of 0.895 (95% confidence interval: 0.829–0.961).

[0005] Patent publication number CN112630311B discloses the use of urine samples to screen for a panel of metabolic markers, including dopamine, serotonin, γ-aminobutyric acid, tryptophan, tyramine, kynurenine, 3,4-dihydroxyphenylacetic acid, epinephrine, and norepinephrine, for the diagnosis of depression. Samples from 200 healthy individuals and 200 patients with depression were collected and tested. The biomarker panel achieved a sensitivity of 93.2% and a specificity of 96.7% for the training set. The area under the curve (AUC) for the validation set data was 0.865, with a sensitivity of 83.6% and a specificity of 89.7%.

[0006] However, most of these existing technologies use metabolite combinations in urine samples to assess depression risk, resulting in relatively modest performance. Therefore, due to the need for more accurate assessment results, a model with improved risk assessment performance is needed to achieve robust and accurate depression risk assessment, which is of great clinical significance. Summary of the Invention

[0007] To address the existing lack of accurate, efficient, and stable models for assessing depression risk, the present invention provides a biomarker combination, reagents and kits for detecting the biomarker combination, and their applications, as well as a method for constructing a model for assessing depression risk, a model constructed thereby, and a system incorporating the model. The biomarker combination has a high correlation with depression, and the model for assessing depression risk has high accuracy, sensitivity, and specificity in assessing depression risk.

[0008] To solve the above technical problems, the present invention provides a technical solution: a biomarker combination, wherein the biomarker combination includes the following metabolites: furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid and L-arginine;

[0009] And / or, the biomarker combination includes the following microorganisms: Eubacterium hallii, Blautia sp. AF19-10LB, Blautia sp. OM07-19, Faecalibacterium sp. AF28-13AC, Roseburia intestinalis, Roseburia faecis, Firmicutes bacterium AF25-13AC, Faecalibacterium prausnitzii and Eubacterium rectale.

[0010] In a preferred embodiment of the present invention, the biomarkers further include one or more of the following metabolites: fructose, phenyllactic acid, L-tyrosine, L-proline, 3-hydroxybutyric acid, 12-hydroxystearic acid, indole-3-acetic acid methyl ester, oleic acid, 10-cis-heptadecenoic acid, palmitoleic acid, acetic acid, dodecanoic acid, nonadecenoic acid (cis-10), 2-hydroxybutyric acid, pentadecanoic acid, myristic acid, cis-5-dodecenoic acid, α-linolenic acid, 3-aminoisobutyric acid, acetylglycine, undecanoic acid, L-carnitine and linoleic acid;

[0011] And / or, the biomarker further includes one or more of the following microorganisms: Ruminococcus gnavus, Dialister sp. Marseille-P5638, Bacteroides vulgatus, Faecalibacterium sp. AF27-11BH, Hungatela hathewayl, Roseburia hominis, Blautia sp. OM05-6, Faecalibacterium sp. AM43-5AT, Ruminococcus torques, Blautia sp. AF19-1, Ruminococcus AM46-18 sp.AM46-18), Ruminococcus sp.AM12-48, Ruminococcus sp.AF33-11BH, Ruminococcus sp.AM23-1LB, Romboutsia timonensis, Dorea formicigenerans, Eggerthella lenta, Gemmiger formicilis, and Blautia obeum.

[0012] To solve the above technical problems, the present invention provides a technical solution: a reagent for detecting the biomarker combination as described in the present invention.

[0013] In a preferred embodiment of the present invention, when the reagent detects a metabolite in the biomarker combination as described in the present invention, the detection is the detection of the serum concentration of the metabolite; when the reagent detects a microorganism in the biomarker combination as described in the present invention, the detection is the detection of the abundance of the microorganism in feces, and the abundance refers to the relative abundance of the microorganism.

[0014] In a specific embodiment of the present invention, the reagent is a gene sequencing reagent and / or an HPLC-MS reagent.

[0015] In the present invention, the relative abundance of microorganisms refers to the ratio of the number of DNA sequences of a certain microorganism (such as bacteria, fungi or viruses) in a sample to the total number of DNA sequences of all microorganisms in the sample. This ratio is usually expressed as a percentage.

[0016] To solve the above technical problems, the present invention provides a technical solution: a kit comprising the biomarker combination or the reagent described in the present invention.

[0017] To solve the above technical problems, the present invention provides a technical solution: use of the biomarker combination, reagent or kit described in the present invention in the preparation of a product for assessing the risk of depression.

[0018] In the present invention, depression refers to depressive disorder, which is characterized by the "three low symptoms" of significant and persistent low mood, slow thinking, and decreased will activity, as well as cognitive impairment and somatic symptoms. It is the main type of mood disorder.

[0019] In the present invention, the clinical diagnosis process of depression generally includes the following steps:

[0020] ① Initial assessment: Diagnosing depression usually begins with an initial consultation between the patient and a healthcare provider. During this consultation, the doctor will ask questions about the patient's symptoms, medical history, family history, and possible life events to understand the patient's situation.

[0021] ②Symptom assessment: The doctor will assess the patient's symptoms, which usually include the following main symptoms of depression:

[0022] Persistent depression, low mood, or feelings of emptiness;

[0023] loss of interest or favorite activities;

[0024] Irritability or irritability;

[0025] Insomnia or excessive sleeping;

[0026] Easy fatigue or lack of energy;

[0027] Feelings of self-guilt or worthlessness;

[0028] difficulty concentrating or making decisions;

[0029] thoughts of death or suicide;

[0030] Rule out other medical conditions: Your doctor will rule out other medical or mental health conditions that may cause similar symptoms, such as thyroid problems, substance abuse, or other mental disorders.

[0031] ③ Severity assessment: Doctors will usually assess the severity of depression, which may help determine treatment options. The severity of depression can range from mild to moderate to severe, depending on the number and severity of symptoms and their impact on daily life.

[0032] Diagnostic criteria: The diagnosis of depression is usually based on international diagnostic criteria such as the Hamilton Depression Rating Scale (HAMD) score, the Diagnostic and Statistical Manual of Mental Disorders (DSM-5), or the International Classification of Diseases (ICD-10 or ICD-11). Doctors will use these criteria to determine whether the diagnosis of depression is met.

[0033] To solve the above technical problems, the present invention provides a technical solution: a method for constructing a model for assessing depression risk, the method comprising:

[0034] The biomarker data in the biomarker database are input into the GB model for machine learning to construct the model for assessing the risk of depression; the biomarker data in the biomarker database are sourced from samples of patients with depression and samples of patients without depression; the biomarkers include the biomarker combination described in the present invention.

[0035] In a preferred embodiment of the present invention, the model for assessing depression risk is validated by an external independent validation dataset.

[0036] In a preferred embodiment of the present invention, in the biomarker data, the metabolite data is the serum concentration of the metabolite, and the microbial data is the abundance of the microorganism in the feces, and the abundance refers to the relative abundance of the microorganism.

[0037] In a preferred embodiment of the present invention, the biomarker data comes from cohort 1 in Table 1, with 70% of the samples used as a training set and 30% of the samples used as a test set; the independent validation data set comes from cohort 2 and cohort 3 in Table 1.

[0038] In a preferred embodiment of the present invention, the metabolite data is obtained by HPLC-MS technology, preferably using a metabolic chip in the patent application with patent application number 201910170381.0 for detection and substance identification.

[0039] In a preferred embodiment of the present invention, the microbial abundance data is obtained by gene sequencing, preferably by sequencing on an Illumina HiSeq X-10 platform.

[0040] In a preferred embodiment of the present invention, the metabolite data are analyzed by the analysis algorithm IP4M (http: / / 47.92.77.13 / ).

[0041] In a preferred embodiment of the present invention, the parameter settings of the GB model include one or more of the following: (1) Data (training set) is set to train; (2) Distribution (the form of the loss function) is set to bernoulli; (3) n.trees (number of iterations) is set to 100; (4) Interaction.depth (the depth of the decision tree) is set to 1; (5) shrinkage (learning rate) is set to 0.03; (6) n.minobsinnode (the minimum number of training set samples in a node to start splitting) is set to 4.

[0042] In a specific embodiment of the present invention, the procedure of the GB model is:

[0043] Training set: data

[0044] Validation set: data

[0045] data4<-read.csv("D: / R_sang / 2r_csv / HF_data4.csv")

[0046] data5<-read.csv("D: / R_sang / 2r_csv / HF_data5.csv")

[0047] idx<-

[0048] c("age","Plt","ALT","AST","Sex","bmi","FBG","TC","TG","GGT","urea","HDLch",

[0049] "LDLch","APOA","TBil","Ferritin","albumin","DM.IFG","AST.ALT","AST.PLT")

[0050] Y<-data4$group_1

[0051] pm_disc<-data4[,idx]

[0052] data_disc<-cbind.data.frame(Y,pm_disc)

[0053] Y<-data5$group_1

[0054] pm_vld<-data5[,idx]

[0055] data_vld<-cbind.data.frame(Y,pm_vld)

[0056] rslt<-matrix(nrow=1,ncol=12)

[0057] colnames(rslt)<-c("accuracy","P_a","P_b","R_a","R_b","F_a","F_b","auROC",

[0058] "auPR","threshold","specificity","sensitivity")

[0059] dataG<-data_disc

[0060] data1G<-data_vldset.seed(1234)

[0061] model_GB<-gbm(Y~.,data=dataG,n.trees=400,distribution="bernoulli",

[0062] interaction.depth=1,shrinkage=0.03,n.minobsinnode=4)

[0063] pred_GB<-predict(model_GB,data1G,n.trees=400,type='link')

[0064] roc<-roc(data1G$Y,pred_GB)

[0065] auROC<-auc(roc)

[0066] best_cut<-as.matrix(coords(roc,'best',transpose=FALSE))[1,]

[0067] #some_cut<-as.matrix(coords(roc,0.8,"specificity",transpose=FALSE))

[0068] #best is the best threshold point, there may be two, when filling in the result of permutations and combinations, only the first one is selected to avoid errors

[0069] The #coords() function can not only select the best cutoff value, but also find the corresponding point or specificity or sensitivity related results according to requirements

[0070] lg<-as.numeric(length(pred_GB))

[0071] for(s in 1:lg){

[0072] if(pred_GB[s] <best_cut[1]){data1G$pred_Y[s]<-0}else{data1G$pred_Y[s]<-1}

[0073] }

[0074] tab<-table(data1G$Y,data1G$pred_Y)

[0075] accu<-sum(diag(prop.table(tab)))

[0076] P_a<-diag(prop.table(tab,2))[1]

[0077] P_b<-diag(prop.table(tab,2))[2]

[0078] R_a<-diag(prop.table(tab,1))[1]#Specificity

[0079] R_b<-diag(prop.table(tab,1))[2]#Sensitivity

[0080] F1_a<-2*P_a*R_a / (P_a+R_a)

[0081] F1_b<-2*P_b*R_b / (P_b+R_b)

[0082] pr<-pr.curve(pred_GB[data1G$Y==1]%>%na.omit, pred_GB[data1G$Y==0]%>%na.omit)

[0083] auPR<-pr$auc.integral

[0084] row<-1

[0085] rslt[row,1]<-accu

[0086] rslt[row,2]<-P_a

[0087] rslt[row,3]<-P_b

[0088] rslt[row,4]<-R_a

[0089] rslt[row,5]<-R_b

[0090] rslt[row,6]<-F1_a

[0091] rslt[row,7]<-F1_b

[0092] rslt[row,8]<-auROC

[0093] rslt[row,9]<-auPR

[0094] rslt[row,10:12]<-best_cut

[0095] write.csv(rslt,"D: / R_sang / 4r_result / ???.csv").

[0096] To solve the above technical problems, the present invention provides a technical solution: a model for assessing the risk of depression, wherein the model for assessing the risk of depression is constructed by the method for constructing a model for assessing the risk of depression as described in the present invention.

[0097] To solve the above technical problems, the present invention provides a technical solution: a method for assessing the depression risk of a sample, the method comprising inputting the biomarker data of the sample to be tested into a model for assessing depression risk as described in the present invention, to obtain an assessment result of the depression risk of the sample; the biomarker comprises the biomarker combination described in the present invention.

[0098] In a preferred embodiment of the present invention, the method is for non-diagnostic purposes.

[0099] In a preferred embodiment of the present invention, in the biomarker data, the metabolite data is the serum concentration of the metabolite, and the microbial data is the abundance of the microorganism in the feces, and the abundance refers to the relative abundance of the microorganism.

[0100] In the present invention, the sample includes a serum sample and / or a stool sample.

[0101] In a preferred embodiment of the present invention, the metabolite data is obtained by HPLC-MS technology, preferably using a metabolic chip in the patent application with patent application number 201910170381.0 for detection and substance identification.

[0102] In a preferred embodiment of the present invention, the microbial abundance data is obtained by gene sequencing, preferably by sequencing on an Illumina HiSeq X-10 platform.

[0103] In a preferred embodiment of the present invention, the metabolite data are analyzed by the analysis algorithm IP4M (http: / / 47.92.77.13 / ).

[0104] In a preferred embodiment of the present invention, the judgment criteria for the evaluation results are: when the output result of the model constructed based on the biomarker combination (a continuous variable between 0 and 1) is greater than or equal to the threshold of the model, the output evaluation result is "higher risk of depression"; when it is less than the model threshold, the output evaluation result is "lower risk of depression".

[0105] In a specific embodiment of the present invention, the judgment standard of the evaluation result is: when the biomarker combination consists of furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid, L-arginine, fructose, phenyllactic acid, L-tyrosine, L-proline, 3-hydroxybutyric acid, 12-hydroxystearic acid, indole-3-acetic acid methyl ester, oleic acid, 10-cis-heptadecenoic acid, palmitoleic acid, acetic acid, dodecanoic acid, nonadecenoic acid (cis-10), 2-hydroxybutyric acid, pentadecanoic acid, myristic acid, cis-5-dodecenoic acid, α-linolenic acid, 3-aminoisobutyric acid, acetylglycine, undecanoic acid, L-carnitine and linoleic acid, when the output result of the model constructed based on the biomarker combination is greater than or equal to 0.577, the output evaluation result is "higher risk of depression"; when it is less than 0.577, the output evaluation result is "lower risk of depression".

[0106] In a specific embodiment of the present invention, the judgment standard of the evaluation result is: when the biomarker combination consists of furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid and L-arginine, when the output result of the model constructed based on the marker combination is greater than or equal to 0.4, the output evaluation result is "higher risk of depression"; when it is less than 0.4, the output evaluation result is "lower risk of depression".

[0107] In a specific embodiment of the present invention, the judgment standard of the evaluation result is: when the biomarker combination consists of furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid, L-arginine, Eubacterium hallii, Blautia sp.AF19-10LB, Blautia sp.OM07-19, Faecalibacterium sp.AF28-13AC, Roseburia intestinalis, Roseburia faecis, Firmicutes bacterium AF25-13AC, Faecalibacterium prausnitzii, prausnitzii) and Eubacterium rectale (Eubacterium rectale), when the output result of the model constructed based on the marker combination is greater than or equal to 0.45, the output evaluation result is "high risk of depression"; when it is less than 0.45, the output evaluation result is "low risk of depression".

[0108] In a specific embodiment of the present invention, the judgment standard of the evaluation result is: when the biomarker combination consists of Eubacterium hallii, Blautia sp. AF19-10LB, Blautia sp. OM07-19, Faecalibacterium sp. AF28-13AC, Roseburia intestinalis, Roseburia faecis, Firmicutes bacterium AF25-13AC, Faecalibacterium prausnitzii, prausnitzii) and Eubacterium rectale (Eubacterium rectale), when the output result of the model constructed based on the marker combination is greater than or equal to 0.39, the output evaluation result is "high risk of depression"; when it is less than 0.39, the output evaluation result is "low risk of depression".

[0109] In a specific embodiment of the present invention, the judgment standard of the evaluation result is: when the biomarker combination consists of furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid, L-arginine, fructose, phenyllactic acid, L-tyrosine, L-proline, 3-hydroxybutyric acid, 12-hydroxystearic acid, indole-3-acetic acid methyl ester, oleic acid, 10-cis-heptadecenoic acid, palmitoleic acid, acetic acid, dodecanoic acid, nonadecenoic acid (cis-10), 2-hydroxybutyric acid, pentadecanoic acid, myristic acid, cis-5-dodecenoic acid, α-linolenic acid, 3-aminoisobutyric acid, acetylglycine, undecanoic acid, L-carnitine, linoleic acid, Eubacterium hallii (Eubacterium hallii) hallii), Blautia sp. AF19-10LB, Blautia sp. OM07-19, Faecalibacterium sp. AF28-13AC, Roseburia intestinalis, Roseburia faecis, Firmicutes bacterium AF25-13AC, Faecalibacterium prausnitzii and Eubacterium rectale, when the output result of the model constructed based on the marker combination is greater than or equal to 0.54, the output evaluation result is "higher risk of depression"; when it is less than 0.54, the output evaluation result is "lower risk of depression".

[0110] In a specific embodiment of the present invention, the judgment standard of the evaluation result is: when the biomarker combination consists of furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid, L-arginine, fructose, phenyllactic acid, L-tyrosine, L-proline, 3-hydroxybutyric acid, 12-hydroxystearic acid, indole-3-acetic acid methyl ester, oleic acid, 10-cis-heptadecenoic acid, palmitoleic acid, acetic acid, dodecanoic acid, nonadecenoic acid (cis-10), 2-hydroxybutyric acid, pentadecanoic acid, myristic acid, cis-5-dodecenoic acid, α-linolenic acid, 3-aminoisobutyric acid, acetylglycine, undecanoic acid, L-carnitine, linoleic acid, Eubacterium hallii (Eubacterium hallii) hallii), Blautia sp. AF19-10LB, Blautia sp. OM07-19, Faecalibacterium sp. AF28-13AC, Roseburia intestinalis, Roseburia faecis, Firmicutes bacterium AF25-13AC, Faecalibacterium prausnitzii, Eubacterium recale, Ruminococcus gnavus, Dialister sp. Marseille-P5638, Bacteroides vulgatus), Faecalibacterium sp.AF27-11BH, Hungatela hathewayl, Roseburia hominis, Blautia sp.OM05-6, Faecalibacterium sp.AM43-5AT, Ruminococcus torques, Blautia sp.AF19-1, Ruminococcus sp.AM46-18, Ruminococcus sp.AM12-48, Ruminococcus Sp.AF33-11BH, and Ruminococcus sp.When the biomarker combination is composed of: AF33-11BH), Ruminococcus Sp.AM23-1LB, Romboutsiatimonensis, Doreaformicigenerans, Eggerthella lenta, Gemmiger formicilis, and Blautia obeum, if the output of the model constructed based on the biomarker combination is greater than or equal to 0.3, the output evaluation result is "higher risk of depression"; if it is less than 0.3, the output evaluation result is "lower risk of depression."

[0111] To solve the above technical problems, the present invention provides a technical solution: a system for assessing depression risk, the system comprising:

[0112] a data receiving module, configured to receive or input biomarker data in a sample, wherein the biomarkers include the biomarker combination described in the present invention;

[0113] The judgment and output module is used to output the assessment result of the depression risk of the individual in the sample after the reception or input is completed through the model for assessing the depression risk as described in the present invention; the judgment standard of the assessment result is: when the output result of the model (a continuous variable between 0 and 1) is greater than or equal to the threshold of the model, the output assessment result is "higher risk of depression"; when it is less than the model threshold, the output assessment result is "lower risk of depression".

[0114] In a specific embodiment of the present invention, the judgment criteria of the evaluation results are:

[0115] When the biomarker combination consists of furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid, L-arginine, fructose, phenyllactic acid, L-tyrosine, L-proline, 3-hydroxybutyric acid, 12-hydroxystearic acid, indole-3-acetic acid methyl ester, oleic acid, 10-cis-heptadecenoic acid, palmitoleic acid, acetic acid, dodecanoic acid, nonadecenoic acid (cis-10), 2-hydroxybutyric acid, pentadecanoic acid, myristic acid, cis-5-dodecenoic acid, α-linolenic acid, 3-aminoisobutyric acid, acetylglycine, undecanoic acid, L-carnitine and linoleic acid, when the output result of the model constructed based on the biomarker combination is greater than or equal to 0.577, the output evaluation result is "higher risk of depression"; when it is less than 0.577, the output evaluation result is "lower risk of depression".

[0116] In a specific embodiment of the present invention, the judgment standard of the evaluation result is: when the biomarker combination consists of furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid and L-arginine, when the output result of the model constructed based on the marker combination is greater than or equal to 0.4, the output evaluation result is "higher risk of depression"; when it is less than 0.4, the output evaluation result is "lower risk of depression".

[0117] In a specific embodiment of the present invention, the judgment standard of the evaluation result is: when the biomarker combination consists of furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid, L-arginine, Eubacterium hallii, Blautia sp.AF19-10LB, Blautia sp.OM07-19, Faecalibacterium sp.AF28-13AC, Roseburia intestinalis, Roseburia faecis, Firmicutes bacterium AF25-13AC, Faecalibacterium prausnitzii prausnitzii) and Eubacterium rectale (Eubacterium rectale), when the output result of the model constructed based on the marker combination is greater than or equal to 0.45, the output evaluation result is "high risk of depression"; when it is less than 0.45, the output evaluation result is "low risk of depression".

[0118] In a specific embodiment of the present invention, the judgment standard of the evaluation result is: when the biomarker combination consists of Eubacterium hallii, Blautia sp. AF19-10LB, Blautia sp. OM07-19, Faecalibacterium sp. AF28-13AC, Roseburia intestinalis, Roseburia faecis, Firmicutes bacterium AF25-13AC, Faecalibacterium prausnitzii, prausnitzii) and Eubacterium rectale (Eubacterium rectale), when the output result of the model constructed based on the marker combination is greater than or equal to 0.39, the output evaluation result is "high risk of depression"; when it is less than 0.39, the output evaluation result is "low risk of depression".

[0119] In a specific embodiment of the present invention, the judgment standard of the evaluation result is: when the biomarker combination consists of furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid, L-arginine, fructose, phenyllactic acid, L-tyrosine, L-proline, 3-hydroxybutyric acid, 12-hydroxystearic acid, indole-3-acetic acid methyl ester, oleic acid, 10-cis-heptadecenoic acid, palmitoleic acid, acetic acid, dodecanoic acid, nonadecenoic acid (cis-10), 2-hydroxybutyric acid, pentadecanoic acid, myristic acid, cis-5-dodecenoic acid, α-linolenic acid, 3-aminoisobutyric acid, acetylglycine, undecanoic acid, L-carnitine, linoleic acid, Eubacterium hallii (Eubacterium hallii) hallii), Blautia sp. AF19-10LB, Blautia sp. OM07-19, Faecalibacterium sp. AF28-13AC, Roseburia intestinalis, Roseburia faecis, Firmicutes bacterium AF25-13AC, Faecalibacterium prausnitzii and Eubacterium rectale, when the output result of the model constructed based on the marker combination is greater than or equal to 0.54, the output evaluation result is "higher risk of depression"; when it is less than 0.54, the output evaluation result is "lower risk of depression".

[0120] In a specific embodiment of the present invention, the judgment standard of the evaluation result is: when the biomarker combination consists of furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid, L-arginine, fructose, phenyllactic acid, L-tyrosine, L-proline, 3-hydroxybutyric acid, 12-hydroxystearic acid, indole-3-acetic acid methyl ester, oleic acid, 10-cis-heptadecenoic acid, palmitoleic acid, acetic acid, dodecanoic acid, nonadecenoic acid (cis-10), 2-hydroxybutyric acid, pentadecanoic acid, myristic acid, cis-5-dodecenoic acid, α-linolenic acid, 3-aminoisobutyric acid, acetylglycine, undecanoic acid, L-carnitine, linoleic acid, Eubacterium hallii (Eubacterium hallii) hallii), Blautia sp. AF19-10LB, Blautia sp. OM07-19, Faecalibacterium sp. AF28-13AC, Roseburia intestinalis, Roseburia faecis, Firmicutes bacterium AF25-13AC, Faecalibacterium prausnitzii, Eubacterium recale, Ruminococcus gnavus, Dialister sp. Marseille-P5638, Bacteroides vulgatus), Faecalibacterium sp.AF27-11BH, Hungatela hathewayl, Roseburia hominis, Blautia sp.OM05-6, Faecalibacterium sp.AM43-5AT, Ruminococcus torques, Blautia sp.AF19-1, Ruminococcus sp.AM46-18, Ruminococcus sp.AM12-48, Ruminococcus Sp.AF33-11BH, and Ruminococcus sp.When the biomarker combination is composed of: AF33-11BH), Ruminococcus Sp.AM23-1LB, Romboutsiatimonensis, Dorea formicigenerans, Eggerthella lenta, Gemmiger formicilis, and Blautia obeum, if the output of the model constructed based on the biomarker combination is greater than or equal to 0.3, the output assessment result is "high risk of depression"; when it is less than 0.3, the output risk result is "low risk of depression."

[0121] In a preferred embodiment of the present invention, the system further comprises a data processing module for collecting biomarker data.

[0122] In a preferred embodiment of the present invention, in the biomarker data, the metabolite data is the serum concentration of the metabolite, and the microbial data is the abundance of the microorganism in the feces, and the abundance refers to the relative abundance of the microorganism.

[0123] In a preferred embodiment of the present invention, the metabolite data is obtained by HPLC-MS technology, preferably using a metabolic chip in the patent application with patent application number 201910170381.0 for detection and substance identification.

[0124] In a preferred embodiment of the present invention, the microbial abundance data is obtained by gene sequencing, preferably by sequencing on an Illumina HiSeq X-10 platform.

[0125] In a preferred embodiment of the present invention, the metabolite data are analyzed by the analysis algorithm IP4M (http: / / 47.92.77.13 / ).

[0126] To solve the above technical problems, the present invention provides a technical solution: a computer-assisted method for assessing depression risk, the method comprising the following steps:

[0127] Step 1: receiving or inputting biomarker data in a sample, wherein the biomarker data includes data of the biomarker combination described in the present invention;

[0128] Step 2: Input the biomarker data received or input in step 1 into the model for assessing depression risk as described in the present invention, and output an assessment result of the depression risk of the individual in the sample, wherein the judgment criteria of the assessment result are:

[0129] When the output result of the model constructed based on the biomarker combination is greater than or equal to the model threshold, the output evaluation result is "higher risk of depression"; when it is less than the model threshold, the output evaluation result is "lower risk of depression".

[0130] In the present invention, the threshold value can be easily determined by those skilled in the art and does not change with changes in the training or validation database.

[0131] In a preferred embodiment of the present invention, the method for assessing the risk of depression further comprises step 0: collecting biomarker data from tissue samples.

[0132] In a preferred embodiment of the present invention, the metabolite data is obtained by HPLC-MS technology, preferably using a metabolic chip in the patent application with patent application number 201910170381.0 for detection and substance identification.

[0133] In a preferred embodiment of the present invention, the microbial abundance data is obtained by gene sequencing, preferably by sequencing on an Illumina HiSeq X-10 platform.

[0134] In a preferred embodiment of the present invention, the metabolite data are analyzed by the analysis algorithm IP4M (http: / / 47.92.77.13 / ).

[0135] To solve the above technical problems, the present invention provides a technical solution: a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the functions of the system as described in the present invention, or implement the steps of the method for assessing depression risk as described in the present invention.

[0136] To solve the above technical problems, the present invention provides a technical solution: an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the functions of the system as described in the present invention, or the steps of the method for assessing depression risk as described in the present invention.

[0137] On the basis of conforming to the common sense in this field, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present invention.

[0138] The reagents and raw materials used in the present invention are commercially available.

[0139] The positive progress effect of the present invention is:

[0140] The present invention provides a method for constructing a model for assessing depression risk, a model for assessing depression risk, and a system and application for assessing depression risk. The technical solution of the present invention can assess depression risk based on the Hamilton Depression Rating Scale and clinical judgment, demonstrating excellent risk assessment capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0141] Figure 1 The abundance fold changes of 28 bacteria.

[0142] Figure 2 Comparison between the GB model and other models.

[0143] Figure 3 This is the ROC curve of the “37 points” training results for cohort 1 serum samples.

[0144] Figure 4 This is the ROC curve of the self-test results of model 1 in male and female serum samples of cohort 1.

[0145] Figure 5 ROC curve for the internal validation results of model 1

[0146] Figure 6 The ROC curve of the results of model 1 in the validation set of cohort 2 and cohort 3 samples.

[0147] Figure 7 ROC curve for a single metabolite.

[0148] Figure 8 Traversing the combinatorial graph for 31 metabolites.

[0149] Figure 9 This is the ROC curve of the training results of model 2.

[0150] Figure 10 ROC curve for the internal validation results of Model 2.

[0151] Figure 11 The ROC curve of the results of model 2 in the validation set of cohort 2 and cohort 3 samples.

[0152] Figure 12 ROC curves for model 3 training results and internal validation results.

[0153] Figure 13 The ROC curve of the results of model 3 in the validation set of cohort 2 and cohort 3 samples.

[0154] Figure 14 Traverse the combined graph for metabolites and bacterial communities together.

[0155] Figure 15 This is the ROC curve of the training results of model 4.

[0156] Figure 16 ROC curve for the internal validation results of Model 4.

[0157] Figure 17 The ROC curve of the results of model 4 in the validation set of cohort 2 and cohort 3 samples.

[0158] Figure 18 This is the ROC curve of the training results of model 5.

[0159] Figure 19 ROC curve for the internal validation results of model 5.

[0160] Figure 20 The ROC curve of the results of model 5 in the validation set of cohort 2 and cohort 3 samples.

[0161] Figure 21 This is the ROC curve of the training results of model 6.

[0162] Figure 22 ROC curve for the internal validation results of Model 6.

[0163] Figure 23 The ROC curve of the results of model 6 in the validation set of cohort 2 and cohort 3 samples.

[0164] Figure 24 Schematic diagram of the system for assessing depression risk.

[0165] Figure 25 A schematic diagram of the structure of an electronic device. DETAILED DESCRIPTION

[0166] The present invention is further illustrated by way of examples below, but the present invention is not limited to the scope of the examples. Experimental methods in the following examples where specific conditions are not specified were performed according to conventional methods and conditions, or selected according to the product specifications.

[0167] Example 1 Collection of samples

[0168] Diagnostic criteria for depression: The diagnosis of depression is usually based on international diagnostic criteria such as the Hamilton Depression Rating Scale (HAMD) score, the Diagnostic and Statistical Manual of Mental Disorders (DSM-5), or the International Classification of Diseases (ICD-10 or ICD-11). Doctors will use these criteria to determine whether the diagnosis of depression is met.

[0169] Specifically: The clinical diagnosis process of depression in this application includes the following steps:

[0170] ① Initial assessment: Diagnosing depression usually begins with an initial consultation between the patient and a healthcare provider. During this consultation, the doctor will ask questions about the patient's symptoms, medical history, family history, and possible life events to understand the patient's situation.

[0171] ②Symptom assessment: The doctor will assess the patient's symptoms, which usually include the following main symptoms of depression:

[0172] Persistent depression, low mood, or feelings of emptiness;

[0173] loss of interest or favorite activities;

[0174] Irritability or irritability;

[0175] Insomnia or excessive sleeping;

[0176] Easy fatigue or lack of energy;

[0177] Feelings of self-guilt or worthlessness;

[0178] difficulty concentrating or making decisions;

[0179] thoughts of death or suicide;

[0180] Rule out other medical conditions: Your doctor will rule out other medical or mental health conditions that may cause similar symptoms, such as thyroid problems, substance abuse, or other mental disorders.

[0181] ③ Severity assessment: Doctors will usually assess the severity of depression, which may help determine treatment options. The severity of depression can range from mild to moderate to severe, depending on the number and severity of symptoms and their impact on daily life.

[0182] Depressed and non-depressed healthy individuals were collected. All samples were collected with informed consent from the subjects, and the collection process was standardized. This experiment was approved by the ethics committee. Depressed and non-depressed individuals were determined based on clinical diagnosis by practicing physicians. Samples were obtained from the Department of Psychiatry and Psychology, First Affiliated Hospital of Shanxi Medical University. All samples were divided into two groups, each containing both healthy and depressed individuals. Details are shown in Table 1.

[0183] Table 1: Sample details table

[0184]

[0185]

[0186] Example 2 Detection of fecal flora abundance

[0187] Total microbial genomic DNA was extracted from stool samples of clinical cohort 1 using the DNeasy PowerSoil kit (QIAGEN, Valencia, CA, USA). The quantity and quality of the extracted DNA were measured using a nano drop ND-1000 spectrophotometer (Thermo Fisher Scientific, Waltham, MA, USA) and electrophoresis gels. DNA samples for quality control were used to construct metagenomic shotgun libraries using the Illumina TruSeq Nano DNA LT Library Preparation Kit. The PE150 strategy for each library was sequenced on an Illumina HiSeq X-10 platform (Illumina, San Diego, CA, USA) and Personal Biotechnology Co. (Shanghai, China). Fastqc was used as the raw analytical read quality control. Adapter classifications were removed from the classification reads using Cutadapt (V1.2.1) (NBIS, Uppsala, Sweden). For low-quality reads, ambiguous bases were removed using a sliding window algorithm. Sequencing reads were compared to the host genome, and host genome reads were removed. The quality-filtered reads were then reassembled to construct metagenomes for each sample. Scaffold / Scaftigs sequences longer than 300 bp were selected for each sample and constructed using megahit (https: / / hku-bal.github.io / megabox / ).

[0188] Calculation method of bacterial abundance:

[0189] 1. Data acquisition: Obtain sample DNA and perform whole-genome sequencing, usually using high-throughput sequencing technology (such as Illumina sequencing).

[0190] 2. Quality control and filtering: Perform quality control on sequencing data, including removing low-quality sequences, removing adapters, removing contaminants, etc., to ensure high-quality sequences.

[0191] 3. Sequence alignment: Use alignment tools to align the cleaned sequence with a known bacterial genome database. This can be done using tools such as Bowtie2 or BWA-MEM.

[0192] 4. Read Counts: After aligning each genome, calculate the read counts of each bacterial genome, that is, the number of sequences in the sequencing data that match the genome.

[0193] 5. Abundance calculation: Convert read counts to relative abundance. This usually involves dividing the read counts of each bacterium by the total read counts to obtain relative abundance.

[0194] Example 3 Detection of serum metabolites

[0195] The metabolic chip, which is patent pending with patent number 201910170381.0, was used for detection and substance identification. The specific method is as follows:

[0196] Serum metabolites were determined using the Q300 kit (Huiyun Biotechnology, Shenzhen, China) according to the following protocol:

[0197] (1) Remove the calibrator I, calibrator II, internal standard I working solution, internal standard II, derivatization reagent, derivatization reagent diluent, and EDC dispersant from the kit and rewarm them at room temperature for at least 30 minutes.

[0198] (2) Take out calibrator I and internal standard II and centrifuge them at 8000g for 15 minutes.

[0199] (3) Prepare 10 mL of 75% methanol solution and 40 mL of 50% methanol solution; take 100 mL of methanol for later use.

[0200] (4) Prepare 3 liquid storage tanks (optional, used when adding samples using a discharge gun).

[0201] (5) Preparation of calibrators: Add 200 μL of methanol to calibrator I and shake at 1200 rpm for 20 minutes at room temperature. Then add 50 μL of calibrator II and shake at 1200 rpm for 20 minutes at room temperature (to ensure that the calibrator is fully dissolved). Take 100 μL of the reconstituted calibrator mixture and dilute it with 75% methanol solution as follows: 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 160, 1 / 320. (A total of 8 concentration gradients of calibrator mixtures are obtained)

[0202] (6) Preparation of internal standard II working solution: Take internal standard II and add 1.2 mL of 50% methanol solution. Oscillate at 1200 rpm at room temperature for 20 minutes to obtain internal standard II working solution.

[0203] (7) Take out a 350 μL 96-well V-bottom plate, add 20 μL of deionized water to well A1, add 20 μL of calibrator mixed solution to wells A2 to A9 in descending order of concentration, and add 20 μL of serum (or plasma) sample to the other wells.

[0204] (8) Add 120 μL of internal standard I working solution to each well, cover with aluminum film, and place on a thermomixer at 10°C and 650 rpm for 20 minutes.

[0205] (9) Centrifuge the plate at 4000 g for 20 minutes using a microplate centrifuge.

[0206] (10) Preparation of derivatization reagent working solution: Take 2.4 mL of derivatization reagent diluent, add it to the derivatization reagent bottle, shake to dissolve, and set aside (ultrasound can be used to assist dissolution if necessary).

[0207] (11) Preparation of EDC working solution: Weigh 110 mg of EDC powder, add 2.4 mL of EDC diluent, shake to dissolve, and set aside (EDC working solution needs to be prepared and used immediately)

[0208] (12) After centrifugation in step 9, carefully remove the aluminum film to prevent the liquid in the 96-well plate from splashing.

[0209] (13) Carefully pipette 30 μL of supernatant from the septum of the above 96-well plate into a new 1 mL 96-well round-bottom plate, and add 20 μL of derivatization reagent working solution and 20 μL of EDC working solution in sequence.

[0210] (14) Cover with aluminum film and place the 96-well plate in a constant temperature oscillator at 30°C and 1200 rpm for 60 minutes.

[0211] (15) After the reaction was completed, 330 μL of 50% methanol solution was added to each well and centrifuged at 4000 g at 4°C for 30 minutes.

[0212] (16) Take a new 350 μL 96-well V-bottom plate and add 10 μL of internal standard II working solution to each well.

[0213] (17) Transfer 140 μL of the supernatant to the microplate pre-filled with the internal standard II working solution, mix at 650 rpm for 5 minutes, and cover with a film.

[0214] (18) The microplate is placed in the automatic sampler of the ultra-high performance liquid chromatography tandem mass spectrometry system for testing.

[0215] Example 4 Screening Method for Metabolites in Serum

[0216] Unidimensional and multidimensional statistical analyses of metabolite and microbial abundance data are important tools for studying the relationship between microbial composition and metabolites in ecosystems and organisms. These analyses can provide information about potential associations between microbial communities and metabolites within an organism. The following are the basic principles and purposes of unidimensional and multidimensional statistical analyses:

[0217] Principles of unidimensional statistical analysis:

[0218] 1) Univariate analysis: Univariate statistical analysis focuses on the properties and variation of a single variable (e.g., a single metabolite or a specific bacterial group). Common methods include t-tests, analysis of variance, and nonparametric tests.

[0219] 2) Correlation analysis: Correlation analysis can be used to assess the linear or nonlinear relationship between metabolites and bacterial abundance. Commonly used methods include the Pearson correlation coefficient and the Spearman rank correlation coefficient.

[0220] 3) Differential analysis: Differential analysis is used to detect significant differences in metabolite or bacterial abundance between two or more groups. In metabolomics, methods such as t-test, ANOVA, and Wilcoxon can be used.

[0221] Purpose of unidimensional statistical analysis:

[0222] 1) Searching for biomarkers: Through univariate analysis, we can find biomarkers that show significant differences at the metabolite or microbiome levels. These biomarkers may be associated with the physiological state or disease state in the organism.

[0223] 2) Understanding the relationship between variables: Through correlation analysis, the potential relationship between metabolites and microbiota can be identified, thereby inferring their interactions in physiological processes.

[0224] Principles of multidimensional statistical analysis:

[0225] 1) Multivariate statistical techniques: Multidimensional statistical analysis focuses on the relationships between multiple variables. Commonly used methods include principal component analysis (PCA), partial least squares regression (PLS-R), cluster analysis, and redundancy analysis (RDA).

[0226] 2) Ecological statistical methods: These methods can simultaneously consider multiple metabolites and bacterial communities, revealing their structure and dynamics in complex ecosystems.

[0227] Purpose of multidimensional statistical analysis:

[0228] 1) Holistic interpretation of data: Multidimensional statistical analysis can help to holistically interpret the composition of metabolites and microbiota and reveal the relationships between them.

[0229] 2) Prediction model: For modeling the relationship between metabolites and microbiota, for example, through PLS-R, a prediction model can be constructed to predict the information of one variable (such as metabolites) based on another variable (such as microbiota).

[0230] 3) Discover potential ecological patterns: Through cluster analysis and ecological statistical methods, potential ecological patterns between metabolites and microbial communities can be discovered, revealing the ecosystem dynamics between microbial composition and metabolites in organisms.

[0231] After accurately quantifying metabolite concentrations in serum from Cohorts 1 and 2 using the Q300 assay, unidimensional and multidimensional statistical analyses of metabolites were performed using the data analysis system based on the gas / liquid chromatography-mass spectrometry platform described in patent CN109061020A (i.e., the analysis algorithm IP4M (http: / / 47.92.77.13 / )). This included calculating significant differences in metabolites between the healthy control and depression groups and VIP values from PLSDA analysis. For Cohort 1, the union of substances with significant differences and VIP values greater than 1 was selected, and the corresponding substances were obtained using the same method for Cohort 2. The intersection of substances with consistent change trends in Cohorts 1 and 2, totaling 31 substances, was selected as target metabolites (see Table 2).

[0232] Table 2: Chinese-English comparison table of 31 metabolites

[0233]

[0234]

[0235] Example 5 Screening method for fecal flora

[0236] 146 clinical stool samples were collected to obtain the corresponding metagenomic microbiome data. Due to the large variety of microbiome species, the top 0.01% bacteria were selected (see Figure 1 ), as the research object, according to the existing public analysis website (https: / / www.microbiomeanalyst.ca / MicrobiomeAnalyst / home.xhtml), the species level Lefse analysis of the healthy control group and the depression group was performed (see Figure 1 and Table 3 ), to obtain differential bacteria.

[0237] Table 3: LDA values of various bacterial species

[0238]

[0239]

[0240] Example 6 Establishment of a model where the marker is a metabolite

[0241] Serum data for 31 metabolites were used as feature variables (biomarker combination 1), and the serum data from cohort 1 in Example 1 was used as the modeling data. 70% of the cohort 1 serum data served as the training set, and 30% served as the test set. The training set included 109 depression samples and 63 healthy controls; the test set included 46 depression samples and 27 healthy controls. A seven-fold cross-validation method was used to build the model (Model 1). The model algorithm was the Gradient Descent (GB) model (key parameters are shown in Table 5). Cohort 2 and cohort 3 serum data were used as independent validation datasets for validation, assessing the risk of depression and evaluating model performance.

[0242] In addition, other models were used to build models and evaluate model performance. The key parameters of other models are shown in Table 4. The GB model has a higher AUC and better performance than other models (see Figure 2 ).

[0243] Table 4: Key parameters of other models

[0244]

[0245] Program code of GB model:

[0246]

[0247]

[0248] #best is the best threshold point, there may be two, when filling in the result of permutations and combinations, only the first one is selected to avoid errors

[0249] The #coords() function can not only select the best cutoff value, but also find the corresponding point or specificity or sensitivity related results according to requirements

[0250] lg<-as.numeric(length(pred_GB))

[0251] for(s in 1:lg){

[0252] if(pred_GB[s] <best_cut[1]){data1G$pred_Y[s]<-0}else{data1G$pred_Y[s]<-1}

[0253] }

[0254] tab<-table(data1G$Y,data1G$pred_Y)

[0255] accu<-sum(diag(prop.table(tab)))

[0256] P_a<-diag(prop.table(tab,2))[1]

[0257] P_b<-diag(prop.table(tab,2))[2]

[0258] R_a<-diag(prop.table(tab,1))[1]#Specificity

[0259] R_b<-diag(prop.table(tab,1))[2]#Sensitivity

[0260] F1_a<-2*P_a*R_a / (P_a+R_a)

[0261] F1_b<-2*P_b*R_b / (P_b+R_b)

[0262] pr<-pr.curve(pred_GB[data1G$Y==1]%>%na.omit, pred_GB[data1G$Y==0]%>%na.omit)

[0263] auPR<-pr$auc.integral

[0264] row<-1

[0265] rslt[row,1]<-accu

[0266] rslt[row,2]<-P_a

[0267] rslt[row,3]<-P_b

[0268] rslt[row,4]<-R_a

[0269] rslt[row,5]<-R_b

[0270] rslt[row,6]<-F1_a

[0271]

[0272] Table 5: Parameters of GB model

[0273] Key parameters Setting value significance Data train training set Distribution bernoulli The form of the loss function is, n.trees 100 Number of iterations Interaction.depth 1 Depth of decision tree shrinkage 0.03 Learning rate n.minobsinnode 4 The minimum number of training set samples in a node to start splitting

[0274] The results show that the training set result of model 1 has an AUC of 0.85, which is superior (see Figure 3 and Figure 4), the internal validation set result AUC=0.81 (see Figure 5 ), the independent validation set result of cohort 2 serum data was AUC = 0.8, indicating that the model has good robustness and strong generalization ability (see Figure 6 ). The independent validation set result of cohort 3 serum data is AUC = 0.73, although the performance has decreased (see Figure 6 ), but considering that the serum data of cohort 3 are data after medication, the model performance is still good in general (generally speaking, various drugs have a certain impact on human metabolism, and the model's risk assessment performance for patients taking medication is often reduced, but the model of the present invention is still effective for risk assessment of patients taking medication). The model constructed by combining 31 metabolites has a larger AUC value and better risk assessment performance than the model constructed by each metabolite alone (also using the serum data of cohort 1 as the modeling data) (see Figure 7 ).

[0275] Judgment criteria for model 1: When the output result of the model constructed based on the biomarker combination 1 (a continuous variable between 0 and 1) is greater than or equal to 0.577, the output evaluation result is "higher risk of depression"; when it is less than 0.577, the output evaluation result is "lower risk of depression".

[0276] In model 1, 31 metabolites were selected and 8 metabolites were selected as the best combination of markers (see Figure 8 ), including furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamic acid, indolepropionic acid, homovanillic acid, and L-arginine. 70% of the samples in the serum data of cohort 1 were used as the training set, and 30% of the samples were used as the test set. The training set included 109 depression samples and 64 healthy controls; the internal validation set included 46 depression samples and 27 healthy controls. The risk assessment performance of model 2 for these 8 metabolites (marker combination 2) is shown in Figure 9 and Figure 10 The AUC value of the training set was 0.88, and the AUC value of the internal validation set was 0.83. The serum data of cohorts 2 and 3 were used as external validation sets for verification. The results are as follows Figure 11 As shown, the AUC value of the validation set results for the serum data of cohort 2 was 0.80, and the AUC value of the validation set results for the serum data of cohort 3 was 0.72. The judgment criteria of this model are that when the output result (a continuous variable between 0 and 1) of the model constructed based on the marker combination 2 is greater than or equal to 0.4, the output evaluation result is "higher risk of depression"; when it is less than 0.4, the output evaluation result is "lower risk of depression".

[0277] Example 7 Establishment of a model with metabolites and bacterial flora as markers

[0278] 146 samples (75 healthy people and 71 depressed people) from cohort 1 with both microbiome and serum data were used as modeling data. 70% of the samples were used as training sets and 30% of the samples were used as validation sets. The model algorithm was the GB model (gradient descent algorithm, with the same parameters as model 1). In the training set, there were 53 healthy people and 50 depressed people; in the internal validation set, there were 22 healthy people and 21 depressed people. 31 metabolites and 28 bacteria (Table 6) were used as biomarker combination 3 feature variables (Model 3). The AUC of the 146 sample training set was 0.89, and the AUC of the internal validation set was 0.81, indicating excellent model performance (see Figure 12 ). The data of cohort 2 and cohort 3 samples were used as external validation sets for verification, and the results are as follows Figure 13 As shown, the AUC value of the validation set of sample data from cohort 2 was 0.77, and the AUC value of the validation set of sample data from cohort 3 was 0.70. The judgment criteria of this model are: when the output result (a continuous variable between 0 and 1) of the model constructed based on the biomarker combination 3 is greater than or equal to 0.3, the output evaluation result is "higher risk of depression"; when it is less than 0.3, the output evaluation result is "lower risk of depression."

[0279] Table 6: Chinese-English comparison table of 28 bacteria

[0280] Chinese name English name Eubacterium rectale Eubacterium rectale Faecalibacterium prausnitzii Faecalibacterium prausnitzii Ruminococcus activus Ruminococcus gnavus Roseburia intestinalis Roseburia intestinalis Bacillus sp. Marseille-P5638 Dialister sp. Marseille-P5638 Roseola faecalis Roseburia faecis Bacteroides vulgaris Bacteroides vulgatus Faecalibacterium AF27-11BH Faecalibacterium sp.AF27-11BH Hungatela hathewayl (no Chinese name) Hungatella hathewayl Roseburia hominis Roseburia hominis Firmicutes bacteria AF25-13AC Firmicutes bacterium AF25-13AC Eubacterium mucosum sp.OM05-6 Blautia sp.OM05-6 Faecalibacterium AM43-5AT Faecalibacterium sp.AM43-5AT Ruminococcus contortus Ruminococcus torques Eubacterium mucosum sp.AF19-1 Blautia sp.AF19-1 Ruminococcus AM46-18 Ruminococcus sp.AM46-18 Ruminococcus AM12-48 Ruminococcus sp.AM12-48 Faecalibacterium AF28-13AC Faecalibacterium sp.AF28-13AC Ruminococcus Sp.AF33-11BH Ruminococcus Sp.AF33-11BH Ruminococcus Sp.AM23-1LB Ruminococcus Sp.AM23-1LB Romboutsia timonensis (no Chinese name) Romboutsia timonensis Eubacterium muciniphilum sp.OM07-19 Blautia sp.OM07-19 Dorea formicigenerans (no Chinese name) Dorea formicigenerans Eggertella tarda Eggerthella lenta Eubacterium mucilaginosum sp.AF19-10LB Blautia sp.AF19-10LB Formic acid budding bacteria Gemmiger formicilis Eubacterium hallii Eubacterium hallii Blautia ovata Blautia obeum

[0281] The 28 bacteria in biomarker combination 3 and the 8 metabolites in biomarker combination 2 in Example 6, a total of 36 variables, were traversed together and 9 bacteria were selected, that is, a total of 17 variables (8 metabolites and 9 bacteria) as the best combination of markers (biomarker combination 4, see Figure 14 ), bacteria include Eubacterium hallii, Eubacterium viscidum sp.AF19-10LB, Eubacterium viscidum sp.OM07-19, Faecalibacterium genus AF28-13AC, intestinal Roseburia, fecal Roseburia, Firmicutes bacteria AF25-13AC, Faecalibacterium prausnitzii and Eubacterium rectum; increasing or decreasing the markers will lead to a decrease in model performance. Consistent with the modeling data of model 3, the microbiome and serum data in cohort 1 were used as modeling data, 70% of the samples were used as training sets, and 30% of the samples were used as test sets. Among them, in the training set, there were 109 healthy people and 64 cases of depression; in the internal validation set, there were 27 healthy people and 46 cases of depression. The risk assessment performance of model 4 with 8 metabolites + 9 bacteria (biomarker combination 4) is shown in Figure 15 and Figure 16 The AUC of the training set was 0.95, and the AUC of the internal validation set was 0.85. The data of cohort 2 and cohort 3 samples were used as external validation sets for verification. The results are as follows Figure 17As shown, the AUC value of the validation set of sample data from cohort 2 is 0.81, and the AUC value of the validation set of sample data from cohort 3 is 0.73. The judgment standard of this model is that when the output result of the model constructed based on the marker combination 4 (a continuous variable between 0 and 1) is greater than or equal to 0.45, the output evaluation result is "higher risk of depression"; when it is less than 0.45, the output evaluation result is "lower risk of depression".

[0282] The training set consisted of 109 healthy individuals and 64 depressed individuals; the internal validation set consisted of 27 healthy individuals and 46 depressed individuals. The risk assessment performance of Model 5 using 9 bacteria (biomarker combination 5, i.e., the 9 bacteria in biomarker combination 4) is shown in Table 5. Figure 18 and Figure 19 The AUC of the training set was 0.84, and the AUC of the internal validation set was 0.76. The data of cohort 2 and cohort 3 samples were used as external validation sets for verification. The results are as follows Figure 20 As shown, the AUC value of the validation set of sample data from cohort 2 was 0.76, and the AUC value of the validation set of sample data from cohort 3 was 0.71. The judgment criteria of this model are that when the output result (a continuous variable between 0 and 1) of the model constructed based on the biomarker combination 5 is greater than or equal to 0.39, the output evaluation result is "higher risk of depression"; when it is less than 0.39, the output evaluation result is "lower risk of depression".

[0283] The training set included 109 healthy subjects and 64 depressed subjects; the internal validation set included 27 healthy subjects and 46 depressed subjects. The risk assessment performance of model 6 for 31 metabolites and 9 bacteria (biomarker combination 6, including all metabolite and biomarker combinations 5 in Table 2) is shown in Figure 21 and Figure 22 The AUC of the training set was 0.91, and the AUC of the internal validation set was 0.82. The data of cohort 2 and cohort 3 samples were used as external validation sets for verification. The results are as follows Figure 23 As shown, the AUC value of the validation set of sample data from cohort 2 was 0.79, and the AUC value of the validation set of sample data from cohort 3 was 0.71. The judgment standard of this model is that when the output result (a continuous variable between 0 and 1) of the model constructed based on the biomarker combination 6 is greater than or equal to 0.54, the output evaluation result is "higher risk of depression"; when it is less than 0.54, the output evaluation result is "lower risk of depression".

[0284] Example 8 Clinical Sample Evaluation

[0285] The methods described in Examples 2 and 3 were used to detect the abundance of fecal flora and the concentration of metabolites in serum of clinical patients to obtain marker data for the clinical patients. The obtained data were then input into the models described in Examples 6 or 7, and the output was compared with the cutoff values of each model to obtain a depression risk assessment result for the clinical patient. The depression risk assessment results for each clinical patient using the models described in Examples 6 and 7 were consistent with the actual clinical diagnosis results.

[0286] Example 9: System for Assessing Depression Risk

[0287] The system 61 for assessing depression risk includes a data receiving module 52 and a judgment and output module 53, and preferably also includes a data processing module 51 (see Figure 24 ).

[0288] The data processing module 51 is used to collect data on a biomarker combination (any of the biomarker combinations 1-6 described above) in a patient's sample to be tested and transmit the data to the data processing module.

[0289] The data receiving module 52 is configured to analyze the received or input biomarker combination data according to the data analysis method described in Example 6 or Example 7 to obtain a calculation result. The biomarker combination data can be collected by the data collecting module 51 or obtained from other sources.

[0290] The judgment and output module 53 is used to judge whether the calculation result meets the preset judgment conditions, that is, when the output result of the model (a continuous variable between 0 and 1) is greater than or equal to the threshold of the model (the point with the highest sensitivity + specificity is determined as the threshold and remains unchanged), the output evaluation result is "high risk of depression"; when it is less than the model threshold, the output evaluation result is "low risk of depression".

[0291] Example 10 Electronic device

[0292] This embodiment provides an electronic device, which can be expressed in the form of a computing device (for example, a server device), including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for assessing the risk of depression in Examples 6 and 7 of the present invention can be implemented.

[0293] Figure 25 The hardware structure diagram of this embodiment is shown. The electronic device 9 specifically includes:

[0294] At least one processor 91, at least one memory 92, and a bus 93 for connecting different system components (including the processor 91 and the memory 92), wherein:

[0295] The bus 93 includes a data bus, an address bus, and a control bus.

[0296] The memory 92 includes a volatile memory, such as a random access memory (RAM) 921 and / or a cache memory 922 , and may further include a read-only memory (ROM) 923 .

[0297] Memory 92 also includes a program / utility 925 having a set (at least one) of program modules 924, such program modules 924 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0298] The processor 91 executes various functional applications and data processing by running computer programs stored in the memory 92 .

[0299] The electronic device 9 can further communicate with one or more external devices 94 (e.g., a keyboard, pointing device, etc.). Such communication can be performed via an input / output (I / O) interface 95. Furthermore, the electronic device 9 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 96. The network adapter 96 communicates with other modules of the electronic device 9 via a bus 93. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device 9, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.

[0300] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the present application, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0301] Example 11 Computer-readable storage medium

[0302] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the method for assessing depression risk in embodiment 6 or embodiment 7 of the present invention are implemented.

[0303] The readable storage medium may include, but is not limited to, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0304] In a possible embodiment, the present invention can also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of the method for assessing depression risk in Example 6 or Example 7 of the present invention.

[0305] The program code for executing the present invention may be written in any combination of one or more programming languages, and may be executed entirely on the user device, partially on the user device, as an independent software package, partially on the user device and partially on a remote device, or entirely on the remote device.

Claims

1. A biomarker combination, characterized in that: The biomarker panel includes the following metabolites: furanoic acid, L-phenylalanine, 5-hydroxytryptamine, L-tryptophan, L-glutamate, indolepropionic acid, homovanillic acid, and L-arginine; and / or the biomarker combination includes the following microorganisms: Eubacterium hallii, Blautia sp. AF19-10LB, Blautia sp. OM07-19, Faecalibacterium sp. AF28-13AC, Roseburia intestinalis, Roseburia faecis, Firmicutes bacterium AF25-13AC, Faecalibacterium prausnitzii, and Eubacterium rectale; Preferably, the biomarkers further include one or more of the following metabolites: fructose, phenyllactic acid, L-tyrosine, L-proline, 3-hydroxybutyric acid, 12-hydroxystearic acid, indole-3-acetic acid methyl ester, oleic acid, 10-cis-heptadecenoic acid, palmitoleic acid, acetic acid, dodecanoic acid, nonadecenoic acid (cis-10), 2-hydroxybutyric acid, pentadecanoic acid, myristic acid, cis-5-dodecenoic acid, α-linolenic acid, 3-aminoisobutyric acid, acetylglycine, undecanoic acid, L-carnitine and linoleic acid; And / or, the biomarker further includes one or more of the following microorganisms: Ruminococcus gnavus, Dialister sp. Marseille-P5638, Bacteroides vulgatus, Faecalibacterium sp. AF27-11BH, Hungatela hathewayl, Roseburia hominis, Blautia sp. OM05-6, Faecalibacterium sp. AM43-5AT, Ruminococcus torques, Blautia sp. AF19-1, Ruminococcus AM46-18 sp.AM46-18), Ruminococcus sp.AM12-48, Ruminococcus sp.AF33-11BH, Ruminococcus sp.AM23-1LB, Romboutsia timonensis, Dorea formicigenerans, Eggerthella lenta, Gemmiger formicilis, and Blautia obeum.

2. A reagent for detecting the biomarker combination according to claim 1.

3. A kit, characterized in that The kit comprises the biomarker combination according to claim 1 or the reagent according to claim 2.

4. Use of the biomarker combination according to claim 1, the reagent according to claim 2, or the kit according to claim 3 in the preparation of a product for assessing the risk of depression.

5. A method for constructing a model for assessing the risk of depression, characterized in that: The construction method comprises: Inputting biomarker data in a biomarker database into a GB model for machine learning to construct the model for assessing depression risk; the biomarker data in the biomarker database are derived from samples of patients with depression and samples of patients without depression; the biomarkers include the biomarker combination according to claim 1; Preferably, in the biomarker data, the metabolite data is the serum concentration of the metabolite, and the microbial data is the abundance of the microorganism in feces.

6. A model for assessing the risk of depression, characterized in that: The model for assessing depression risk is constructed by the construction method as described in claim 5.

7. A method for assessing the risk of depression in a sample, characterized in that: The method comprises inputting biomarker data of a sample to be tested into a model for assessing depression risk as described in claim 6 to obtain an assessment result of the sample's depression risk; the biomarker comprises the biomarker combination as described in claim 1; and / or the method is for non-diagnostic purposes; Preferably, in the biomarker data, the metabolite data is the serum concentration, and the microbial data is the abundance of microorganisms in feces; More preferably, the judgment standard of the result is: when the output result of the model constructed based on the biomarker combination is greater than or equal to the threshold of the model, the output evaluation result is "high risk of depression"; When it is less than the model threshold, the output evaluation result is "low risk of depression".

8. A system for assessing the risk of depression, characterized in that The system comprises: a data receiving module, configured to receive or input biomarker data in a sample, wherein the biomarker comprises the biomarker combination according to claim 1; a judgment and output module, configured to output, after the receiving or inputting is completed, an assessment result of the depression risk of the individual in the sample using the depression risk assessment model according to claim 6; wherein the judgment standard of the assessment result is: when the output result of the model is greater than or equal to the model threshold, the assessment result is output as "high depression risk"; when it is less than the model threshold, the assessment result is output as "low depression risk"; Preferably, the system further comprises a data processing module for collecting the biomarker data; preferably, in the biomarker data, the metabolite data is the serum concentration of the metabolite, and the microbial data is the abundance of the microorganism in feces.

9. A computer-assisted method for assessing the risk of depression, characterized in that: The method for assessing the risk of depression comprises the following steps: Step 1: receiving or inputting biomarker data in a sample, wherein the biomarker data includes data of the biomarker combination according to claim 1; Step 2: Input the biomarker data received or input in step 1 into the model for assessing depression risk as described in claim 6, and output an assessment result of the depression risk of the individuals in the sample; when the output result of the model is greater than or equal to the threshold of the model, the output assessment result is "higher risk of depression"; when it is less than the model threshold, the output assessment result is "lower risk of depression".

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement the functions of the system according to claim 8 or the steps of the method for assessing depression risk according to claim 9.

11. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to execute the computer program to implement the functions of the system according to claim 8 or the steps of the method for assessing depression risk according to claim 9.

Citation Information

Patent Citations

  • Data analysis system based on gas / liquid chromatogram and mass spectrum platform

    CN109061020A

  • Metabolite batch quantification software system

    CN109918056A

  • Metabolic markers and kits for detecting mood disorders and their usage

    CN112630311B