A combination of microbial markers, method, system, device and medium for diagnosing or predicting depression
Through Mendel's randomization analysis, the intestinal microbial marker combinations related to depression were screened out, and a machine learning model was constructed, which solved the problem of poor correlation of characteristic values of depression prediction models in the existing technology, and achieved accurate diagnosis and risk prediction of depression.
Patent Information
- Application Number
- CN202411305521.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-09-19
AI Technical Summary
The lack of effective combination of intestinal microbial markers in the prior art leads to poor correlation between the characteristic values of the depression prediction model and depression, resulting in a low causal relationship and inability to effectively explain the relationship between depression.
Through Mendel's randomization analysis, the combination of intestinal microbial markers that have a causal relationship with depression was screened out, including Actinomycetes, Prevotella, Bifidobacteria, etc., and a machine learning model was constructed for diagnosis and prediction.
Accurate diagnosis and risk prediction of depression are achieved, the training set AUC reaches 0.8684, the verification set AUC reaches 0.8219, and the independent verification set AUC reaches 0.8172, which has high accuracy and clinical application value.
Smart Images

Figure CN118813841B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of microbial markers, and specifically, relates to a microbial marker combination, method, system, device and medium for diagnosing or predicting depression. Background Art
[0002] Depression is a serious mental illness with unclear pathogenesis and lack of objective diagnostic biomarkers. Studies have found that intestinal microbiome disorders are a potential new cause of depression. The microbiome represents all microorganisms living in / in the human body. The intestinal microbiome is the largest and most direct external environment of the human body. It has the property of interacting with host genetics to shape the phenotype and plays an indispensable role in maintaining the health of the body. About half of the small molecules in the blood are produced by microorganisms or regulated by intestinal microorganisms. The interactions between microorganisms and the host significantly affect human health.
[0003] More and more evidence shows that intestinal microorganisms can affect the function of the central nervous system through the interaction of multiple mechanisms such as neural, endocrine or immune signaling. Therefore, the concept of the "microbe-gut-brain" axis was formally proposed in 2012. That is, the increase of pathogenic microorganisms and the decrease of protective microorganisms in the intestinal microorganisms can affect brain function, behavior and the occurrence of brain diseases through the "gut-brain" axis, making this field a frontier hotspot in global life science research today.
[0004] Numerous studies have shown that intestinal flora can synthesize neurotransmitters. For example, Lactobacillus, Bifidobacterium, and Enterobacter can all produce gamma-aminobutyric acid; Escherichia coli can produce serotonin and dopamine; and Lactobacillus can produce acetylcholine. Studies have also found that the lack of intestinal flora is associated with a significant decrease in the levels of neurotransmitters in the intestine, such as serotonin, norepinephrine, and gamma-aminobutyric acid, and that normal neurotransmitter concentrations can be reestablished through colonization of the flora.
[0005] Therefore, screening intestinal microbial markers and predicting the risk of depression through machine learning models can help humans better prevent depression. At present, there is no good prediction model for depression prediction based on intestinal microbial markers. The feature values of some research prediction models are poorly correlated with depression and have a low causal relationship, which means that this feature often cannot explain the relationship with depression. Summary of the invention
[0006] In order to solve at least one of the above technical problems, the inventors obtained intestinal microbial markers that are causally related to depression through Mendelian randomization analysis, and further explored their causal relationship with depression. The markers obtained are all intestinal microorganisms that are causally related to depression, which can better diagnose or predict.
[0007] The first aspect of the present invention provides a microbial marker combination for diagnosing or predicting depression, wherein the microbial marker combination comprises Actinomyces ( Actinomyces ), Pseudo-Prevotella ( Alloprevotella ), Bifidobacterium spp. Bifidobacterium ), Streptobacillus ( Catenibacterium ), Collinsella spp. Collinsella ), Coprococcus ( Coprobacter )、Flavone-degrading bacteria( Flavonifractor ), Haemophilus ( Haemophilus ), Lachnospiraceae ( Lachnospira ), Lactobacillus spp. Lactobacillus ), Marvinburya ( Marvinbryantia ), Methanobrevibacterium ( Methanobrevibacter ), Parabacteroides ( Parabacteroides ), Streptococcus ( Streptococcus ), Geobacillus ( Terrisporobacter )、Food Grain Pseudomonas( Victivallis ), Adler Kreutzella ( Adlercreutzia ), by Eubacterium mucosum ( Blautia ), Desulfovibrio ( Desulfovibrio ), Dialisterella spp. Dialister ), Egeria ( Eggerthella ), Faecalibacterium prausnitzii ( Faecalibacterium )、Gordonella parafaciens( Gordonibacter )、Clostridium family Hungatella Genus ( Hungatella )、Intestinal core bacteria( Lachnoclostridium ), Oscillatory Spiral ( Oscillibacter ), Peptococcus ( Peptococcus ) and Veillonella ( Veillonella ).
[0008] By detecting each microorganism in the above-mentioned microbial marker combination, its abundance data is obtained, and further by constructing a machine learning model to diagnose whether the subject suffers from depression or predict whether the subject has a risk of depression.
[0009] A second aspect of the present invention provides a method for diagnosing or predicting depression, comprising the following steps:
[0010] Constructing a prediction model in a computer using the abundance data of the combination of microbial markers described in the first aspect of the present invention in biological samples of a population, wherein the population includes patients with depression and non-depressed subjects;
[0011] The abundance data of the combination of microbial markers in the subject's biological sample is input into the prediction model to obtain a diagnosis or risk prediction result of depression.
[0012] It is worth noting that all steps of the method for diagnosing or predicting depression of the present invention are implemented by a computer or other device.
[0013] In some embodiments of the present invention, the abundance data of each microbial marker is obtained based on qPCR, 16S RNA sequencing or metagenomic sequencing.
[0014] In some specific embodiments of the present invention, it refers to the method obtained by 16S RNA sequencing, specifically including:
[0015] Extracting genomic DNA from the biological sample and performing 16S RNA sequencing to obtain sequencing data;
[0016] The original reads of the sequencing data are preprocessed, filtered to obtain high-quality reads, compared with the 16S RNA gene reference database, and chimera sequences are removed. Finally, the filtered sequences are clustered in a certain manner (including but not limited to NanoCLUST) to obtain multiple sequence clustering operational taxonomy units (OTUs), and the OTUs are annotated with taxonomy to obtain the abundance data of each microorganism.
[0017] In some embodiments of the invention, the abundance is a relative abundance.
[0018] In the present invention, the biological sample includes but is not limited to feces, intestinal lavage fluid, anal swab samples, preferably, a feces sample.
[0019] In some embodiments of the present invention, constructing a prediction model comprises the following steps:
[0020] S1, randomly dividing the abundance data of the microbial marker combination in the biological samples of the group into two groups, one group is a training set, and the other group is a test set;
[0021] S2, using the training set data, builds a prediction model based on the machine learning algorithm and performs multi-fold cross validation;
[0022] S3, verify the obtained prediction model in the test set.
[0023] In some embodiments of the present invention, stratified random sampling is used for grouping, and the grouping ratio may be 7:3, 4:1, etc.
[0024] In some embodiments of the present invention, the machine learning algorithm is selected from any one of the following algorithms: logistic regression algorithm, linear regression algorithm, random forest algorithm, neural network algorithm, support vector machine algorithm, Bayesian classification algorithm, gradient boosting algorithm, K nearest neighbor algorithm and decision tree algorithm.
[0025] In some specific embodiments of the present invention, the machine learning algorithm is a decision tree algorithm, and the maximum depth of the prediction model obtained based on the decision tree algorithm is 6, and the maximum number of leaf nodes is 10.
[0026] The third aspect of the present invention provides a computer device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any method described in the second aspect of the present invention when executing the computer program.
[0027] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any method described in the second aspect of the present invention are implemented.
[0028] A fifth aspect of the present invention provides a system for diagnosing or predicting depression, comprising the following modules:
[0029] A data input module, used to input the abundance data of the microbial marker combination described in the first aspect of the present invention in the biological sample of the subject obtained;
[0030] A database storage module, for storing the abundance data of the microbial marker combination in biological samples of a population, wherein the population includes depression patients and non-depression subjects;
[0031] A disease prediction module is connected to the data input module and the database storage module, respectively, and is used to construct a prediction model using the abundance data of the microbial marker combination in the biological samples of the population, and to diagnose whether the subject suffers from depression or predict whether the subject has a risk of depression based on the abundance data of the microbial marker combination of the subject obtained from the data input module.
[0032] The sixth aspect of the present invention provides use of the abundance detection reagent of the microbial marker combination described in the first aspect of the present invention in the preparation of a kit for diagnosing or predicting depression.
[0033] In some preferred embodiments of the present invention, the abundance detection reagent refers to a high-throughput sequencing reagent, including a nucleic acid extraction, amplification reagent and / or purification reagent.
[0034] In some embodiments of the present invention, the abundance detection reagent includes primers and / or probes. Further, the probes are prepared into a chip.
[0035] Compared with the prior art, the present invention has achieved the following beneficial effects:
[0036] By using the microbial markers of the present invention and establishing a machine learning model, the AUC reached 0.8684 in the training set, 0.8219 in the validation set, and 0.8172 in the independent validation set, indicating that the microbial markers of the present invention can be used to accurately diagnose whether a subject suffers from depression or predict whether a subject has a risk of depression, achieve accurate prediction, enable timely medical intervention, and have great clinical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram of a decision tree constructed in Example 2 of the present invention is shown.
[0038] Figure 2 The performance evaluation of the machine learning model in Example 2 of the present invention in the training set and the validation set is shown.
[0039] Figure 3 The performance evaluation of the machine learning model in Example 3 of the present invention in the external validation set is shown. DETAILED DESCRIPTION
[0040] Unless otherwise indicated, implied from the context, or customary in the prior art, all parts and percentages in this application are based on weight, and the tests and characterization methods used are all synchronized with the filing date of this application. Where applicable, the contents of any patent, patent application or disclosure involved in this application are fully incorporated herein by reference, and their equivalent patent families are also introduced as references, especially the definitions of relevant terms in the art disclosed in these documents. If the definition of a specific term disclosed in the prior art is inconsistent with any definition provided in this application, the definition of the term provided in this application shall prevail.
[0041] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments.
[0042] The following examples are used to demonstrate preferred embodiments of the present invention. It will be appreciated by those skilled in the art that the techniques disclosed in the following examples represent techniques discovered by the inventors that can be used to implement the present invention and therefore can be considered as preferred embodiments of the present invention. However, it will be appreciated by those skilled in the art based on this specification that many modifications may be made to the specific embodiments disclosed herein and still achieve the same or similar results without departing from the spirit or scope of the present invention.
[0043] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs, and the disclosure and materials cited therein are hereby incorporated by reference.
[0044] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many technical equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the claims.
[0045] The experimental methods in the following examples are conventional methods unless otherwise specified. The instruments and equipment used in the following examples are conventional laboratory instruments and equipment unless otherwise specified; the experimental materials used in the following examples are purchased from conventional biochemical reagent stores unless otherwise specified.
[0046] Example 1 Screening of intestinal microbial markers for depression
[0047] This example screens intestinal microbial markers associated with depression, and the specific screening steps are as follows:
[0048] 1. Get the dataset
[0049] Single nucleotide polymorphisms (SNPs) related to gut microbiome (GM) were selected as instrumental variables (IVs) for exposure. The genetic variation data of these SNPs came from the MiBioGen Consortium, which analyzed the 16SrRNA gene sequencing data of 18,340 participants.
[0050] The genetic variation data of depression comes from the IEU OpenGWAS project, which contains 29 data sets, and the data numbers are:
[0051] EBI databases (ebi-a-GCST90038650, ebi-a-GCST90018833, ebi-a-GCST90013878, ebi-a-GCST90013928, ebi-a-GCST90013909, ebi-a-GCST90013959, ebi-a-GCST90013910, ebi-a-GCST90013960 and ebi-a-GCST90018613), FinnGen Biobank (finn-b-F5_DEPRESSIO, finn-b-ANTIDEPRESSANTS and finn-b-O15_POSTPART_DEPR), UKB Biobank (finn-b-F5_DEPRESSIO, finn-b-ANTIDEPRESSANTS and finn-b-O15_POSTPART_DEPR), kb-e-2100_CSA, ukb-e-2090_CSA, ukb-e-2100_AFR, ukb-e-2090_AFR, ukb-e-20126_p1_CSA, ukb-e-20126_p4_CSA, ukb-e-20126_p5_CSA, ukb-e-20126_p1_AFR, u kb-e-2090_EAS, ukb-e-4609_CSA, ukb-e-2100_MID, ukb-e-2090_MID, ukb-e-460 9_AFR, ukb-e-4620_CSA, ukb-e-4620_AFR, ukb-e-20446_AFR and ukb-e-20510_AFR).
[0052] 2. Select instrumental variables
[0053] After excluding 15 unknown genera from the 211 intestinal microorganisms, a total of 196 intestinal microorganisms were included in this example, including 9 phylum-level microorganisms, 16 class-level microorganisms, 20 order-level microorganisms, 32 family-level microorganisms, and 119 genus-level microorganisms.
[0054] Instrumental variable inclusion criteria: The SNP significance threshold was set at P < 1.0 × 10 -5 , the linkage disequilibrium standard is r 2 <0.001, genetic distance 10000 kb, and highly correlated SNPs were removed to ensure that the included SNPs were independent of each other.
[0055] In order to eliminate the influence of weak instrumental variable bias on the estimation of association effect, the F value is used to test the strength of the instrumental variable. The calculation formula is as follows:
[0056] F=[R 2 / (R 2 -1)]×[(NK-1) / K]
[0057] Among them, N represents the sample size of the gut microbiome GWAS study, K represents the number of instrumental variables, and R 2 represents the extent to which the SNP explains the exposure, R 2 The calculation formula is: 2 =2×β 2 × (1-MAF) × MAF, MAF is the minor allele frequency, and β is the effect size of SNP on exposure.
[0058] F>10 is considered to be able to effectively avoid the bias caused by weak instrumental variables, so SNPs with F<10 were eliminated. A total of 1519 SNPs were obtained as instrumental variables.
[0059] 3. Statistical Analysis
[0060] Statistical analysis was performed using the TwoSampleMR package and the MR-PRESSO package in R. Inverse-variance weighted analysis (IVW) was mainly used to calculate the OR value and 95% confidence interval (95% CI) to evaluate the potential causal association between intestinal microbial abundance and the risk of depression, combined with MR-Egger regression, weighted median (WME), and weighted mode (WM) as supplementary methods of IVW.
[0061] The IVW method was used to screen the eigenvalues, and P < 0.05 was considered statistically significant.
[0062] 4. Sensitivity and pleiotropy analysis
[0063] In terms of sensitivity, this example calculates Cochran's Q statistic by IVW and MR-Egger regression. If P>0.05, it means there is no obvious heterogeneity. At the same time, the inventors also eliminated the included SNPs one by one by leaving one out method to observe whether it affects the analysis results, and draw a forest map. If P>0.05 is obtained after eliminating a certain SNP, it is considered that the SNP will not have a significant impact on the results.
[0064] In terms of pleiotropy, the inventors simultaneously used the intercept term of MR-Egger regression and the Mendelian randomization pleiotropy residual sum and outlier (MR-PRESSO) to test the horizontal pleiotropy of the included SNPs. In MR-Egger regression, if the intercept tends to 0, it can be considered that there is no horizontal pleiotropy. The horizontal pleiotropy screening retains the characteristic value of P>0.05.
[0065] 5. Summary Analysis
[0066] IVW is considered the standard method for MR summary data, which uses the Wald ratio method to calculate the causal effect estimate for each included instrumental SNP and then performs a weighted summary analysis.
[0067] The results of the inverse variance weighted analysis and weighted median analysis of the above 29 data in this embodiment are shown in Tables 1 to 29.
[0068] Table 1 IVW analysis results of ebi-a-GCST90038650 data
[0069]
[0070] As can be seen from Table 1, in ebi-a-GCST90038650, Veillonella (Veillonella spp.), Ruminiclostridium 6 (Clostridium rumenii 6), Eubacterium ventriosum group (Eubacterium venosus group), Erysipelatoclostridium (Erysipelothrix spp.), Ruminococcus gnavus group (active rumen coccus group), Ruminococcaceae UCG 013 (Ruminococcaceae UCG 013), Lachnospiraceae UCG 010 (Lachnospirillum UCG 010) and Turicibacte The causal relationship between (spore-forming intestinal bacteria) and depression reached a statistically significant level (P<0.05). Veillonella , Ruminiclostridium 6 , Eubacterium ventriosum group, Erysipelatoclostridium , Ruminococcus gnavus group and Turicibacte It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Ruminococcaceae UCG 013 and Lachnospiraceae UCG 010 was negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0071] Table 2 IVW analysis results of ebi-a-GCST90018833 data
[0072]
[0073] As can be seen from Table 2, in ebi-a-GCST90018833, Lachnoclostridium (Core bacteria in the intestine), Tyzzerella3 (Tezerella spp.), Gordonibacter (Gordonella parafaciens) and Ruminococcaceae UCG 004The causal relationship between (Ruminococcaceae UCG 004) and depression reached a statistically significant level (P<0.05). Tyzzerella 3 It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Lachnoclostridium , Gordonibacter and Ruminococcaceae UCG 004 It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0074] Table 3 IVW analysis results of ebi-a-GCST90013878 data
[0075]
[0076] From Table 3, we can see that in ebi-a-GCST90013878, Eubacterium ventriosum group, Ruminococcus gnavus group and Ruminiclostridium 6 The causal relationship with depression reached a statistically significant level (P<0.05). Eubacterium ventriosum group, Ruminococcus gnavus group and Ruminiclostridium 6 They are all positively correlated with the occurrence of depression, suggesting that they are risk factors for depression.
[0077] Table 4 IVW analysis results of ebi-a-GCST90013928 data
[0078]
[0079] It can be seen from Table 4 that in ebi-a-GCST90013928, Eubacterium ventriosum group, Ruminococcus gnavus group and Ruminiclostridium 6 The causal relationship with depression reached a statistically significant level (P<0.05). All three bacteria were positively correlated with the occurrence of depression, suggesting that they are risk factors for depression.
[0080] Table 5 IVW analysis results of ebi-a-GCST90013909 data
[0081]
[0082] From Table 5, we can see that in ebi-a-GCST90013909, Butyricicoccus (Butyric acid bacteria), Faecalibacterium (Faecalibacterium prausnitzii), Enterorhabdus (Enterobacter spp.), Oxalobacter (Oxalobacter spp.), Ruminiclostridium 6 , Ruminococcaceae UGC 013 ((Ruminococcaceae UCG 013)), Eubacterium brachy group (Brevibacterium group), Coprococcus 3 (Coprococcus spp. 3) and Blautia The causal relationship between Eubacterium mucosum and depression reached a statistically significant level (P<0.05). Faecalibacterium , Enterorhabdus , Oxalobacter , Ruminiclostridium 6 , Eubacterium brachy group and Coprococcus 3 It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Butyricicoccus , Ruminococcaceae UGC 013 and Blautia It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0083] Table 6 IVW analysis results of ebi-a-GCST90013959 data
[0084]
[0085] From Table 6, we can see that in ebi-a-GCST90013959, Eubacterium brachy group , Ruminiclostridium 6 , Enterorhabdus , Blautia , Oxalobacter , Ruminococcaceae UGC 013 (Ruminococcaceae UCG 013), Coprococcus 3 , Faecalibacterium and Butyricicoccus The causal relationship with depression reached a statistically significant level (P<0.05). Eubacterium brachy group , Ruminiclostridium 6 , Enterorhabdus , Oxalobacter , Coprococcus 3 and Butyricicoccus It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Blautia , Ruminococcaceae UGC 013 and Butyricicoccus It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0086] Table 7 IVW analysis results of ebi-a-GCST90013910 data
[0087]
[0088] From Table 7, we can see that in ebi-a-GCST90013910, Ruminococcaceae UCG 005 (Ruminococcaceae UCG 005), Tyzzerella 3 , Desulfovibrio (Desulfovibrio spp.) and Coprococcus 3 The causal relationship with depression reached a statistically significant level (P<0.05). Desulfovibrio and Coprococcus 3 It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Ruminococcaceae UCG 005 and Tyzzerella 3 It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0089] Table 8 IVW analysis results of ebi-a-GCST90013960 data
[0090]
[0091] It can be seen from Table 8 that in ebi-a-GCST90013960, Desulfovibrio , Coprococcus 3 , Ruminococcaceae UCG 005 and Tyzzerella 3 The causal relationship with depression reached a statistically significant level (P<0.05). Desulfovibrio and Coprococcus 3 It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Ruminococcaceae UCG 005 and Tyzzerella 3 It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0092] Table 9 IVW analysis results of ebi-a-GCST90018613 data
[0093]
[0094] It can be seen from Table 9 that in ebi-a-GCST90018613, Family XIII AD3011 group and Lachnospiraceae UCG 010 The causal relationship with depression reached statistical significance (P<0.05). Lachnospiraceae UCG 010 It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Family XIII AD3011 group It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0095] Table 10 IVW analysis results of finn-b-F5_DEPRESSIO data
[0096]
[0097] From Table 10, we can see that in finn-b-F5_DEPRESSIO, Dialister (Diarlistella spp.), Veillonella , Slackia (Slackia spp.), Ruminococcaceae UGC 011 (Ruminococcaceae UCG 011), Roseburia (Rosenbergia spp.), Catenibacterium (Streptobacillus spp.) and Lachnospiraceae FCS020 group The causal relationship between Lachnospiraceae FCS020 group and depression reached a statistically significant level (P<0.05). Veillonella , Slackia , Roseburia , Catenibacterium and Lachnospiraceae FCS020 group It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Roseburia and Ruminococcaceae UGC 011 It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0098] Table 11 IVW analysis results of finn-b-ANTIDEPRESSANTS data
[0099]
[0100] From Table 11, we can see that in finn-b-ANTIDEPRESSANTS, Hungatella (Clostridium family Hungatella genus), Bifidobacterium (Bifidobacterium), Dialister , Alloprevotella (PseudoPrevotella spp.) and Desulfovibrio The causal relationship with depression reached a statistically significant level (P<0.05). Hungatella It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Bifidobacterium , Dialister , Alloprevotella and Desulfovibrio It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0101] Table 12 IVW analysis results of finn-b-O15_POSTPART_DEPR data
[0102]
[0103] From Table 12, we can see that in finn-b-O15_POSTPART_DEPR, Ruminococcaceae UCG 010 (Ruminococcaceae UCG 010), Roseburia , Slackia , Family XIIIAD3011 group and Ruminococcaceae UCG 011 The causal relationship between (Ruminococcaceae UCG 011) and depression reached a statistically significant level (P<0.05). Roseburia , Slackia , Family XIIIAD3011 group and Ruminococcaceae UCG 011 It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Ruminococcaceae UCG 010 It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0104] Table 13 IVW analysis results of ukb-e-2100_CSA data
[0105]
[0106] It can be seen from Table 13 that in ukb-e-2100_CSA, Coprococcus 1 (Coprococcus 1), Tyzzerella 3 , Roseburia and Methanobrevibacter The causal relationship between Methanobacterium and depression reached a statistically significant level (P<0.05). Tyzzerella 3 and Methanobrevibacter It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Coprococcus and Roseburia It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0107] Table 14 IVW analysis results of ukb-e-2090_CSA data
[0108]
[0109] From Table 14, we can see that in ukb-e-2090_CSA, Adlercreutzia (Adler's Kreutzeria spp.), Ruminococcaceae UCG 013 , Flavonifractor (Flavone-degrading bacteria), Haemophilus (Haemophilus spp.) and Coprococcus 1 The causal relationship with depression reached a statistically significant level (P<0.05). Ruminococcaceae UCG 013 , Flavonifractor and Haemophilus It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Adlercreutzia and Coprococcus 1 It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0110] Table 15 IVW analysis results of ukb-e-2100_AFR data
[0111]
[0112] From Table 15, we can see that in ukb-e-2100_AFR, Lachnospiraceae FCS020 group , Eggerthella (Eggella spp.), Collinsella (Collinsella spp.), Rikenellaceae RC9 gutgroup (Venkenellaceae RC9 enteric group), Ruminococcaceae UCG 004 and Oxalobacter The causal relationship with depression reached a statistically significant level (P<0.05). Lachnospiraceae FCS020 group , Collinsella and Rikenellaceae RC9 gutgroup It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Eggerthella and Oxalobacter It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0113] Table 16 IVW analysis results of ukb-e-2090_AFR data
[0114]
[0115] From Table 16, we can see that in ukb-e-2090_AFR, Ruminococcus torques group (Ruminococcus torque group), Enterorhabdus , Eubacterium brachy group , Adlercreutzia and Lachnospiraceae UCG 008 The causal relationship between (Lachnospirillum UCG 008) and depression reached a statistically significant level (P<0.05). Ruminococcus torques group It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Enterorhabdus , Eubacterium brachy group , Adlercreutzia and Lachnospiraceae UCG 008 It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0116] Table 17 IVW analysis results of ukb-e-20126_p1_CSA data
[0117]
[0118] From Table 17, we can see that in ukb-e-20126_p1_CSA, Adlercreutzia , Erysipelotrichaceae UCG 003 (Erysipelothrixaceae UCG 003), Eubacterium eligensgroup (selective Eubacterium group), Eisenbergiella (Eisenbergia spp.), Fusicatenibacter (Spindle-like rod genus), Prevotella 7 (Prevotella genus 7), Terrisporobacter (Geobasidiomyces), Family XIIIAD3011 group , Ruminococcaceae UCG 004 and Oscillibacter The causal relationship between Oscillatory Spirochete and depression reached a statistically significant level (P<0.05). Adlercreutzia , Erysipelotrichaceae UCG 003 , Eisenbergiella , Family XIIIAD3011 group , Ruminococcaceae UCG 004 and Oscillibacter It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Eubacterium eligens group , Fusicatenibacter , Prevotella 7 and Terrisporobacter It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0119] Table 18 IVW analysis results of ukb-e-20126_p4_CSA data
[0120]
[0121] From Table 18, we can see that in ukb-e-20126_p4_CSA, Lachnoclostridium , Eubacterium ventriosum group and Adlercreutzia The causal relationship with depression reached a statistically significant level (P<0.05). Lachnoclostridium and Eubacterium ventriosum group It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Adlercreutzia It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0122] Table 19 IVW analysis results of ukb-e-20126_p5_CSA data
[0123]
[0124] From Table 19, we can see that in ukb-e-20126_p5_CSA, Marvinbryantia (Marvinberry), Intestinimonas (Oscillospiraceae Intestinimonas genus), Rikenellaceae RC9 gutgroup and Oscillibacter The causal relationship with depression reached a statistically significant level (P<0.05). Marvinbryantia and Rikenellaceae RC9 gutgroup It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Intestinimonas and Oscillibacter It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0125] Table 20 IVW analysis results of ukb-e-20126_p1_AFR data
[0126]
[0127] From Table 20, we can see that in ukb-e-20126_p1_AFR, Lachnoclostridium The causal relationship with depression reached a statistically significant level (P<0.05). Lachnoclostridium It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression.
[0128] Table 21 IVW analysis results of ukb-e-2090_EAS data
[0129]
[0130] From Table 21, we can see that in ukb-e-2090_EAS, Victivallis (Food Pseudomonas), Fusicatenibacter , Coprococcus 3 and Parabacteroides The causal relationship between (Parabacteroides) and depression reached a statistically significant level (P<0.05). Victivallis , Coprococcus 3 and Parabacteroides It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Fusicatenibacter It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0131] Table 22 IVW analysis results of ukb-e-4609_CSA data
[0132]
[0133] From Table 22, we can see that in ukb-e-4609_CSA, Eubacterium fissicatena grou p (Eubacterium fragmentum group), Enterorhabdus , Lachnospiraceae UCG 001 (Lachnospirillum UCG 001) and Gordonibacter The causal relationship with depression reached a statistically significant level (P<0.05). Lachnospiraceae UCG 001 It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Eubacterium fissicatena grou p. Enterorhabdus and Gordonibacter It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0134] Table 23 IVW analysis results of ukb-e-2100_MID data
[0135]
[0136] From Table 23, we can see that in ukb-e-2100_MID, Coprobacter (Coprococcus), Streptococcus (Streptococcus), Clostridium innocuum group (Clostridium innocua group) and Lachnospiraceae NK4A136 group The causal relationship between Lachnospiraceae and depression reached a statistically significant level (P<0.05). Clostridiuminnocuum group and Lachnospiraceae NK4A136 group It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Coprobacter and Streptococcus It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0137] Table 24 IVW analysis results of ukb-e-2090_MID data
[0138]
[0139] From Table 24, we can see that in ukb-e-2090_MID, Fusicatenibacter and Lachnospira The causal relationship with depression reached a statistically significant level (P<0.05). Lachnospira It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Fusicatenibacter It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0140] Table 25 IVW analysis results of ukb-e-4609_AFR data
[0141]
[0142] From Table 25, we can see that in ukb-e-4609_AFR, Ruminococcaceae UCG 005 , Eisenbergiella , Butyricicoccus , Oscillibacter , Lactobacillus (Lactobacillus spp.) and Ruminiclostridium 6 The causal relationship with depression reached a statistically significant level (P<0.05). Ruminococcaceae UCG 005 , Eisenbergiella and Ruminiclostridium 6 It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Butyricicoccus , Oscillibacter and Lactobacillus It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0143] Table 26 ukb-e-4620_CSA data IVW analysis results
[0144]
[0145] From Table 26, we can see that in ukb-e-4620_CSA, Desulfovibrio , Erysipelotrichaceae UCG 003 , Actinomyces (Actinomyces), Ruminococcus 2 (Ruminococcus spp.), Lachnospiraceae UCG 001 and Lachnoclostridium The causal relationship with depression reached a statistically significant level (P<0.05). All bacteria were positively correlated with the occurrence of depression, suggesting that they are all risk factors for depression.
[0146] Table 27 IVW analysis results of ukb-e-4620_AFR data
[0147]
[0148] From Table 27, we can see that in ukb-e-4620_AFR, Ruminococcaceae NK4A214 group (Ruminococcaceae NK4A214), Hungatella , Slackia and Eggerthella The causal relationship with depression reached a statistically significant level (P<0.05). Ruminococcaceae NK4A214 group and Hungatella It is positively correlated with the occurrence of depression, suggesting that it is a risk factor for depression; Slackia and Eggerthella It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0149] Table 28 IVW analysis results of ukb-e-20446_AFR data
[0150]
[0151] From Table 28, we can see that in ukb-e-20446_AFR, Peptococcus The causal relationship between (Peptococcus) and depression reached a statistically significant level (P<0.05). Peptococcus It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0152] Table 29 IVW analysis results of ukb-e-20510_AFR data
[0153]
[0154] From Table 29, we can see that in ukb-e-20510_AFR, Peptococcus The causal relationship with depression reached a statistically significant level (P<0.05). Peptococcus It is negatively correlated with the occurrence of depression, suggesting that it is a protective factor for depression.
[0155] Some of the above exposure factors (microorganisms) only appear in one data set, totaling 29, as shown in Table 30:
[0156] Table 30 29 intestinal microorganisms that only appear in one data set
[0157]
[0158] For those intestinal microorganisms that only appear in one data set, the intestinal microorganisms with OR≈1.0 are filtered out, and the remaining intestinal microorganisms are retained, thereby filtering out Erysipelatoclostridium and Turicibacter , retaining the remaining 27 intestinal microorganisms, which are: Actinomyces , Alloprevotella , Bifidobacterium , Catenibacterium , Clostridium innocuum group , Collinsella , Coprobacter , Flavonifractor , Haemophilus , Intestinimonas , Lachnospira , Lachnospiraceae NK4A136 group , Lachnospiraceae UCG 008 , Lactobacillus , Marvinbryantia , Methanobrevibacter , Parabacteroides , Prevotella7 , Ruminococcaceae NK4A214 group , Ruminococcaceae UCG 010 , Ruminococcus 2 , Streptococcus , Terrisporobacter , Victivallis , Eubacterium eligens group , Eubacterium fissicatena group and Ruminococcus torques group .
[0159] There were also some intestinal microorganisms that appeared in multiple data sets, totaling 36. For these intestinal microorganisms that appeared in multiple data sets, Meta-analysis was used to merge the data and calculate the OR value. 2 When >50% and P<0.05, the random effect model was used, otherwise the fixed effect model was used. The results are shown in Table 31:
[0160] Table 31 36 intestinal microorganisms appearing in multiple data sets
[0161]
[0162] In this way, the data of 36 exposure factors were merged and the exposure factors with OR=1.0 were filtered out, including: Eubacterium brachy group , Oxalobacter , Roseburia , Ruminococcaceae UCG 004 , Ruminococcaceae UCG 005 , Ruminococcaceae UCG 013 , Tyzzerella 3 and Slackia , retaining the remaining 28 intestinal microorganisms, namely Adlercreutzia , Blautia , Butyricicoccus , Coprococcus 1 , Coprococcus 3 , Desulfovibrio , Dialister , Eggerthella , Eisenbergiella , Enterorhabdus , Erysipelotrichaceae UCG 003 , Eubacterium ventriosum group , Faecalibacterium , Family XIIIAD3011 group , Fusicatenibacter , Gordonibacter , Hungatella , Lachnoclostridium , Lachnospiraceae FCS020 group , Lachnospiraceae UCG 001 , Lachnospiraceae UCG 010 , Oscillibacter , Peptococcus , Rikenellaceae RC9 gutgroup , Ruminiclostridium 6 , Ruminococcaceae UCG 011 , Ruminococcus gnavus group and Veillonella .
[0163] Example 2 Depression Prediction Machine Learning Model
[0164] This embodiment uses the DecisionTreeClassifier function in the sklearn module in Python to perform decision tree modeling on the data.
[0165] This example constructs a prediction model for calculating the risk of depression in a host based on the abundance values of intestinal microbial markers screened in Example 1. The specific steps are as follows:
[0166] S1, Dataset acquisition: 555 sample data were downloaded from the MicroPhenoDB database, including 300 cases in the control group and 255 cases in the case group.
[0167] S2, the StandardScaler function of the klearn.preprocessing module standardizes the data, and then uses the recursive feature elimination method for feature selection. The recursive feature elimination method uses a base model for multiple rounds of training. After each round of training, several unimportant features are eliminated, and then the next round of training is based on the new feature set. It uses model accuracy to identify which attribute combinations contribute the most to the prediction of the target attribute, and then eliminates useless features. Sklearn provides two recursive feature elimination methods, namely, recursive feature elimination (RFE) and cross-recursive feature elimination (RFECV). The present invention uses a random forest classifier as the base model and uses 5-fold cross validation to screen the RFECV recursive feature elimination method. 55 intestinal microbial features were finally screened and 28 were retained, which are: Actinomyces , Alloprevotella , Bifidobacterium , Catenibacterium , Collinsella , Coprobacter , Flavonifractor , Haemophilus , Lachnospira , Lactobacillus , Marvinbryantia , Methanobrevibacter , Parabacteroides , Streptococcus , Terrisporobacter , Victivallis , Adlercreutzia , Blautia , Desulfovibrio , Dialister , Eggerthella , Faecalibacterium , Gordonibacter , Hungatella , Lachnoclostridium , Oscillibacter , Peptococcus and Veillonella The 28 intestinal microbial characteristics will be used for subsequent model construction and prediction.
[0168] S3, the train_test_split function divides the data set into a training set and an internal validation set. In this embodiment, 25% of the samples are used as the internal validation set.
[0169] S4, a decision tree model is established by cross-validation, and the model is continuously optimized. max_depth is the maximum depth of the decision tree, and max_leaf_nodes is the maximum number of leaf nodes of the model. This embodiment uses parameter grid search to select the most suitable parameter combination. When max_depth=6 and max_leaf_nodes=10, the model prediction effect is the best.
[0170] The root node has 416 samples, gini=0.498, which represents impurity. The lower the score, the better. The classification value indicates the number of categories. The classification value = [222,194] means that category 1 has 222 samples and category 2 has 194 samples. Lachnoclostridium Is it <=0.047 to divide the left branch and the right branch, the left branch is Bifidobacterium <=0.357, the right branch is Collinsella <=0.023 is divided in sequence. Figure 1 shown.
[0171] The ROC curves of the training set and the validation set are as follows: Figure 2 As shown, the accuracy of the training set is 82.69%, and the area under the curve (AUC) is 0.8684; the accuracy of the validation set is 79.14%, AUC=0.8219.
[0172] Example 3: Evaluation of model performance in independent validation set
[0173] The inventors applied the depression prediction model constructed using Example 2 to an independent validation set.
[0174] The independent validation set mainly comes from real physical examination sample data. The sample type is feces, with a total of 108 samples, including 58 cases with depression and 50 healthy controls. Nanopore 16S RNA sequencing analysis was used. The specific analysis steps are as follows:
[0175] (1) Sequencing
[0176] DNA extraction: DNA was extracted using a fecal sample bacterial genome extraction kit, and the extracted DNA concentration and purity were tested using Nanodrop;
[0177] 16S rRNA amplification: Targeted amplification of the V4 region of the bacterial 16S rRNA gene region was performed using a 16S sequencing library construction kit for 30 cycles. The reaction conditions were: 95°C, 1 min, 94°C, 30 s, 62°C, 30 s, 62°C, 2 min, 62°C, 5 min; then maintained at 4°C.
[0178] Recovery of PCR products: Purification of target fragments was performed using the AMPure XP Nucleic Acid Purification Kit (Beckman) and quantification by Qubit;
[0179] Adapter connection: Mix the above barcoded libraries according to the concentration ratio into a 10μL system, so that the total amount of DNA after mixing is 200ng. Add 1μL Rapid Sequencing Adapter (RAP) to the mixed sample, blow gently with the tip of the pipette to mix, and react at room temperature for 5min.
[0180] Finally, the constructed library was sequenced using the GridION nanopore sequencer.
[0181] (2) Bioinformatics analysis
[0182] First, NanoStat software was used to quality control, filter, and remove sequencing adapters of the raw data to obtain high-quality reads. Minimap2 software was used for global alignment to find chimeras, and then yacrd (Yet Another Chimeric Read Detector) software was used to remove chimera sequences. Filtlong software was then used to filter data of specific length and Q value. Finally, the filtered sequences were clustered in a certain way (NanoCLUST in this embodiment) to obtain multiple sequence clustering operational taxonomy units (OTUs), and OTUs were annotated with taxonomy to obtain species abundance information.
[0183] (3) Model prediction
[0184] The abundance information of intestinal microbial markers in the case group and the control group was collated, and then model prediction and risk level classification were performed.
[0185] The model constructed in Example 2 was used to perform the prediction analysis of the above independent validation set, and the results were as follows: Figure 3 As shown, AUC=0.8172, accuracy rate 75%. The above results show that the model constructed by using Example 2 can distinguish depression patients from healthy people, and can be used clinically to assist in diagnosing whether a subject suffers from depression.
[0186] All documents mentioned in the present invention are cited as references in this application, just as each document is cited as reference individually. In addition, it should be understood that after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the claims attached to this application.
Claims
1. A combination of microbial markers for diagnosing or predicting depression, characterized in that: The microbial marker combination includes Actinomyces ( Actinomyces ), Pseudo-Prevotella ( Alloprevotella ), Bifidobacterium spp. Bifidobacterium ), Streptobacillus ( Catenibacterium ), Collinsella spp. Collinsella ), Coprococcus ( Coprobacter )、Flavone-degrading bacteria( Flavonifractor ), Haemophilus ( Haemophilus ), Lachnospiraceae ( Lachnospira ), Lactobacillus spp. Lactobacillus ), Marvinburya ( Marvinbryantia ), Methanobrevibacterium ( Methanobrevibacter ), Parabacteroides ( Parabacteroides ), Streptococcus ( Streptococcus ), Geobacillus ( Terrisporobacter )、Food Grain Pseudomonas( Victivallis ), Adler Kreutzeria ( Adlercreutzia ), Eubacterium mucosum ( Blautia ), Desulfovibrio ( Desulfovibrio ), Dialisterella spp. Dialister ), Egeria ( Eggerthella ), Faecalibacterium prausnitzii ( Faecalibacterium )、Gordonella parvum( Gordonibacter )、Clostridium family Hungatella Genus ( Hungatella ), intestinal core bacteria Lachnoclostridium , Oscillatory Spirillum Oscillibacter ), Peptococcus ( Peptococcus ) and Veillonella ( Veillonella ).
2. A method for constructing a prediction model for diagnosing or predicting depression, characterized in that: The following steps are involved: S1, randomly dividing the abundance data of the combination of microbial markers described in claim 1 in biological samples of a population into two groups, one group being a training set and the other group being a test set, wherein the population includes patients with depression and non-depressed subjects; S2, using the training set data, builds a prediction model based on the machine learning algorithm and performs multi-fold cross validation; S3, verify the obtained prediction model in the test set.
3. The method according to claim 2, characterized in that The abundance data of the microbial marker combination is obtained based on qPCR, 16S RNA sequencing or metagenomic sequencing.
4. The method according to claim 2, characterized in that: The machine learning algorithm is selected from any one of the following algorithms: Logistic regression algorithm, linear regression algorithm, random forest algorithm, neural network algorithm, support vector machine algorithm, Bayesian classification algorithm, gradient boosting algorithm, K nearest neighbor algorithm and decision tree algorithm.
5. The method according to claim 4, characterized in that The machine learning algorithm is a decision tree algorithm. The maximum depth of the prediction model obtained based on the decision tree algorithm is 6, and the maximum number of leaf nodes is 10.
6. A computer device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method according to any one of claims 2 to 5 when executing the computer program.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 2 to 5 are implemented.
8. A system for diagnosing or predicting depression, characterized in that: Includes the following modules: A data input module, used to input the abundance data of the microbial marker combination according to claim 1 obtained in the subject's biological sample; A database storage module, for storing the abundance data of the microbial marker combination in biological samples of a population, wherein the population includes depression patients and non-depression subjects; A disease prediction module is connected to the data input module and the database storage module, respectively, and is used to construct a prediction model using the abundance data of the microbial marker combination in the biological samples of the population, and to diagnose whether the subject suffers from depression or predict whether the subject has a risk of depression based on the abundance data of the microbial marker combination of the subject obtained from the data input module.
9. Use of the abundance detection reagent of the microbial marker combination according to claim 1 in the preparation of a kit for diagnosing or predicting depression.