Depression risk assessment method and system based on intestinal flora characteristics

By using gut microbiota characteristic assessment methods, combined with functional gene abundance and metabolite concentration data, a neuroactive metabolic function profile is constructed, which solves the problem of insufficient objectivity in the diagnosis of depression in existing technologies and achieves high-precision depression risk assessment and personalized management.

CN121460180APending Publication Date: 2026-02-03ANSHAN (TIANJIN) BIOTECHNOLOGY CO LTD

Patent Information

Application Number
CN202511648676.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Current technologies lack objective and quantitative biological biomarkers in the diagnosis of depression, resulting in poor diagnostic consistency, inability to achieve early warning, difficulty in predicting efficacy, and blind selection of drugs. Furthermore, single 16S rRNA gene sequencing analysis cannot reveal the specific molecular mechanisms by which the microbiome affects the host brain, and its predictive accuracy and reliability are limited.

Method used

By obtaining individual gut microbiota samples, performing nucleic acid sequencing and metabolite detection, a neuroactive metabolic functional profile is constructed. Combined with functional gene abundance and metabolite concentration data, the data are input into a pre-trained risk assessment model to output a depression risk score.

Benefits of technology

It achieves high-precision mapping from complex biomarkers to clinical phenotypes, significantly improving the accuracy and reproducibility of depression risk assessment, and has the potential for early screening, dynamic monitoring, and personalized health management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121460180A_ABST
    Figure CN121460180A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical assistance, in particular to a depression risk assessment method and system based on intestinal flora characteristics. The method comprises the following steps: acquiring an individual enteric microorganism sample, performing nucleic acid sequencing on the enteric microorganism sample, and determining abundance data of functional genes related to a neural activity metabolic pathway in enteric microorganisms; carrying out metabolite detection on the intestinal microorganism sample, and determining concentration data of metabolites related to the neural activity metabolic pathway in the intestinal microorganisms; based on the abundance data of the functional genes and the concentration data of the metabolites, data fusion processing is carried out, a neural activity metabolism function spectrum is constructed, and the neural activity metabolism function spectrum is a multi-dimensional feature vector; and inputting the neural activity metabolic function spectrum into a pre-trained risk assessment model, and outputting a depression risk score of the individual. The accuracy and repeatability of evaluation are remarkably improved, and the method has great potential for early screening, dynamic monitoring and personalized health management.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical assistance, in particular to a depression risk assessment method and system based on intestinal flora characteristics. BACKGROUND

[0002] Depression is a high-incidence and high-disability mental illness with a large number of patients worldwide. At present, the clinical diagnosis highly depends on the subjective symptom description of patients and the scale (such as PHQ-9, HAMD) used by doctors, lacking objective and quantitative biological markers. This leads to poor diagnostic consistency, inability to achieve early warning, difficulty in predicting efficacy, and blindness in drug selection.

[0003] Intestinal flora plays a key role in the pathogenesis of depression through the "gut-brain axis". Existing technologies are mostly focused on analyzing the correlation between the species composition of flora and depression through 16S rRNA gene sequencing (for example, finding changes in the abundance of Prevotella or Faecalibacterium in patients with depression). However, this method has fundamental limitations: function and species are disconnected: changes in species abundance are not entirely equivalent to changes in their functional activity. The same bacterial species may have significant differences in functional genes and metabolic capacity in different individuals or environments. Mechanism explanation is ambiguous: only staying at the correlation level of "who is more and who is less", it cannot reveal the specific molecular mechanisms of how the microbiome affects the host brain (i.e. "how to act"), which belongs to "black box" correlation, with limited prediction accuracy and reliability. Inconsistent results: different studies find different species with poor reproducibility, making it difficult to translate into a universal diagnostic tool. SUMMARY

[0004] The present application provides a depression risk assessment method based on intestinal flora characteristics to solve the above technical problems.

[0005] In a first aspect, the present application provides a depression risk assessment method based on intestinal flora characteristics, the method comprising: S1, obtaining an intestinal microbial sample of an individual, performing nucleic acid sequencing on the intestinal microbial sample, and determining the abundance data of functional genes related to a neural activity metabolic pathway in intestinal microorganisms; S2, performing metabolite detection on the intestinal microbial sample, and determining the concentration data of metabolites related to the neural activity metabolic pathway in intestinal microorganisms; S3, based on the abundance data of the functional genes and the concentration data of the metabolites, performing data fusion processing, and constructing a neural activity metabolic function spectrum, wherein the neural activity metabolic function spectrum is a multi-dimensional feature vector; S4, inputting the neural activity metabolic function spectrum into a pre-trained risk assessment model, and outputting a depression risk score of the individual.

[0006] By the technical solution, the "neuroactive metabolic function spectrum" is constructed through multi-omics data fusion (functional gene abundance and metabolite concentration), which overcomes the limitation of single data source. The functional gene data reveals the potential metabolic capacity of intestinal flora in the key neural pathway (such as tryptophan-serotonin pathway), and the metabolite data directly reflects the level of actual metabolites. This "capacity" and "output" double verification provides a more comprehensive and true portrait of the functional state of intestinal flora, laying a reliable biological foundation for risk assessment. In the data preprocessing, standardization and feature selection are strictly implemented, effectively eliminating technical biases such as sequencing depth and detection batch, ensuring the comparability of data from different sources. Subsequently, through feature weighting fusion, key signals are amplified according to biological importance and statistical significance, thereby refining the core feature set most relevant to the pathogenesis of depression. The high-precision mapping from complex biomarkers to clinical phenotypes is achieved, successfully transforming the traditional assessment method relying on subjective questionnaires into quantitative analysis based on objective biomarkers, which not only significantly improves the accuracy and repeatability of the assessment, but also has great potential for early screening, dynamic monitoring and personalized health management.

[0007] Optionally, the neuroactive metabolic pathway includes at least one of a tryptophan-serotonin / kynurenine pathway, a short-chain fatty acid synthesis pathway, a GABAergic system, and a choline and acetylcholine metabolic pathway; the abundance data of the functional gene corresponds to the relative content of the functional gene involved in the neuroactive metabolic pathway; the bioinformatics analysis includes sequence quality control, sequence assembly, open reading frame prediction, and functional gene alignment.

[0008] Optionally, for the tryptophan-serotonin / kynurenine pathway, the sequences of tryptophan hydroxylase gene, tryptophan pyrrolase gene, kynurenine enzyme gene, and indoleamine-2,3-dioxygenase homologous gene are extracted; for the short-chain fatty acid synthesis pathway, the sequences of butyryl-CoA:acetyl-CoA transferase gene and butyrate kinase gene are extracted; for the GABAergic system pathway, the sequence of glutamate decarboxylase gene is extracted; for the choline and acetylcholine metabolic pathway, the sequence of choline trimethylamine lyase gene is extracted; in the extraction process of the functional gene, sequence alignment and annotation are performed using bioinformatics tools, and the functional gene abundance is controlled, including removing low-quality sequences and correcting sequencing bias.

[0009] Optionally, a list of targeted metabolites corresponding to the neural activity metabolic pathways is selected, including at least tryptophan, 5-hydroxytryptophan, serotonin, kynurenine, quinolinic acid and kynurenic acid in the tryptophan-serotonin / kynurenine pathway, butyric acid, propionic acid and acetic acid in the short-chain fatty acid synthesis pathway, gamma-aminobutyric acid in the GABAergic system pathway, and trimethylamine and trimethylamine oxide in the choline and acetylcholine metabolism pathway; a targeted metabolomics strategy is used for metabolite extraction and detection of the intestinal microbial sample, wherein metabolite extraction uses organic solvents and adds internal standards to correct extraction efficiency; the concentration of the metabolites is quantified by chromatography-mass spectrometry technology, and the raw concentration data is standardized, including internal standard correction, batch effect correction and concentration normalization, to eliminate experimental variations and obtain comparable metabolite concentration data.

[0010] Optionally, in the sample pretreatment stage, an appropriate amount of fecal sample is weighed, methanol-water extraction solution containing internal standards is added, metabolite release is promoted by vortex mixing and ultrasonic treatment, and then the supernatant is taken by centrifugation as the detection sample; in the chromatographic separation stage, a hydrophilic interaction chromatographic column is used for liquid chromatographic separation, and the mobile phase gradient is optimized to achieve efficient separation of target metabolites; in the mass spectrometry detection stage, the metabolites are quantified in the mass spectrometer by multiple reaction monitoring mode, and the absolute concentration or relative concentration is calculated by comparing the standard curve and the internal standard response value; in the data post-processing stage, the raw concentration is corrected by internal standard, and the batch-to-batch variation is corrected by statistical method, and finally the concentration data is converted to logarithmic scale or Z-score to improve the distribution characteristics.

[0011] Optionally, the abundance data of the functional genes and the concentration data of the metabolites are standardized to eliminate technical bias caused by sequencing depth, sample size or detection method, wherein the standardization includes converting the data to a distribution with a mean of 0 and a variance of 1 using the Z-score standardization method; feature selection is performed to select features significantly related to depression risk from the standardized data, wherein the feature selection is based on the importance score of random forest to select the top 15%-20% features; feature weighting fusion is performed according to the biological importance of each feature in the depression pathogenesis mechanism, wherein the weight is automatically determined based on prior knowledge or feature importance in the model training process; the weighted features are combined into a multi-dimensional feature vector, i.e. the neural activity metabolic function spectrum, wherein the dimension of the multi-dimensional feature vector is consistent with the number of selected features, and each dimension represents a standardized and weighted feature value; the neural activity metabolic function spectrum is used to represent the functional state of the intestinal microorganisms of the individual in the neural activity metabolic pathway.

[0012] Optionally, the pre-trained risk assessment model is a classification or regression model trained based on a training sample set with known depression status, using a machine learning algorithm selected from random forest, XGBoost, support vector machine, or logistic regression. Before model application, the input neuroactive metabolic functional profile is pre-processed and feature-scaled consistently with the training phase to ensure data compatibility. The processed functional profile is input into the model, which calculates a depression risk probability value or a continuous score based on the learned feature weights. The model output is post-processed, including mapping the probability value to a risk level or generating a visualized report, where the risk score is positively correlated with the severity of depression.

[0013] Optionally, the training sample set is constructed by collecting gut microbiota samples with known depression status, obtaining their functional gene abundance data and metabolite concentration data, and constructing the corresponding neuroactive metabolic functional profile. In the feature engineering phase, feature selection is performed on the training set to optimize the feature subset, and feature scaling is applied to improve model convergence. In the model training phase, a machine learning algorithm is used to optimize hyperparameters through grid search or cross-validation, and a classifier is trained with depression status as the label. In the model validation phase, the data is divided into training and test sets using stratified sampling, and the model performance is evaluated using five-fold or ten-fold cross-validation, with performance metrics including AUC, accuracy, and F1 score. The model generalization ability is tested on an independent validation set to ensure stability and accuracy on unseen data.

[0014] Optionally, the individual's gut microbiota sample is collected, and the sample type is fecal sample or intestinal content sample. The sample is aliquoted and stored, immediately frozen at -80°C or placed in nucleic acid / metabolite stabilizing solution to prevent degradation. In the nucleic acid extraction phase, commercial DNA extraction kits are used for genomic DNA extraction, and DNA concentration and purity are detected. In the metabolite extraction phase, organic solvents are used for metabolite extraction, and internal standards are added for subsequent quantitative correction. After sample pre-treatment, metagenomic sequencing and targeted metabolomics detection are performed respectively to ensure consistency and comparability of data sources.

[0015] Through the above technical solutions, through standardized sample collection, storage and pre-treatment processes, the coefficient of variation introduced by technical operation between samples is significantly reduced. This not only significantly improves the accuracy and comparability of metagenomic functional gene abundance data and targeted metabolite concentration data, but also lays a solid foundation for subsequent construction of high-quality training sample sets and obtaining reliable depression risk scores. The unified process enables seamless integration of sample data collected at different times and different locations, greatly enhancing the transferability and clinical application value of the risk assessment model.

[0016] In a second aspect, the present application provides a depression risk assessment system based on intestinal flora characteristics, comprising: a sample analysis module configured to obtain an intestinal microorganism sample of an individual, perform nucleic acid sequencing on the intestinal microorganism sample, and determine abundance data of functional genes related to a neuroactive metabolic pathway in the intestinal microorganism; a metabolic analysis module configured to perform metabolite detection on the intestinal microorganism sample, and determine concentration data of metabolites related to the neuroactive metabolic pathway in the intestinal microorganism; a multi-dimensional feature analysis module configured to perform data fusion processing based on the abundance data of the functional genes and the concentration data of the metabolites, and construct a neuroactive metabolic function spectrum, wherein the neuroactive metabolic function spectrum is a multi-dimensional feature vector; a depression risk analysis module configured to input the neuroactive metabolic function spectrum into a pre-trained risk assessment model, and output a depression risk score of the individual. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0018] Figure 1 An application scenario schematic diagram provided by an embodiment of the present application; Figure 2 A flowchart of a depression risk assessment method based on intestinal flora characteristics provided by an embodiment of the present application; Figure 3 A structural schematic diagram of a depression risk assessment system based on intestinal flora characteristics provided by an embodiment of the present application. DETAILED DESCRIPTION

[0019] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0020] In addition, the term "and / or" used in this document merely describes an associated relationship between associated objects, which means that there can be three relationships, for example, A and / or B can represent three cases of A alone, A and B together, and B alone. In addition, the character " / " in this document generally represents an "or" relationship between the front and rear associated objects unless otherwise specified.

[0021] The embodiments of the present application will be further described in detail below with reference to the accompanying drawings of the specification.

[0022] The evaluation results of the prior art have subjectivity and hysteresis, and the core problem lies in the inability to early and accurately predict the depression risk of individuals through deep fusion and intelligent analysis of multi-source data. The lack of such evaluation capability makes prevention and intervention measures often based on late clinical symptoms, easily missing the best intervention window, thereby restricting the accuracy and efficiency of mental health management and increasing the risk of disease progression and medical burden.

[0023] Based on this, the present application provides a depression risk assessment method and system based on intestinal flora characteristics, which integrates functional gene abundance and metabolite concentration two types of multi-omics data to break through the analysis bottleneck of a single data type. Functional gene information reflects the metabolic potential of intestinal flora in key neural pathways such as tryptophan-serotonin, while metabolite concentration directly presents the abundance of actual metabolites. The mutual confirmation of this "metabolic potential" and "metabolic output" can more completely and truly depict the neural activity function state of intestinal flora, providing a solid biological basis for subsequent risk assessment. In the data preprocessing stage, through standardization process and feature screening, technical interference such as sequencing depth and batch effect is effectively controlled, ensuring the comparability and consistency of cross-sample data. Further, based on biological significance and statistical significance, the features are weighted and fused to strengthen key signals and select a core feature set closely related to depression mechanism. This method realizes accurate mapping from complex biomarkers to clinical symptoms, changes the previous evaluation mode mainly relying on subjective questionnaire to quantitative analysis based on objective biomarkers, greatly improves the accuracy and repeatability of evaluation, and also shows broad application prospects in early warning, dynamic tracking and personalized health intervention.

[0024] Figure 1 An application scenario provided by the present application is shown in the figure. When evaluating the depression risk of intestinal flora characteristics, the method provided by the present application realizes high-precision mapping from complex biomarkers to clinical phenotypes, successfully changing the existing evaluation method relying on subjective questionnaire to quantitative analysis based on objective biomarkers.

[0025] Specifically, the method provided by the application is applied to any server, the server interacts with an individual to be evaluated, and an intestinal microorganism sample provided by the individual to be evaluated is obtained, a more comprehensive and true intestinal flora functional state portrait is provided, technical biases such as sequencing depth and detection batch are effectively eliminated, the comparability of data from different sources is ensured, a depression risk score is generated and output to an evaluator, the accuracy and repeatability of evaluation are significantly improved, and the method has great potential for early screening, dynamic monitoring and personalized health management.

[0026] The specific implementation can refer to the following embodiments.

[0027] Figure 2 A flowchart of a depression risk assessment method based on intestinal flora characteristics provided by an embodiment of the application, the method of the embodiment can be applied to the server in the above scenarios. As shown in the method includes: Figure 2 S1, obtaining an intestinal microorganism sample of an individual, performing nucleic acid sequencing on the intestinal microorganism sample, and determining the abundance data of functional genes related to a neural active metabolic pathway in the intestinal microorganism. S1, obtaining an intestinal microorganism sample of an individual, performing nucleic acid sequencing on the intestinal microorganism sample, and determining the abundance data of functional genes related to a neural active metabolic pathway in the intestinal microorganism.

[0028] First, an intestinal microorganism sample of an individual is collected, and the sample type is a fecal sample or an intestinal content sample. The sample is immediately divided after collection and stored in a -80°C ultra-low temperature refrigerator or placed in a nucleic acid stabilizing liquid (such as RNAlater) to prevent DNA degradation. Before nucleic acid extraction, a commercial DNA extraction kit (such as QIAamp DNA Stool Mini Kit) is used for genomic DNA extraction, and the DNA concentration and purity (such as A260 / A280 ratio should be between 1.8-2.0) are detected by ultraviolet spectrophotometry or fluorometry to ensure that the DNA quality meets the sequencing requirements.

[0029] The gut microbiome samples are sequenced using metagenomic sequencing technology (e.g. Illumina NovaSeq platform) with a recommended sequencing depth of no less than 10G raw data per sample. The raw sequencing data are quality controlled (using FastQC tool to remove low quality sequences and adapter contamination), then subjected to sequence assembly (using MEGAHIT or SPAdes software) and open reading frame (ORF) prediction (using Prodigal tool). The abundance data of functional genes are quantified by aligning to a custom-built database of genes involved in neuroactive metabolic pathways (including tryptophan-serotonin / kynurenine pathway, short-chain fatty acid synthesis pathway, GABAergic system, and choline and acetylcholine metabolism pathway), using BLAST or DIAMOND tool for sequence alignment, and calculating the relative abundance of functional genes (e.g. by RPKM or TPM normalization method). For example, for the tryptophan-serotonin pathway, the sequences of tryptophan hydroxylase gene (TPH) and indoleamine-2,3-dioxygenase homologous gene (IDO) are extracted and their abundance in the sample is calculated.

[0030] S2, performing metabolite detection on the gut microbiome sample to determine the concentration data of metabolites related to the neuroactive metabolic pathway in the gut microbiome.

[0031] The same gut microbiome sample is subjected to metabolite analysis using targeted metabolomics strategy. First, an appropriate amount of fecal sample (e.g. 100 mg) is weighed and added into methanol-water extraction solution (ratio usually 80:20, v / v) containing internal standards (e.g. stable isotope-labeled metabolite standards), and metabolite release is promoted by vortex mixing (5 minutes) and ultrasonic treatment (10 minutes, 4°C), followed by centrifugation (12000 rpm, 10 minutes, 4°C) to obtain the supernatant as the detection sample. Metabolite detection uses liquid chromatography-mass spectrometry (LC-MS) technology, and the chromatographic separation uses a hydrophilic interaction chromatography column (HILIC column, such as Waters ACQUITY UPLC BEH Amide column), and the mobile phase gradient is optimized as follows: 0-5 minutes, 95% acetonitrile; 5-15 minutes, acetonitrile ratio from 95% to 50% to achieve efficient separation of target metabolites. Mass spectrometry detection is performed in a mass spectrometer (e.g. Sciex QTRAP 6500+) using multiple reaction monitoring (MRM) mode for quantification, and the absolute or relative concentration of metabolites is calculated by comparing the pre-established standard curve and internal standard response value.

[0032] The raw concentration data were corrected by internal standard (to correct extraction efficiency variation) and batch effect (using statistical methods such as ComBat algorithm). Finally, the concentration data were transformed to log-scale or Z-score (mean = 0, variance = 1) to improve distribution properties and remove experimental variation. The targeted metabolite list includes tryptophan, 5-hydroxytryptophan, serotonin, kynurenine, quinolinic acid and kynurenic acid in tryptophan-serotonin / kynurenine pathway; butyric acid, propionic acid and acetic acid in short-chain fatty acid synthesis pathway; gamma-aminobutyric acid in GABAergic system pathway; and trimethylamine and trimethylamine-N-oxide in choline and acetylcholine metabolism pathway.

[0033] S3, based on the abundance data of the functional genes and the concentration data of the metabolites, performing data fusion processing to construct a neural activity metabolic function spectrum, wherein the neural activity metabolic function spectrum is a multi-dimensional feature vector.

[0034] First, the abundance data of the functional genes and the concentration data of the metabolites are standardized to eliminate technical bias caused by sequencing depth, sample size or detection method. Standardization uses Z-score method to convert each feature to a distribution with mean 0 and variance 1. For example, the functional gene abundance data is normalized by dividing the total sequencing depth, and then the Z-score is calculated; the metabolite concentration data is directly converted by Z-score.

[0035] Features significantly associated with depression risk are selected from the standardized data. Feature selection is based on the importance score of the random forest algorithm, and the top 15-20% features are selected (for example, 15-20 key features are selected from the initial 100 features). The importance score is calculated by training a random forest model (using a training set with known depression status), and the feature importance is indicated by the Gini index or the average accuracy drop.

[0036] Each feature is assigned a weight according to its biological importance in the pathogenesis of depression. The weight can be determined automatically based on prior knowledge (such as the importance of metabolic pathways reported in the literature) or the importance of features in the model training process. For example, features related to the tryptophan-serotonin pathway (such as serotonin concentration) may be assigned a higher weight (such as 1.2), while other features have a weight of 1.0. The weighted feature values are combined into a multi-dimensional feature vector through linear combination.

[0037] The weighted features are combined into a multi-dimensional feature vector with the same number of dimensions as the number of selected features (for example, a 15-dimensional vector), and each dimension represents a standardized and weighted feature value. This functional spectrum is used to represent the functional state of the individual's gut microbiota in the neural activity metabolic pathway and as input for machine learning models.

[0038] S4, inputting the neural metabolic function profile into a pre-trained risk assessment model, outputting a depression risk score of the individual.

[0039] The pre-trained risk assessment model is a classification or regression model trained by machine learning algorithm based on a training sample set with known depression status (e.g. containing 500 depression patients and 500 healthy controls). The machine learning algorithm can be selected from random forest, XGBoost, support vector machine or logistic regression. Before applying the model, the input neural metabolic function profile is pre-processed and feature scaled consistently with the training phase (e.g. Z-score standardization using the mean and variance of the training set) to ensure data compatibility.

[0040] The processed function profile is input into the model, which calculates the depression risk probability value (between 0 and 1) or continuous score (e.g. 0-100 points) based on the learned feature weights. For example, when using a random forest model, the risk score is output by majority voting or average probability; when using a logistic regression model, the probability is output by a sigmoid function.

[0041] The model output is post-processed, including mapping the probability value to a risk level (e.g. low risk: <0.3, medium risk: 0.3-0.7, high risk: >0.7) or generating a visual report (e.g. risk curve graph or text description). The risk score is positively correlated with the severity of depression, with a high score indicating a high risk of depression. The performance of the model is evaluated on an independent validation set to ensure that the AUC (area under the curve) is greater than 0.85 and the accuracy is greater than 80%.

[0042] Beneficial effects: By integrating multi-omics data (functional gene abundance and metabolite concentration) to construct a "neural metabolic function profile", the limitations of a single data source are overcome. Functional gene data reveals the potential metabolic capacity of gut microbiota in key neural pathways (such as the tryptophan-serotonin pathway), while metabolite data directly reflects the levels of actual metabolites. This dual verification of "capacity" and "output" provides a more comprehensive and realistic portrait of the functional state of gut microbiota, laying a solid biological foundation for risk assessment. Strict standardization and feature selection are implemented in data preprocessing, effectively eliminating technical biases such as sequencing depth and detection batch, ensuring the comparability of data from different sources. Subsequently, through feature weighting fusion, key signals are amplified based on biological importance and statistical significance, thus refining a core feature set most relevant to the pathogenesis of depression. This achieves high-precision mapping from complex biomarkers to clinical phenotypes, successfully transforming traditional assessment methods relying on subjective questionnaires into quantitative analysis based on objective biomarkers, significantly improving the accuracy and reproducibility of assessment, and possessing great potential for early screening, dynamic monitoring and personalized health management.

[0043] In some embodiments, the nucleic acid sequencing is metagenomic sequencing, which is used to obtain the genomic information of all microorganisms in the intestinal microbial sample, and the abundance data of functional genes related to the neural active metabolic pathways are quantified by bioinformatics analysis; the neural active metabolic pathways include at least one of the tryptophan-serotonin / kynurenine pathway, the short-chain fatty acid synthesis pathway, the GABAergic system, and the choline and acetylcholine metabolic pathway; the abundance data of the functional genes correspond to the relative content of the functional genes involved in the neural active metabolic pathways; the bioinformatics analysis includes sequence quality control, sequence assembly, open reading frame prediction, and functional gene alignment.

[0044] Shotgun Metagenomic Sequencing can comprehensively capture the genomic information of intestinal microorganisms by unbiased sequencing of all microbial DNA in the sample, avoiding the limitation of 16S rRNA sequencing which only provides species classification information. Intestinal microorganisms participate in neural active metabolic pathways by encoding specific functional genes, and the abundance of these genes directly reflects the synthesis potential of microbial communities for neurotransmitter precursors such as tryptophan and GABA. For example, the tryptophan hydroxylase gene (TPH) in the tryptophan-serotonin pathway is responsible for catalyzing the conversion of tryptophan to 5-hydroxytryptophan, which is a key step in serotonin synthesis; the butyryl-CoA:acetyl-CoA transferase gene (but) in the short-chain fatty acid synthesis pathway is involved in butyric acid production, and butyric acid can affect neuroinflammation by passing through the blood-brain barrier. By bioinformatics analysis, the reads obtained by sequencing are mapped to the functional genes in the KEGG (Kyoto Encyclopedia of Genes and Genomes) or MetaCyc database, the relative abundance of these genes can be quantified, and the potential functional activity of the neural active metabolic pathways can be evaluated. This method converts microbial genomic information into quantifiable functional indicators, providing a molecular basis for depression risk assessment.

[0045] Total genomic DNA is extracted from 200 mg of fecal sample using a commercial DNA extraction kit (such as QIAamp PowerFecal Pro DNA Kit). The extraction process includes mechanical lysis (bead-beating), enzyme digestion (lysozyme treatment), and column purification to ensure DNA integrity. The extracted DNA is detected by Nanodrop for concentration (≥20 ng / μL) and purity (A260 / A280 ratio 1.8-2.0), and verified by agarose gel electrophoresis for no degradation.

[0046] Illumina NovaSeq 6000 platform was used for paired-end sequencing (PE150), and at least 10 Gb of raw data was generated for each sample. TruSeq DNA PCR-Free Kit was used for library construction to avoid amplification bias. The sequencing depth ensured coverage of the diversity of the intestinal microbial community (average sequencing depth ≥ 10x).

[0047] Fastp (v0.23.0) was used for raw read quality control, including removing low-quality reads (quality value < Q20), adapter sequences, and contaminated reads (such as host DNA), and outputting high-quality clean reads. After quality control, the read length was ≥ 100 bp, and the Q30 proportion was ≥ 90%.

[0048] MEGAHIT (v1.2.9) was used for de novo assembly of clean reads, with parameters set to k-min 21, k-max 141, to generate contigs. After assembly, the N50 was ≥ 1 kb, and short sequences with a length < 500 bp were filtered.

[0049] MetaGeneMark (v3.38) was used to predict ORFs for assembled contigs, identify coding sequences (CDS), and output ORF nucleotide and amino acid sequences.

[0050] DIAMOND (v2.0.15) was used to align ORF amino acid sequences to KEGG KO database (version 2023-01) using blastp mode, with an e-value threshold ≤ 1e-5 and coverage ≥ 70%. For the neurotransmitter metabolism pathway, specific KO numbers were screened for functional genes, such as: Tryptophan-serotonin / kynurenine pathway: KO number K00522 (tryptophan hydroxylase, TPH), K00453 (indoleamine-2,3-dioxygenase, IDO1), K01817 (kynureninease, KYNU).

[0051] Short-chain fatty acid synthesis pathway: KO number K01034 (butyryl-CoA:acetyl-CoA transferase, but), K00929 (butyrate kinase, buk).

[0052] GABAergic system: KO number K01580 (glutamate decarboxylase, gad).

[0053] Choline and acetylcholine metabolism pathway: KO number K14154 (choline trimethylamine lyase, cutC).

[0054] Based on the alignment results, the relative abundance of each functional gene was calculated using a custom Python script. The abundance was expressed in counts per million (CPM) with the formula: CPM = (number of aligned reads for a gene / total number of quality-controlled reads) x 10 6 . For low-abundance genes (CPM < 1), they were considered as noise and removed. The final output was a functional gene abundance matrix for each sample, with rows representing samples and columns representing functional genes.

[0055] Beneficial effects: metagenomic sequencing covers all microbial genomes, can detect low-abundance functional genes, significantly improves the resolution of neuroactive metabolic pathways, uniformity in sequencing depth calculation makes data comparable across batches and laboratories, directly quantifies functional gene abundance, complements metabolite concentration data, enhances the mechanistic interpretation of the depression risk model, automated bioinformatics pipeline supports large-scale sample analysis, can handle hundreds of samples in a single run, and takes less time.

[0056] In some embodiments, for the tryptophan-serotonin / kynurenine pathway, the sequences of tryptophan hydroxylase genes, tryptophan pyrrolase genes, kynurenine enzyme genes, and indoleamine-2, 3-dioxygenase homologous genes are extracted; for the short-chain fatty acid synthesis pathway, the sequences of butyryl-CoA:acetyl-CoA transferase genes and butyrate kinase genes are extracted; for the GABAergic system pathway, the sequence of glutamate decarboxylase gene is extracted; for the choline and acetylcholine metabolism pathway, the sequence of choline trimethylamine lyase gene is extracted; in the extraction process of functional genes, bioinformatics tools are used for sequence alignment and annotation, and the functional gene abundance is quality controlled, including removing low-quality sequences and correcting sequencing bias.

[0057] The functional status of neuroactive metabolic pathways is directly determined by the activity of its key catalytic enzymes, which in turn is highly correlated with the presence and abundance of their encoding genes. Therefore, precisely locating and quantifying these core functional genes is the key to assess the functional contribution of gut microbiota in the "gut-brain axis". For example, in the tryptophan-serotonin pathway, tryptophan hydroxylase (TPH) is the rate-limiting enzyme for the conversion of tryptophan to 5-hydroxytryptophan (5-HTP) by microorganisms, and its gene abundance directly affects the synthesis potential of serotonin precursor; while indoleamine-2,3-dioxygenase (IDO) dominates the conversion of tryptophan into kynurenine pathway, the activation of which is associated with neuroinflammation and depression. Short-chain fatty acids (such as butyric acid) are synthesized by enzymes encoded by butyryl-CoA:acetyl-CoA transferase (but) and butyrate kinase (buk) genes, and butyric acid has anti-inflammatory and maintains the integrity of the blood-brain barrier. GABA (gamma-aminobutyric acid) is an important inhibitory neurotransmitter, and its microbial synthesis depends on the glutamate decarboxylase (gad) gene. Choline trimethylamine lyase (cutC) converts choline to trimethylamine (TMA), and TMA is further oxidized to TMAO, which is associated with the risk of cardiovascular and nervous system diseases. By specifically extracting the sequences of these key genes from metagenomic data and accurately quantifying their abundance, a high-value feature set can be constructed that directly reflects the potential of neuroactive metabolism, providing input with clear biological significance for subsequent risk assessment models.

[0058] The specific workflow of the present embodiment includes the following core steps: Database construction of functional genes: All reference protein sequences related to the four neuroactive metabolic pathways were downloaded from KEGG, MetaCyc and UniProt databases. For each target enzyme, such as tryptophan hydroxylase (K00522), indoleamine-2,3-dioxygenase (K00453), kynurenine enzyme (K01817), butyryl-CoA:acetyl-CoA transferase (K01034), butyrate kinase (K00929), glutamate decarboxylase (K01580), choline trimethylamine lyase (K14154), a local reference sequence database was constructed. This database contains homologous sequences from different microbial species to cover the diversity of gut microorganisms.

[0059] High-confidence sequence alignment and extraction: DIAMOND (v2.0.15) was used to align the amino acid sequences predicted from metagenomic ORFs against the custom reference database described above. Blastp mode was adopted with a stringent e-value threshold of <1e-10, sequence identity of >70%, and query coverage of >80%. Only the alignments that met all three conditions were kept to ensure the extracted functional gene sequences were highly accurate and specific. For each ORF, the matched KO number, enzyme name, and corresponding reference sequence were recorded.

[0060] Sequence-specific verification: To further exclude false positives, HMMER (v3.3.2) software was used. For each target enzyme family, the hidden Markov model (HMM profile) was obtained from its Pfam or TIGRFAM database. The hmmsearch command was used to scan the preliminarily screened ORF sequences with an e-value threshold of <1e-5. Through the alignment of the HMM model, the protein domain could be more reliably identified, ensuring that the extracted sequences indeed belonged to the target functional gene family.

[0061] Abundance calculation and normalization: The original abundance of functional genes was calculated by counting the number of high-quality reads aligned to the unique identifier of the gene (such as the KO number). To eliminate the influence of sequencing depth differences between different samples, the "median of ratios" method of the DESeq2 package was used for normalization to generate standardized abundance values. This method could effectively compensate for the differences in library size between samples and reduce the influence of high-abundance species dominance. At the same time, the count per million (CPM) was calculated as an auxiliary indicator. For subsequent analysis, low-abundance or sporadically expressed genes with CPM <1 in more than 90% of samples were removed to reduce noise.

[0062] Functional gene abundance quality control: Sequence quality control: Before assembly, Fastp had been used to quality control the original reads. In the gene extraction stage, FastQC report was used to evaluate the quality distribution and base composition of contigs after assembly. Removal of PCR duplicates: The MarkDuplicates function of Picard Tools was used to identify and mark duplicate sequences that may be introduced by PCR amplification, and exclude these duplicate reads in abundance calculation. Correction of sequencing bias: The Seqtk (v1.3) tool was used to randomly downsample the original reads, making the sequencing depth of all samples uniform to the minimum depth of the batch of samples, to eliminate the abundance estimation bias caused by uneven sequencing depth. At the same time, the influence of GC content preference on gene abundance calculation was checked and corrected.

[0063] Beneficial effects: By focusing on the key enzyme genes of the four major neuroactive metabolic pathways, the biological interpretability of the functional characteristics is significantly improved; compared with whole genome indiscriminate counting, the precise extraction of target genes significantly improves the feature contribution of the risk model.

[0064] In some embodiments, a list of targeted metabolites corresponding to the neuroactive metabolic pathways is selected, including at least tryptophan, 5-hydroxytryptophan, serotonin, kynurenine, quinolinic acid and kynurenic acid in the tryptophan-serotonin / kynurenine pathway, butyric acid, propionic acid and acetic acid in the short-chain fatty acid synthesis pathway, gamma-aminobutyric acid in the GABAergic system pathway, and trimethylamine and oxidized trimethylamine in the choline and acetylcholine metabolism pathway; a targeted metabolomics strategy is used for metabolite extraction and detection of the intestinal microbial sample, wherein metabolite extraction uses organic solvents and adds internal standards to correct extraction efficiency; the concentrations of the metabolites are quantified by chromatography-mass spectrometry technology, and the raw concentration data are standardized, including internal standard correction, batch effect correction and concentration normalization, to eliminate experimental variations and obtain comparable metabolite concentration data.

[0065] Metabolites are the end products of functional gene expression, and their concentrations directly reflect the real-time biochemical activity and flux of neuroactive metabolic pathways. Unlike the "potential" information provided by gene abundance, metabolite concentration provides direct evidence of the "current state". For example, in the tryptophan-serotonin pathway, even if the tryptophan hydroxylase gene abundance is high, if the tryptophan substrate concentration is low or the kynurenine pathway is highly competitive, the final concentration of serotonin may be insufficient, which is closely related to depressive symptoms. Short-chain fatty acids (SCFAs) such as butyric acid, produced by intestinal microbes fermenting dietary fiber, not only provide energy for intestinal epithelial cells, but also enter the circulatory system and cross the blood-brain barrier to exert anti-neuroinflammatory and promote brain-derived neurotrophic factor expression effects. GABA is the main inhibitory neurotransmitter in the brain, and its peripheral concentration is associated with central nervous system status. Choline metabolite trimethylamine (TMA) and its oxidation product TMAO are associated with systemic inflammation and endothelial dysfunction, which may indirectly affect brain health. The targeted metabolomics strategy specifically monitors these pre-defined metabolites with clear biological significance, using the high sensitivity and high selectivity of liquid chromatography-mass spectrometry technology to achieve accurate quantification in complex biological matrices (such as feces). By introducing stable isotope-labeled internal standards and performing systematic data standardization, technical errors introduced during sample pretreatment and instrument analysis can be minimized, ensuring that metabolite concentration data obtained from different batches and different time points are highly comparable and biologically reliable.

[0066] The specific workflow of the present embodiment includes the following core steps: Targeted metabolite list and internal standard selection: Metabolite panel: Based on four major neuroactive metabolic pathways, a targeted list containing at least 15 core metabolites was determined: Tryptophan-serotonin / kynurenine pathway: L-tryptophan, 5-hydroxytryptophan, 5-hydroxytryptamine, L-kynurenine, quinolinic acid, kynurenic acid.

[0067] Short-chain fatty acid synthesis pathway: Acetate, propionate, butyrate.

[0068] GABAergic system pathway: Gamma-aminobutyric acid.

[0069] Choline and acetylcholine metabolism pathway: Choline, trimethylamine, trimethylamine-N-oxide.

[0070] Internal standard: To correct for variations in extraction and ionization efficiency, a stable-isotope-labeled internal standard was provided for each metabolite or each class of chemically similar metabolites. For example: d5-tryptophan, d4-serotonin, 13C4-butyrate, d6-GABA, d9-choline, d9-TMAO, etc. The internal standard was added at a known concentration before sample extraction.

[0071] Metabolite extraction: 100 mg (±5 mg) of homogenized frozen fecal sample was accurately weighed into a 2 mL centrifuge tube. 1 mL of pre-cooled extraction solution (methanol:water = 4:1, v / v, containing the above-mentioned mixed internal standards) was added, and vortexed immediately for 30 seconds. Ultrasonic-assisted extraction (ice water bath, power 300 W, duration 5 minutes) was performed to break the cells and release the metabolites. Centrifugation was performed at 14,000 rpm for 15 minutes at 4°C. 800 μL of supernatant was carefully aspirated and transferred to a new injection vial or 96-well plate, ready for LC-MS / MS analysis. If not immediately analyzed, it was stored at -80°C.

[0072] Chromatography-mass spectrometry detection: Liquid chromatography conditions: Instrument: ultra-high performance liquid chromatography system. Chromatographic column: for polar metabolites, a hydrophilic interaction chromatographic column (HILIC, such as Waters ACQUITY UPLC BEH Amide, 2.1 mm x 100 mm, 1.7 μm) was selected. Mobile phase: A phase (95% acetonitrile / 5% water containing 10 mM ammonium formate), B phase (50% acetonitrile / 50% water containing 10 mM ammonium formate). Gradient program: 0-2 min, 100% A; 2-8 min, 100% A → 40% A; 8-9 min, 40% A; 9-9.1 min, 40% A → 100% A; 9.1-12 min, 100% A (column equilibration). Flow rate: 0.3 mL / min; column temperature: 40°C; injection volume: 2 μL.

[0073] Mass spectrometry conditions: Instrument: Triple quadrupole mass spectrometer (e.g. Sciex QTRAP 6500+) equipped with an electrospray ion source. Ionization mode: Positive ion mode (ESI+) is used for most metabolites (tryptophan, serotonin, GABA, choline, TMAO, etc.), some metabolites (e.g. SCFAs) might need negative ion mode (ESI-) or detection after derivatization. Scan mode: Multiple reaction monitoring (MRM). Precursor ion, product ion, declustering voltage and collision energy are optimized for each target metabolite and its internal standard. Source parameters: Ion source temperature 500 °C, ionization voltage 5500 V (ESI+), curtain gas, nebulizer gas and auxiliary gas are set according to instrument optimization.

[0074] Data standardization: Raw data acquisition and integration: MRM signals are acquired using instrument software (e.g. Sciex Analyst or Agilent MassHunter), and chromatographic peak identification and integration are performed. Internal standard correction: The ratio of peak area of each target metabolite to that of its corresponding internal standard is calculated. This ratio can effectively correct the difference in extraction recovery and mass spectrometry ionization efficiency. Standard curve and quantification: A series of standard samples with known concentrations are analyzed simultaneously with internal standards, and a standard curve (usually linear regression) is drawn. The peak area ratio of metabolite / internal standard in the sample is substituted into the standard curve to calculate its absolute concentration (unit such as ng / g feces). For metabolites without absolute standards, relative concentrations are reported. Batch effect correction: QC samples made by mixing all samples are inserted in each analysis batch. Statistical methods (e.g. R package loess or ComBat) are used to correct batch-to-batch variation in the entire data set based on signal drift of QC samples. Concentration normalization: To eliminate the difference in water content or total amount of matrix between samples, the concentration after internal standard correction is divided by the wet weight of the sample, and finally the concentration data per unit mass is obtained. To further improve the data distribution to meet the requirements of subsequent statistical analysis and model input, the concentration data can be log2 transformed or Z-score standardized.

[0075] Beneficial effects: The targeted MRM mode effectively avoids interference in complex matrix, combined with isotope internal standard correction, making the accuracy of metabolite quantification high and precision good; through systematic internal standard correction, batch effect correction and concentration normalization, experimental technical variation is significantly reduced, making the data obtained by different laboratories and different time points reliable integrated and compared, laying the foundation for multi-center research; the obtained metabolite concentration data provides direct downstream biochemical verification for functional gene abundance data, realizes the full-link evaluation from "genetic potential" to "functional output", greatly enhances the biological rationality and persuasiveness of the depression risk assessment model, and the chromatography-mass spectrometry technology can detect metabolites in the fecal sample with a concentration range spanning several orders of magnitude (from nM to μM level), ensuring that low-abundance but functionally important metabolites are not missed.

[0076] In some embodiments, the targeted metabolomics strategy includes: a sample pretreatment stage, weighing an appropriate amount of fecal sample, adding an internal standard-containing methanol-water extraction solution, promoting metabolite release through vortex mixing and ultrasonic treatment, then centrifuging to obtain the supernatant as the detection sample; a chromatographic separation stage, using a hydrophilic interaction chromatographic column for liquid chromatographic separation, optimizing the mobile phase gradient to achieve efficient separation of target metabolites; a mass spectrometry detection stage, using a multiple reaction monitoring mode to quantify metabolites in the mass spectrometer, calculating absolute or relative concentrations by comparing standard curves and internal standard response values; a data post-processing stage, performing internal standard correction on the original concentration, and correcting batch-to-batch variation through statistical methods, and finally converting the concentration data to log scale or Z-score to improve the distribution characteristics.

[0077] The core of the targeted metabolomics strategy is to achieve high-precision and high-reproducible quantification of specific metabolites through a standardized process. This strategy breaks down the complex biological sample analysis into four interlocking stages: sample pretreatment aims to maximize the reproducible extraction of target metabolites while removing interfering substances such as proteins and lipids; chromatographic separation stage utilizes the difference in distribution coefficients between stationary phase and mobile phase of different metabolites to separate the co-extracted complex mixture in time dimension, thereby reducing ion suppression effect during mass spectrometric detection and ensuring independent detection window for each metabolite; mass spectrometric detection stage, especially the multiple reaction monitoring mode, gives the method extremely high selectivity and sensitivity by two-stage mass screening (selecting specific parent ions and their characteristic daughter ions), enabling accurate quantification of low-abundance metabolites in complex fecal matrix; the data post-processing stage is the key to convert the original instrument signal into reliable and comparable biological data, through internal standard correction to offset the fluctuations of extraction and ionization efficiency, and through batch correction to eliminate instrument performance drift and reagent differences. The final data transformation makes the data distribution more consistent with the assumptions of subsequent statistical analysis and machine learning models. The design philosophy of the entire strategy is "control variation and ensure comparability", ensuring that the final metabolite concentration data truly reflects biological differences rather than technical noise.

[0078] The specific workflow of this embodiment includes the following four precisely controlled stages: Sample pre-treatment stage: Accurate weighing and internal standard addition: Take the fecal sample from -80 °C and thaw on ice. Accurately weigh (100.0 ± 1.0) mg of homogenized feces into a 2 mL pre-cooled grinding tube using a precision balance. Immediately add 1.0 mL of pre-cooled extraction solution (80% methanol in water, v / v, containing known concentrations of mixed stable isotope internal standards such as ^13C6-tryptophan, d4-serotonin, ^13C4-butyric acid, etc.). Internal standards are added at the beginning of extraction for whole process monitoring and correction. Homogenization and extraction: Homogenize using a tissue grinder (e.g. MP Biomedicals FastPrep-24) at a speed of 6.0 m / s for 60 seconds to ensure complete sample disruption. Subsequently, place the centrifuge tube in an ice water bath and perform ultrasonic-assisted extraction (ultrasonic power 300 W, work for 2 seconds, intermittent for 3 seconds, total duration 5 minutes) to further release intracellular metabolites. Deproteinization and clarification: Centrifuge at 14,000 x g for 15 minutes in a 4 °C low-temperature centrifuge. After centrifugation, the sample is divided into three layers: the lower layer precipitate (bacterial residue, undigested food fiber), the middle thin layer (protein, lipid), and the upper layer clear extract. Supernatant collection: Carefully pipette 800 μL of supernatant and transfer to a clean 1.5 mL injection vial or 96-well plate. If chromatographic analysis cannot be performed immediately, seal with parafilm and store at -80 °C to avoid repeated freeze-thawing (no more than 2 times). Chromatographic separation stage: Chromatographic system and column: An ultra-high performance liquid chromatography system is used. For the polar metabolites in this protocol, a hydrophilic interaction chromatography column (HILIC) is selected. The column oven is set to 40 °C. Mobile phase and gradient elution: Phase A: 95% acetonitrile / 5% water containing 10 mM ammonium formate, pH 3.0. Phase B: 50% acetonitrile / 50% water containing 10 mM ammonium formate, pH 3.0. Gradient program: 0-1.0 min, 100% A; 1.0-8.0 min, 100% A → 60% A; 8.0-8.5 min, 60% A; 8.5-8.6 min, 60% A → 100% A; 8.6-12.0 min, 100% A (for column equilibration). Flow rate and injection: The flow rate is constant at 0.35 mL / min. The automatic injector temperature is maintained at 4 °C, and the injection volume is 3.0 μL. Random needle washing and needle cleaning programs are used to prevent cross-contamination.

[0079] Mass spectrometry detection stage: Mass spectrometer and ion source: A triple quadrupole mass spectrometer (e.g. Sciex QTRAP 6500+) equipped with an electrospray ion source is used. Ion source parameters: for positive ion mode (ESI+), the following settings are used: ion source temperature 550 °C, ionization voltage 5500 V, gas curtain gas 35 psi, nebulizer gas 50 psi, auxiliary gas 50 psi. Multiple reaction monitoring: at least one pair of MRM ion transitions (precursor ion → product ion) is optimized and set for each target metabolite and its internal standard, and the de-clustering voltage and collision energy are optimized. Mass spectrometry is performed with MRM scanning during the entire chromatographic gradient, with a dwell time set to 20-50 ms to ensure sufficient data points (typically > 12 points) for each chromatographic peak for accurate quantification. Quantitative calculation: using pure standard analytes, a standard curve is prepared with 5-8 concentration points analyzed under the same conditions. The chromatographic peak areas are automatically integrated by the instrument software (e.g. Sciex Analyst or MultiQuant), and the peak area ratio of the metabolite to the internal standard in the sample is substituted into the standard curve to calculate the absolute concentration (unit: ng / mL or mM). For metabolites without commercial standards, the relative abundance is reported relative to the internal standard or total signal.

[0080] Data post-processing stage: Internal standard correction: based on the ratio of the measured peak area of the internal standard in each sample to the theoretical / average peak area, the concentration of the corresponding target metabolite is corrected to compensate for individual differences in extraction and ionization efficiency. Batch effect correction: in each analysis batch, a quality control sample mixed from all samples is inserted at intervals. Using the loess function or a special package (e.g. MetaboAnalystR) in R language, based on the signal trend of the QC sample, the retention time and peak intensity of the entire batch of data are non-linearly corrected to eliminate instrument drift. Data transformation and standardization: first, the concentration after internal standard and batch correction is divided by the sample wet weight to obtain the concentration per unit mass (e.g. ng / g). Then, to improve the normal distribution of the data and stabilize the variance, all concentration data are log2 transformed. Finally, to further unify the dimension to adapt to subsequent machine learning models, the log2 transformed data can be Z-score standardized in the full sample range, with a mean of 0 and a standard deviation of 1.

[0081] Beneficial effects: Each step from sample processing to data output has strict SOPs and quality control, ensuring high repeatability and comparability of the data, effectively separating and specifically detecting target metabolites, significantly reducing matrix effects, making quantitative results accurate and reliable, and through data post-processing, the final metabolite concentration data distribution is better, and the scale of functional gene abundance data is more matched, directly improving the efficiency and performance of subsequent data fusion and machine learning model training, and improving the convergence speed of subsequent models.

[0082] In some embodiments, the data fusion processing comprises: performing standardization processing on the abundance data of the functional genes and the concentration data of the metabolites to eliminate technical bias caused by sequencing depth, sample size or detection method, wherein the standardization processing comprises converting the data into a distribution with a mean of 0 and a variance of 1 using a Z-score standardization method; performing feature selection to screen out features significantly related to depression risk from the standardized data, wherein the feature selection is based on the importance score of random forest to select the top 15%-20% features; performing feature weighting fusion to assign weights according to the biological importance of each feature in the depression pathological mechanism, wherein the weights are automatically determined based on prior knowledge or feature importance in the model training process; combining the weighted features into a multi-dimensional feature vector, i.e., the neural activity metabolic function spectrum, wherein the dimension of the multi-dimensional feature vector is consistent with the number of selected features, and each dimension represents a standardized and weighted feature value; and using the neural activity metabolic function spectrum to represent the functional state of the intestinal microorganisms of the individual in the neural activity metabolic pathway.

[0083] The mechanism by which the intestinal microbiome affects depression risk is multi-faceted and interconnected, and a single type of data (such as only gene abundance or only metabolite concentration) is difficult to fully capture this complexity. Functional gene abundance reflects the genetic potential of what the gut microbiota "can do", while metabolite concentration reveals the functional output of what they "actually do". Integrating these two data types allows for a closed-loop analysis from genetic potential to functional performance, thereby constructing a risk assessment model with more biological significance and predictive power. The core working principle is that, through machine learning algorithms, the feature pattern that can most effectively distinguish between high-risk and low-risk individuals is learned from the high-dimensional integrated feature set. For example, the model may learn that the combination of "low tryptophan hydroxylase gene abundance" + "low serotonin concentration" + "high kynurenine concentration" is a joint biomarker pattern that is more indicative of depression risk than any single indicator. Data preprocessing ensures that data from different sources and with different dimensions can be fairly compared and modeled; cross-validation and strict performance evaluation ensure that the model is not overfit to the training data, but has good generalization ability and can be applied to new, unseen individuals.

[0084] The functional gene abundance matrix (samples x genes) and the metabolite concentration matrix (samples x metabolites) from the preceding examples were horizontally merged by sample ID to form a unified integrated data matrix (samples x (genes + metabolites)). For missing values, the following strategy was adopted: if a feature (gene or metabolite) was missing in more than 20% of the samples, the feature was removed. For sporadic missing values in the remaining features, K-Nearest Neighbors algorithm was used to fill in the missing values (using R package impute, k = 10) to estimate the missing values using the similarity between samples in other features. Although the gene and metabolite data had been normalized in the previous stage, to eliminate the dimensional differences of the integrated features, all features were Z-score standardized to have a mean of 0 and a standard deviation of 1. This step was crucial for distance-based machine learning algorithms.

[0085] The random forest algorithm was preferred for model training. Its advantages are: it can handle high-dimensional features; it is not sensitive to multicollinearity between features; it can provide feature importance ranking, enhancing the interpretability of the model. The integrated feature set was used as the input variable (X), and the clinically evaluated depression status (e.g., PHQ-9 score ≥ 10 points as the case group, < 5 points as the healthy control group) as the binary output variable (Y). Using the RandomForestClassifier in the scikit-learn library of Python, the initial parameters were set, such as the number of decision trees in the forest n_estimators = 500, and the remaining parameters used the default values, and the model was fitted on the training set.

[0086] Five-fold cross-validation was used to optimize hyperparameters within the training set. Grid search or random search was used to optimize key parameters, with the average AUC under cross-validation as the optimization target. The optimized model was tested on an independent validation set that did not participate in the training process at all. The following core performance indicators were calculated: area under the curve: the receiver operating characteristic curve was drawn, and the AUC value was calculated to evaluate the overall discrimination ability of the model, the target was generally AUC > 0.75; accuracy: (true positives + true negatives) / total sample number; precision: true positives / (true positives + false positives), measuring the proportion of true positives among samples predicted as positive; recall: true positives / (true positives + false negatives), measuring the proportion of all true positive samples that are successfully predicted; F1 score: harmonic mean of precision and recall; performance threshold: set the minimum standard for model deployment, for example, AUC ≥ 0.80 on the independent validation set, and accuracy, precision, and recall are all ≥ 0.70.

[0087] The final model that meets the performance criteria (including its structure, parameters, and feature scaler) is saved in a serialized format (e.g., using Python's pickle format). When deployed, for a new individual sample, its functional gene abundance and metabolite concentration data are extracted following the exact same process, preprocessed and feature-scaled in the same way, and then the processed integrated features are input into the saved model to output its depression risk probability score.

[0088] Benefits: By integrating gene ("cause") and metabolite ("effect") data, the model can capture more complete biological pathway information. Cross-validation and independent validation set evaluation ensure that the model has good generalization ability and avoids overfitting. The feature importance ranking provided by the random forest model can intuitively reveal which functional genes and metabolites contribute most to distinguishing between risk groups, providing strong clues and hypothesis directions for subsequent mechanism research.

[0089] In some embodiments, the pre-trained risk assessment model is a classification or regression model trained based on a training sample set with known depression status, using a machine learning algorithm selected from random forest, XGBoost, support vector machine, or logistic regression. Before model application, the input neuroactive metabolite functional profile is preprocessed and feature-scaled in the same way as in the training phase to ensure data compatibility. The processed functional profile is input into the model, which calculates the depression risk probability value or continuous score based on the learned feature weights. Post-processing of the model output includes mapping the probability value to a risk level or generating a visual report, where the risk score is positively correlated with the severity of depression.

[0090] The core idea of this embodiment is to realize the closed loop from "digital risk assessment" to "physical intervention action". The gut microbiome is a highly plastic system, and its influence on the brain through neuroactive metabolic pathways can be adjusted through external intervention. However, the intervention must be targeted, otherwise it may be ineffective or even have adverse effects. For example, for individuals at high risk due to insufficient serotonin synthesis, supplementing precursor substances that produce 5-HTP or serotonin-producing probiotics is key; while for individuals at risk due to overactivation of the kynurenine pathway, intervention measures to suppress this pathway are needed. The working principle of this scheme is to use the explainable results (feature importance) of the machine learning model as a "navigation map" to accurately locate the dysfunctional biological pathway nodes. Subsequently, based on existing scientific research evidence on the microbiome-gut-brain axis, a knowledge base is constructed that connects "biological disorder types" with "evidence-based interventions". Finally, through logical matching, a multi-modal intervention plan is "tailored" for the individual, aiming to correct the identified intestinal microecological dysfunction at the root and thereby reduce the risk of depression.

[0091] The detailed workflow of this embodiment includes the following core steps: Model interpretation and key biomarker identification: risk probability interpretation: the model outputs the depression risk probability of individual i as P_i (e.g., P_i = 0.85). Set a risk threshold (such as 0.65), when P_i > 0.65, determine that the individual is high risk, and trigger the intervention scheme generation.

[0092] Contribution factor analysis: call the feature importance ranking provided by the random forest model (such as based on Gini impurity reduction or permutation importance). Select the top N (e.g., N = 10) features as the "key biomarkers" that contribute most to the risk of this individual. These markers may include: [gene] Kynureninase (KYNU) - high abundance; [metabolite] Quinolinic acid - high concentration; [metabolite] Serotonin - low concentration; [gene] Tryptophan hydroxylase (TPH) - low abundance.

[0093] Functional pathway mapping and imbalance type diagnosis: pathway mapping: map the above key biomarkers to the neural activity metabolic pathways they are directly involved in. For example, KYNU and quinolinic acid are highly expressed, while TPH and serotonin are low, which strongly indicates an imbalance type of "tryptophan metabolism tilting towards kynurenine pathway, insufficient serotonin synthesis".

[0094] Imbalance type classification: several common core functional imbalance types are predefined, such as: Type A: insufficient serotonin synthesis type; Type B: kynurenine pathway overactivation / neurotoxicity type; Type C: short-chain fatty acid (especially butyric acid) deficiency type; Type D: GABA synthesis deficiency type; Type E: choline-TMAO metabolic disorder / pro-inflammatory type; Type F: mixed type (combination of multiple types described above).

[0095] Intervention strategy knowledge base matching: Knowledge base structure: this is a pre-set rule database based on scientific literature and clinical guidelines. It is indexed by "functional imbalance type" and associated with a series of interventions. For example: For "Type B: Kynurenine Pathway Overactivation": Probiotics / prebiotics: Recommend supplementation with a probiotic combination rich in Lactobacillus and Bifidobacterium strains, as they have been shown to modulate tryptophan metabolism and reduce kynurenine production. Also suggest supplementation with prebiotics like inulin to promote the growth of these beneficial bacteria. Dietary adjustments: Recommend increasing foods rich in vitamin B6 (like tuna, chicken, chickpeas), as vitamin B6 is a cofactor inhibitor of kynurenine aminotransferase (KYNU) and promotes the production of kynurenic acid, a neuroprotective metabolite. Reduce single, large intakes of tryptophan-rich foods that are not balanced. Lifestyle tweaks: Recommend regular aerobic exercise, as studies have shown that exercise can lower peripheral kynurenine levels.

[0096] For "Type C: Butyrate Deficiency": Probiotics / prebiotics: Recommend supplementation with butyrate-producing bacteria like Clostridium butyricum or Blautia producta. Strongly suggest increasing the intake of dietary fibers like resistant starches (e.g., cooled potatoes, bananas), arabinoxylan (found in whole grains). Dietary adjustments: The diet recommendation is explicitly "high-fiber diet" with a specific list of foods (e.g., legumes, oats, nuts).

[0097] Individualized protocol report generation: Integration of baseline information: The matched general interventions are further personalized based on individual baseline information (e.g., age, gender, food allergy history, vegetarian preference, current medication). For example, for individuals with lactose intolerance, avoid recommending yogurt with Lactobacillus and suggest using a capsule formulation instead.

[0098] Report structured output: Generate an easy-to-understand personalized report containing: First section: Risk assessment summary: Display the risk score and the main imbalanced pathways. Second section: Key findings: Show the most contributing biomarkers in graphical form. Third section: Personalized action plan: Probiotic supplementation recommendations: e.g., recommended strains: Lactobacillus plantarum PS128, Bifidobacterium longum BB536. One capsule per day, taken after meals. Core dietary adjustments: e.g., ensure a daily intake of 30 grams of dietary fiber, with a focus on resistant starches (e.g., one slightly green banana per day). Eat Omega-3 rich fish (e.g., salmon) three times a week. Lifestyle tweaks: e.g., recommend 150 minutes of moderate-intensity aerobic exercise per week, such as brisk walking or jogging. Fourth section: Mechanistic explanation: Explain "why" these measures are taken in layman's terms, linking them to the detected biological imbalances, enhancing adherence.

[0099] Beneficial effects: change the "one-size-fits-all" supplement mode, make the intervention directly act on the identified problem pathway, significantly improve the expected intervention efficiency, provide personalized reports with clear biological explanation, users can understand the internal relationship between intervention measures and their own health status, and improve user compliance.

[0100] In some embodiments, the training and verification process of the pre-trained risk assessment model includes: training sample set construction, collecting intestinal microbial samples with known depression status, obtaining functional gene abundance data and metabolite concentration data, and constructing corresponding neural activity metabolic function spectrum; feature engineering stage, performing feature selection on the training set to optimize the feature subset, and applying feature scaling to improve model convergence; model training stage, using machine learning algorithms, optimizing hyperparameters through grid search or cross-validation to train the classifier with depression status as the label; model verification stage, using stratified sampling to divide the data into training set and test set, and using five-fold or ten-fold cross-validation to evaluate the model performance, the performance indicators include AUC, accuracy and F1 score; test the generalization ability of the model on the independent verification set to ensure the stability and accuracy of the model on unseen data.

[0101] The reliability and generalization ability of the pre-trained risk assessment model is the core of the depression risk assessment system. Machine learning models learn the difference patterns between depression and non-depression samples from high-dimensional feature vectors, but the model performance is highly dependent on the quality of training data, the rationality of feature engineering, and the rigor of verification strategy. The training sample set needs to cover a variety of depression states (such as mild, moderate, and severe depression) and healthy controls to capture the variability in the real world. Feature engineering reduces noise and improves model convergence efficiency by selecting key features and unifying scales. Model training adjusts the algorithm structure through hyperparameter optimization to prevent overfitting or underfitting. The verification stage uses stratified sampling and cross-validation to ensure unbiased evaluation, while the independent verification set test evaluates the stability of the model on unknown data, avoiding the situation that the model only performs well on training data. The whole process ensures that the model has reliable prediction ability and migratory in clinical application.

[0102] Collect gut microbiome samples from individuals with known depression status, from multiple sources including multi-center clinical cohorts (e.g. hospitals, communities) to ensure sample diversity. Depression status is confirmed by clinical diagnostic tools (e.g. PHQ-9 score, MINI International Neuropsychiatric Interview), and labeled into depression group (PHQ-9≥10) and healthy control group (PHQ-9<5). Sample size is usually ≥400 (half-half for depression and control), and each sample is subjected to both metagenomic sequencing and targeted metabolomics detection. Functional gene abundance data is obtained by metagenomic analysis pipeline (e.g. Fastp quality control, MEGAHIT assembly, DIAMOND alignment KEGG KO), represented in CPM (counts per million mapped reads); metabolite concentration data is obtained by LC-MS / MS targeted detection, represented in ng / g after internal standard correction and batch effect correction. Based on these data, construct the neuroactive metabolic functional profile: first, Z-score standardize gene abundance and metabolite concentration respectively, then use random forest feature importance score to select top 15% features (e.g. tryptophan hydroxylase gene, serotonin concentration, etc.), finally, weighted fusion according to biological importance (e.g. tryptophan metabolic pathway weight 1.2) and model adaptive weight (based on Gini importance), form a multi-dimensional feature vector with dimension N (N≈20). The sample set is randomly divided into training set and test set according to the ratio of 7:3, ensuring consistent class proportions.

[0103] Perform feature optimization on the training set, including: use random forest algorithm (n_estimators=500) to calculate the importance score of each feature, select the top 15%-20% features (e.g. key genes and metabolites involving tryptophan-serotonin pathway and GABAergic system), and remove low importance features to reduce dimensionality. Alternative methods include LASSO regression (alpha=0.01) or recursive feature elimination (RFE), and the optimal feature subset is determined by cross-validation. Apply Z-score standardization (use StandardScaler) to convert features to a distribution with mean 0 and variance 1, to improve model convergence speed; for skewed distribution data, optional log transformation or QuantileTransformer. Feature engineering is realized by Python pipeline (e.g. scikit-learn) to ensure compatibility with subsequent model training.

[0104] A classifier is trained using machine learning algorithms with depression status as the binary label (0 = healthy, 1 = depressed). Common algorithms include Random Forest, XGBoost, or Support Vector Machine. Hyperparameter optimization is performed through GridSearchCV or cross-validation (e.g., 5-fold cross-validation). The training process is performed on the training set, with early stopping (e.g., early_stopping_rounds=50 for XGBoost) to prevent overfitting, and the optimal model is saved as.pkl or ONNX format.

[0105] The data is divided into training and test sets (7:3 ratio) using stratified sampling to ensure consistent class proportions. Model performance is evaluated through 5-fold or 10-fold cross-validation: the training set is divided into 5 or 10 subsets, and one subset is used as the validation set, while the remaining subsets are used as the training set. The average performance indicators, including AUC (Area Under Curve), accuracy, F1 score, and sensitivity and specificity, are calculated. For example, the Random Forest model has a target AUC ≥ 0.90 and accuracy ≥ 85% in cross-validation. Finally, the model's generalization ability is tested on an independent validation set (e.g., 100 samples provided by an external agency), and the evaluation indicators are consistent with cross-validation, ensuring that the model's AUC does not decrease by more than 0.05 on unseen data.

[0106] Beneficial effects: Through a systematic training and validation process, the reliability and clinical applicability of the depression risk assessment model are significantly improved. Feature engineering optimizes the feature subset, increasing model training efficiency by about 30% and reducing the risk of overfitting; cross-validation and independent validation ensure the stability of the model in diverse populations, with high AUC in internal cross-validation and high AUC in independent validation. The entire process supports the deployment of the model in multiple centers, providing a standardized and repeatable solution for early screening of depression.

[0107] In some embodiments, the individual's intestinal microbiota sample is collected, and the sample type is a fecal sample or intestinal content sample; the sample is aliquoted and stored, immediately frozen at -80°C or placed in a nucleic acid / metabolite stabilizing solution to prevent degradation; in the nucleic acid extraction stage, commercial DNA extraction kits are used for genomic DNA extraction, and DNA concentration and purity are detected; in the metabolite extraction stage, organic solvents are used for metabolite extraction, and internal standards are added for subsequent quantitative correction; after sample pretreatment, macro-genome sequencing and targeted metabolomics detection are performed respectively to ensure consistency and comparability of data sources.

[0108] Sample collection and pre-processing are the cornerstone of ensuring the accuracy of depression risk assessment results. Intestinal microbial communities and their metabolites are extremely susceptible to changes after leaving the body environment. Rapid microbial reproduction, death, and the continuous activity of metabolic enzymes can significantly alter the original state of the sample, leading to bias in subsequent metagenomic and metabolomic data. Therefore, standardized collection, preservation, and pre-processing procedures are crucial for "freezing" the biological true state of the sample. Rapid cryogenic freezing (-80°C) can effectively inhibit enzyme activity and microbial metabolism; the use of nucleic acid / metabolite stabilizing liquid can prevent the degradation of biological macromolecules through chemical means. Unified DNA extraction and metabolite extraction procedures ensure high comparability of data between different samples, providing high-quality input for subsequent multi-omic data fusion and model construction, controlling technical variation from the source, and improving the robustness and repeatability of the entire risk assessment process.

[0109] Use disposable sterile fecal collection tubes (such as SAS-100) or endoscopic intestinal content samples. The subject collects or is operated by medical personnel, ensuring that the sample is processed within 30 minutes of excretion. For fecal samples, use the built-in sampling spoon to take about 2-3 grams of the core part, avoiding contact with the toilet or air for too long. The sample is immediately transferred to a pre-cooled 5mL cryogenic tube that has been pre-added with 1mL of methanol-water stabilizing liquid containing 1% protease inhibitor and 10µM antioxidant (such as ascorbic acid) to instantaneously inhibit metabolic activity. Subsequently, in a biological safety cabinet, the sample is quickly aliquoted into multiple 0.5g / tube screw-capped cryogenic tubes to reduce repeated freeze-thawing. After aliquoting is complete, all sample tubes are transferred to a -80°C ultra-low temperature freezer within 1 hour for long-term storage, and detailed sample metadata (collection time, storage time points, etc.) are recorded.

[0110] Take the sample out of -80°C freezer and thaw on ice. Use commercial fecal genomic DNA extraction kit (e.g. QIAamp PowerFecal Pro DNA Kit) which combines mechanical disruption (bead beating) and chemical lysis to effectively break the cell wall of Gram-positive bacteria. Detailed procedure: weigh 0.25 gram of thawed fecal sample, add lysis buffer and proteinase K provided in the kit, vortex and heat at 65°C for 10 minutes; then use a tissue homogenizer for high-speed bead beating (6.5 m / s, 2 x 45 s); after centrifugation, take the supernatant and perform DNA binding, washing (AW1 / AW2 buffer) and elution (EB buffer) through a silica gel membrane column, finally obtain about 50-100 µL of genomic DNA. Use a microspectrophotometer (e.g. NanoDrop One) to detect the DNA concentration (requirement ≥ 20 ng / µL) and purity (A260 / A280 ratio between 1.8-2.0, A260 / A230 > 2.0), and use a fluorometer (e.g. Qubit) for accurate quantification. The qualified DNA samples are stored at -20°C or directly enter the metagenomic sequencing library construction process.

[0111] Parallel to nucleic acid extraction, another aliquot of sample (0.3 gram) is taken for metabolite extraction. After thawing on ice, quickly transfer the sample into a pre-cooled 2 mL centrifuge tube, add 1 mL of pre-cooled extraction solvent (methanol:water = 4:1, v / v, containing a known concentration of isotope internal standard mixture, such as ^13C6-tryptophan, ^13C4-butyric acid, D4-serotonin, etc.). Use a tissue homogenizer to homogenize at 4°C for 60 seconds, then vortex for 2 minutes and ultrasonic treatment in an ice water bath for 10 minutes (power 300W, work 2s, intermittent 3s) to fully release intracellular metabolites. Centrifuge at 12,000 rpm for 15 minutes at 4°C, carefully aspirate 800 µL of supernatant, filter through a 0.22 µm organic phase microporous filter, and collect the filtrate in a sample injection bottle. This extract can be directly used for LC-MS / MS analysis. The addition of internal standards can correct the quantitative deviation caused by extraction efficiency, ion suppression, etc. at the beginning of extraction.

[0112] To ensure batch consistency, quality control samples are set up for each batch of sample processing (from collection to extraction): Blank control: a tube containing only extraction solvent, used to monitor background contamination. Process control: use commercial standard fecal matrix (e.g. SeraCon) or mixed pool sample, which undergoes the exact same process as the sample to be tested, to monitor the stability of extraction and detection. Internal standard recovery rate: by calculating the ratio of the measured concentration of internal standard to the added concentration, evaluate the extraction efficiency, the recovery rate should be between 85%-115%. All quality control data are recorded for subsequent data correction and batch effect removal.

[0113] Beneficial effects: By standardizing the sample collection, preservation and pre-processing procedures, the coefficient of variation introduced by technical operations between samples is significantly reduced. This not only significantly improves the accuracy and comparability of metagenomic functional gene abundance data and targeted metabolite concentration data, but also lays a solid foundation for subsequent construction of high-quality training sample set and obtaining reliable depression risk score. The unified procedure enables seamless integration of sample data collected at different times and different places, greatly enhancing the transferability and clinical application value of the risk assessment model.

[0114] Figure 3 The structure diagram of a depression risk assessment system based on intestinal flora characteristics provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the depression risk assessment system 300 based on intestinal flora characteristics of the embodiment comprises a sample analysis module 301, a metabolite analysis module 302, a multi-dimensional feature analysis module 303 and a depression risk analysis module 304. Figure 3

[0115] The sample analysis module 301 is configured to obtain an intestinal microbial sample of an individual, perform nucleic acid sequencing on the intestinal microbial sample, and determine the abundance data of functional genes related to a neural active metabolic pathway in intestinal microorganisms. The metabolite analysis module 302 is configured to perform metabolite detection on the intestinal microbial sample, and determine the concentration data of metabolites related to the neural active metabolic pathway in intestinal microorganisms. The multi-dimensional feature analysis module 303 is configured to perform data fusion processing based on the abundance data of the functional genes and the concentration data of the metabolites, and construct a neural active metabolic function spectrum, wherein the neural active metabolic function spectrum is a multi-dimensional feature vector. The depression risk analysis module 304 is configured to input the neural active metabolic function spectrum into a pre-trained risk assessment model, and output a depression risk score of the individual.

[0116] Optionally, when the nucleic acid sequencing is metagenomic sequencing, the sample analysis module 301 is configured to obtain the genomic information of all microorganisms in the intestinal microbial sample, and when quantifying the abundance data of functional genes related to the neural active metabolic pathway through bioinformatics analysis, the sample analysis module 301 is specifically configured to: the neural active metabolic pathway comprises at least one of the tryptophan-serotonin / kynurenine pathway, the short-chain fatty acid synthesis pathway, the GABA system and the choline and acetylcholine metabolic pathway; the abundance data of the functional genes corresponds to the relative content of the functional genes involved in the neural active metabolic pathway; and the bioinformatics analysis comprises sequence quality control, sequence assembly, open reading frame prediction and functional gene alignment.

[0117] ​Optionally, the sample analysis module 301, in the extraction of functional genes, is specifically used for: for the tryptophan-serotonin / kynurenine pathway, extracting the sequences of tryptophan hydroxylase gene, tryptophan pyrrolase gene, kynurenine enzyme gene and indoleamine-2, 3-dioxygenase homolog gene; for the short-chain fatty acid synthesis pathway, extracting the sequences of butyryl-CoA: acetyl-CoA transferase gene and butyric acid kinase gene; for the GABAergic system pathway, extracting the sequence of glutamate decarboxylase gene; for the choline and acetylcholine metabolism pathway, extracting the sequence of choline trimethylamine lyase gene; in the extraction process of functional genes, bioinformatics tools are used for sequence alignment and annotation, and the abundance of functional genes is controlled, including removing low-quality sequences and correcting sequencing bias.

[0118] Optionally, the metabolic analysis module 302, in the determination of the concentration data of metabolites in the intestinal microorganism related to the neuroactive metabolic pathway, is specifically used for: selecting a list of target metabolites corresponding to the neuroactive metabolic pathway, the target metabolites at least including tryptophan, 5-hydroxytryptophan, serotonin, kynurenine, quinolinic acid and kynurenic acid in the tryptophan-serotonin / kynurenine pathway, butyric acid, propionic acid and acetic acid in the short-chain fatty acid synthesis pathway, γ-aminobutyric acid in the GABAergic system pathway, and trimethylamine and oxidized trimethylamine in the choline and acetylcholine metabolism pathway; adopting a targeted metabolomics strategy to extract and detect metabolites from the intestinal microbial sample, wherein organic solvents are used for metabolite extraction and internal standards are added to correct extraction efficiency; quantifying the concentration of the metabolites by chromatography-mass spectrometry and standardizing the raw concentration data, including internal standard correction, batch effect correction and concentration normalization, to eliminate experimental variation and obtain comparable metabolite concentration data.

[0119] Optionally, the metabolic analysis module 302, the targeted metabolomics strategy, is specifically used for: in the sample pretreatment stage, weighing an appropriate amount of fecal sample, adding methanol-water extract containing internal standard, promoting metabolite release by vortex mixing and ultrasonic treatment, then centrifuging to obtain supernatant as detection sample; in the chromatographic separation stage, using a hydrophilic interaction chromatographic column for liquid chromatographic separation, optimizing the mobile phase gradient to realize efficient separation of target metabolites; in the mass spectrometry detection stage, using multiple reaction monitoring mode to quantify metabolites in mass spectrometer, calculating absolute concentration or relative concentration by comparing standard curve and internal standard response value; in the data post-processing stage, correcting the raw concentration with internal standard, and correcting batch-to-batch variation by statistical method, and finally converting the concentration data to logarithmic scale or Z-score to improve the distribution characteristics.

[0120] Optionally, the multi-dimensional feature analysis module 303, in the data fusion process, is specifically used for: standardizing the abundance data of the functional genes and the concentration data of the metabolites to eliminate technical bias caused by sequencing depth, sample size or detection method, wherein the standardization process includes converting the data into a distribution with a mean of 0 and a variance of 1 using the Z-score standardization method; performing feature selection to filter out features significantly related to depression risk from the standardized data, wherein the feature selection is based on the importance score of the random forest to select the top 15%-20% features; performing feature weighted fusion according to the biological importance of each feature in the depression pathological mechanism, wherein the weight is automatically determined based on prior knowledge or feature importance in the model training process; combining the weighted features into a multi-dimensional feature vector, i.e., the neural activity metabolic function spectrum, wherein the dimension of the multi-dimensional feature vector is consistent with the number of selected features, and each dimension represents a standardized and weighted feature value; and the neural activity metabolic function spectrum is used to represent the functional state of the intestinal microorganisms in the neural activity metabolic pathway of the individual.

[0121] Optionally, the depression risk analysis module 304, in the process of inputting the neural activity metabolic function spectrum into the pre-trained risk assessment model and outputting the depression risk score of the individual, is specifically used for: the pre-trained risk assessment model is a classification or regression model trained by a machine learning algorithm based on a known depression state training sample set, and the machine learning algorithm is selected from random forest, XGBoost, support vector machine or logistic regression; before model application, the input neural activity metabolic function spectrum is preprocessed and feature scaled in accordance with the training phase to ensure data compatibility; the processed function spectrum is input into the model, and the model calculates the depression risk probability value or continuous score based on the learned feature weight; and the model output is post-processed, including mapping the probability value to a risk level or generating a visual report, wherein the risk score is positively correlated with the severity of depression.

[0122] Optionally, the depression risk analysis module 304 is specifically configured to: construct a training sample set, collect intestinal microbiota samples with known depression status, obtain functional gene abundance data and metabolite concentration data, and construct a corresponding neural activity metabolic function spectrum; in a feature engineering stage, perform feature selection on the training set to optimize the feature subset, and apply feature scaling to improve model convergence; in a model training stage, use a machine learning algorithm, optimize hyperparameters through grid search or cross-validation, and train a classifier with depression status as a label; in a model validation stage, divide the data into a training set and a test set using stratified sampling, and use five-fold or ten-fold cross-validation to evaluate the model performance, and the performance indicators include AUC, accuracy and F1 score; test the model generalization ability on an independent validation set to ensure the stability and accuracy of the model on unseen data.

[0123] Optionally, the system 300 further comprises a sample quality control module 305, which is specifically configured to: collect intestinal microbiota samples of individuals, and the sample types are fecal samples or intestinal content samples; package and store the samples, immediately freeze them at -80°C or place them in nucleic acid / metabolite stabilizing liquid to prevent degradation; in the nucleic acid extraction stage, use a commercial DNA extraction kit to extract genomic DNA, and detect the DNA concentration and purity; in the metabolite extraction stage, use organic solvents to extract metabolites, and add internal standards for subsequent quantitative correction; after sample pretreatment, perform metagenomic sequencing and targeted metabolomics detection respectively to ensure the consistency and comparability of data sources.

[0124] The system of the embodiment can be used to execute the method of any of the above embodiments, and has similar implementation principles and technical effects, which will not be described here again.

Claims

1. A method for assessing depression risk based on gut microbiota characteristics, characterized in that, include: S1. Obtain an individual's gut microbiota sample, perform nucleic acid sequencing on the gut microbiota sample, and determine the abundance data of functional genes related to neuroactive metabolic pathways in the gut microbiota; S2. Perform metabolite detection on the gut microbial sample to determine the concentration data of metabolites related to the neuroactive metabolic pathway in the gut microbial sample; S3. Based on the abundance data of the functional genes and the concentration data of the metabolites, perform data fusion processing to construct a neuroactive metabolic function profile, wherein the neuroactive metabolic function profile is a multidimensional feature vector. S4. Input the neuroactive metabolic function spectrum into the pre-trained risk assessment model and output the individual's depression risk score.

2. The method according to claim 1, characterized in that, In step S1, the nucleic acid sequencing is metagenomic sequencing, which is used to obtain the genomic information of all microorganisms in the gut microbiota sample and to quantify the abundance data of functional genes related to neuroactive metabolic pathways through bioinformatics analysis. The neuroactive metabolic pathways include at least one of the following: the tryptophan-serotonin / kynurenine pathway, the short-chain fatty acid synthesis pathway, the GABAergic system, and the choline and acetylcholine metabolic pathway. The abundance data of the functional genes correspond to the relative content of the functional genes involved in the neural active metabolic pathway; The bioinformatics analysis includes sequence quality control, sequence assembly, open reading frame prediction, and functional gene alignment.

3. The method according to claim 2, characterized in that, The extraction of the functional genes includes: Sequences of tryptophan hydroxylase gene, tryptophan pyrrolase gene, kynurenase gene and indoleamine-2,3-dioxygenase homolog gene were extracted for the tryptophan-serotonin / kynurenine pathway. Sequences of butyryl-CoA:acetyl-CoA transferase gene and butyrate kinase gene were extracted for the short-chain fatty acid synthesis pathway. Sequence of the glutamate decarboxylase gene was extracted targeting the GABAergic system pathway; Sequence of the choline trimethylamine lyase gene was extracted targeting the choline and acetylcholine metabolic pathways. During the extraction of functional genes, bioinformatics tools were used for sequence alignment and annotation, and the abundance of functional genes was quality controlled, including the removal of low-quality sequences and correction of sequencing bias.

4. The method according to claim 1, characterized in that, In step S2, determining the concentration data of metabolites related to the neuroactive metabolic pathway in the gut microbiota includes: Select a list of target metabolites corresponding to the neuroactive metabolic pathways, wherein the target metabolites include at least tryptophan, 5-hydroxytryptophan, serotonin, kynurenine, quinolinic acid and kynurenic acid in the tryptophan-serotonin / kynurenine pathway, butyric acid, propionic acid and acetic acid in the short-chain fatty acid synthesis pathway, γ-aminobutyric acid in the GABAergic system pathway, and trimethylamine and trimethylamine oxide in the choline and acetylcholine metabolic pathway; A targeted metabolomics strategy was used to extract and detect metabolites from the gut microbiota samples. The metabolite extraction used organic solvents and added internal standards to correct the extraction efficiency. The concentrations of the metabolites were quantified using chromatography-mass spectrometry (GC-MS), and the raw concentration data were standardized, including internal standard correction, batch effect correction, and concentration normalization, to eliminate experimental variability and obtain comparable metabolite concentration data.

5. The method according to claim 4, characterized in that, The targeted metabolomics strategy includes: In the sample pretreatment stage, an appropriate amount of fecal sample was weighed, and methanol-water extract containing internal standard was added. The release of metabolites was promoted by vortex mixing and ultrasonic treatment. Then, the supernatant was collected by centrifugation as the test sample. In the chromatographic separation stage, a hydrophilic interaction chromatographic column is used for liquid chromatography separation, and the mobile phase gradient is optimized to achieve efficient separation of target metabolites. In the mass spectrometry detection stage, the metabolites are quantified in the mass spectrometer using multiple reaction monitoring mode, and the absolute or relative concentration is calculated by comparing the response values ​​of the standard curve and the internal standard. In the data post-processing stage, the original concentrations are corrected using internal standards, and batch-to-batch variation is corrected using statistical methods. Finally, the concentration data are converted to logarithmic scale or Z-score to improve distribution characteristics.

6. The method according to claim 5, characterized in that, The data fusion process includes: The abundance data of the functional genes and the concentration data of metabolites are standardized to eliminate technical biases caused by sequencing depth, sample size or detection method. The standardization process includes using the Z-score standardization method to convert the data into a distribution with a mean of 0 and a variance of 1. Feature selection was performed to screen out features that were significantly associated with depression risk from the standardized data. The feature selection was based on the importance score of random forest, which selected the top 15% to 20% of features. Feature weighted fusion is performed, and weights are assigned to each feature based on its biological importance in the pathological mechanism of depression. The weights are automatically determined based on prior knowledge or the importance of features during model training. The weighted features are combined into a multidimensional feature vector, namely the neuroactive metabolic function spectrum, wherein the dimension of the multidimensional feature vector is consistent with the number of selected features, and each dimension represents a standardized and weighted feature value. The neuroactive metabolic functional profile is used to represent the functional status of the individual's gut microbiota in neuroactive metabolic pathways.

7. The method according to claim 1, characterized in that, In step S5, inputting the neuroactive metabolic function profile into the pre-trained risk assessment model and outputting the individual's depression risk score includes: The pre-trained risk assessment model is a classification or regression model trained using a training sample set of known depressive states through a machine learning algorithm, which is selected from random forest, XGBoost, support vector machine or logistic regression. Before applying the model, the input neural activity metabolic function spectrum is preprocessed and feature scaled in the same way as during the training phase to ensure data compatibility. The processed functional spectrum is input into the model, which calculates the depression risk probability value or continuous score based on the learned feature weights. The model output is post-processed, including mapping probability values ​​to risk levels or generating visualization reports, where the risk score is positively correlated with the severity of depression.

8. The method according to claim 5, characterized in that, The training and validation process of the pre-trained risk assessment model includes: Training sample set construction: collect gut microbiota samples of known depressive states, obtain their functional gene abundance data and metabolite concentration data, and construct the corresponding neuroactive metabolic functional profile; In the feature engineering phase, feature selection is performed on the training set to optimize the feature subset, and feature scaling is applied to improve model convergence. During the model training phase, machine learning algorithms are used to optimize hyperparameters through grid search or cross-validation, and a classifier is trained with the depressive state as the label. During the model validation phase, stratified sampling is used to divide the data into training and test sets, and five-fold or ten-fold cross-validation is used to evaluate the model performance. Performance metrics include AUC, accuracy, and F1 score. Test the model's generalization ability on an independent validation set to ensure that the model remains stable and accurate on unseen data.

9. The method according to claim 5, characterized in that, The method further includes a sample collection and preprocessing step, which is performed before S1, including: Collect gut microbiota samples from individuals; the sample type is either a fecal sample or a gut contents sample. Aliquot and store the samples immediately, freezing them at -80°C or placing them in a nucleic acid / metabolite stabilizing solution to prevent degradation; During the nucleic acid extraction stage, commercial DNA extraction kits were used to extract genomic DNA, and the DNA concentration and purity were tested. In the metabolite extraction stage, organic solvents were used for metabolite extraction, and internal standards were added for subsequent quantitative correction. After sample preprocessing, metagenomic sequencing and targeted metabolomics detection were performed to ensure the consistency and comparability of data sources.

10. A depression risk assessment system based on gut microbiota characteristics, characterized in that, The method applied to any one of claims 1-9 includes: The sample analysis module is used to obtain gut microbial samples from individuals, perform nucleic acid sequencing on the gut microbial samples, and determine the abundance data of functional genes related to neuroactive metabolic pathways in the gut microbiota. The metabolic analysis module is used to detect metabolites in the gut microbial sample and determine the concentration data of metabolites related to the neuroactive metabolic pathway in the gut microbial sample. The multidimensional feature analysis module is used to perform data fusion processing based on the abundance data of the functional genes and the concentration data of the metabolites to construct a neuroactive metabolic function profile, wherein the neuroactive metabolic function profile is a multidimensional feature vector. The depression risk analysis module is used to input the neuroactive metabolic function spectrum into the pre-trained risk assessment model and output the individual's depression risk score.

Citation Information

Patent Citations

  • Application of bacteria in development assessment and treatment of children

    CN115835875A

  • Method for identifying disease microorganism-metabolite combined marker based on host-flora co-metabolism model

    CN120375920A

  • Compositions for modulating gut microflora populations, treatment of dysbiosis and disease prevention, and methods for making and using same

    WO2024182434A2

Cited By

  • Machine learning-based clinical mass spectrum risk prediction method and equipment

    CN122084815A