Methods and systems for detecting metabolites and microbiomes associated with anxiety and / or depression in irritable bowel syndrome patients
The integration of gut metagenome data from bacteriome, mycobiome, and virome, along with metabolic features, addresses the limitations of single-domain microbiome studies in IBS, enabling precise detection and personalized treatment of anxiety and depression.
Patent Information
- Application Number
- PCT/CN2025/093680
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-09
- Filing Date
- 2025-05-09
- Publication Date
- 2025-11-13
Smart Images

Figure CN2025093680_13112025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR DETECTING METABOLITES AND MICROBIOMES ASSOCIATED WITH ANXIETY AND / OR DEPRESSION IN IRRITABLE BOWEL SYNDROME PATIENTSInventors: Zhaoxiang BIAN; and Qin LIUCross-Reference to Related Applications:
[0001] The present application claims priority from U.S. provisional patent application serial number 63 / 644,512 filed May 9th, 2024, and the disclosure of which is incorporated herein by reference in its entirety.Field of the Invention:
[0002] The present invention generally relates to the medical and bioinformatic fields. More specifically the present invention relates to methods and systems for detecting anxiety and depression markers in irritable bowel syndrome patients using multi-kingdom gut microbiome and metabolome profiles.Background of the Invention:
[0003] Irritable bowel syndrome (IBS) is a chronic functional gastrointestinal disorder, with an estimated prevalence of 6.6%in Hong Kong. The frequent co-occurrence of mental health conditions with IBS is well documented-over one-quarter of individuals with IBS experience depressive symptoms, while more than one-third present with symptoms of anxiety. These rates are significantly higher than those observed in the general population. Globally, anxiety disorders affect more than 260 million people and rank as the sixth leading cause of disability. Meanwhile, depression, identified by the World Health Organization as the leading contributor to global disability, affects over 300 million individuals and is associated with approximately 800,000 suicide deaths annually. The burden of mental disorders continues to rise, with many individuals experiencing both anxiety and depression concurrently.
[0004] There is a critical need for biological markers that reflect the underlying physiological mechanisms of psychiatric comorbidities in IBS. Gut dysbiosis, or imbalance in the gut microbiota, has been implicated in IBS, particularly among patients with psychiatric symptoms. However, most existing studies focus narrowly on bacterial populations, neglecting the broader microbiome landscape that includes fungi, viruses, and other microorganisms, as well as their genetic content. Recent research supports the use of multi-kingdom and functional microbiome markers as promising diagnostic and therapeutic tools across various diseases. This suggests that prior microbiome studies in IBS with comorbid mental disorders may have overlooked key contributors to disease development. A more integrative, multi-biome approach is essential to accurately capture the in vivo state and to identify more comprehensive and meaningful biomarkers for psychiatric conditions in IBS. We hypothesize that incorporating multi-kingdom microbial data will reveal intricate interactions within the microbiome and allow for improved stratification of disease, which cannot be achieved by studying single microbial domains alone.
[0005] The identification of robust biomarkers for mental disorders in IBS could significantly advance clinical care and public health by: (1) enabling earlier and more accurate diagnosis; (2) facilitating stratification of patients into subgroups for targeted, personalized therapies; (3) empowering patients with a better understanding of their condition, thereby enhancing engagement and treatment adherence; and (4) informing public health strategies by identifying high-risk groups and optimizing the distribution of mental health resources, particularly within the IBS population in Hong Kong.
[0006] Biomarker discovery in the context of IBS-associated mental disorders holds the potential to transform patient care. It offers a pathway to more precise diagnostics, individualized treatment plans, and a deeper understanding of the gut-brain axis. Such advancements could lead to earlier interventions, improved health outcomes, and reduced burden of mental illness in affected populations. Therefore, the present invention addresses this need.Summary of the Invention:
[0007] It is an objective of the present invention to provide a method or system to solve the aforementioned technical problems.
[0008] In accordance with a first aspect of the present invention, a method for detecting anxiety and / or depression metabolic and microbiome markers in a patient with IBS is provided. The method includes the following steps: obtaining a serum sample and a fecal sample from the patient; extracting a gut metagenome from the fecal sample including a bacteriome, a mycobiome, and a virome of the patient; extracting metabolic features from the serum sample and fecal sample; and detecting whether a marker set of microbial species and metabolites for depression and anxiety is present in the gut metagenome and the metabolic features by utilizing a classifier to screen the gut metagenome and the metabolic features and locate the presence of the marker set of microbial species and metabolites.
[0009] In accordance with one embodiment of the present invention, the marker set of microbial species and metabolites for depression and anxiety has one or more metabolite selected from the group consisting of linoleic acid, triglycerides, glycolithocholic acid (GLCA) , isolithocholic acid, ω-muricholic acid (wMCA) , oleic acid, palmioleic acid, 6-keto lithocholic (6_KLCA) , palmitic acid, cholic acid (CA) , hyocholic acid, oleamide, and 2-arachidonylglycerol.
[0010] In accordance with one embodiment of the present invention, the marker set of microbial species and metabolites for depression and anxiety includes one or more microbial species selected from bacteria, viruses, or fungi, including but not limited to: Escherichia coli, Enterobacteria phage cdtI, Enterobacteria phage mEp460, Shigella phage fII, Staphylococcus phage SPbeta-like, Bacteroides uniformis, Escherichia phage TL 2011b, Malassezia vespertilionis, Janthinobacterium phage vB JliM Donnerlittchen, Escherichia phage D6, Komagataella phaffii, Enterobacteria phage SfI, Firmicutes bacterium CAG: 83, Alcaligenes phage vB_AfaP_QDWS595, Microbacterium phage Pumpernickel, Shigella phage SfIV, Enterobacteria phage SfV, Nakaseomyces glabratus, Synechococcus phage S-H9-1, Ruminococcus gnavus, Ustilago maydis, Bacillus phage vB_BanS_Nate, Streptomyces phage Success, Faecalibacterium prausnitzii, Thermothelomyces thermophilus, Streptomyces phage Faust, Escherichia phage vB_EcoM_Ro157c2YLVW, Bacteroides xylanisolvens, ruthenibacterium lactatiformans, triglyceride glucose (TyG) index, Aspergillus chevalieri, Fusarium fujikuroi, Schizosaccharomyces pombe, Fusarium graminearum, Fusarium venenatum, Streptomyces phage Wakanda, Ruminococcus bromii, Fusarium falciforme, Synechococcus phage S-SCSM1, Salmonella phage SEN22, Yarrowia lipolytica, Brettanomyces bruxellensis, Fusarium oxysporum, Psilocybe cubensis, Colletotrichum lupini, Aeromonas phage LAh10, Cercospora beticola, Akanthomyces muscarius, Pyricularia grisea, Roseburia inulinivorans, Kosakonia phage Kc263, Kluyveromyces lactis, oleic acid, Intestinimonas butyriciproducens, Aspergillus fumigatus, Trichoderma breve, Pyricularia oryzae, Bacteroides intestinalis, Bacteroides thetaiotaomicron, Salmonella phage JD01, Alistipes shahii, Clostridium leptum, Alistipes putredinis, Naumovozyma castellii, Enterobacteria phage IME10, Lachancea thermotolerans, Schizosaccharomyces osmophilus, Saccharomycodes ludwigii, Ustilaginoidea virens, Candida dubliniensis, Saccharomyces kudriavzevii, Parabacteroides distasonis, Synechococcus phage S-SZBM1, Enterobacteria phage mEp237, Naumovozyma dairenensis, Purpureocillium takamizusanense, Rhizoctonia solani, Pyricularia pennisetigena, Brettanomyces nanus, Enterobacteria phage P4, 2-Arachidonylglycerol, Blautia wexlerae, Bacteroides vulgatus, Bacteroides clarus, Eubacterium sp. CAG: 38, Bacteroides nordii, Bacteroides cellulosilyticus, Marasmius oreades, and Enterobacteria phage phiP27.
[0011] In accordance with another embodiment of the present invention, the GLCA, isolithocholic acid, wMCA, 6_KLCA, CA, hyocholic acid are detected in the serum sample.
[0012] In accordance with one embodiment of the present invention, the gut metagenome is extracted using weighted similarity network fusion (WSNF) , wherein the weighting is based on the number of species in each bacteriome, mycobiome, and virome.
[0013] In accordance with one embodiment of the present invention, the method further includes a step of calculating a triglyceride glucose (TyG) index using the formula: ln [Fasting TG (mg / dL) ×Fasting Glucose (mg / dL) / 2] .
[0014] In accordance with one embodiment of the present invention, the classifier is a machine learning classifier trained to detect multi-class disease phenotypes, including IBS with anxiety and IBS with depression.
[0015] In accordance with one embodiment of the present invention, the classifier includes a random forest model trained using gut multi-omics data and validated using nested cross-validation.
[0016] In accordance with a second aspect of the present invention, a system for detecting anxiety and / or depression metabolites and microbiomes in IBS patients is introduced. The system includes the following components: a sample processing module configured to analyze serum and fecal samples to generate gut metagenome and metabolome data; a fusion module configured to integrate bacteriome, mycobiome, and virome data using WSNF; a classification module implementing a trained machine learning model configured to classify IBS patients with or without comorbid anxiety or depression based on a comparison with a marker set of microbial species and metabolites for depression and anxiety; and an output interface configured to display the predicted mental health status of the subject based on the classifier output.
[0017] In accordance with one embodiment of the present invention, the fusion module is configured to: assign weights to each microbiome sub-dataset based on species richness; integrate bacteriome, mycobiome, and virome data using WSNF; and generate a composite similarity network for spectral clustering of patient profiles.
[0018] In accordance with one embodiment of the present invention, the classification module is further configured to: implement a random forest model trained on a labeled dataset comprising multi-omics features of IBS patients with and without comorbid anxiety or depression; compute a classification score representing the likelihood of mental disorder; and output model confidence scores or probabilities associated with the prediction.Brief Description of the Drawings:
[0019] Embodiments of the invention are described in more details hereinafter with reference to the drawings, in which:
[0020] FIGs. 1A-1F depict the stratification of IBS-D by WSNF using multi-kingdom microbiome, in which FIG. 1A is a heatmap illustrating WSNF similarity scores of IBS patients stratified by spectral clustering, FIG. 1B shows the comparison of Bristol score between the two identified patient clusters (cluster 1, n = 118; cluster 2, n = 166) according to integrated multi-biome profiles, derived from n = 284 biologically independent samples, FIG. 1C depicts a principal coordinate analysis (PCoA) of bacterial microbiome based on Bray-Curtis dissimilarity illustrating patient clusters and health control, FIG. 1D demonstrates the Bray-Curtis principal coordinates analysis of R. gnavus, FIG. 1E illustrates the microbial species associated with IBS clusters and FIG. 1F illustrates the microbial species associated with non-IBS clusters;
[0021] FIGs. 2A-2C depict the multi-omics based machine learning for identifying mental disorder in IBS, in which FIG. 2A illustrates the area under the receiver operating characteristic curve (AUROC, centre for the error bands is median) , FIG. 2B shows the feature importance summary of top 30 features from random forest classifier, and FIG. 2C demonstrates the microbial species associated with health status or different disease phenotypes;
[0022] FIGs. 3A-3B depict the prediction properties of feature numbers in the random forest model, in which FIG. 3A shows the receiver operating curves with area under the curve (AUC) for prediction performances for top 100 ranked multi-omics features; and
[0023] FIG. 4 depicts a pipeline for multi-biome data integration.Detailed Description:
[0024] In the following description, systems and / or methods of detecting metabolites and microbiomes associated with anxiety and / or depression and the likes are set forth as preferred examples. It will be apparent to those skilled in the art that modifications, including additions and / or substitutions may be made without departing from the scope and spirit of the invention. Specific details may be omitted so as not to obscure the invention; however, the disclosure is written to enable one skilled in the art to practice the teachings herein without undue experimentation.
[0025] The term “irritable bowel syndrome (IBS) ” used herein refers to a chronic functional gastrointestinal disorder characterized by recurrent abdominal pain or discomfort associated with altered bowel habits, such as diarrhea, constipation, or a combination of both, in the absence of detectable structural abnormalities. It is a multifactorial condition influenced by gut-brain axis dysregulation, gastrointestinal motility disturbances, visceral hypersensitivity, immune activation, and microbiome alterations. IBS is commonly classified into subtypes based on predominant stool patterns: IBS with constipation (IBS-C) , IBS with diarrhea (IBS-D) , mixed IBS (IBS-M) , and unclassified IBS (IBS-U) . The condition significantly impacts patients' quality of life and is often associated with psychological comorbidities such as anxiety and depression. Although the exact cause remains unclear, IBS is typically diagnosed based on symptom-based criteria, such as the Rome IV criteria, and managed through a combination of dietary modifications, pharmacological treatments, and psychological interventions.
[0026] The term “weighted similarity network fusion (WSNF) ” used herein refers to a data integration method that combines multiple similarity networks-each representing pairwise relationships between samples derived from different data types-into a single unified network. Unlike traditional SNF, which treats all input networks equally, WSNF assigns a specific weight to each network, reflecting its relative importance, reliability, or biological relevance. This weighting allows the fusion process to prioritize more informative or higher-quality data sources, producing a composite network that more accurately captures the underlying structure or patterns across heterogeneous datasets. WSNF is commonly used in fields such as multi-omics integration, patient stratification, and biomarker discovery.
[0027] In accordance with a first aspect of the present invention, a method for detecting metabolites and microbiomes associated with anxiety and / or depression in patients diagnosed with IBS is provided.
[0028] The method begins with the collection of biological specimens from the patient, specifically a serum sample and a fecal sample. These two types of biological material allow for the comprehensive assessment of both systemic metabolic profiles and localized gut microbial communities.
[0029] From the fecal sample, a gut metagenome is extracted. The metagenomic data encompasses three distinct subcomponents of the microbiome: the bacteriome, the mycobiome, and the virome. These collectively represent the bacterial, fungal, and viral populations residing within the patient’s gastrointestinal tract. The diversity and abundance of microbial species in each of these domains are captured using high-throughput sequencing technologies, allowing for high-resolution profiling of the gut microbiota.
[0030] Simultaneously, metabolic features are extracted from both the serum and the fecal sample. These features include small molecules and biochemical compounds that may be indicative of underlying physiological or pathological states. These metabolic readouts provide complementary information to the gut microbiome and are critical for identifying metabolic signatures associated with mood disorders.
[0031] Following the extraction of both the microbiome and metabolome profiles, the method employs a trained classifier to analyze these profiles. The classifier is configured to screen the patient’s gut metagenomic and metabolic features to detect the presence of a marker set of microbial species and metabolites. If these signature features are identified in the patient’s data, it is inferred that the patient may exhibit underlying or comorbid anxiety and / or depressive symptoms.
[0032] In some embodiments, the marker set of microbial species and metabolites has both microbial and metabolic markers. Metabolites in the panel include linoleic acid, triglycerides, glycolithocholic acid (GLCA) , isolithocholic acid, ω-muricholic acid (wMCA) , oleic acid, palmioleic acid, 6-keto lithocholic (6_KLCA) , palmitic acid, cholic acid (CA) , hyocholic acid, oleamide, and 2-arachidonylglycerol. Particularly, GLCA, isolithocholic acid, wMCA, 6_KLCA, CA, hyocholic acid are preferably detected in the serum sample.
[0033] On the microbial side, the marker set includes a diverse set of bacteria, viruses, and fungi, such as Escherichia coli, Enterobacteria phage cdtI, Enterobacteria phage mEp460, Shigella phage fII, Staphylococcus phage SPbeta-like, Bacteroides uniformis, Escherichia phage TL 2011b, Malassezia vespertilionis, Janthinobacterium phage vB JliM Donnerlittchen, Escherichia phage D6, Komagataella phaffii, Enterobacteria phage SfI, Firmicutes bacterium CAG: 83, Alcaligenes phage vB_AfaP_QDWS595, Microbacterium phage Pumpernickel, Shigella phage SfIV, Enterobacteria phage SfV, Nakaseomyces glabratus, Synechococcus phage S-H9-1, Ruminococcus gnavus, Ustilago maydis, Bacillus phage vB_BanS_Nate, Streptomyces phage Success, Faecalibacterium prausnitzii, Thermothelomyces thermophilus, Streptomyces phage Faust, Escherichia phage vB_EcoM_Ro157c2YLVW, Bacteroides xylanisolvens, ruthenibacterium lactatiformans, triglyceride glucose (TyG) index, Aspergillus chevalieri, Fusarium fujikuroi, Schizosaccharomyces pombe, Fusarium graminearum, Fusarium venenatum, Streptomyces phage Wakanda, Ruminococcus bromii, Fusarium falciforme, Synechococcus phage S-SCSM1, Salmonella phage SEN22, Yarrowia lipolytica, Brettanomyces bruxellensis, Fusarium oxysporum, Psilocybe cubensis, Colletotrichum lupini, Aeromonas phage LAh10, Cercospora beticola, Akanthomyces muscarius, Pyricularia grisea, Roseburia inulinivorans, Kosakonia phage Kc263, Kluyveromyces lactis, oleic acid, Intestinimonas butyriciproducens, Aspergillus fumigatus, Trichoderma breve, Pyricularia oryzae, Bacteroides intestinalis, Bacteroides thetaiotaomicron, Salmonella phage JD01, Alistipes shahii, Clostridium leptum, Alistipes putredinis, Naumovozyma castellii, Enterobacteria phage IME10, Lachancea thermotolerans, Schizosaccharomyces osmophilus, Saccharomycodes ludwigii, Ustilaginoidea virens, Candida dubliniensis, Saccharomyces kudriavzevii, Parabacteroides distasonis, Synechococcus phage S-SZBM1, Enterobacteria phage mEp237, Naumovozyma dairenensis, Purpureocillium takamizusanense, Rhizoctonia solani, Pyricularia pennisetigena, Brettanomyces nanus, Enterobacteria phage P4, 2-Arachidonylglycerol, Blautia wexlerae, Bacteroides vulgatus, Bacteroides clarus, Eubacterium sp. CAG: 38, Bacteroides nordii, Bacteroides cellulosilyticus, Marasmius oreades, and Enterobacteria phage phiP27.
[0034] The method may also include calculating the triglyceride-glucose (TyG) index using the formula: ln [Fasting TG (mg / dL) ×Fasting Glucose (mg / dL) / 2] as an additional metabolic risk marker. To evaluate whether the subject exhibits comorbid anxiety or depression symptoms, the integrated microbiome data and the metabolite features are input into a machine learning classifier trained to recognize such mental health conditions among IBS patients.
[0035] In one embodiment, the classifier includes a random forest model trained using gut multi-omics data-specifically, combined microbiome and metabolome features-and validated using a nested cross-validation procedure. The classifier is capable of performing multi-class prediction to differentiate between IBS alone, IBS with anxiety, IBS with depression, and IBS with both conditions.
[0036] Through this integrative analysis, the method enables early and precise detection of potential psychiatric comorbidities in IBS patients by leveraging rich and diverse data from the gut microbiome and host metabolic state, combined with advanced computational modeling.
[0037] In accordance with a second aspect of the present invention, a comprehensive system for detecting potential anxiety and depression symptoms in individuals diagnosed with IBS is introduced. The system leverages gut microbiome and metabolome data in combination with machine learning-based classification and includes several coordinated modules working together to process biological data, integrate multi-omic datasets for generating accurate diagnostic predictions.
[0038] At the core of the system is a sample processing module, which is configured to receive and analyze both serum and fecal samples collected from IBS patients. The sample processing module performs biological extractions and analytical assessments to generate two types of data: the gut metagenome dataset and the gut metabolome dataset. The metagenome dataset includes sub-datasets corresponding to different microbial domains-specifically, the bacteriome (bacterial data) , mycobiome (fungal data) , and virome (viral data) . In parallel, the metabolome dataset captures various chemical markers from the same subjects, including serum and fecal metabolites relevant to host-microbiome interactions and potential psychiatric symptoms.
[0039] The system further includes a fusion module, which is designed to perform integration of the bacteriome, mycobiome, and virome data using a technique known as WSNF. To accurately reflect the diversity and contribution of each microbial domain, the fusion module assigns weights to each microbiome sub-dataset based on its species richness-that is, the number of distinct species detected within the sub-dataset. Using these weights, the WSNF process generates a composite similarity network, enabling spectral clustering of patient profiles based on microbiome similarities across all domains. This multi-dimensional integration enhances the precision of downstream classification, particularly for complex phenotypes like IBS with comorbid mental health disorders.
[0040] Once the integrated data are prepared, they are passed to a classification module, which implements a trained machine learning model-specifically, a random forest classifier. This classifier has been trained on a labeled dataset containing multi-omic profiles of IBS patients previously diagnosed with or without anxiety and / or depression. Based on this training, the model is capable of recognizing complex patterns of microbial and metabolic variation that correlate with mental health comorbidities in IBS.
[0041] Upon receiving the patient’s integrated microbiome data and metabolomic features, the classification module calculates a classification score that represents the likelihood of comorbid anxiety or depression in the IBS patient. The output of this module includes not only the binary or multi-class prediction but also the model confidence score-expressed as probabilities-to indicate the reliability of the classification. This probabilistic output allows clinicians to interpret results with greater nuance and to consider borderline cases with appropriate clinical judgment.
[0042] The system includes an output interface, which is configured to present the final prediction in an accessible format. The interface displays the predicted mental health status of the subject based on the classifier’s output and may include visualization tools for confidence intervals, feature importance, and links to the associated multi-omic biomarkers contributing to the prediction.
[0043] EXAMPLES
[0044] Sample collection
[0045] Adults meeting the Rome IV criteria for irritable bowel syndrome (IBS) are prospectively recruited from two Chinese medicine clinics affiliated with the School of Chinese Medicine at Hong Kong Baptist University. Specifically, participants are eligible for inclusion if they meet the following criteria: (1) fulfillment of the Rome IV diagnostic criteria, including recurrent abdominal pain occurring, on average, at least one day per week over the past three months; (2) an IBS Symptom Severity Scale (IBS-SSS) score greater than 75 at baseline; (3) age between 18 and 65 years; (4) normal colonic evaluation within the past five years confirmed by colonoscopy or barium enema; and (5) provision of written informed consent. Participants are excluded if they meet any of the following criteria: (1) pregnancy or breastfeeding; (2) history of inflammatory bowel disease (IBD) , carbohydrate malabsorption, hormonal disorders, known allergies to food additives, or other serious medical conditions; (3) surgical history involving gallbladder removal, gastrointestinal (GI) tract, or cerebral cranium; (4) evidence of parasitic infection; or (5) current use of medications known to influence gastrointestinal function, blood pressure, or lipid metabolism.
[0046] Disease classification is based on predominant bowel habits reported on days with abnormal bowel movements, as assessed using the Bristol Stool Form Scale and defecation frequency. Age-and sex-matched healthy controls without any medical history of neurodegenerative, cardiovascular, metabolic, or gastrointestinal diseases, and without prior surgeries involving the gallbladder, GI tract, or cerebral cranium, are also recruited from the same clinical centers. All participants are instructed to provide first-morning stool samples and fasting blood samples on the same day for biochemical assessments and multi-omics analyses. The use of antibiotics, probiotics, prebiotics, or other microbiota-influencing supplements is discontinued for at least three weeks prior to stool collection. Collected serum and stool samples are transported on dry ice and stored at -80℃ until further treatments.
[0047] For instance, feces (100 mg) are completely homogenized with five-fold volume of ice-cold distilled water. After high-speed centrifugation (13,000 rpm for 15 min at 4 ℃) , water extractions are transferred to a new 2 mL tube. Subsequently, another five-fold volume (500 μL) of methanol is added into the pellet sample. The mixture is completely homogenized and centrifuged again. Methanol extractions are combined with the previous water extractions. For serum sample, serum (50 μL) is prepared with four volumes of cold methanol for protein precipitation, and metabolite extracts are obtained after vortex and centrifugation. The 200 μL of fecal or serum supernatant is dried and redissolved in the same volume of solvent consisting of water and acetonitrile (98: 2, v / v) . Meanwhile, quality control (QC) samples pooling all samples are individually prepared using the same protocol. P-chlorophenylalanine (5 μg / mL) is added as an internal standard.
[0048] Referring to FIG. 4, a method for multi-biome data integration in according to one embodiment of the present invention is depicted. Workflow shows the multi-omics data (bacteriome, mycobiome, virome and metabolome) integration approach, and machine learning. The ultimate goal of data integration is to discover disease biomarkers, confirm phenotypic spectrum.
[0049] Briefly, for the microbiome arm, shotgun metagenomic sequencing data is first subjected to rigorous quality control. This includes read filtering using FastQC and Trimmomatic to remove low-quality reads and adapter sequences. To eliminate potential contamination from host DNA, the data is further processed using Bowtie2 in conjunction with KneadData, ensuring that only microbial reads are retained for taxonomic profiling. Following preprocessing, the filtered microbial reads are analyzed across three microbial domains. The bacteriome is profiled using MetaPhlAn3, which enables high-resolution taxonomic classification. To capture the fungal components, the mycobiome is characterized using Kraken2, a comprehensive metagenomic classifier. For viral populations, the virome is identified using a combination of VirSorter2 and Diamond, which detect both known and novel viral sequences based on similarity searches. These individual domain-specific features are then fused into a comprehensive microbial feature matrix, with weighted integration applied to preserve relative domain contributions. This produces a single, unified dataset representing the integrated microbiome profile for each individual. In parallel, metabolomic data is acquired using a high-throughput analytical platform, likely liquid chromatography-mass spectrometry (LC-MS) . The raw spectral data undergoes feature extraction, including peak detection, alignment, and normalization, resulting in a metabolite abundance matrix. Subsequent feature selection identifies key metabolites relevant to disease classification or subtype differentiation. Together, the microbial and metabolomic datasets are prepared for multi-modal data integration. These fused and preprocessed features form the input to the data analysis and machine learning pipeline for classification and biomarker discovery.
[0050] A total of 173 individuals with IBS and 84 healthy controls are recruited for metagenomic sequencing analysis. For each participant, 24 intrinsic physiological and clinical parameters and 45 questionnaire-based lifestyle and behavioral factors are collected. The 24 intrinsic factors include: gender, age, body mass index (BMI) , IBS subtype (IBS-D, IBS-C, IBS-M, or IBS-U) , serum total bile acids (TBA, μmol / g) , fecal TBA (μmol / g) , alkaline phosphatase (ALP, U / L) , alanine transaminase (ALT, U / L) , aspartate transaminase (AST, U / L) , urea (mmol / L) , creatinine (μmol / L) , total cholesterol (TC, mmol / L) , fasting glucose (mmol / L) , triglycerides (TG, mmol / L) , serum 7α-hydroxy-4-cholesten-3-one (C4, ng / mL) , fibroblast growth factor 19 (FGF19, pg / mL) , stool frequency (per day) , Bristol stool score, Hamilton depression rating scale (HAMD) , Zung self-rating anxiety scale (SAS) , Zung self-rating depression scale (SDS) , and three components of the IBS symptom severity scale (IBS-SSS) : pain, distention, and total score.
[0051] The 45 questionnaire-derived factors encompass: (1) basic information (education level, marriage and monthly income) ; (2) dietary habits (rice products quantity and frequency, flour products quantity and frequency, coarse cereals quantity and frequency, viscera blood products quantity and frequency, red meat quantity and frequency, white meat quantity and frequency, egg products quantity and frequency, vegetables quantity and frequency, tuber products quantity and frequency, bean products quantity and frequency, fruit quantity and frequency, vegetable oil quantity and frequency, animal oil quantity and frequency) ; (3) beverage consumption (tea quantity and frequency, coffee quantity and frequency, alcohols quantity and frequency, milk products quantity and frequency) ; (4) physical exercise habits (physical activity quantity, exercise quantity and frequency) ; (5) sleeping status (sleep quality and sleep time) ; (6) smoking (smoke duration) ; (7) stress status (spirit quantity and frequency) . IBS patients with comorbid depression and anxiety are identified by experienced physicians based on self-reported SDS scores (≥53 indicating depression) and SAS scores (≥50 indicating anxiety) . This is approved by the Ethics Committee on the Use of Human & Animal Subjects in Teaching & Research (Approval no. HASC / 15-16 / 0300 and HASC / 16-17 / 0027) . Written informed consent is obtained from each participant prior to sample collection.
[0052] Example 1. Bioinformatics evaluation
[0053] Samples are subjected to both metagenomic and metabolomic analysis. To extract microbial DNA from a wide range of organisms-including bacteria, fungi, and viruses-samples may undergo mechanical and chemical lysis. For consistency and high yield, commercial kits such as the QIAamp PowerFecal DNA Kit are commonly used. For virome-specific analysis, additional steps such as filtration (e.g., using 0.22 μm filters) and DNase treatment may be employed to enrich for viral particles prior to lysis.
[0054] Extracted DNA is quantified using fluorometric assays (e.g., Qubit) and spectrophotometry (e.g., NanoDrop) , and its integrity is confirmed via agarose gel electrophoresis. Sequencing libraries are prepared through enzymatic fragmentation, adapter ligation, and PCR amplification. Shotgun metagenomic sequencing is then performed using Illumina platforms (e.g., NovaSeq) to obtain high-quality paired-end reads. It is worth noting that viral sequence identification poses a challenge in metagenomic analysis due to the absence of a universal viral marker-unlike bacterial 16S rRNA. As a result, reference-based read mapping is hindered by the limited number of annotated viral genomes. To address this, an optimized pipeline capable of de novo viral contig assembly and retrieval from shotgun reads is employed.
[0055] Raw sequence quality is assessed using FASTQC and filtered utilizing Trimmomatic using the following parameters; SLIDINGWINDOW: 4: 20, MINLEN: 60 HEADCROP 15; CROP 225. Contaminating human reads are filtering using Kneaddata (Reference database: GRCh38 p12) with default parameters. Megahit, with default parameters, is chosen to assemble the reads into contigs per sample. Assemblies are subsequently pooled and retained if longer than 1 kb. Bacterial contamination is removed by using an extensive set of inclusion criteria to select viral sequences only.
[0056] Briefly, contigs are required to fulfill one of the following criteria; 1) Categories 1–6 from VirSorter when run with default parameters and Refseqdb (–db) (1) positive, (2) circular, (3) greater than 3 kb with no BLASTn alignments to the NT database (January ‘19) (e-value threshold: 1e-10) , (4) a minimum of 2 pVogs with at least 3 per 1 kb, (5) BLASTn alignments to viral RefSeq database (v. 89) (e-value threshold: 1e-10) , and (6) less than three ribosomal proteins as predicted using the COG database.
[0057] HMMscan is used to search the pVOGs hmm profile database using predicted protein sequences on VLS with an e-value filter of 1e-5, retaining the top hit in each case. Afterward, a fasta file combining viral contigs is compiled. The redundant sequences are eliminated by CD-HIT-EST provided from CD-HIT 4.8.1. This viral database includes the viral contigs recovered by the screening criteria from the bulk metagenomic assemblies. Then the paired reads are mapped to the viral contig database with BWA, using default parameters. The viral operational taxonomic unit (OTU) table of viral abundance is pulled from BWA sam output files by script, and normalized by the number of metagenomic reads and the OTU sequence length. The contigs are analyzed according to their open reading frames (ORFs) . The ORFs on the contigs are predicted using MetaProdigal v2.6.3 with the metagenomics procedure (-p meta) . To annotate the predicted ORFs, the amino acid sequences of the ORFs are queried by Diamond against the viral RefSeq protein (v84) with an E-value <10-5 and a bitscore >50. The viral Refseq proteins with the top closest homologies (E-value <10-5 and bitscore >50) are considered for each ORF, analogous to a previously reported method.
[0058] Example 2. Integration and clustering analysis of multi-biome data
[0059] The patients are stratified to provide evidence for a clear link between gut microbiota and IBS. It is observed that the integrative microbiome can capture microbiome interactions and allow better stratification of disease, which cannot be appreciated by the study of a single microbial group. Thus, it is further explored whether integrative multi-kingdom microbiome of IBS provide a novel framework for understanding IBS and identification of treatable traits.
[0060] For each biome dataset, microbes prevalent in at least 5%of patients (that is, n ≥ 7) with an average abundance of 1%are kept for analysis. Integration of bacterial, fungal, and viral community data is performed by weighted SNF (WSNF) using online software. In short, when a biome contains more taxa (e.g., bacteria >viruses) , it is more likely to influence the multi-biome as a whole-a factor overlooked in standard (unweighted) SNF. Thus, bacterial and viral profiles should not be treated equally (as in conventional unweighted SNF) , since the bacterial community’s higher taxonomic richness provides more information. Consequently, weighting each biome by its taxonomic richness during data integration is justified. Briefly, the respective weights of each biome are assigned based on the richness of the data, as demonstrated by the number of species present in each biome. Then, a Bray-Curtis similarity matrix is created for each biome dataset using vegan package, which is subsequently integrated using WSNF analysis pipeline. The optimal number of clusters (n = 2) is determined by WSNF using the eigengap method and the value of K nearest neighbors, which is set based on the optimal silhouette width.
[0061] Therefore, bacterial, fungal, and viral microbiome profiles are generated for all of the patients to assess to more holistic microbiome in each individual. Compositional profiles of each single microbial group are integrated by WSNF approach. Spectral clustering of the resultant similarity matrix identifies two patient clusters, including C1 cluster (IBS cluster) and C2 cluster (non-IBS cluster) (FIG 1A and FIG. 1B) . Principal coordinate analysis demarcates two distinct subgroups (FIG. 1C, PERMANOVA, p value < 0.05) . Bray-Curtis principal coordinates analysis of R. gnavus is also conducted (FIG. 1D) The differential compositional taxa associations are performed using microbiome multivariable association with linear models (MaAsLin2) with adjustment for multiple comparisons. Notably, Ruminococcus gnavus, Escherichia coli and Fusobacterium varium are more abundant in C2 cluster than C1 cluster (P < . 05) (FIG. 1E and FIG. 1F) .
[0062] Example 3. Multi-class machine learning for detecting IBS syndrome encompassing depression and anxiety
[0063] Multi-class model is implemented by Python 3.6.7 using standard libraries that are publicly available: pandas (0.23.4) , numpy (1.14.5) , scikit-learn (1.1) , and matplotlib (2.2.3) . For each subgroup, samples are randomly divided into a training set (70%of samples) and a test set for independent evaluation (remaining 30%) . Random forests (RF) are used as classifier models for the detection of anxiety and depression by using multi-omics profiles. The RF multi-class classifier with the following modifications is implemented to the default SciKit-learn settings: n_estimaters = 2000 and class_weight = balanced. Hyperparameters are tuned using a grid search for balanced accuracy, which end up selecting n_estimaters as 2000. Other settings are kept as default. A nested cross-validation procedure is applied to calculate within-training set accuracy by splitting data into training and test sets for 20-times repeated, fivefold-stratified cross-validation (balancing class proportions across folds) . The mean AUROC and AUPR value are calculated accordingly for the visualization of results. The highly ranked and frequently selected multi-omics features are considered predictive signatures for further interpretation. The optimal models selected based on cross-validated results are evaluated in the withheld evaluation dataset. This process repeats 10 times to obtain a distribution of random forest prediction evaluations on the validation set, and the mean AUROC and AUPR value are calculated accordingly for the visualization of results.
[0064] The AUROC is included to characterize the model performance as the models initially provided outputs of probabilities for each disease phenotype, and these predicted probabilities are then used to estimate the risk of disease occurrence or absence, which forms a binary status that is analyzed to provide an AUROC value. The AUROC is a widely applied metric that considers the trade-offs between sensitivity and specificity at all possible thresholds for comparing the performance across various classifiers with a baseline value of 0.5 for a random classifier. AUPR is provided as a complimentary assessment, which considers the trade-offs between precision (or positive predictive value) and recall (or sensitivity) with a baseline that equals the proportion of positive disease cases in all samples.
[0065] Therefore, using gut multi-kingdom microbiome and metabolites data, random forest machine learning model is trained to differentiate IBS with mental disorder from normal IBS. Machine learning model are trained on the training set (70%, n=180) with fivefold cross-validation, and then are applied to the test set (30%, n = 77) for validation. The random forest model achieved a mean AUROC of 0.74-0.83 (FIG. 2A) , suggesting that multi-class disease classification based on the faecal microbiome is feasible. Applying the variable selection strategy based on random forest, important features are identified based one the importance score. Particularly, important features of random forests in C2 cluster are distinct with health control (FIG. 2C) . The top 30 multi-omic features contributing to the random forest multi-class classifier are clustered using hierarchical clustering (FIG. 2B) . Associations are coloured by direction of effect (red, positive; purple, negative; p < 0.05) , with associations significant at FDR < 0.05 marked with a plus (positive correlations) or minus (negative correlations) , respectively. The nominal significance (p-value) of associations is calculated by MaAsLin2, and the false discovery rate (FDR) is computed by Benjamini–Hochberg correction.
[0066] To further test whether a simplified panel of biomarkers could be used as biomarker to distinguish the mental disorder in IBS according to the multi-omics data, predictions are performed using top ranked features contributing to the random forest classifier (FIG. 3A) . It is tested how many of the features representatives of the mental disorder in IBS are necessary to achieve the comparable predictive performance by training the classifier model with different number of top-ranking features that are chosen based on the mean decrease in GINI from the classifier trained with the full set of features. The results show that using as few as 100 features (FIG 3B, Table 1) achieves best average AUC 0.82 for all groups.
[0067] Table 1. Top 100 ranked features from the random forest classifier
[0068] The foregoing description of the present invention has been provided for the purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations will be apparent to the practitioner skilled in the art.
[0069] The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, thereby enabling others skilled in the art to understand the invention for various embodiments and with various modifications that are suited to the particular use contemplated.
Claims
1.A method for detecting anxiety and / or depression metabolic and microbiome markers in a patient with irritable bowel syndrome (IBS) , comprising:obtaining a serum sample and a fecal sample from the patient;extracting a gut metagenome from the fecal sample including a bacteriome, a mycobiome, and a virome of the patient;extracting metabolic features from the serum sample and fecal sample; anddetecting whether a marker set of microbial species and metabolites for depression and anxiety is present in the gut metagenome and the metabolic features by utilizing a classifier to screen the gut metagenome and the metabolic features and locate the presence of the marker set of microbial species and metabolites.2.The method of claim 1, wherein the marker set of microbial species and metabolites for depression and anxiety comprises one or more metabolite selected from the group consisting of linoleic acid, triglycerides, glycolithocholic acid (GLCA) , isolithocholic acid, ω-muricholic acid (wMCA) , oleic acid, palmioleic acid, 6-keto lithocholic (6_KLCA) , palmitic acid, cholic acid (CA) , hyocholic acid, oleamide, and 2-arachidonylglycerol.3.The method of claim 1, wherein the marker set of microbial species and metabolites for depression and anxiety comprises one or more microbial species selected from bacteria, viruses, or fungi, including but not limited to: Escherichia coli, Enterobacteria phage cdtI, Enterobacteria phage mEp460, Shigella phage fII, Staphylococcus phage SPbeta-like, Bacteroides uniformis, Escherichia phage TL 2011b, Malassezia vespertilionis, Janthinobacterium phage vB JliM Donnerlittchen, Escherichia phage D6, Komagataella phaffii, Enterobacteria phage SfI, Firmicutes bacterium CAG: 83, Alcaligenes phage vB_AfaP_QDWS595, Microbacterium phage Pumpernickel, Shigella phage SfIV, Enterobacteria phage SfV, Nakaseomyces glabratus, Synechococcus phage S-H9-1, Ruminococcus gnavus, Ustilago maydis, Bacillus phage vB_BanS_Nate, Streptomyces phage Success, Faecalibacterium prausnitzii, Thermothelomyces thermophilus, Streptomyces phage Faust, Escherichia phage vB_EcoM_Ro157c2YLVW, Bacteroides xylanisolvens, ruthenibacterium lactatiformans, triglyceride glucose (TyG) index, Aspergillus chevalieri, Fusarium fujikuroi, Schizosaccharomyces pombe, Fusarium graminearum, Fusarium venenatum, Streptomyces phage Wakanda, Ruminococcus bromii, Fusarium falciforme, Synechococcus phage S-SCSM1, Salmonella phage SEN22, Yarrowia lipolytica, Brettanomyces bruxellensis, Fusarium oxysporum, Psilocybe cubensis, Colletotrichum lupini, Aeromonas phage LAh10, Cercospora beticola, Akanthomyces muscarius, Pyricularia grisea, Roseburia inulinivorans, Kosakonia phage Kc263, Kluyveromyces lactis, oleic acid, Intestinimonas butyriciproducens, Aspergillus fumigatus, Trichoderma breve, Pyricularia oryzae, Bacteroides intestinalis, Bacteroides thetaiotaomicron, Salmonella phage JD01, Alistipes shahii, Clostridium leptum, Alistipes putredinis, Naumovozyma castellii, Enterobacteria phage IME10, Lachancea thermotolerans, Schizosaccharomyces osmophilus, Saccharomycodes ludwigii, Ustilaginoidea virens, Candida dubliniensis, Saccharomyces kudriavzevii, Parabacteroides distasonis, Synechococcus phage S-SZBM1, Enterobacteria phage mEp237, Naumovozyma dairenensis, Purpureocillium takamizusanense, Rhizoctonia solani, Pyricularia pennisetigena, Brettanomyces nanus, Enterobacteria phage P4, 2-Arachidonylglycerol, Blautia wexlerae, Bacteroides vulgatus, Bacteroides clarus, Eubacterium sp. CAG: 38, Bacteroides nordii, Bacteroides cellulosilyticus, Marasmius oreades, and Enterobacteria phage phiP27.4.The method of claim 2, wherein the GLCA, isolithocholic acid, wMCA, 6_KLCA, CA, hyocholic acid are detected in the serum sample.5.The method of claim 1, wherein the gut metagenome is extracted using weighted similarity network fusion (WSNF) , wherein the weighting is based on the number of species in each bacteriome, mycobiome, and virome.6.The method of claim 1, further comprising calculating a triglyceride glucose (TyG) index using the formula: ln [Fasting TG (mg / dL) ×Fasting Glucose (mg / dL) / 2] .7.The method of claim 1, wherein the classifier is a machine learning classifier trained to detect multi-class disease phenotypes, including IBS with anxiety and IBS with depression.8.The method of claim 7, wherein the classifier comprises a random forest model trained using gut multi-omics data and validated using nested cross-validation.9.A system for detecting anxiety and / or depression metabolites and microbiomes in IBS patients, comprising:a sample processing module configured to analyze serum and fecal samples to generate gut metagenome and metabolome data;a fusion module configured to integrate bacteriome, mycobiome, and virome data using weighted similarity network fusion;a classification module implementing a trained machine learning model configured to classify IBS patients with or without comorbid anxiety or depression based on a comparison with a marker set of microbial species and metabolites for depression and anxiety; andan output interface configured to display the predicted mental health status of the subject based on the classifier output.10.The system of claim 9, wherein the fusion module is configured to:assign weights to each microbiome sub-dataset based on species richness;integrate bacteriome, mycobiome, and virome data using WSNF; andgenerate a composite similarity network for spectral clustering of patient profiles.11.The system of claim 9, wherein the classification module is further configured to:implement a random forest model trained on a labeled dataset comprising multi-omics features of IBS patients with and without comorbid anxiety or depression;compute a classification score representing the likelihood of mental disorder; andoutput model confidence scores or probabilities associated with the prediction.
Citation Information
Patent Citations
Irritable bowel syndrome related flora marker and kit thereof
CN110838365A
Methods of diagnosing disease
US20220128556A1
Methods of diagnosing irritable bowel syndrome
WO2022058552A1
Systems and methods for identifying microbial signatures
WO2023278352A1