USCG-based intestinal flora probiotic identification method

By constructing a USCG reference database and reconstructing MACG, combined with phylogenetic tree analysis, the problem of low resolution in fecal microbiota transplantation donor screening was solved, and efficient screening of low-abundance strains in the gut microbiota and prediction of treatment effects were achieved.

CN121054091APending Publication Date: 2025-12-02SUZHOU MICRODIMENSIONAL BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511137843.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing technologies for fecal microbiota transplantation (FMT) donor screening suffer from low resolution, high cost, and difficulty in accurately screening for beneficial strains, especially in intraspecific diversity research where there is a lack of efficient and low-cost methods.

Method used

A probiotic identification method based on highly conserved single-copy genes (USCG) was adopted. A USCG reference database was constructed by intestinal metagenomic sequencing, MACG was reconstructed and a phylogenetic tree was built. Association analysis was performed in combination with clinical outcome data to screen out candidate probiotics.

Benefits of technology

It enables efficient and low-cost strain-level analysis in the gut microbiota, accurately screens low-abundance strains, improves the resolution and accuracy of intraspecific sequence reconstruction, and is suitable for accurate screening of fecal microbiota transplantation donors and prediction of treatment effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121054091A_ABST
    Figure CN121054091A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of microbiology and bioinformatics, and particularly relates to a USCG-based intestinal flora probiotic recognition method, which is characterized in that a ucgMLST technology is utilized, metagenome data is analyzed to realize strain-level resolution, and 39 candidate probiotics are screened out in combination with species retention rate and clinical outcome data; the method can realize accurate detection of low-abundance species, and is especially suitable for research and scenes such as fecal fungi transplantation (FMT) donor screening, receptor treatment effect prediction, probiotic development and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of microbiology and bioinformatics, specifically relating to a method for identifying probiotics in the gut microbiota based on ultra-conserved single-copy genes (USCG). Background Technology

[0002] The gut microbiota plays a vital role in human health, closely related to metabolism, immunity, and diseases such as ulcerative colitis, obesity, and cancer. Fecal microbiota transplantation (FMT), as a potential therapeutic approach for modulating the gut microbiota, has received widespread attention in recent years and has been approved for the treatment of recurrent Clostridium difficile infection (rCDI). Although FMT has shown potential in adjuvant therapy for other diseases, its clinical application has been delayed due to issues such as inconsistent donor selection criteria. Currently, FMT donor selection only excludes specific pathogenic microorganisms (such as Clostridium difficile) or multidrug-resistant bacteria (such as methicillin-resistant Staphylococcus aureus (MRSA), using donor microbial diversity as a positive evaluation criterion. The resolution is only down to the species level, but intraspecific diversity may be a significant factor contributing to variations in FMT efficacy. Therefore, estimating the intraspecific diversity of microorganisms in FMT donor samples, reconstructing multi-strain sequences, and further screening for potentially beneficial strains while excluding potentially harmful strains have significant clinical application value and safety implications.

[0003] As research depth and technological methods advance, methods for studying intraspecific diversity are constantly evolving. However, existing methods have limitations in terms of applicability and resolution. Specifically: community fingerprinting methods, such as denaturing gradient gel electrophoresis (DGGE), restriction fragment length polymorphism (TRFLP), and automated ribosomal spacer analysis (ARISA), can be used for culture-free intraspecific diversity studies, but their low throughput and low resolution limitations make them unsuitable for the current research environment involving large datasets and high throughput. 16S rRNA amplicon sequencing-based methods, such as oligotyping and amplicon sequence diversity (ASV), have been applied in intraspecific diversity studies, but their limited detection site coverage restricts their resolution in intraspecific analyses. Metagenomic sequencing provides more information than amplicon sequencing, such as conserved genes or whole-genome sequences. Based on this, many methods have been developed for intraspecific diversity research, including marker-based methods. Metagenomic sequencing offers various methods for analyzing diversity in gene or single-species reference genome sequences, comparing similarity among pedigree-level reference genomes, sequence typing, and comparing gene components. However, these methods rely on limited reference sequences or metagenomically assembled sequences, limiting the potential of metagenomic sequencing to uncover vast amounts of unknown species diversity. Furthermore, compared to single-genome assembly, metagenomic assembly suffers from higher fragmentation and lower integrity, making it less comprehensive for intraspecific analyses. Newer technologies, such as species-specific enrichment before sequencing and the generation of single-amplified genomes (SAGs) after single-cell sequencing, can compensate for these limitations. Additionally, combining multi-omics approaches, such as transcriptomics, proteomics, and metabolomics, with metagenomics allows for analysis by integrating phenotypic and sequence differences. However, under current technological conditions, these approaches still face significant cost disadvantages compared to metagenomics. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings and limitations of the prior art and provide a method for applying metagenomic sequences in the field of intraspecific diversity research, enabling efficient, low-cost, and accurate strain-level analysis and tracking in complex gut microbiota communities (such as fecal samples), especially for tracking low-abundance strains.

[0005] To achieve the above objectives, embodiments of this application provide a method for identifying probiotics in gut microbiota based on USCG, characterized by the following steps: Step S1: Perform metagenomic sequencing of gut samples to obtain a metagenomic sequencing dataset composed of sequencing reads; Step S2: Construct a USCG reference database, which includes USCG and corresponding species annotation information; Step S3: Construct a strain-level metagenomic analysis workflow ucgMLST, including comparing the sequencing reads with the USCG database to obtain a USCG sequence profile, and reconstructing the core genotype (MACG) and phylogenetic tree of metagenomic assembly based on the USCG sequence profile; Step S4: Calculate the species survival rate based on the phylogenetic tree; Step S5: Perform correlation analysis by combining the species survival rate with clinical outcome data; Step S6: Screen candidate probiotics.

[0006] In some embodiments, the metagenomic sequencing of the intestinal sample in step S1 includes: step S11, extracting DNA from a fecal sample, such as an FMT donor and recipient; step S12, performing metagenomic sequencing on the DNA using a high-throughput sequencer to generate a predetermined number of paired end reads; and step S13, performing quality control to remove low-quality sequences and host-derived reads from the paired end reads to obtain the sequencing reads as metagenomic sequencing data in the metagenomic sequencing dataset.

[0007] In some embodiments, step S2, constructing the USCG database, includes: Step S21, in 556,157 genomes covering bacteria, archaea, and eukaryotes, all genomes are subdivided into 116,687 single-stranded clusters with a genetic distance less than a first distance, and a representative set of genomes is randomly selected; the first distance is 0.02; Step S22, for the sequences in the above-mentioned representative genome set, a hidden Markov model is used to search for sequences of all potential coding frames and generate a hidden Markov model spectrum of this genome set, which is compared with a set of homologous genes that are commonly found in the tree of life, including bacteria, identified from the BUSCO database, i.e., a set of homologous genes is screened from the 556,157 genomes; the set of homologous genes is screened, and the final screened result is defined as the USCG; finally, 480 genes are screened; Step S23, a reference genome is extracted from the NCBI RefSeq database to construct a USCG database containing gene sequences and species / strain annotations.

[0008] In some embodiments, step S3, constructing the strain-level metagenomic analysis workflow ucgMLST processing specifically includes: S31, a spectrum processing step for performing species classification analysis and reconstructing MACG based on the USCG reference database; S32, a phylogenetic tree processing step for integrating genomic data and performing USCG phylogenetic analysis.

[0009] In some embodiments, the spectrum processing performs the following operations: S311, aligning each metagenomic sequencing data to a USCG reference database constructed by the database processing, iteratively evaluating the presence of USCGs for each representative genome in the metagenomic data, and summarizing the alignment scores of all reads aligned with each representative genome in each iteration; S312, sorting all representative genomes according to the result scores, selecting the representative genome with the highest score, and deleting the reads that match it in the next iteration; S313, terminating the spectrum processing step when the number of reads in the metagenomic data aligned with USCGs in the database is less than a preset value; S314, the spectrum processing predicts the concordant sequences of all reads aligned with each selected representative USCG sequence, these concordant sequences are called the metagenome-assembled core genotype. Genotypes (MACG) are USCG sequences reconstructed from metagenomic data. To minimize the impact of random sequencing errors, when the read coverage of each site on the MACG is ≥3 and the base type consistency is ≥80%, an uppercase base character (A, C, G, or T) is assigned to that site; otherwise, a lowercase base character is assigned to indicate that it is a low-quality base.

[0010] In some embodiments, the phylogenetic tree processing loads the MACG generated in the spectrum processing step and compares it with all representative genomes in the database; for this purpose, the reconstructed MACG and the USCG sequences of all representative genomes in the database are aligned with selected reference sequences; under default parameters, only MACGs covering more than 20% of the reference USCG sequence are retained, and MACGs below this threshold are filtered out; the maximum likelihood phylogenetic tree of tandem USCG sequences is calculated to resolve the phylogenetic relationships between USCGs.

[0011] In some embodiments, low-quality sequences are ignored by default during the analysis process, or the "--risky" parameter is set in the phylogenetic tree processing to include low-quality MACG sequences.

[0012] In some embodiments, step S4, calculating the species retention rate includes: step S41, calculating the genetic distance between any two strains on the phylogenetic tree, for example, when the genetic similarity of strains reaches 98% or more, using 0.001 as the threshold, and the genetic distance exceeding the threshold is considered as different strains; step S42, calculating the species retention rate: the species retention rate within a certain time period = the number of identical strains / (the number of identical strains + the number of different strains), and classifying species with a retention rate greater than a preset value as high retention based on the species retention rate.

[0013] In some embodiments, step S5, performing association analysis in conjunction with clinical outcome data, includes: step S51, collecting clinical outcome data; step S52, using statistical methods to analyze the association between strain clusters and the clinical outcome data, and identifying health-related strains and disease-related strains.

[0014] In some embodiments, the screening of probiotics in step S6 includes screening probiotics by comprehensively considering retention rate factors and health / disease-related phenotypes. The conditions for screening probiotics are: the species to which the strain belongs has a high retention rate; and the strain is a health-related strain within the species.

[0015] This application provides a method for identifying probiotics in gut microbiota based on highly conserved housekeeping genes. The method includes: gut metagenomic detection; constructing a USCG reference database, including extracting a reference genome and annotations; quality-controlled sequence alignment with the USCG database; parsing and filtering the alignment results; species annotation of reads; reconstructing the MACG based on the alignment results; constructing a phylogenetic tree based on the reconstructed MACG; calculating species retention rate; performing association analysis in conjunction with clinical outcome data; and screening out the candidate probiotics.

[0016] The method proposed in this invention can effectively reconstruct the biological genetic sequences in samples. Furthermore, by using a reference database constructed by combining NCBI RefSeq (NCBI Reference Sequence Database) and USCG, the specificity of the results is ensured, thereby overcoming the excessive negative and false positive results of species annotation and improving the comprehensiveness, resolution, and accuracy of intraspecific sequence reconstruction. In some embodiments, downstream analysis further incorporates clinical outcome data and intraspecific microbial heterogeneity to screen for probiotic candidates related to treatment efficacy. In some embodiments, accurate screening of fecal microbiota transplantation donors, prediction of treatment efficacy, and development of probiotics are achieved. Attached Figure Description

[0017] Figure 1 The present invention provides a flowchart and process diagram of ucgMLST according to an embodiment of the present application.

[0018] Figure 2This diagram illustrates the superior species identification capabilities of ucgMLST compared to other methods (accurately identifying 29 streptococcal species with no false positives).

[0019] Figure 3 This is a schematic diagram illustrating the superior performance of ucgMLST at ultra-low sequencing depth compared to other methods.

[0020] Figure 4 This is a schematic diagram illustrating the high correlation of pairwise genetic distances based on USCG and core genome calculations.

[0021] Figure 5 A schematic diagram illustrating how ucgMLST enables accurate tracking of clinical samples.

[0022] Figure 6 This diagram illustrates the classification of species retention across multiple cohorts. High-retention bacteria (retention rate ≥ 66%, right side), medium-retention bacteria (retention rate 34%-66%, middle), or low-retention bacteria (retention rate ≤ 34%, left side) are shown.

[0023] Figure 7 This is a schematic diagram illustrating the correlation between intraspecific heterogeneity of Enterococcus faecalis and disease.

[0024] Figure 8 A schematic diagram illustrating potential probiotic species.

[0025] Figures 9A to 9E This is a flowchart of a USCG-based gut microbiota probiotic identification method according to an embodiment of this application. Detailed Implementation

[0026] The specific embodiments of this application will now be described in detail with reference to the accompanying drawings. A complete description of each step included in the method of this invention.

[0027] like Figure 1 and Figure 9A As shown, the main process of the USCG-based gut microbiota probiotic identification method according to an embodiment of the present invention may include the following steps:

[0028] Step S1: Perform metagenomic sequencing on the intestinal sample to obtain a metagenomic sequencing dataset consisting of sequencing reads;

[0029] Step S2: Construct the USCG reference database;

[0030] Step S3: Construct a strain-level metagenomic analysis workflow, namely Universal Core Genome Multilocus Sequence Typing (ucgMLST), to align the sequencing reads with the USCG reference database and reconstruct the MACG and phylogenetic tree;

[0031] Step S4: Calculate the species retention rate;

[0032] Step S5: Perform correlation analysis by combining clinical outcome data;

[0033] Step S6: Screen out the probiotics.

[0034] The specific steps for each major step are as follows:

[0035] Step S1, metagenomic sequencing of the intestinal sample, yields a metagenomic sequencing dataset composed of sequencing reads; this may specifically include:

[0036] Step S11: Extract DNA from fecal samples (such as FMT donors and recipients);

[0037] Step S12: Use a high-throughput sequencer (such as Illumina NovaSeq 6000) to perform metagenomic sequencing and generate paired end reads of 150 bp in length;

[0038] Step S13, optionally, involves quality control to remove low-quality sequences and host-derived reads to obtain metagenomic sequencing data, thereby forming a metagenomic sequencing dataset.

[0039] Step S2, constructing a USCG reference database, which includes USCGs and corresponding species annotation information; specifically, it may include:

[0040] Step S21: Among the 556,157 genomes covering bacteria, archaea, and eukaryotes, all genomes are subdivided into 116,687 single-stranded clusters with a genetic distance less than a first distance, for example, 0.02. A representative set of genomes is selected by randomly selecting a sequence from each single-stranded cluster.

[0041] Step S22: For the sequences in the representative genome set, a Hidden Markov Model (HMM) is used to search for all potential coding frames and compare them with a set of 573 homologous genes commonly found in the Tree of Life, including bacteria, identified from the BUSCO database. This establishes an HMM profile derived from the representative genome set, i.e., a set of homologous genes is selected from 556,157 genomes. Further screening of this set of homologous genes is performed, with the following criteria: these genes are single copies in at least 70% of the genomes of at least one of the kingdoms Bacteria, Archaea, Fungi, Protozoa, and Invertebrates. Finally, a certain number of genes, such as 480 genes, are selected and defined as highly conserved single-copy genes (USCGs). For example, these 480 genes are defined as highly conserved single-copy genes (USCGs) and used as data in the USCG reference database.

[0042] Step S23: Extract the reference genome from the NCBI RefSeq database and construct the USCG reference database containing gene sequences and species / strain annotations.

[0043] Step S3: Construct the strain-level metagenomic analysis workflow ucgMLST, such as... Figure 1 and Figure 9B As shown, it can specifically include two processing steps:

[0044] Step S31, spectral processing steps to reconstruct MACG; and

[0045] Step S32, phylogenetic tree processing step based on MACG.

[0046] Specifically, in the spectral processing step S31, as follows: Figure 9C As shown, it may include:

[0047] Step S311: Apply the minimap2 program (minimap2 program source: https: / / github.com / lh3 / minimap2, v2.27) to align each metagenomic sequencing data in the metagenomic sequencing dataset to the database to establish the USCG reference database constructed in step S2;

[0048] Step S312: Perform segment annotation based on the species information in the USCG reference database;

[0049] Step S313: Iteratively evaluate the presence of USCG in metagenomic sequencing data for each representative genome;

[0050] like Figure 9D As shown, step S313 may further include

[0051] Step S3131: In each iteration, summarize the alignment scores of all reads aligned with each representative genome.

[0052] Step S3132: Sort all representative genomes according to the alignment scores;

[0053] Step S3133: Select the representative genome with the highest alignment score and delete the matching reads in the next iteration. The spectrum establishment process is terminated when the number of reads in the metagenomic sequencing data aligned to USCGs in the database is less than a predetermined number, such as 3.

[0054] The spectrum processing step predicts consistent sequences for all reads aligned with each selected representative USCG sequence. These consistent sequences are called metagenome-assembled core genotypes (MACGs), which are USCG sequences reconstructed from metagenomic sequencing data. To minimize the impact of random sequencing errors, in some embodiments, when the read coverage at each site on the MACG is ≥3 and the base type consistency is ≥80%, an uppercase base character (A, C, G, or T) is assigned to that site; otherwise, a lowercase base character is assigned to indicate low-quality bases.

[0055] like Figure 9E As shown, the phylogenetic tree processing step S32 may specifically include:

[0056] Step S321: First, load the MACG generated in the spectral processing step S32;

[0057] Step S322: Compare it with all representative genomes in the USCG reference database. To do this, align the reconstructed MACG and the USCG sequences of all representative genomes in the USCG reference database with the selected reference sequence. Under default parameters, only MACGs covering more than 20% of the reference USCG sequence are retained, and MACGs below this threshold are filtered out.

[0058] Step S323: Use the IQTree program (v2.2.2.6) to calculate the maximum likelihood phylogenetic tree of the tandem USCG sequences, thereby resolving the phylogenetic relationships between the USCG sequences.

[0059] It is worth noting that, in this embodiment, low-quality sequences, i.e., sequences with lowercase base characters, can be ignored by default during the phylogenetic tree processing step to prevent overestimation of MACG-specific single nucleotide polymorphisms (SNPs). To meet the requirement of including low-quality MACG sequences, the "--risky" parameter is provided as an option in the phylogenetic tree processing step.

[0060] Step S4, calculating the species survival rate, can specifically include:

[0061] Step S41: Calculate the genetic distance between any two strains on the phylogenetic tree. A threshold of 0.001 can be used. If the genetic distance exceeds this threshold, the strains are considered to be different strains.

[0062] Step S42: Calculate the species retention rate over a certain period of time. The species retention rate over a certain period of time is calculated as: (Number of identical strains) / (Number of identical strains + Number of different strains). Based on the species retention rate classification rules, species are categorized into three types: high-retention bacteria (e.g., retention rate ≥ 66%), medium-retention bacteria (e.g., retention rate 34%-66%), and low-retention bacteria (e.g., retention rate ≤ 34%). In this embodiment, all species are divided into three categories based on their bacterial species retention rate to identify high-retention bacteria for subsequent selection.

[0063] Step S5, combining clinical outcome data for correlation analysis, can specifically include:

[0064] Step S51: Collect clinical outcome data, such as improvement in FMT receptor symptoms, changes in inflammatory markers, etc.; and

[0065] Step S52: Using statistical methods, such as Fisher's exact test (q < 0.05), analyze the association between the bacterial clusters and the clinical outcome data to identify health-related and disease-related bacterial strains.

[0066] Step S6, screening the candidate probiotics may include comprehensively considering the retention rate factor and the health / disease-related phenotype to screen candidate probiotics. The screening criteria are: 1) the species to which the strain belongs has the high retention rate, that is, the retention rate in the above embodiment is ≥66%; 2) the strain is a health-related strain within the species.

[0067] It should be understood that the above processing steps can also be implemented as program modules. Typically, the USCG reference database processing can be implemented as a database module (Base), the spectral processing as a spectral module (Profile), and the phylogenetic tree processing as a phylogenetic tree module (Phylo). The reference database establishes the USCG reference database; the spectral processing step is used for species classification analysis and reconstruction of the core genotype (MACG); and the phylogenetic tree processing step is used for phylogenetic analysis.

[0068] Experimental Example

[0069] Experimental Example 1: Comparison of ucgMLST with other methods

[0070] Experimental environment: Metagenomic sequencing data containing 1 million metagenomic reads was generated from the genomes of 29 streptococcal species to test the species discrimination capability of the metagenomic analysis process; real metagenomic sequencing data was also used to test performance at ultra-low sequencing depth (0.1× coverage).

[0071] Experimental materials:

[0072] 1) Simulated dataset: 1 million metagenomic reads from 29 streptococcal genomes;

[0073] 2) Real-world datasets: Real-world metagenomic sequencing datasets used for testing ultra-low sequencing depth;

[0074] 3) Clinical samples: used to verify the applicability of ucgMLST in tracking probiotic strain colonization; the samples are longitudinal.

[0075] Experimental procedure:

[0076] 1) Species identification capability test: 1 million simulated metagenomic reads were input into metagenomic analysis tools such as ucgMLST, Kraken2, and MetaPhlAn4, and their performance in terms of species identification accuracy and correlation with simulated relative abundance was compared.

[0077] 2) Ultra-low sequencing depth test: Using real-world metagenomic datasets, the performance of ucgMLST was evaluated at an ultra-low sequencing depth of 0.1× coverage.

[0078] 3) Comparison with whole genome: The practicality of USCG, i.e. the genome segment used by ucgMLST, with the whole genome in disease outbreak identification is assessed by constructing a phylogenetic tree and calculating pairwise genetic distances based on USCG and the core genome.

[0079] 4) Clinical validation: In longitudinal samples, ucgMLST was used to track the colonization of probiotic strains, observe their performance at multiple time points, and distinguish individual-specific colonization events.

[0080] Experimental results:

[0081] 1) Species identification capability: ucgMLST accurately identified all 29 streptococcal species without false positives. It outperformed Kraken2 and MetaPhlAn4 in both species identification accuracy and correlation with simulated relative abundance. Regarding the correlation with relative abundance, ucgMLST had a correlation coefficient R² of 0.98 with simulated relative abundance, which was superior to Kraken2 (R² = 0.83) and MetaPhlAn4 (R² = 0.90). Figure 2 As shown.

[0082] 2) Ultra-low sequencing depth performance: ucgMLST can detect results with a minimum coverage of 0.026× coverage. Under the same data quality conditions, ucgMLST significantly outperforms leading tools such as StrainPhlAn in terms of the number of strains detected. Figure 3 As shown.

[0083] 3) Comparison with the whole genome: USCG accurately reconstructs phylogenetic clusters identical to those in the whole genome. A high correlation exists between pairwise genetic distances calculated based on USCG and the core genome, with a slope of 0.61 and R0. 2 A value of 0.95 indicates that it has sufficient resolution for strain tracking. For example... Figure 4 As shown.

[0084] 4) Clinical validation: ucgMLST can reliably track the colonization of probiotic strains in longitudinal samples and distinguish individual-specific colonization events. For example... Figure 5 As shown.

[0085] Experiment Example 2: Application of ucgMLST in multi-cohort data across diseases

[0086] Experimental materials: Seven metagenomic datasets (project numbers: PRJEB38984, PRJNA643353, PRJNA741822, PRJNA345144, PRJEB36140, PRJEB23524, PRJNA672867) were downloaded from NCBI SRA (http: / / www.ncbi.nlm.nih.gov / sra) using the "fastq-dump" tool. These datasets included healthy adults (HA), a diabetes cohort (T2D), cardiovascular disease (CAD), an infant cohort (Infants), melanoma (MEL), recurrent Clostridium difficile infection (rCDI), and inflammatory bowel disease (IBS). Reference genomes were downloaded from GenBank based on the GTDB classification method, and data integrity was ensured through quality control.

[0087] Experimental procedure: As mentioned above, a phylogenetic tree of species was constructed based on ucgMLST, which includes seven public databases corresponding to seven metagenomic datasets. The retention rate was calculated and intraspecific heterogeneity was analyzed.

[0088] Experimental results:

[0089] 1. Based on multi-cohort data, all species were classified into three categories according to their survival rate: high survival rate (≥66%), medium survival rate (34%-66%), and low survival rate (≤34%). The results are as follows: Figure 6 As shown. Figure 6 In this context, HA stands for healthy adults; T2D stands for type 2 diabetes; CAD stands for coronary artery disease; Infants stands for infants; rCDl stands for recurrent Clostridium difficile infection; MEL stands for melanoma; and IBS stands for irritable bowel syndrome.

[0090] 2. Several species exhibited significant intraspecific subgroup classifications, which were significantly correlated with different host states. Taking *Faecalibacterium duncaniae* as an example, we observed significant intraspecific heterogeneity, with subgroup A associated with healthy status and subgroup B associated with disease status. The results are as follows: Figure 7 As shown.

[0091] 3. Considering both high retention rate and health relevance, 39 probiotics were selected as candidates, including *Phocaeicola vulgatus* and *Faecalibacterium duncaniae*. The results are as follows... Figure 8 As shown. Figure 8 In this context, HA stands for healthy adults; T2D stands for type 2 diabetes; CAD stands for coronary artery disease; Infants stands for infants; rCDl stands for recurrent Clostridium difficile infection; MEL stands for melanoma; and IBS stands for irritable bowel syndrome.

[0092] This application develops a method for identifying probiotics in the gut microbiota based on USCG and ucgMLST technologies, enabling accurate detection of low-abundance species. USCG achieves strain-level resolution in metagenomic data. By combining species retention rate and clinical outcome data, probiotic candidates significantly associated with treatment efficacy are screened. The efficient and low-cost analytical workflow proposed in this invention is suitable for research on complex gut microbiota.

Claims

1. A method for identifying probiotic gut microbiota based on USCG, characterized in that: Including steps Step S1 involves performing metagenomic sequencing on intestinal samples to obtain a metagenomic sequencing dataset consisting of sequencing reads; Step S2: Construct a USCG reference database, which includes USCG and corresponding species annotation information; Step S3: Construct the strain-level metagenomic analysis workflow ucgMLST, which includes aligning the sequencing reads with the USCG database to obtain the USCG sequence profile, and reconstructing the core genotype (MACG) and phylogenetic tree of the metagenomic assembly based on the USCG sequence profile; Step S4: Calculate the species retention rate based on the phylogenetic tree; Step S5: Perform a correlation analysis by combining the species survival rate with clinical outcome data; Step S6: Select the candidate probiotics.

2. The method for identifying probiotic gut microbiota based on USCG according to claim 1, characterized in that: The metagenomic sequencing of the intestinal sample in step S1 includes: Step S11: Extract DNA from fecal samples, such as FMT donors and recipients; Step S12: Use a high-throughput sequencer to perform metagenomic sequencing on the DNA to generate a predetermined number of paired end reads; Step S13: Perform quality control, remove low-quality sequences and host-origin reads from the paired end reads to obtain the sequencing reads as metagenomic sequencing data in the metagenomic sequencing dataset.

3. The method for identifying probiotic gut microbiota based on USCG according to claim 1, characterized in that: Step S2, constructing the USCG database, includes: Step S21: Among the 556,157 genomes covering bacteria, archaea, and eukaryotes, all genomes are subdivided into 116,687 single-stranded clusters with a genetic distance less than a first distance, and a representative set of genomes is randomly selected; the first distance is 0.

02. Step S22: For the sequences in the above representative genome set, use a Hidden Markov Model (HMM) to search for sequences of all potential coding frames and generate an HMM spectrum for this genome set. Compare this spectrum with a set of homologous genes that are commonly found in the Tree of Life, including bacteria, identified from the BUSCO database. That is, screen a set of homologous genes from 556,157 genomes. Screen this set of homologous genes and define the final screening result as the USCG. Finally, 480 genes are screened. Step S23: Extract the reference genome from the NCBI RefSeq database and construct the USCG database containing gene sequences and species / strain annotations.

4. The method for identifying probiotic gut microbiota based on USCG according to claim 1, characterized in that: Step S3, constructing the strain-level metagenomic analysis workflow, specifically includes the ucgMLST process: S31, the spectrum processing step is used to perform species classification analysis and reconstruct MACG based on the USCG reference database; S32, the phylogenetic tree processing step, is used to integrate genomic data and perform USCG phylogenetic analysis.

5. The method for identifying probiotic gut microbiota based on USCG according to claim 6, characterized in that: The spectral processing performs the following operations: S311, each metagenomic sequencing data is aligned to the USCG reference database constructed by the database processing, the presence of USCG for each representative genome in the metagenomic data is evaluated iteratively, and the alignment scores of all reads aligned with each representative genome are summarized in each iteration. S312: Sort all representative genomes according to the result scores, select the representative genome with the highest score, and delete the reads that match it in the next iteration; S313, when the number of reads in the metagenomic data aligned to the USCG reference database is less than a preset value, the spectrum processing step is terminated. S314, spectrum processing predicts the consistent sequences of all reads aligned with each selected representative USCG sequence. These consistent sequences are called the core genotype of metagenomic assembly (MACG), which is the USCG sequence reconstructed from metagenomic data. To minimize the impact of random sequencing errors, when the read coverage of each site on the MACG is ≥3 and the base type consistency is ≥80%, an uppercase base character (A, C, G, or T) is assigned to that site; otherwise, a lowercase base character is assigned to indicate that it is a low-quality base.

6. The method for identifying probiotic gut microbiota based on USCG according to claim 6, characterized in that: The phylogenetic tree processing loads the MACG generated in the spectrum processing step and compares it with all representative genomes in the USCG reference database. To do this, the reconstructed MACG and the USCG sequences of all representative genomes in the USCG reference database are aligned with selected reference sequences. Under default parameters, only MACGs covering more than 20% of the reference USCG sequences are retained, and MACGs below this threshold are filtered out. The maximum likelihood phylogenetic tree of tandem USCG sequences is calculated to resolve the phylogenetic relationships between USCG sequences.

7. The method for identifying probiotic gut microbiota based on USCG according to claim 7, characterized in that: The low-quality MACG sequences are ignored by default during the analysis process, or parameters are set in the phylogenetic tree processing to include low-quality MACG sequences.

8. The method for identifying probiotic gut microbiota based on USCG according to claim 1, characterized in that: Step S4, calculating the species retention rate includes: Step S41: Calculate the genetic distance between any two strains on the phylogenetic tree. For example, if the genetic similarity of strains reaches 98% or more, use 0.001 as the threshold. If the genetic distance exceeds the threshold, they are considered different strains. Step S42, calculate the species retention rate: the species retention rate within a certain period of time = the number of identical strains / (the number of identical strains + the number of different strains). Based on the species retention rate, species with a retention rate greater than the preset value are classified as high retention.

9. The method for identifying probiotic gut microbiota based on USCG according to claim 1, characterized in that: Step S5, performing correlation analysis based on clinical outcome data, includes: Step S51: Collect clinical outcome data; Step S52: Statistical methods are used to analyze the association between bacterial clusters and the clinical outcome data to identify health-related and disease-related bacterial strains.

10. The method for identifying probiotic gut microbiota based on USCG according to claim 1, characterized in that: Step S6 involves screening probiotics by comprehensively considering retention rate and health / disease-related phenotypes. The screening criteria for probiotics are: the species to which the strain belongs has a high retention rate; and... The strain is a species-specific health-related strain.