Chronic disease and multi-disease coexistence detection method driven by tongue coating microbiota
By using metagenomic sequencing and machine learning algorithms of the tongue microbiota, we can identify microbial biomarkers that are significantly associated with chronic diseases and coexistence of multiple diseases, establish a gender-specific risk assessment model, overcome the shortcomings of non-invasive and precise screening in existing technologies, and achieve efficient and personalized health risk assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-04
- Publication Date
- 2026-04-03
AI Technical Summary
Existing methods lack non-invasive, precise early screening technologies with species-level resolution for chronic diseases and coexisting conditions, and fail to effectively consider the moderating effects of demographic factors such as gender, thus limiting the application of personalized health screening and precision intervention.
By collecting tongue coating samples, obtaining microbial DNA using metagenomic sequencing technology, and combining statistical tests and machine learning algorithms to screen for significantly relevant microbial biomarkers, a gender-specific risk assessment model was established, and a disease risk classification and early warning system based on microbial characteristics and population heterogeneity was constructed.
It enables non-invasive and highly accurate risk assessment of chronic diseases and multiple coexisting diseases, improves the sensitivity and specificity of personalized health screening, significantly enhances the accuracy and applicability of gender-specific risk assessment, and provides a microbiological basis for individualized health management.
Smart Images

Figure CN121789770A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biological detection and risk assessment, and in particular to a method for detecting chronic diseases and coexisting diseases driven by the tongue coating microbiota. Background Technology
[0002] With the increasing global burden of non-communicable diseases and the growing prevalence of co-existence of multiple diseases, public health and sustainable socioeconomic development are under serious threat. Studies have shown that the human microbiome, especially the oral microbiome, plays a crucial role in the occurrence and development of chronic diseases. The tongue coating, as an important carrier of "internal symptoms manifesting externally" in Traditional Chinese Medicine, has a microbial composition closely related to systemic diseases such as hypertension, diabetes, and dyslipidemia. Furthermore, its non-invasive sampling and ease of operation make it a potential biological window for disease early warning and health monitoring.
[0003] However, existing methods still have significant limitations. Most studies are limited to a single chronic disease type or a higher classification level, lacking a systematic analysis of microbial species-level changes in coexisting multiple diseases. Furthermore, existing methods do not adequately address the moderating role of demographic factors such as host sex in the association between microbes and diseases, failing to establish risk assessment models with population heterogeneity, thus limiting their application in personalized health screening and precision intervention. Therefore, based on these challenges, this invention proposes a tongue coating microbiota-driven method for detecting chronic diseases and coexisting multiple diseases. Summary of the Invention
[0004] Purpose of the invention To address the aforementioned issues, the present invention aims to provide a tongue coating microbiome-driven method for detecting chronic diseases and coexisting multiple diseases. This method addresses the lack of non-invasive, accurate, and species-level resolution early screening technologies in existing chronic disease and coexisting disease risk assessments. By systematically analyzing the dynamic changes of the tongue coating microbiome under disease states and incorporating host factors such as gender for stratified modeling, the invention provides a microbial biomarker identification and risk warning tool applicable to large-scale population screening and personalized health management.
[0005] Technical solution To achieve the above objectives, this invention provides a tongue coating microbiota-driven method for detecting chronic diseases and coexisting multiple diseases. This method involves standardized collection of tongue coating samples and extraction of microbial DNA; obtaining species-level species composition profiles using metagenomic sequencing technology; screening for microbial biomarkers significantly associated with chronic diseases and coexisting multiple diseases through statistical tests and machine learning algorithms; establishing a microbial change model that reflects the disease progression gradient; and further constructing a gender-specific risk assessment sub-model, ultimately forming a disease risk classification and early warning system that integrates microbial characteristics and population heterogeneity.
[0006] In a first aspect, the present invention provides a method for detecting chronic diseases and multiple coexisting diseases driven by the tongue coating microbiota, comprising: Tongue samples were collected from the subjects, and microbial DNA was extracted. The DNA was sequenced using high-throughput metagenomic sequencing technology; Constructing microbial species-level community composition profiles based on sequencing results; The samples were grouped by chronic disease status, including no disease, single disease, and comorbidity; Statistical methods and machine learning models were used to identify characteristic bacterial species that were significantly associated with chronic diseases and coexistence of multiple diseases. The identification steps of the characteristic bacterial species include constructing a multimodal composite screening index by fusing statistical significance, machine learning predictive importance, and biological gradient discrimination ability, and thereby screening out microbial biomarkers that are significantly related to disease status and have clear indications of disease progression from the community composition spectrum at the microbial species level.
[0007] The method enables non-invasive, high-precision microbiome analysis with species-level resolution, significantly improving the sensitivity and specificity of early screening for chronic diseases.
[0008] Furthermore, the characteristic bacterial species include a set of pre-determined marker species whose abundance changes in chronic disease states are statistically significant and can effectively distinguish different disease states and disease progression gradients.
[0009] Furthermore, the marker species include species that exhibit abundance gradients or non-gradient changes in healthy states, single chronic diseases, and coexisting multiple diseases.
[0010] Furthermore, the gradient change is manifested as at least one species gradually decreasing in abundance as the severity of the disease increases, and / or at least another species gradually increasing in abundance as the severity of the disease increases.
[0011] Furthermore, the non-gradient changes include bacterial species that are significantly reduced in the diseased group and bacterial species that are significantly enriched in the diseased group.
[0012] Furthermore, it includes a gender stratification analysis step to identify male- or female-specific microbial biomarkers and establish gender-specific risk assessment sub-models accordingly, thereby improving the accuracy and applicability of personalized health risk assessment.
[0013] Furthermore, the sample collection steps include standardized fasting sampling, aseptic rinsing, and multi-swab rotation scraping to ensure the representativeness and consistency of the samples; the sequencing steps adopt a paired-end sequencing mode and undergo strict quality control and host genome decontamination to ensure data reliability.
[0014] Furthermore, the identification step of the characteristic bacterial species further includes: A species interaction network of the tongue coating microbiota is constructed, and the network centrality index of each microbial species in the interaction network is calculated. Based on the network centrality index, the discrimination weights of the characteristic bacterial species identified by the statistical method and machine learning model are dynamically adjusted to screen out microorganisms that play a key pivot role in the interaction network as an optimized set of characteristic bacterial species. The construction of species interaction networks is based on the abundance correlation among microbial species, and the network centrality index includes at least one of degree centrality, betweenness centrality, and compactness centrality.
[0015] In a second aspect, the present invention also provides a tongue coating microbiota-driven detection system for chronic diseases and multiple coexisting diseases, the system being based on the method described in the first aspect above, comprising: The tongue coating sample processing module is used to standardize the processing of tongue coating samples and extract microbial DNA. A species abundance detection unit is used to detect the relative abundance of the marker species; The risk prediction unit integrates LEfSe analysis and random forest algorithm to output the risk level of individual chronic diseases and multiple co-occurrence of diseases based on microbial abundance data.
[0016] The system can quantitatively assess the association between microbial characteristics and disease states, and generate visualized risk stratification results.
[0017] Furthermore, the risk prediction unit includes a gender-specific analysis module for processing the microbial data of male and female subjects respectively, and providing corresponding risk prediction and interpretation instructions.
[0018] Thirdly, the present invention also provides a combination of tongue coating microbial markers for risk assessment of chronic diseases and coexistence of multiple diseases, comprising at least one of the following microbial species: TM7_phylum_sp_oral_taxon_348, Streptococcus_sanguinis, Streptococcus_salivarius, Rothia_aeria, Prevotella_vespertina, Prevotella_koreensis, Prevotella_jejuni, Prevotella_histicola, Ottowia_sp_Marseille_P4747, Neisseria_sp_oral_taxon _014, Lautropia_mirabilis, Lachnoanaerobaculum_saburreum, Dialister_invisus, Candidatus_Nanosynsacchari_sp_TM7_ANC_ 38_39_G1_1, Candidatus_Nanosynbacter_SGB96076, Actinomyces_sp_ICM58, Actinomyces_graevenitzii, Actinomyces_dentalis.
[0019] Fourthly, the present invention also provides the application of tongue coating microbiome in the preparation of a risk assessment kit for chronic diseases and multiple coexisting diseases, including: The tongue coating microbiome contains at least one of the following microbial species: TM7_phylum_sp_oral_taxon_348, Streptococcus_sanguinis, Streptococcus_salivarius, Rothia_aeria, Prevotella_vespertina, Prevotella_koreensis, Prevotella_jejuni, Prevotella_histicola, Ottowia_sp_Marseille_P4747, Neisseria_sp_oral_taxon _014, Lautropia_mirabilis, Lachnoanaerobaculum_saburreum, Dialister_invisus, Candidatus_Nanosynsacchari_sp_TM7_ANC_ 38_39_G1_1, Candidatus_Nanosynbacter_SGB96076, Actinomyces_sp_ICM58, Actinomyces_graevenitzii, Actinomyces_dentalis; The kit includes specific primers, probes, or sequencing components for detecting the tongue microbiome.
[0020] Furthermore, the kit also includes instructions for differentiating and interpreting microbial biomarker data from subjects of different genders to support personalized medical decision-making.
[0021] Fifthly, the present invention also provides a non-transitory computer-readable medium storing computer-executable instructions, which, when executed by a processor, implement the method described in the first aspect or control the operation of the system described in the second aspect; the medium includes, but is not limited to, solid-state memory, optical storage devices, or cloud storage systems, ensuring that the method and system can be stably implemented on various hardware platforms.
[0022] This invention relates to a non-invasive method for assessing the risk of chronic diseases and coexisting diseases based on species-level feature analysis of tongue coating microbial communities. It involves standardized collection of tongue coating samples and acquisition of microbial composition data using high-throughput metagenomic sequencing technology; construction of species-level abundance spectra based on bioinformatics analysis; screening for microbial biomarkers significantly correlated with disease status and progression gradients by combining statistical tests and machine learning algorithms; and further, the introduction of a gender-stratified analysis framework to establish a risk assessment model that reflects population heterogeneity, ultimately achieving individualized disease risk stratification and early warning.
[0023] This method provides a high-resolution, high-specificity non-invasive screening tool for chronic disease risk. It can not only reveal the dynamic changes of the microbial community as the disease progresses, but also significantly improve the accuracy of risk prediction and population applicability through the construction of a gender-specific model. It provides a reliable microbiological basis and innovative technical approach for community health monitoring, individualized intervention and precision public health practices.
[0024] Beneficial effects By implementing the tongue coating microbiota-driven detection method for chronic diseases and multiple coexisting diseases provided by the present invention, the following technical effects are achieved: (1) In the field of chronic disease and coexistence of multiple diseases, a systematic and precise analysis and characteristic map construction of the tongue coating microbial community at the species level has been completed. This breakthrough has broken through the bottleneck that previous studies were mostly limited to higher taxonomic levels, and a large number of specific species-level biomarkers have been discovered, providing unprecedented microbiological resolution and technical basis for non-invasive and precise early disease screening and risk warning.
[0025] (2) By introducing a gender-stratified analysis framework, significant gender heterogeneity was revealed in the association patterns between microorganisms and chronic diseases and the coexistence of multiple diseases. Based on this, a gender-specific risk assessment sub-model was constructed. This overcomes the limitations of the "one-size-fits-all" model, greatly improves the accuracy and applicability of individualized health risk assessment, and lays the core theoretical foundation for moving towards gender-specific precision prevention and intervention strategies.
[0026] (3) By constructing a composite screening index that integrates statistical significance, machine learning predictive importance, and biological gradient discrimination ability, the omission and bias problems existing in traditional single-method screening of biomarkers are systematically solved. It can select microbial species that are not only statistically significant but also have clear indications of disease progression, thereby significantly enhancing the reliability, interpretability, and generalization ability of the final risk assessment model.
[0027] (4) By introducing a dynamic weight adjustment mechanism for microbial interaction networks, the biological rationality and robustness of the risk assessment model are significantly enhanced from a systems ecology perspective. This model effectively identifies species that play key pivotal roles in the microbial community. These species not only have their abundance changes correlated with disease status but also significantly impact the stability of their respective ecological networks. Therefore, compared to traditional methods that only consider species-independent effects, the set of biomarkers selected by this model better represents the overall functional state of the microbial community, thereby improving the model's ability to interpret disease physiological mechanisms and providing a more solid biological basis for risk prediction results. Attached Figure Description
[0028] To make the above-described method for detecting chronic diseases and multiple coexisting diseases driven by tongue microbiota of the present invention more obvious and understandable, the accompanying drawings used in the specific embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0029] Figure 1 This is a flowchart illustrating the method described in this application; Figure 2 This is a schematic diagram illustrating the principle of this application; Figure 3 This indicates the changing trends of characteristic bacterial species in the non-chronic disease group, the single chronic disease group, and the multiple disease coexistence group. Detailed Implementation
[0030] Example 1: A method for detecting chronic diseases and multiple coexisting diseases driven by the tongue coating microbiota is provided, the procedure of which is as follows: Figure 1 As shown, the principle is as follows Figure 2 As shown, it specifically includes: Tongue samples were collected from the subjects, and microbial DNA was extracted. The DNA was sequenced using high-throughput metagenomic sequencing technology; Constructing microbial species-level community composition profiles based on sequencing results; The samples were grouped by chronic disease status, including no disease, single disease, and comorbidity; Statistical methods and machine learning models were used to identify characteristic bacterial species that were significantly associated with chronic diseases and coexistence of multiple diseases. We will construct an individual risk scoring model for population chronic disease risk classification and early warning.
[0031] The specific characteristic bacterial species include 18 key bacteria species for chronic diseases, namely: TM7_phylum_sp_oral_taxon_348 (TM7 phylum bacteria (oral taxonomic unit 348)), Streptococcus sanguinis (blood streptococcus), Streptococcus salivarius (salivarius streptococcus), Rothia aeria (aerobic streptococcus), Prevotella vespertina (prevotella vespertina), Prevotella koreensis (prevotella koreensis), Prevotella jejuni (prevotella jejuni), Prevotella histicola (prevotella histiocytola), and Ottowia_sp_Marseille_P4747 (Ottowia sp. Marseille P4747). *Tortel* (Marseille-P4747 strain), *Neisseria* sp. oral taxon 014, *Lautropia mirabilis*, *Lachnoanaerobaculum saburreum*, *Dialister invisus*, *Candidatus Nanosynsacchari* sp. TM7 ANC 38 39 G1 1, *Candidatus Nanosynbacter* SGB96076, *Actinomyces* sp. ICM58, *Actinomyces graevenitzii*, and *Actinomyces dentalis*. As shown in the table below, the statistical tests are based on nonparametric methods, which are robust to data distribution.
[0032] Table 1. 18 Differentially Occurring Bacterial Species in the Tongue Flora of the Chronic Disease Group and the Non-Chronic Disease Group
[0033] The changing trends of characteristic bacterial species in the non-chronic disease group, the single chronic disease group, and the multiple disease coexistence group are as follows: Figure 3 As shown, except for Candidatus Nanosynsacchari sp. TM7 ANC 38 39 G1 1, the expression levels of the other 17 bacterial species differed significantly in the three groups of no chronic disease, single chronic disease, and multiple coexisting diseases, and their abundance changes showed a gradient or non-gradient pattern.
[0034] The gradient change pattern was characterized by a gradual decrease in Candidatus Nanosynbacter SGB96076 with increasing disease severity, and a gradual increase in Neisseria sp. oral taxon 014 with increasing disease severity; the non-gradient change pattern was characterized by a significant decrease in Prevotella vespertina, Actinomyces graevenitzii, and Actinomyces sp. ICM58 in the diseased group, while Neisseria sp. oral taxon 014, Ottowia sp. Marseille-P4747, Lautropia mirabilis, and Prevotella koreensis were significantly enriched in the diseased group.
[0035] It also includes a gender stratification analysis step to identify male- or female-specific microbial biomarkers; The male-specific biomarkers included 16 bacterial species, 15 of which showed significant variations among the three groups. Bergeyella cardium was enriched in the coexisting disease group, while Prevotella histicola, Megasphaera micronucifo, Actinomyces sp ICM58, and Gemella SGB47116 were reduced in the coexisting disease group. Female-specific biomarkers included 31 bacterial species, of which 23 showed significant variations across the three groups. *Slackia exigua*, *Rothia mucilaginosa*, *Porphyromonas gingivalis*, and *Prevotella buccae* were enriched in the co-occurrence group, while *Campylobacter showae*, *GGB2672_SGB3600*, and *Alloprevotella_SGB1463* were reduced in the co-occurrence group.
[0036] Venn diagram analysis identified 49 species associated with chronic diseases, primarily Prevotella and Actinomyces. Among them, Actinomyces_sp_ICM58, Prevotella_jejuni, and Prevotella_histicola showed overlap in both overall and male-specific analyses; Actinomyces_SGB17168 and Prevotella_koreensis were shared in both overall and female-specific datasets; and GGB10485_SGB49305 was a differentially expressed species between males and females.
[0037] Example 2: Based on the aforementioned embodiments, a standardized process for collecting and processing tongue coating samples is provided.
[0038] Before sampling, subjects must fast for ≥8 hours and rinse their mouths with sterile water at least 3 times. To ensure sample integrity, subjects must not have used glucocorticoids or antibiotics within the past 3 months, have no history of long-term medication use, and have no oral diseases.
[0039] Standardized sampling procedures include: Step 1: Using a sterile swab, scrape the inside of the subject's mouth 30 times in a single direction from the base of the tongue to the tip of the tongue; Step 2: A total of 6 swabs are used during the scraping process, and each swab is rotated 5 times to ensure coating adhesion; Step 3: After sampling, immediately transfer the swab to a cryovial and freeze it in liquid nitrogen to -80°C for long-term storage.
[0040] Example 3: Based on the aforementioned embodiments, a microbial DNA sequencing and data analysis workflow is provided.
[0041] The DNA sequencing platform used was BGI-DIPSEQ, with a read length of 100bp paired-end sequencing and a data volume of ≥15Gb / sample. Host decontamination was performed by Bowtie2 alignment with the human genome GRCh38.
[0042] The quality control standards for data preprocessing in the data analysis were: Phred quality value ≥20, read length ≥30bp, and use of fastp v0.23.2 to filter low-quality data. Species annotation used MetaPhlAn4 species spectrum analysis for 481 species with relative abundance ≥0.1%.
[0043] Example 4: Based on the aforementioned embodiments, a process for screening and validating microbial biomarkers is provided.
[0044] The diversity calculation included α-diversity and β-diversity. For α-diversity, the Shannon index was 3.27 for the chronic disease group and 3.30 for the healthy group (P = 0.012). Microbial evenness was reduced in the chronic disease group (Simpson index P = 0.003, reverse Simpson index P = 0.003, Pielou evenness index P = 0.004). For β-diversity, there were significant differences in microbial composition between the healthy and chronic disease groups (PERMANOVA analysis P = 0.001).
[0045] Differential species screening methods included: Wilcoxon rank-sum test, identifying 200 differentially expressed bacterial species from the chronic disease group and the healthy group; LEfSe analysis to screen out 50 differentially expressed bacterial species with an LDA score ≥ 2; and random forest model, sorting the 50 species by their average precision decrease value and finally taking the intersection.
[0046] Example 5: This embodiment provides a specific implementation method for constructing a risk stratification model for chronic diseases and multiple coexisting diseases based on tongue coating microbial biomarkers.
[0047] First, in the overall population analysis, the least absolute shrinkage and selection operator regression method was used to select variables for the 18 candidate microbial taxa previously screened, and finally seven key microbial biomarkers were identified: Candidatus Nanosynbacter_SGB96076, Ottowia_sp_Marseille_P4747, Actinomyces_sp_ICM58, TM7_phylum_sp_oral_taxon_348, Prevotella_koreensis, Dialister_invisus, and Streptococcus_salivarius.
[0048] The relative abundance data of the seven key microbial biomarkers were combined with the participants' demographic covariates to construct a multivariate logistic regression model to assess an individual's risk of developing chronic diseases and multiple co-occurring diseases. Based on the risk score calculated by this logistic regression model, all participants were divided into three risk levels: low-risk group (risk score ≤ 20th percentile), medium-risk group (risk score between 21st and 79th percentiles), and high-risk group (risk score ≥ 80th percentile).
[0049] As shown in Table 2, the model validation results show that, compared with the low-risk group, the odds ratio of chronic diseases in the medium-risk group was significantly higher (OR=4.21, 95% CI: 2.91–6.35), while the odds ratio of the high-risk group was significantly higher (OR=10.20, 95% CI: 6.94–15.60), indicating that the model has good risk stratification ability.
[0050] Table 2. Summary of Model Validation Results
[0051] Furthermore, to examine gender heterogeneity, gender-specific risk models were constructed for both male and female subject groups.
[0052] In the male population, minimum absolute contraction and selection operator regression identified eight key taxonomic units for modeling, including: Prevotella_histicola, GGB10485_SGB49305, Megasphaera_micronuciformis, Actinomyces_sp_ICM58, Solobacterium_SGB6828, Veillonella_tobetsuensis, Bergeyella_cardium, and Lachnoanaerobaculum_sp_ICM7. After incorporating these biomarkers and covariates into the model, risk stratification results showed that compared to the low-risk reference group, the odds ratio (OR) for intermediate-risk individuals was 2.99 (95% CI: 1.97–4.73), and the OR for high-risk individuals was 6.17 (95% CI: 3.98–9.96).
[0053] In the female population, minimum absolute contraction and selection operator regression also identified eight key taxonomic units: GGB10485_SGB49305, Leptotrichia_hongkongensis, Prevotella_oris, Leptotrichia_wadei, Actinomyces_graevenitzii, Slackia_exigua, Streptococcus_australis, and Granulicatella_SGB8244. The corresponding sex-specific risk model showed an odds ratio (OR) of 3.25 (95% CI: 1.97–5.70) for intermediate-risk participants and 5.96 (95% CI: 3.46–10.80) for high-risk participants.
[0054] This embodiment demonstrates that the risk prediction model based on tongue coating microbial biomarkers can effectively stratify the risk of chronic diseases and multiple co-occurring diseases in the overall population and different gender subgroups, providing a reliable microbiological basis for individualized risk assessment.
[0055] Example 6: Building upon the aforementioned embodiments, to address the potential omissions or biases that may arise from relying on single statistical or machine learning methods to screen biomarkers in traditional approaches, an integrated screening framework is proposed. This framework organically integrates the results of different screening methods by constructing a composite index called "Multimodal Discriminant Weighted Score." This index not only considers the significance of differences between groups and the importance of the model, but also introduces a gradient weight factor that reflects the consistency of its changing trends and discriminative ability across different disease states. This allows for a more comprehensive and reliable identification of core biomarkers that are statistically and biologically valuable.
[0056] Using metagenomic sequencing data, a list of differentially expressed microbial species was obtained through LEfSe analysis and random forest algorithm, respectively.
[0057] For each species in the initial screening list, calculate its two base scores.
[0058] The formula for calculating the statistical significance score is:
[0059] In the formula, The statistical significance score; This is the probability value calculated in a statistical test.
[0060] The model importance score is obtained by Min-Max normalization of the average precision decrease value provided by the random forest, so that its range is [0,1].
[0061] The mean abundance of this species in the healthy group, the single-disease group, and the comorbidity group was analyzed, and its gradient discriminant weights were calculated:
[0062] In the formula, Gradient discrimination weights; This represents the average relative abundance of this microbial species in the multi-disease coexistence group; This represents the average relative abundance of the microbial species in the healthy group; This represents the average relative abundance of the microbial species in a single chronic disease group; It is a very small positive number used to prevent the denominator of the formula from being zero, and can be ignored in actual calculations.
[0063] This formula quantifies the ratio of the overall change in species abundance from healthy to comorbid to the degree of fluctuation in the intermediate process. The larger the value, the more consistent and monotonous the change trend of the species, and the stronger its ability to identify the gradient of disease progression.
[0064] Integrating the above indicators, a final comprehensive score is generated for each species:
[0065] In the formula, Weighted score for multimodal discrimination; The weighting coefficients assigned to the statistical significance score; These are the weighting coefficients assigned to the model's importance score; Assign importance score to the model; These are the weight coefficients assigned to the gradient discrimination weights.
[0066] All species were sorted from highest to lowest MDWS score, and the top N species were selected as the final set of core microbial biomarkers for constructing the risk assessment model.
[0067] This method overcomes the limitations of single-method approaches by integrating statistical significance, machine learning importance, and disease gradient discriminative ability. It ensures that the selected biomarkers not only exhibit significant differences but also demonstrate clear biological changes related to disease progression, significantly enhancing the interpretability and reliability of the system. Results show that this multimodal discriminative weighted ensemble screening model, compared to traditional methods relying solely on statistical significance or machine learning importance, can more comprehensively and reliably identify core microbial biomarkers with high discriminative value. By integrating multi-dimensional information such as statistical significance, model-predicted importance, and biological gradient discriminative ability, this model effectively avoids the omission or misjudgment of potentially high-value biomarkers due to the limitations of a single screening criterion. The resulting set of biomarkers not only possesses excellent statistical discrimination ability but also exhibits clear and consistent biological changes, significantly enhancing the interpretability, stability, and generalization ability of the final risk assessment model.
[0068] Example 7: Building upon the aforementioned embodiments, this embodiment addresses the issue that traditional methods primarily rely on abundance variations of individual microbial species while neglecting the complex interactions between species within the microbial community. These interactions may exert synergistic or antagonistic effects on disease states. Based on network science theory, this embodiment constructs an interaction network of the tongue coating microbiota, calculates the centrality index of each species within the network, and thus quantifies its functional importance within the entire community. By fusing the network centrality score with existing multimodal discriminative weighted scores, the weight of each species in risk assessment is dynamically adjusted, giving higher weights to species that occupy key positions in the interaction network and significantly impact community stability. This not only captures the overall dynamics of the microbial community but also improves the biological rationale for risk prediction.
[0069] First, a relative abundance matrix at the microbial species level was obtained from metagenomic sequencing data of tongue coating samples. The data underwent quality control and host decontamination processes to retain species with a relative abundance ≥0.1%. Logarithmic transformation was performed on the abundance data to reduce the impact of skewed distribution.
[0070] Calculate the Spearman correlation coefficients among all microbial species to form a correlation matrix. Set a correlation threshold to retain only significant correlations, and construct an undirected weighted network. Here, nodes represent microbial species, edges represent the interaction strength between species, and edge weights are the absolute values of the correlation coefficients.
[0071] For each species node, calculate the following centrality metric: Degree centrality (DC): the number of connections a node has to other nodes, normalized to [0,1].
[0072] Betweenness centrality (BC): the mediating role of a node on the shortest path in the network, normalized to [0,1].
[0073] Compactness centrality (CC): The inverse of the average distance from a node to other nodes, normalized to [0,1].
[0074] Then, calculate the overall network importance score for each species:
[0075] In the formula, The network importance score for species i; , and For the weighting coefficients, satisfying The default value is set to , , ; The degree centrality is normalized. For normalized betweenness centrality; This represents the normalized compactness centrality.
[0076] The network importance score is combined with the multimodal discriminative power weighted score described in the above embodiments to generate a comprehensive risk score, as shown in the following formula:
[0077] In the formula, The overall risk score for species i; The weighted score for calculating the multimodal discriminative power; The minimum network importance score for all species is used to normalize the network importance score to the range [0,1]. The maximum value of the network importance score for all species is used to normalize the network importance score to the range [0,1]. This parameter is used to adjust the contribution of the network importance; the default value is 0.2, and it is optimized through cross-validation.
[0078] All species were ranked in descending order based on their comprehensive risk scores, and the top-ranked species were selected as the final key microbial biomarkers. These species not only have high statistical and machine learning importance, but also play crucial roles in microbial interaction networks.
[0079] Key species selected were input into a random forest model to train a risk classifier for chronic diseases and coexisting diseases. A network importance analysis module was added to the risk prediction unit to calculate the microbial interaction network and overall risk score of new samples in real time, outputting individualized risk levels. Simultaneously, in the sex-specific analysis, interaction networks for males and females were constructed separately to capture sex heterogeneity.
[0080] This model significantly improves the performance of risk assessment models by integrating microbial interaction networks, more comprehensively reflects the ecological dynamics of microbial communities, reduces the bias of single-species analysis, and thus improves the robustness and clinical applicability of the model.
Claims
1. A method for detecting chronic diseases and multiple coexisting diseases driven by tongue coating microbiota, characterized in that, include: Tongue samples were collected from the subjects, and microbial DNA was extracted. The DNA was sequenced using high-throughput metagenomic sequencing technology; Constructing microbial species-level community composition profiles based on sequencing results; The samples were grouped according to their chronic disease status; Statistical methods and machine learning models were used to identify characteristic bacterial species that were significantly associated with chronic diseases and coexistence of multiple diseases. The identification steps of the characteristic bacterial species include constructing a multimodal composite screening index by fusing statistical significance, machine learning predictive importance, and biological gradient discrimination ability, and thereby screening out microbial biomarkers that are significantly related to disease status and have clear indications of disease progression from the community composition spectrum at the microbial species level.
2. The method according to claim 1, characterized in that: The characteristic bacterial species include a pre-determined set of marker species whose abundance changes are statistically significant in chronic disease states.
3. The method according to claim 2, characterized in that: The marker species include species that exhibit either gradient or non-gradient abundance changes in healthy states, single chronic diseases, and states with multiple coexisting diseases.
4. The method according to claim 3, characterized in that: The gradient change is characterized by a gradual decrease in the abundance of at least one species as the severity of the disease increases, and / or a gradual increase in the abundance of at least another species as the severity of the disease increases.
5. The method according to claim 3, characterized in that: The non-gradient changes include bacterial species that are significantly reduced in the diseased group and bacterial species that are significantly enriched in the diseased group.
6. The method according to claim 1, characterized in that, The identification step of the characteristic bacterial species further includes: A species interaction network of the tongue coating microbiota is constructed, and the network centrality index of each microbial species in the interaction network is calculated. Based on the network centrality index, the discrimination weights of the characteristic bacterial species identified by the statistical method and machine learning model are dynamically adjusted to screen out microorganisms that play a key pivot role in the interaction network as an optimized set of characteristic bacterial species. The construction of species interaction networks is based on the abundance correlation among microbial species, and the network centrality index includes at least one of degree centrality, betweenness centrality, and compactness centrality.
7. A tongue coating microbiome-driven detection system for chronic diseases and multiple coexisting diseases, characterized in that: The system is implemented based on the method of any one of claims 1-6, including: The tongue coating sample processing module is used to standardize the processing of tongue coating samples and extract microbial DNA. A species abundance detection unit is used to detect the relative abundance of the marker species; The risk prediction unit integrates LEfSe analysis and random forest algorithm to output the risk level of individual chronic diseases and multiple co-occurrence of diseases based on microbial abundance data.
8. The system according to claim 7, characterized in that: The risk prediction unit further includes a gender-specific analysis module, which processes the microbial data of male and female subjects separately and provides corresponding risk prediction and interpretation instructions.
9. A combination of tongue coating microbial characteristics for assessing the risk of chronic diseases and coexistence of multiple diseases, characterized in that, The characteristic bacterial species combination includes at least one of the following microbial species: TM7_phylum_sp_oral_taxon_348, Streptococcus_sanguinis, Streptococcus_salivarius, Rothia_aeria, Prevotella_vespertina, Prevotella_koreensis, Prevotella_jejuni, Prevotella_histicola, Ottowia_sp_Marseille_P4747, Neisseria_sp_oral_taxon _014, Lautropia_mirabilis, Lachnoanaerobaculum_saburreum, Dialister_invisus, Candidatus_Nanosynsacchari_sp_TM7_ANC_ 38_39_G1_1, Candidatus_Nanosynbacter_SGB96076, Actinomyces_sp_ICM58, Actinomyces_graevenitzii, Actinomyces_dentalis.
10. The application of a tongue coating microbiome in the preparation of a risk assessment kit for chronic diseases and multiple co-existing diseases, characterized in that: The tongue coating microbiome contains at least one of the following microbial species: TM7_phylum_sp_oral_taxon_348, Streptococcus_sanguinis, Streptococcus_salivarius, Rothia_aeria, Prevotella_vespertina, Prevotella_koreensis, Prevotella_jejuni, Prevotella_histicola, Ottowia_sp_Marseille_P4747, Neisseria_sp_oral_taxon _014, Lautropia_mirabilis, Lachnoanaerobaculum_saburreum, Dialister_invisus, Candidatus_Nanosynsacchari_sp_TM7_ANC_ 38_39_G1_1, Candidatus_Nanosynbacter_SGB96076, Actinomyces_sp_ICM58, Actinomyces_graevenitzii, Actinomyces_dentalis; The kit includes specific primers, probes, or sequencing components for detecting the tongue microbiome.
Citation Information
Patent Citations
Tumor prediction system and method based on tongue coating microorganisms and application of tumor prediction system and method
CN115083600A
Application of tumor prediction system based on tongue coating microorganisms
CN116936074A
Tumor prediction method based on tongue coating microorganisms
CN117292814A
Compositions and methods for detecting pathogen infection
US20050277181A1
Methods for predicting response to a therapy for a disorder through core microbiome guilds
WO2024226805A2