Marker screening method, medium and equipment for judging damp-heat syndrome of atrophic gastritis
By performing microbiome and metabolome analysis on tongue coating samples and integrating tongue coating flora and metabolite data, a predictive model was constructed, which solved the problem of inaccurate diagnosis of damp-heat syndrome in atrophic gastritis in existing technologies and achieved more efficient disease identification.
Patent Information
- Application Number
- CN202511687554.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-13
AI Technical Summary
Existing technologies fail to effectively integrate microbial flora and metabolite information when identifying damp-heat syndrome in atrophic gastritis, resulting in insufficient diagnostic accuracy and reliability. Furthermore, the cumbersome process of collecting stool samples has a weak correlation with traditional Chinese medicine diagnosis.
By obtaining tongue coating samples from subjects, we conduct microbiome and non-targeted metabolomics analysis, integrate tongue coating microbiota marker data and metabolite marker data, construct a fused feature vector, and use machine learning algorithms to build a predictive model to output the risk level or classification result of atrophic gastritis with damp-heat syndrome.
It significantly improves the accuracy and reliability of identifying damp-heat syndrome in atrophic gastritis, reduces the rate of misdiagnosis and missed diagnosis, and is suitable for real-time clinical diagnosis.
Smart Images

Figure CN121528495A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of software performance testing technology, specifically to a method, medium, and device for screening biomarkers to identify damp-heat syndrome in atrophic gastritis. Background Technology
[0002] In existing technologies, some progress has been made in using metabolomics or microbiome technologies for disease diagnosis and biomarker screening. Some approaches focus on analyzing single types of omics data; for example, some studies have constructed disease prediction models by detecting specific metabolites in tongue coating samples alone. However, such methods analyze only from the single dimension of metabolites, failing to integrate information such as the microbial community. This makes it difficult to comprehensively capture the complex biological state within the human body, potentially leading to the omission of crucial information and thus limiting their accuracy and reliability in identifying subtle physiological and pathological changes (such as different syndromes in Traditional Chinese Medicine).
[0003] Other existing technologies attempt to employ multi-omics integrated analysis strategies, but their biological samples are mostly derived from feces. Although feces can reflect the state of the gut, the collection process is relatively cumbersome, and the composition of its flora and metabolites has a weak correlation with the core of traditional Chinese medicine diagnostic theories (such as "tongue diagnosis"). Because these methods fail to utilize tongue coating samples, which are closely integrated with traditional Chinese medicine syndrome differentiation, they cannot directly and effectively serve the objective and standardized diagnosis of traditional Chinese medicine syndromes.
[0004] Therefore, there is an urgent need in this field for a technical solution that can overcome the above-mentioned defects. This solution can utilize tongue coating samples that are highly consistent with the diagnostic theories of traditional Chinese medicine and achieve more accurate and reliable identification of TCM syndromes such as damp-heat syndrome of atrophic gastritis by deeply integrating multi-omics information. Summary of the Invention
[0005] In view of the above problems, this application provides a technical solution for screening biomarkers to identify damp-heat syndrome in atrophic gastritis, in order to solve the technical problem that the existing technology is not accurate enough in identifying damp-heat syndrome in atrophic gastritis.
[0006] To achieve the above objectives, in a first aspect, this application provides a method for screening biomarkers to identify damp-heat syndrome in atrophic gastritis, the method comprising the following steps:
[0007] S1: Obtain a tongue coating sample from the subject;
[0008] S2: Perform microbiome analysis and non-targeted metabolome analysis on tongue coating samples from the same subject to obtain tongue coating microbiota marker data and tongue coating metabolite marker data for the subject, respectively.
[0009] S3: The tongue flora marker data and the tongue metabolite marker data are fused to construct a fused feature vector for the prediction model input. The fusion process specifically includes:
[0010] S31: Based on a predefined microbial-metabolite interaction knowledge base, identify and calculate the features of microbial-metabolite pairings with known biological associations, and generate the first type of fusion features;
[0011] S32: Based on the tongue microbiota marker data and tongue metabolite marker data, calculate the association strength between all microbiota markers and metabolite markers, and construct a microbial-metabolite association network based on the association strength. Extract topological features from the association network as a second type of fusion feature. The topological features include: degree centrality representing the local importance of nodes, feature vector centrality representing the global influence, and average path length representing the density of the network.
[0012] S33: Combine the first type of fusion feature, the second type of fusion feature, and the tongue coating microbiota marker data and tongue coating metabolite marker data to generate the fusion feature vector;
[0013] S4: Input the fused feature vector into the prediction model and output the prediction result; wherein, the prediction model learns from the training set data through a machine learning algorithm, the training set data includes fused feature data of tongue flora and metabolites known to be patients with chronic atrophic gastritis and damp-heat syndrome, and the prediction result is used to indicate the risk level or classification of the subject having chronic atrophic gastritis and damp-heat syndrome.
[0014] Furthermore, step S31 specifically includes: identifying microbial-metabolite pairs that have a synthetic, degradative, or transformative relationship in the metabolic pathway from the knowledge base; for each microbial-metabolite pair, performing numerical estimation on the corresponding microbial abundance data and metabolite concentration data to simulate the interaction strength and generate the first type of fusion feature;
[0015] Step S32 specifically includes: calculating the association strength between all microbial community markers and metabolite markers using Spearman's rank correlation coefficient; determining that the relationship pairs with an absolute value of the correlation coefficient greater than a preset correlation coefficient threshold and a significance probability less than a preset probability threshold are significantly associated; constructing the microbial-metabolite association network using all microbial community-metabolite pairs with significant associations; and extracting topological features from the association network as the second type of fusion features.
[0016] Furthermore, inputting the fused feature vector into the prediction model includes:
[0017] S41: Input the tongue coating microbiota marker data into the first feature extraction sub-network of the prediction model and output the high-order feature representation of the microbiota; and input the tongue coating metabolite marker data into the second feature extraction sub-network of the prediction model and output the high-order feature representation of the metabolites.
[0018] S42: The higher-order feature representation of the microbial community and the higher-order feature representation of the metabolites are concatenated and input into the attention layer; the attention layer dynamically calculates and outputs the attention weight of each feature unit in the higher-order feature of the microbial community and the higher-order feature of the metabolites;
[0019] S43: Based on the attention weights, the concatenated higher-order features are weighted and summed to obtain a weighted fusion feature representation, which is then input into the final classifier to generate a prediction result for atrophic gastritis with damp-heat syndrome; at the same time, the attention weights are output as the contribution of the corresponding tongue flora markers and tongue metabolite markers to the prediction result.
[0020] Furthermore, the method also includes:
[0021] Based on the attention weights, obtain the data of several microbial biomarkers and metabolite biomarkers that are ranked first in terms of contribution, which are used by the prediction model in generating the prediction results.
[0022] Based on the data of several microbial biomarkers and metabolite biomarkers ranked first in terms of contribution, the risk level corresponding to the prediction result and the explanation of the key biological pathways of the risk level are determined and output by querying a predefined microbial-metabolite-pathway knowledge base.
[0023] Based on the elucidation of the key biological pathways, at least one personalized intervention recommendation is generated by matching from a pre-set intervention protocol library.
[0024] Furthermore, tongue samples from the same subject include the first tongue sample;
[0025] Microbiome analysis of tongue coating samples from the same subject includes the following steps:
[0026] Total genomic DNA of microorganisms was extracted from the first tongue coating sample;
[0027] PCR amplification and high-throughput sequencing were performed on the variable regions of the pre-defined marker genes to obtain bacterial community sequence data;
[0028] The bacterial community sequence data is annotated to obtain multiple taxonomic units. Based on the set of bacterial community markers associated with damp-heat syndrome of chronic atrophic gastritis, the relative abundance of taxonomic units in the set of bacterial community markers is calculated from the multiple taxonomic units to obtain the tongue coating bacterial community marker data.
[0029] Furthermore, the variable region of the preset marker gene is the V3-V4 variable region of 16S rRNA gene sequencing.
[0030] Furthermore, tongue samples from the same subject included a second tongue sample;
[0031] Untargeted metabolomics analysis of tongue coating samples from the same subject includes the following steps:
[0032] Metabolites were extracted from the second tongue coating sample using an organic solvent extraction solution;
[0033] The metabolite extract was separated by liquid chromatography, and the chromatographic effluent was detected by mass spectrometry to obtain metabolite spectral data.
[0034] Peak extraction, peak alignment, and metabolite identification are performed on the metabolite spectrum data. Based on the identification results, the peak area or relative intensity of each metabolite is determined to obtain the tongue coating metabolite marker data.
[0035] Furthermore, the tongue coating microbiota markers include at least one of the following genera: *Macrococcus*, *Miscanthus*, *Porphyromonas*, and *Peptostreptococcus*; the tongue coating metabolite markers include at least one of the following metabolites: L-palmitoylcarnitine, (4-fluorophenyl)[2-(4-methoxyphenyl)-1H-imidazol-5-yl] methyl ketone, avermectin, 4-amino-5-(4-bromophenyl)-4H-1,2,4-triazol-3-thiol, 2-[(6-amino-9H-purin-8-yl)thio]acetic acid, ethyl trans-caffeate, maltotetraose, and 8-bromo-3,7-dimethyl-3,7-dihydro-1H-purin-2,6-dione.
[0036] In a second aspect, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in the first aspect of this application.
[0037] In a third aspect, this application provides an electronic device having a computer program stored thereon, including a processor and a storage medium, wherein the computer program is stored on the storage medium, and when executed by the processor, the computer program implements the biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in the first aspect of this application.
[0038] Unlike existing technologies, the above-mentioned technical solution involves a method, medium, and device for screening biomarkers to identify damp-heat syndrome in atrophic gastritis. The method includes: obtaining tongue samples from subjects and performing microbiome and non-targeted metabolomics analyses to obtain tongue microbiota biomarker data and tongue metabolite biomarker data, respectively; fusing the two types of data to construct a fused feature vector; and inputting this fused feature vector into a trained prediction model to output a risk level or classification result indicating whether the subject has chronic atrophic gastritis with damp-heat syndrome. This invention, through deep fusion of multi-omics data, effectively captures the complex interactions between microorganisms and metabolites, significantly improving the accuracy and reliability of identifying and screening for damp-heat syndrome in atrophic gastritis.
[0039] The above description of the invention is merely an overview of the technical solution of this application. In order to enable those skilled in the art to better understand the technical solution of this application and to implement it based on the description and drawings, and to make the above-mentioned objectives and other objectives, features and advantages of this application easier to understand, the following description is provided in conjunction with the specific embodiments and drawings of this application. Attached Figure Description
[0040] The accompanying drawings are only used to illustrate the principles, implementation methods, applications, features, and effects of specific embodiments of this application and other related content, and should not be considered as limitations on this application.
[0041] In the accompanying drawings of the instruction manual:
[0042] Figure 1 This is a flowchart of a biomarker screening method for identifying damp-heat syndrome in atrophic gastritis, as described in the first exemplary embodiment of this application.
[0043] Figure 2 A flowchart illustrating the biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in the second exemplary embodiment of this application;
[0044] Figure 3 This is a flowchart of the biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in the third exemplary embodiment of this application;
[0045] Figure 4 This is a flowchart of the biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in the fourth exemplary embodiment of this application;
[0046] Figure 5 This is a flowchart of the biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in the fifth exemplary embodiment of this application;
[0047] Figure 6 This is a schematic diagram of an electronic device according to an exemplary embodiment of this application;
[0048] Figure 7 This is a schematic table illustrating the classification and evaluation indicators of the CAG damp-heat syndrome group, CAG non-damp-heat syndrome group, and healthy group identification model as described in an exemplary embodiment of this application.
[0049] Figure 8 This is an example embodiment of the ROC curves for tongue flora recognition in the CAG damp-heat syndrome group, CAG non-damp-heat syndrome group, and healthy group described in this application.
[0050] Figure 9 This is a schematic table illustrating the classification and evaluation indicators of the CAG damp-heat syndrome group, CAG non-damp-heat syndrome group, and healthy group identification model as described in another exemplary embodiment of this application.
[0051] Figure 10 This is a ROC curve diagram for tongue flora recognition in the CAG damp-heat syndrome group, CAG non-damp-heat syndrome group, and healthy group, as described in another exemplary embodiment of this application.
[0052] Figure 11 This is a schematic table illustrating the classification and evaluation indicators of the CAG damp-heat syndrome group, CAG non-damp-heat syndrome group, and healthy group as described in an exemplary embodiment of this application.
[0053] Figure 12 ROC curve of the combined model of CAG damp-heat syndrome group, CAG non-damp-heat syndrome group and healthy group described in an exemplary embodiment of this application;
[0054] Figure 13 This is a schematic diagram illustrating the gender and age composition of the research subjects as described in an exemplary embodiment of this application.
[0055] Figure 14 For this application Figure 13 The table showing the tongue coating distribution of the research subjects is shown below.
[0056] Figure 15 This is a partial table showing the differences in the horizontal distribution of the healthy group and the CAG damp-heat syndrome group according to an exemplary embodiment of this application;
[0057] Figure 16 As exemplarily described in this application Figure 15 Corresponding remaining table chart;
[0058] Figure 17 This is a comparative table showing the abundance differences between the CAG non-humid heat syndrome group and the CAG humid heat syndrome group as exemplarily described in this application;
[0059] The reference numerals used in the above figures are explained as follows:
[0060] 10. Electronic devices;
[0061] 101. Processor;
[0062] 102. Storage medium. Detailed Implementation
[0063] To explain in detail the possible application scenarios, technical principles, specific feasible solutions, and the objectives and effects that this application can achieve, the following detailed description is provided in conjunction with the listed specific embodiments and accompanying drawings. The embodiments described herein are only used to more clearly illustrate the technical solutions of this application, and are therefore only examples, and should not be used to limit the scope of protection of this application.
[0064] like Figure 1 As shown, in a first aspect, this application provides a method for screening biomarkers to identify damp-heat syndrome in atrophic gastritis, the method comprising the following steps:
[0065] S1: Obtain a tongue coating sample from the subject;
[0066] S2: Perform microbiome analysis and non-targeted metabolome analysis on tongue coating samples from the same subject to obtain tongue coating microbiota marker data and tongue coating metabolite marker data for the subject, respectively.
[0067] S3: The tongue flora marker data and the tongue metabolite marker data are fused to construct a fused feature vector for the prediction model input. The fusion process specifically includes:
[0068] S4: Input the fused feature vector into the prediction model and output the prediction result; wherein, the prediction model learns from the training set data through a machine learning algorithm, the training set data includes fused feature data of tongue flora and metabolites known to be patients with chronic atrophic gastritis and damp-heat syndrome, and the prediction result is used to indicate the risk level or classification of the subject having chronic atrophic gastritis and damp-heat syndrome.
[0069] In this embodiment, the tongue coating microbiota marker data refers to the microbiota characteristic data extracted from the subject's tongue coating sample that is associated with the damp-heat syndrome of atrophic gastritis. It includes the relative abundance information of specific genera, such as Megasphaera and Atopobium, and can be obtained and analyzed by high-throughput DNA sequencing technology.
[0070] Tongue coating metabolite biomarker data refers to the metabolite characteristic data extracted from the tongue coating samples of subjects that can distinguish atrophic gastritis with damp-heat syndrome from other states. It covers the concentration or peak area information of specific metabolites, such as L-palmitoylcarnitine and maltotetrasaccharide, and is obtained by detection using liquid chromatography-mass spectrometry (LC-TOF-MS).
[0071] Fusion feature vectors refer to a unified data carrier formed by integrating tongue microbiota biomarker data and tongue metabolite biomarker data through multi-dimensional feature integration strategies (including knowledge base-guided cross-referencing and network topology feature extraction). This data carrier can be directly input into the prediction model to achieve complementarity and enhancement of the two types of data information, thereby improving the model's ability to distinguish disease states.
[0072] Predictive models refer to disease discrimination models built based on machine learning algorithms. They use fusion feature data of known diagnostic results (patients with atrophic gastritis and damp-heat syndrome / patients without damp-heat syndrome / healthy people) as the training set, and establish a mapping relationship between input and output (risk level / classification result) by learning the data patterns.
[0073] In step S1, the subject's tongue coating sample can be collected as follows: First, instruct the subject to rinse their mouth with 0.9% saline 3-4 times (30mL each time) to remove food residue and bacteria interference in the oral cavity; then, use a sterile pharyngeal swab to scrape the thicker areas of the subject's tongue coating (such as the middle and root of the tongue) 3-4 times with appropriate force to ensure sufficient sample volume; place the scraped sterile pharyngeal swab into a sterile centrifuge tube, and divide each subject into 4 tubes (1 pharyngeal swab per tube), numbered according to the order of enrollment; after collection, immediately place the tube in an ice box and transport it to the laboratory, and after arrival, quickly place it in a -80℃ freezer to prevent sample degradation and lay the foundation for subsequent analysis.
[0074] In this embodiment, the tongue coating samples of the same subject include a first tongue coating sample and a second tongue coating sample.
[0075] In step S2, as Figure 4 As shown, the microbiome analysis of tongue coating samples from the same subject includes the following steps:
[0076] Step S401: Extract total genomic DNA of microorganisms from the first tongue coating sample;
[0077] Step S402: Perform PCR amplification and high-throughput sequencing on the variable region of the preset marker gene to obtain bacterial community sequence data; preferably, the variable region of the preset marker gene is the V3-V4 variable region of the 16S rRNA gene sequencing.
[0078] Step S403: Species annotation is performed on the bacterial community sequence data to obtain multiple taxonomic units. Based on the set of bacterial community markers associated with damp-heat syndrome of chronic atrophic gastritis, the relative abundance of taxonomic units in the set of bacterial community markers is calculated from the multiple taxonomic units to obtain the tongue coating bacterial community marker data.
[0079] In practical applications, 16S rDNA high-throughput sequencing technology is employed. First, total genomic DNA is extracted from tongue coating samples using CTAB or SDS methods. DNA purity and integrity are checked using Nanodrop and agarose gel electrophoresis. Then, specific primers are designed for the V3-V4 variable region of the 16S rRNA gene, and PCR amplification is performed (the reaction system contains DNA template, primers, Taq enzyme, etc.; amplification conditions are: 95℃ pre-denaturation for 5 min, 95℃ denaturation for 30 s, 55℃ annealing for 30 s, 72℃ extension for 30 s, for a total of 30 cycles, with a final extension at 72℃ for 10 min). The amplified products are then purified using an Agilent 2100 bioanalyzer and quantified using an Illumina library quantification kit before high-throughput sequencing. The raw sequencing data is then assembled (using FLASH software) and filtered (removing low-quality reads and chimeric sequences) to obtain high-quality, effective sequences. Based on these effective sequences, ASVs (Amplicon Sequence Variants) are clustered, and the Silva database is used to analyze each ASV. Species annotation was performed to obtain species classification information and abundance distribution; combined with the validated list of bacterial genera associated with atrophic gastritis and damp-heat syndrome (including *Macrococcus*, *Miscanthus*, *Porphyromonas*, *Peptostreptococcus*, etc.), the relative abundance of these genera was calculated, and finally, data on tongue flora markers were generated.
[0080] like Figure 5 As shown, non-targeted metabolomics analysis of tongue coating samples from the same subject includes the following steps:
[0081] S501: Extract metabolites from the second tongue coating sample using an organic solvent extract;
[0082] S502: The metabolite extract is separated by liquid chromatography, and the chromatographic effluent is detected by mass spectrometry to obtain metabolite spectral data;
[0083] S503: Perform peak extraction, peak alignment, and metabolite identification on the metabolite spectrum data, and determine the peak area or relative intensity of each metabolite based on the identification results to obtain the tongue coating metabolite marker data.
[0084] Specifically, in practical applications, liquid chromatography-mass spectrometry (LC-TOF-MS) is used. 25 mg of thawed tongue coating sample is weighed and 500 μL of extraction buffer (methanol:acetonitrile:water = 2:2:1, containing an isotope-labeled internal standard mixture, such as L-valine-d8) is added. The sample is then homogenized using a tissue homogenizer under ice bath conditions. Following this, ultrasonic extraction is performed (ice-water bath, 300 W power, 15 min). After extraction, the sample is placed in a -40℃ freezer for 30 min to promote protein precipitation. The sample is then centrifuged at 12000 rpm for 15 min at 4℃, and the supernatant is collected. The supernatant is filtered through a 0.22 μm organic phase filter membrane and injected into a Vanquish ultra-high performance liquid chromatograph (Thermo Fisher Scientific) for separation. The chromatographic column used is a Waters ACQUITYUPLC BEH Amide (2.1 mm × 10 mm, 1.7 μm). The mobile phase A contains 25 mmol / L ammonium acetate and 25 mmol / L sodium chloride. An aqueous solution of mmol / L ammonia was used, with acetonitrile as phase B. The gradient elution program was: 0-2 min, 95% B; 2-10 min, 95%-65% B; 10-12 min, 65%-40% B; 12-13 min, 40%-95% B; 13-15 min, 95% B. The column temperature was 40℃, the flow rate was 0.3 mL / min, and the injection volume was 5 μL. The chromatographic effluent was detected in a Q Exactive Focus high-resolution mass spectrometer (Thermo Fisher Scientific) using an electrospray ionization (ESI) source in positive ion mode with a scan range of 100-1000, a resolution of 70000, and collision energies of 20 / 40 / 60 eV. The raw mass spectrometry data were converted to mzXML format using Proteo Wizard software, and peak identification, extraction, and alignment were performed using a self-developed R package (with XCMS kernel). The data was compared with Biotree... DB (V2.1) uses a self-built secondary mass spectrometry database to identify metabolites (algorithm score cutoff value set to 0.3); combined with the list of metabolites associated with atrophic gastritis and damp-heat syndrome, the peak area or relative intensity of these metabolites is extracted to form tongue coating metabolite marker data.
[0085] In step S3, as Figure 2 As shown, step S3 includes the following steps:
[0086] S31: Based on a predefined microbial-metabolite interaction knowledge base, identify and calculate the features of microbial-metabolite pairings with known biological associations, and generate the first type of fusion features;
[0087] S32: Based on the tongue microbiota marker data and tongue metabolite marker data, calculate the association strength between all microbiota markers and metabolite markers, and construct a microbial-metabolite association network based on the association strength. Extract topological features from the association network as a second type of fusion feature. The topological features include: degree centrality representing the local importance of nodes, feature vector centrality representing the global influence, and average path length representing the density of the network.
[0088] S33: Combine the first type of fusion feature, the second type of fusion feature, and the tongue coating microbiota marker data and tongue coating metabolite marker data to generate the fusion feature vector.
[0089] Preferably, step S31 specifically includes: identifying microbial-metabolite pairs that are associated in metabolic pathways (such as having a synthetic, degradation, or transformation relationship) from the knowledge base; for each microbial-metabolite pair, performing numerical estimation on the corresponding microbial abundance data and metabolite concentration data to simulate the interaction strength and generate the first type of fusion feature.
[0090] Metabolic pathway associations refer to the direct or indirect interactions between microbiota and metabolites during the metabolic processes in organisms. These interactions include the synthesis (e.g., microbiota generating specific metabolites through enzymatic reactions), degradation (microbiota breaking down metabolites into smaller molecules), or transformation (microbiota converting one metabolite into another).
[0091] From the microbial-metabolite interaction knowledge base, metabolic pathway associations directly related to glossopharyngeal microbiota markers and metabolite markers were screened. For example: ① Megasphaera – L-palmitoylcarnitine: Megasphaera participates in the fatty acid β-oxidation pathway, promoting the conversion of long-chain fatty acids into acylcarnitine, and is directly related to the synthesis of L-palmitoylcarnitine (long-chain acylcarnitine); ② Atopobium – ethyl trans-caffeate: Atopobium contains caffeate esterase, which can degrade ethyl trans-caffeate to trans-caffeic acid, belonging to degradation association; ③ Peptostreptococcus – maltotetrasaccharide: Peptostreptococcus can secrete amylase, which hydrolyzes maltotetrasaccharide (tetrasaccharide) into glucose, belonging to degradation association; finally, 8 core metabolic pathway associations were identified (covering 4 genera and 6 metabolites).
[0092] The simulation calculation process for interaction strength is as follows: For each pair of metabolic pathway associations, the first type of fusion feature is calculated using the formula "relative abundance of microbial community × metabolite peak area × pathway weight coefficient". The pathway weight coefficient is set according to the association strength (referencing the EC number annotations for enzymatic reactions in the KEGG pathway; a strong association coefficient is 1.2, a moderate association is 1.0, and a weak association is 0.8). For example, the pathway weight coefficient for *Macrococcus* and L-palmitoylcarnitine is 1.2, and the coefficient for *Miscanthus* and ethyl trans-caffeate is 1.0. This formula not only integrates the basic data of microbial communities and metabolites but also incorporates the association strength information of metabolic pathways, making the first type of fusion feature more biologically meaningful.
[0093] When generating the first type of fusion features, a microbial-metabolite interaction knowledge base is first established. Specifically, it can integrate verified microbial-metabolite association information from public databases such as KEGG (Kyoto Encyclopedia of Genes and Genomes) and MetaCyc. For example, it can clarify known biological relationships such as "Megacoccus spp. can participate in short-chain fatty acid metabolism and has an indirect relationship with palmitoylcarnitine synthesis" and "Megabacterium spp. can affect amino acid metabolism and is related to the conversion of trans-caffeic acid ethyl ester". Then, microbial-metabolite pairs that match the tongue coating microbial markers (4 genera) and tongue coating metabolite markers (8 metabolites) are selected from the knowledge base, resulting in 12 core pairs. For each pair, the corresponding relative abundance data of microbial communities and metabolite peak area data are numerically calculated (preferably by multiplication to simulate the interaction strength between the two, such as "relative abundance of Megacoccus spp. × peak area of L-palmitoylcarnitine") to generate 12 first-type fusion features.
[0094] Step S32 specifically includes: calculating the association strength between all microbial community markers and metabolite markers using Spearman's rank correlation coefficient; determining that the relationship pairs with an absolute value of the correlation coefficient greater than a preset correlation coefficient threshold and a significance probability less than a preset probability threshold are significantly associated; constructing the microbial-metabolite association network using all microbial community-metabolite pairs with significant associations; and extracting topological features from the association network as the second type of fusion features.
[0095] Spearman's rank correlation coefficients were calculated for the abundance data of the four bacterial community markers and the peak area data of the eight metabolite markers obtained in step S2. For example, the correlation coefficient between "abundance of *Gastropoda*" and "peak area of L-palmitoylcarnitine" was 0.42; the correlation coefficient between "abundance of *Gastropoda*" and "peak area of ethyl trans-caffeate" was -0.38 (a negative correlation indicates that the higher the abundance of the bacterial community, the lower the concentration of the metabolite, which is consistent with the degradation association); and the correlation coefficient between "abundance of *Porphyromonas*" and "peak area of 8-bromo-3,7-dimethyl-3,7-dihydro-1H-purine-2,6-dione" was 0.35, all of which met the condition of "absolute value > 0.3".
[0096] The process for screening significant association pairs is as follows: The correlation coefficients of all 32 bacterial community-metabolite pairs were tested for significance (using t-tests), and multiple correction was performed using the FDR method (to control for false discovery rate). For example, the p-value for *Macrococcus* and L-palmitoylcarnitine was 0.008 (after correction), the p-value for *Miscanthus* and ethyl trans-caffeate was 0.021 (after correction), and the p-value for *Porphyromonas* and 8-bromo-3,7-dimethyl-3,7-dihydro-1H-purine-2,6-dione was 0.035 (after correction), all less than the preset probability threshold of 0.05. Finally, 15 significant association pairs were selected. Compared to the 18 pairs before refinement, 3 false associations with p-values > 0.05 were eliminated, improving network reliability.
[0097] The process of constructing the association network and extracting topological features is as follows: A microbial-metabolite association network is constructed using Cytoscape software. The node size is set according to the abundance of the microbial community / the peak area of the metabolite (the higher the abundance / peak area, the larger the node). The thickness of the edges is set according to the absolute value of the correlation coefficient (the larger the coefficient, the thicker the edge). When extracting topological features from the network, the connection ratio between microbial community nodes and metabolite nodes is calculated as a supplement to the second type of fusion features, further enriching the network structure information.
[0098] Specifically, when generating the second type of fusion feature, Spearman's rank correlation coefficients were calculated for all possible microbial-metabolite pairs (4 genera abundance) and metabolite markers (8 metabolite peak areas) based on the tongue coating microbial biomarker data (abundance of 4 genera) and the tongue coating metabolite biomarker data (peak areas of 8 metabolites). This correlation coefficient effectively reflects the monotonic correlation between the two sets of data and is not affected by the data distribution pattern. A preset correlation coefficient threshold of 0.3 and a preset probability threshold of 0.05 (obtained after FDR multiple test correction) were set, and correlation coefficients with an absolute value > 0.3 were selected. Eighteen significant association pairs with P (i.e., significance probability) < 0.05 were identified. A microbial-metabolite association network was constructed using four microbial community markers and eight metabolite markers as nodes and the 18 significant association pairs as edges. Three types of topological features were extracted from this network: ① Degree centrality (calculating the number of edges directly connecting each node, e.g., "Megacoccus spp. connects to 3 metabolite nodes, with a degree centrality of 3"), representing the local importance of nodes; ② Eigenvector centrality (calculating the degree of connection between each node and other high-importance nodes, e.g., "L-palmitoylcarnitine connects to 2 high-abundance bacterial genera, with a high eigenvector centrality"), representing global influence; ③ Average path length (calculating the average of the shortest paths between all node pairs in the network, reflecting the network's density; a total of 12 nodes, with an average path length of 1.8). Finally, 25 second-type fusion features were generated: 4 (microbial community degree centrality) + 8 (metabolite degree centrality) + 4 (microbial community eigenvector centrality) + 8 (metabolite eigenvector centrality) + 1 (average path length).
[0099] In step S33, the 12 first-class fusion features obtained in S31, the 25 second-class fusion features obtained in S32, the 4 tongue coating microbial biomarker data (relative abundance of bacterial genus) and the 8 tongue coating metabolite biomarker data (metabolite peak area) obtained in step S2 are concatenated to obtain a fusion feature vector of 12+25+4+8=49 dimensions. All features in the vector are standardized (using Z-score standardization, i.e., subtracting the mean of the feature from each feature value and then dividing by the standard deviation) to eliminate dimensional differences and ensure that the weights of each feature are balanced in model training.
[0100] In step S4, the standardized fusion feature vector obtained in step S33 is input into the trained SVM model. The model classifies the samples through the optimal hyperplane (the hyperplane equation is ѡTx+b=0, where ѡ is a 49-dimensional normal vector and b is a bias term); it outputs two types of prediction results: ① risk level (specifically including high risk: predicted probability > 0.8, medium risk: 0.5 < predicted probability ≤ 0.8, low risk: predicted probability ≤ 0.5), ② classification result (directly determined as "atrophic gastritis with damp-heat syndrome" or "non-atrophic gastritis with damp-heat syndrome"), providing quantitative basis for clinical diagnosis.
[0101] The above solution breaks through the limitations of existing technologies that only focus on a single biomarker (microbiota or metabolites). By fusing microbiota and metabolite data, it integrates microbial information with TCM syndrome types, improves the accuracy of diagnosing damp-heat syndrome in atrophic gastritis, and effectively reduces the rate of misdiagnosis and missed diagnosis.
[0102] By screening for first-class fusion features through metabolic pathway associations, we avoid random data combinations without biological basis, making the features more closely resemble real physiological processes. For example, the fusion feature of *Macrococcus* and L-palmitoylcarnitine can directly reflect abnormalities in fatty acid metabolism pathways, which is highly consistent with the "damp-heat accumulation" pathogenesis (fatty acid metabolism disorder leading to damp-heat accumulation) in atrophic gastritis with damp-heat syndrome. Through Spearman's rank correlation coefficient combined with multiple tests for correction, spurious associations are effectively eliminated, reducing the computational burden of model training and improving prediction efficiency, making it suitable for real-time clinical diagnostic scenarios.
[0103] In some embodiments, such as Figure 3 As shown, inputting the fused feature vector into the prediction model includes:
[0104] S41: Input the tongue coating microbiota marker data into the first feature extraction sub-network of the prediction model and output the high-order feature representation of the microbiota; and input the tongue coating metabolite marker data into the second feature extraction sub-network of the prediction model and output the high-order feature representation of the metabolites.
[0105] S42: The higher-order feature representation of the microbial community and the higher-order feature representation of the metabolites are concatenated and input into the attention layer. The attention layer dynamically calculates and outputs the attention weight of each feature unit in the higher-order feature of the microbial community and the higher-order feature of the metabolites.
[0106] S43: Based on the attention weights, the concatenated higher-order features are weighted and summed to obtain a weighted fusion feature representation, which is then input into the final classifier to generate a prediction result for atrophic gastritis with damp-heat syndrome; at the same time, the attention weights are output as the contribution of the corresponding tongue flora markers and tongue metabolite markers to the prediction result.
[0107] In this embodiment, the feature extraction sub-network refers to a sub-module in the prediction model specifically used to extract high-order abstract features from the original biomarker data. It is divided into a first feature extraction sub-network (processing microbial community data) and a second feature extraction sub-network (processing metabolite data). It adopts a deep neural network (DNN) structure, which can capture the nonlinear correlations hidden in the data. Compared with traditional feature engineering, it can more fully mine the data information.
[0108] Higher-order feature representations refer to abstract feature vectors with dimensions higher than the original data, obtained after processing by the feature extraction sub-network. For example, the first feature extraction sub-network transforms the abundance data of four bacterial community markers into a 64-dimensional higher-order feature representation of the bacterial community. This feature not only contains information about individual bacterial genera, but also integrates the interaction patterns between bacterial genera, making it more suitable for the discrimination of complex diseases.
[0109] The attention layer is a model module that simulates the human attention mechanism. By calculating the attention weight of each feature unit, it assigns higher weights to features that are more important to the diagnosis and suppresses the interference of irrelevant features. The larger the weight value, the higher the contribution of the feature to the prediction result, which can achieve the effect of focusing on important features.
[0110] A classifier is the final decision-making module of a prediction model. It receives weighted fused feature representations and outputs classification results. In this embodiment, the classifier is preferably a logistic regression classifier. This classifier has a simple structure, strong interpretability, and outputs probability values (0-1), which are easy to convert into risk levels. At the same time, it avoids the overfitting problem of complex models (such as deep learning).
[0111] In step S41, the first feature extraction sub-network (microbial data processing) adopts a 3-layer fully connected neural network structure. The input layer is 4-dimensional tongue coating microbial biomarker data (standardized relative abundance of bacterial genera), hidden layer 1 contains 32 neurons (activation function is ReLU), hidden layer 2 contains 64 neurons (activation function is ReLU), and the output layer is a 64-dimensional high-order feature representation of the microbial community. During training, the model's classification loss (cross-entropy loss) is used as the optimization objective. The network parameters are updated through the Adam optimizer (learning rate 0.001, decay rate 0.9) so that the high-order features can maximize the differentiation of samples with different diagnostic labels.
[0112] The second feature extraction sub-network (metabolite data processing) adopts a 3-layer fully connected neural network structure symmetrical to the first sub-network. The input layer is 8-dimensional tongue coating metabolite marker data (standardized metabolite peak area), hidden layer 1 contains 32 neurons (ReLU activation function), hidden layer 2 contains 64 neurons (ReLU activation function), and the output layer is a 64-dimensional high-order feature representation of metabolites. During training, it shares optimizer parameters with the first sub-network to ensure that the feature extraction standards of the two types of data are consistent.
[0113] In step S42, feature concatenation involves concatenating the 64-dimensional higher-order feature representations of the microbial community with the 64-dimensional higher-order feature representations of metabolites, resulting in a 128-dimensional concatenated feature vector. The attention layer employs an additive attention mechanism. First, the 128-dimensional concatenated feature vector is input into a fully connected layer (containing 32 neurons with ReLU activation function) to reduce the dimensionality to 32. Then, the 32-dimensional features are normalized using a softmax function, resulting in 32 attention weights (the sum of the weights is 1). These 32 weights are then mapped back to the 128-dimensional concatenated features (each weight corresponds to 4 original concatenated features), ultimately yielding a 128-dimensional attention weight vector. For example, dimensions 1-4 of the microbial community higher-order features correspond to a weight of 0.08, and dimensions 50-53 of the metabolite higher-order features correspond to a weight of 0.12. Higher weight values indicate a greater contribution of the feature to the diagnosis.
[0114] In step S43, the weighted fusion feature representation generation process is as follows: the 128-dimensional concatenated feature vector is multiplied element-wise with the 128-dimensional attention weight vector to obtain the weighted feature vector; the weighted feature vector is summed to obtain a 1-dimensional weighted fusion feature representation, which integrates the information of all higher-order features and highlights the contribution of important features.
[0115] The classification prediction and contribution output process is as follows: The weighted fusion feature representation is input into the logistic regression classifier, and the output sample is the probability value (0-1) of atrophic gastritis with damp-heat syndrome; the risk level (>0.8 is high risk, 0.5-0.8 is medium risk, and <0.5 is low risk) and classification result are determined according to the probability value; at the same time, a 128-dimensional attention weight vector is output. By sorting the weights, the 10-12th dimension of the higher-order features of the microbial community (corresponding to the interaction between *Macrococcus* and other genera) and the 35-37th dimension of the higher-order features of metabolites (corresponding to the association between L-palmitoylcarnitine and maltotetrasaccharide) can be identified as the core contributing features, providing direction for subsequent biological interpretation.
[0116] The above scheme, through a feature extraction sub-network, mines high-order abstract features from the original microbial community and metabolite data, such as synergistic / antagonistic patterns among microbial communities and pathway associations between metabolites, effectively improving the accuracy of distinguishing between damp-heat syndrome and non-damp-heat syndrome in atrophic gastritis. The weight values output by the attention layer directly quantify the contribution of each feature to the prediction result, providing a clear target for subsequent biomarker validation. The combination of the feature extraction sub-network and the attention layer allows the model to adapt to data with different sample sizes (focusing on core features through weight adjustment for small samples, and fully utilizing all features for large samples), improving the model's generalization ability.
[0117] In some embodiments, the method further includes:
[0118] Based on the attention weights, obtain the data of several microbial biomarkers and metabolite biomarkers that are ranked first in terms of contribution, which are used by the prediction model in generating the prediction results.
[0119] Based on the data of several microbial biomarkers and metabolite biomarkers ranked first in terms of contribution, the risk level corresponding to the prediction result and the explanation of the key biological pathways of the risk level are determined and output by querying a predefined microbial-metabolite-pathway knowledge base.
[0120] Based on the elucidation of the key biological pathways, at least one personalized intervention recommendation is generated by matching from a pre-set intervention protocol library.
[0121] In this embodiment, the core biomarkers refer to the tongue microbiota biomarkers and tongue metabolite biomarkers that rank highly in attention weight. These biomarkers contribute the most to the prediction results and are key bioindicators reflecting the pathological state of damp-heat syndrome in atrophic gastritis.
[0122] The microbiome-metabolite-pathway knowledge base is a comprehensive knowledge base that expands upon the microbiome-metabolite interaction knowledge base. It adds information on the association between metabolites and pathways (such as L-palmitoylcarnitine corresponding to the fatty acid metabolism pathway) and the association between pathways and disease pathogenesis (such as abnormal fatty acid metabolism pathway corresponding to the pathogenesis of "damp-heat accumulation" in traditional Chinese medicine). It can achieve a complete mapping from biomarkers to disease mechanisms.
[0123] In some embodiments, the tongue coating microbiota markers include at least one of the following genera: *Macrococcus*, *Miscanthus*, *Porphyromonas*, and *Peptostreptococcus*; the tongue coating metabolite markers include at least one of the following metabolites: L-palmitoylcarnitine, (4-fluorophenyl)[2-(4-methoxyphenyl)-1H-imidazol-5-yl] methyl ketone, avermectin, 4-amino-5-(4-bromophenyl)-4H-1,2,4-triazol-3-thiol, 2-[(6-amino-9H-purin-8-yl)thio]acetic acid, ethyl trans-caffeate, maltotetraose, and 8-bromo-3,7-dimethyl-3,7-dihydro-1H-purin-2,6-dione.
[0124] To determine the effects of different tongue flora markers and tongue metabolite markers on chronic atrophic gastritis (CAG), such as... Figures 13-17 As shown, the following experiments were also conducted to verify this application:
[0125] like Figure 13 As shown, this application included a total of 153 study subjects, including 25 healthy individuals, 55 individuals with chronic atrophic gastritis (CAG) without damp-heat syndrome, and 73 individuals with CAG with damp-heat syndrome. The gender and age composition of the three groups of study subjects is as follows:
[0126] 1. Gender distribution: In the healthy group, there were 11 males (44.00%) and 14 females (56.00%); in the CAG non-damp-heat syndrome group, there were 26 males (47.27%) and 29 females (52.73%); in the CAG damp-heat syndrome group, there were 43 males (58.90%) and 30 females (41.10%). There was no statistically significant difference in gender composition among the three groups (P>0.05), thus excluding the interference of gender factors on subsequent analysis of gut microbiota and metabolites.
[0127] 2. Age distribution: The mean age of the healthy group was 42.68±8.72 years, the mean age of the CAG non-damp-heat syndrome group was 45.58±8.94 years, and the mean age of the CAG damp-heat syndrome group was 46.74±9.64 years. The mean age of the three groups showed a gradual increasing trend, but the difference between the groups was not statistically significant (P>0.05), which is consistent with the clinical characteristics of chronic atrophic gastritis being "more common in middle-aged and elderly people". At the same time, it was ensured that the age factor would not have a systematic impact on the differences in tongue flora and metabolites among the groups.
[0128] In summary, the three groups of subjects showed good balance in baseline characteristics such as gender and age, laying a reliable sample foundation for subsequent comparative analysis of differences in tongue flora and metabolites between CAG patients with damp-heat syndrome and those without, as well as those in healthy states.
[0129] like Figure 14 As shown, statistical analysis of tongue coating types in three groups of subjects revealed a significant correlation between tongue coating distribution and the subjects' health status and CAG TCM syndrome types. The specific patterns are as follows:
[0130] Tongue coating characteristics in the healthy group: All 25 healthy individuals had a thin white tongue coating (100%), with no other tongue coating types. The thin white coating is a typical manifestation of "normal tongue coating" in traditional Chinese medicine theory. Its even distribution further verifies that the healthy group subjects had good gastrointestinal function and overall health status, and can be used as a "normal control benchmark" for subsequent comparison of tongue coating differences between disease groups.
[0131] Tongue coating characteristics of CAG non-damp-heat syndrome group: Among the 55 CAG non-damp-heat syndrome patients, 16 cases (29.09%) had thin white coating, 10 cases (18.18%) had thin yellow coating, 12 cases (21.82%) had white greasy coating, and 17 cases (30.91%) had yellow greasy coating. The tongue coating types showed a diverse distribution, with thin white coating accounting for the highest proportion and yellow greasy coating accounting for the second highest proportion. This reflects that the degree of gastrointestinal dysfunction in CAG non-damp-heat syndrome patients is relatively mild, and the pathogenesis of internal damp-heat is not significant.
[0132] Tongue coating characteristics in the CAG damp-heat syndrome group: Among 73 patients with CAG damp-heat syndrome, 40 cases (54.79%) had a yellow and greasy coating, 13 cases (17.81%) had a thin white coating, 12 cases (16.44%) had a white and greasy coating, and 8 cases (10.96%) had a thin yellow coating. The proportion of yellow and greasy coating was significantly higher than other tongue coating types, and much higher than that in the CAG non-damp-heat syndrome group (30.91%). In traditional Chinese medicine theory, yellow and greasy coating is the core tongue manifestation of "internal damp-heat". This result is highly consistent with the criteria for CAG damp-heat syndrome, which requires a "heat (fire) and dampness" syndrome element score ≥100. This further verifies the clinical observation that "yellow and greasy coating can be used as a typical tongue diagnosis feature of CAG damp-heat syndrome", and provides a theoretical basis in traditional Chinese medicine for subsequent analysis of flora and metabolites based on tongue coating samples.
[0133] In summary, the distribution of tongue coating types can serve as a direct indicator to distinguish between healthy individuals, patients with non-damp-heat syndrome of CAG, and patients with damp-heat syndrome of CAG. Among them, yellow and greasy tongue coating has a high indicative significance for the diagnosis of damp-heat syndrome of CAG.
[0134] like Figure 15 and Figure 16 As shown, 16S rRNA high-throughput sequencing technology was used to analyze the microbial community of tongue coating samples from the healthy group and the CAG damp-heat syndrome group. At the genus level, 15 genera with significant differences were screened (P < 0.05). These differential genera can reflect the disordered structure of the tongue coating microbial community in CAG damp-heat syndrome at the microscopic level. The specific differences are as follows:
[0135] The abundance of bacterial genera significantly increased in the CAG damp-heat syndrome group: Compared with the healthy group, the abundance of seven bacterial genera significantly increased in the CAG damp-heat syndrome group: Leptotrichia (P=0.047), Prevotella (P=0.007), Megasphaera (P=0.001), Selenomonas (P=0.001), TM7x (P=0.001), Atopobium (P=0.004), and Lachnospiraceae_unclassified (P=0.001). Among them, the abundance differences of *Macrococcus*, *Lentinula*, *TM7x*, and unclassified *Trichophyton* were the most significant (P=0.001). For example, the mean abundance of *Macrococcus* in the CAG damp-heat syndrome group (689.00‰) was 10.9 times that of the healthy group (63.00‰), suggesting that the excessive proliferation of these genera may be related to the pathogenesis of "damp-heat accumulation" in CAG damp-heat syndrome. For example, fatty acid metabolism disorders involving *Macrococcus* may lead to the accumulation of metabolic products and aggravate the damp-heat state of the gastrointestinal tract.
[0136] The abundance of bacterial genera significantly decreased in the CAG damp-heat syndrome group: Compared with the healthy group, the abundance of eight bacterial genera was significantly reduced in the CAG damp-heat syndrome group: Amnipila (P=0.008), Escherichia-Shigella (P=0.001), Trichodoccus (P=0.018), Peptostreptococcus (P=0.002), Porphyromonas (P=0.035), Streptococcus (P=0.011), Rothia (P=0.026), and Haemophilus (P=0.008). Among them, the abundance differences of Escherichia coli-Shigella, Peptostreptococcus, and Haemophilus were particularly prominent (P≤0.008). For example, the mean abundance of Haemophilus in the healthy group (7933.00‰) was 2.9 times that in the CAG damp-heat syndrome group (2696.00‰). The decrease in the abundance of these genera may disrupt the balance of the tongue coating microecology, weaken the barrier function of the oral and gastrointestinal mucosa, and create conditions for the occurrence and development of CAG damp-heat syndrome.
[0137] In summary, there were significant and regular differences in the tongue flora between the healthy group and the CAG damp-heat syndrome group at the genus level. These different genera can not only serve as potential flora markers to distinguish between the two groups, but also provide an objective basis for the "damp-heat accumulation" pathogenesis of CAG damp-heat syndrome from a microbiological perspective.
[0138] like Figure 17 As shown, to further screen for tongue flora markers specific to CAG damp-heat syndrome, the differences in tongue flora genera between the CAG non-damp-heat syndrome group and the CAG damp-heat syndrome group were compared and analyzed. The results showed significant differences between the two groups in four genera (P < 0.05), and these differential genera can be used as core flora markers to distinguish between the two CAG syndrome types, as detailed below:
[0139] The genera of bacteria with significantly increased abundance in the CAG damp-heat syndrome group include:
[0140] Megasphaera: The mean abundance of Megasphaera in the CAG damp-heat syndrome group was 689.00‰ (range 63.00-1449.00‰), which was significantly higher than that in the CAG non-damp-heat syndrome group (94.00‰, range 56.00-507.00‰), and the difference was statistically significant (P=0.008). The abundance multiple was 7.3 times, making it the genus with the most significant difference between the two groups.
[0141] *Atopobium*: The mean abundance in the CAG damp-heat syndrome group was 351.00‰ (range 68.50-748.50‰), significantly higher than the 137.00‰ (range 67.00-298.00‰) in the CAG non-damp-heat syndrome group, with a statistically significant difference (P=0.025), representing a 2.56-fold increase in abundance. The genera whose abundance was significantly reduced in the CAG damp-heat syndrome group include:
[0142] Peptostreptococcus: The mean abundance of Peptostreptococcus in the CAG damp-heat syndrome group was 491.00‰ (range 198.50-1134.00‰), which was significantly lower than that in the CAG non-damp-heat syndrome group (1010.00‰, range 323.00-1708.00‰), and the difference was statistically significant (P=0.044). The abundance was only 48.6% of that in the non-damp-heat syndrome group.
[0143] Porphyromonas: The mean abundance of Porphyromonas in the CAG damp-heat syndrome group was 1022.00‰ (range 387.50-2721.00‰), which was significantly lower than that in the CAG non-damp-heat syndrome group (1988.00‰, range 603.00-4821.00‰), and the difference was statistically significant (P=0.013). The abundance was only 51.4% of that in the non-damp-heat syndrome group.
[0144] Will Figure 17 and Figure 15 , Figure 16 Intersection analysis of differentially expressed bacterial genera revealed significant differences in four genera: *Macrococcus*, *Migraine*, *Peptostreptococcus*, and *Porphyromonas*, in both the healthy group vs. the CAG damp-heat syndrome group and the CAG non-damp-heat syndrome group vs. the CAG damp-heat syndrome group. The trends of these differences were consistent (*Macrococcus* and *Migraine* increased in the CAG damp-heat syndrome group, while *Peptostreptococcus* and *Porphyromonas* decreased). This result indicates that these four genera are core biomarkers of the tongue coating microbiota specific to CAG damp-heat syndrome, distinguishing it from the healthy state and the CAG non-damp-heat syndrome. Changes in their abundance can specifically reflect the pathophysiological state of CAG damp-heat syndrome, providing crucial microbiota characteristic data for the subsequent construction of a "microbiota-metabolite joint prediction model."
[0145] In some embodiments, this application further verifies the method involved in this application by setting different experimental groups, and the classification evaluation indexes of the corresponding groups and the ROC curves of the corresponding experimental results are shown in the figure. Figures 7-12 As shown.
[0146] In a second aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in the first aspect of the present invention.
[0147] The computer-readable storage medium may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory.
[0148] The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD ROM); the magnetic surface memory may be a disk storage device or a magnetic tape storage device.
[0149] The volatile memory may be random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synclink dynamic random access memory (SLDRAM), and direct memory bus random access memory (DRRAM). The computer-readable storage media described in the embodiments of the present invention are intended to include these and any other suitable types of memory.
[0150] like Figure 7 As shown, in a third aspect, the present invention provides an electronic device 10, including a processor 101 and a storage medium 102, wherein a computer program is stored on the storage medium, and when the computer program is executed by the processor, it implements the biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in the first aspect of the present invention.
[0151] In some embodiments, the processor may be implemented by software, hardware, firmware, or a combination thereof, and may use at least one of the following: circuit, single or multiple application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), central processing units (CPUs), controllers, microcontrollers, and microprocessors, so that the processor can execute some or all of the steps or any combination of the steps in the biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in the various embodiments of this application.
[0152] Finally, it should be noted that although the above embodiments have been described in the text and drawings of this application, this should not limit the scope of patent protection of this application. Any technical solutions that are based on the essential concept of this application and utilize the content described in the text and drawings of this application, resulting in equivalent structural or procedural substitutions or modifications, as well as the direct or indirect application of the technical solutions of the above embodiments to other related technical fields, are all included within the scope of patent protection of this application.
Claims
1. A method for screening biomarkers to identify damp-heat syndrome in atrophic gastritis, characterized in that, The method includes the following steps: S1: Obtain a tongue coating sample from the subject; S2: Perform microbiome analysis and non-targeted metabolome analysis on tongue coating samples from the same subject to obtain tongue coating microbiota marker data and tongue coating metabolite marker data for the subject, respectively. S3: The tongue flora marker data and the tongue metabolite marker data are fused to construct a fused feature vector for the prediction model input. The fusion process specifically includes: S31: Based on a predefined microbial-metabolite interaction knowledge base, identify and calculate the features of microbial-metabolite pairings with known biological associations, and generate the first type of fusion features; S32: Based on the tongue microbiota marker data and tongue metabolite marker data, calculate the association strength between all microbiota markers and metabolite markers, and construct a microbial-metabolite association network based on the association strength. Extract topological features from the association network as a second type of fusion feature. The topological features include: degree centrality representing the local importance of nodes, feature vector centrality representing the global influence, and average path length representing the density of the network. S33: Combine the first type of fusion feature, the second type of fusion feature, and the tongue coating microbiota marker data and tongue coating metabolite marker data to generate the fusion feature vector; S4: Input the fused feature vector into the prediction model and output the prediction result; wherein, the prediction model learns from the training set data through a machine learning algorithm, the training set data includes fused feature data of tongue flora and metabolites known to be patients with chronic atrophic gastritis and damp-heat syndrome, and the prediction result is used to indicate the risk level or classification of the subject having chronic atrophic gastritis and damp-heat syndrome.
2. The biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in claim 1, characterized in that, Step S31 specifically includes: identifying microbial-metabolite pairs that have a synthesis, degradation, or transformation relationship in the metabolic pathway from the knowledge base; for each microbial-metabolite pair, performing numerical estimation on the corresponding microbial abundance data and metabolite concentration data to simulate the interaction strength and generate the first type of fusion feature; Step S32 specifically includes: calculating the association strength between all microbial community markers and metabolite markers using Spearman's rank correlation coefficient; determining that the relationship pairs with an absolute value of the correlation coefficient greater than a preset correlation coefficient threshold and a significance probability less than a preset probability threshold are significantly associated; constructing the microbial-metabolite association network using all microbial community-metabolite pairs with significant associations; and extracting topological features from the association network as the second type of fusion features.
3. The biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in claim 1, characterized in that, Inputting the fused feature vector into the prediction model includes: S41: Input the tongue coating microbiota marker data into the first feature extraction sub-network of the prediction model and output the high-order feature representation of the microbiota; and input the tongue coating metabolite marker data into the second feature extraction sub-network of the prediction model and output the high-order feature representation of the metabolites. S42: The higher-order feature representation of the microbial community and the higher-order feature representation of the metabolites are concatenated and input into the attention layer. The attention layer dynamically calculates and outputs the attention weight of each feature unit in the higher-order feature of the microbial community and the higher-order feature of the metabolites. S43: Based on the attention weights, the concatenated higher-order features are weighted and summed to obtain a weighted fusion feature representation, which is then input into the final classifier to generate a prediction result for atrophic gastritis with damp-heat syndrome; at the same time, the attention weights are output as the contribution of the corresponding tongue flora markers and tongue metabolite markers to the prediction result.
4. The biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in claim 3, characterized in that, The method further includes: Based on the attention weights, obtain the data of several microbial biomarkers and metabolite biomarkers that are ranked first in terms of contribution, which are used by the prediction model in generating the prediction results. Based on the data of several microbial biomarkers and metabolite biomarkers ranked first in terms of contribution, the risk level corresponding to the prediction result and the explanation of the key biological pathways of the risk level are determined and output by querying a predefined microbial-metabolite-pathway knowledge base. Based on the elucidation of the key biological pathways, at least one personalized intervention recommendation is generated by matching from a pre-set intervention protocol library.
5. The biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in claim 1, characterized in that, Tongue samples from the same subject include the first tongue sample; Microbiome analysis of tongue coating samples from the same subject includes the following steps: Total genomic DNA of microorganisms was extracted from the first tongue coating sample; PCR amplification and high-throughput sequencing were performed on the variable regions of the pre-defined marker genes to obtain bacterial community sequence data; The bacterial community sequence data is annotated to obtain multiple taxonomic units. Based on the set of bacterial community markers associated with damp-heat syndrome of chronic atrophic gastritis, the relative abundance of taxonomic units in the set of bacterial community markers is calculated from the multiple taxonomic units to obtain the tongue coating bacterial community marker data.
6. The biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in claim 5, characterized in that, The variable region of the preset marker gene is the V3-V4 variable region of the 16S rRNA gene sequencing.
7. The method for screening biomarkers for identifying damp-heat syndrome in atrophic gastritis as described in claim 1, characterized in that, Tongue samples from the same subject include a second tongue sample; Untargeted metabolomics analysis of tongue coating samples from the same subject includes the following steps: Metabolites were extracted from the second tongue coating sample using an organic solvent extraction solution; The metabolite extract was separated by liquid chromatography, and the chromatographic effluent was detected by mass spectrometry to obtain metabolite spectral data. Peak extraction, peak alignment, and metabolite identification are performed on the metabolite spectrum data. Based on the identification results, the peak area or relative intensity of each metabolite is determined to obtain the tongue coating metabolite marker data.
8. The method for screening biomarkers for identifying damp-heat syndrome in atrophic gastritis as described in any one of claims 1 to 7, characterized in that, The tongue coating microbiota markers include at least one of the following genera: *Macrococcus*, *Miscanthus*, *Porphyromonas*, and *Peptostreptococcus*; the tongue coating metabolite markers include at least one of the following metabolites: L-palmitoylcarnitine, (4-fluorophenyl)[2-(4-methoxyphenyl)-1H-imidazol-5-yl] methyl ketone, avermectin, 4-amino-5-(4-bromophenyl)-4H-1,2,4-triazol-3-thiol, 2-[(6-amino-9H-purin-8-yl)thio]acetic acid, ethyl trans-caffeate, maltotetraose, and 8-bromo-3,7-dimethyl-3,7-dihydro-1H-purin-2,6-dione.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in any one of claims 1 to 8.
10. An electronic device having a computer program stored thereon, characterized in that, The device includes a processor and a storage medium, wherein a computer program is stored on the storage medium, and when executed by the processor, the computer program implements the biomarker screening method for identifying damp-heat syndrome in atrophic gastritis as described in any one of claims 1 to 8.