A multi-source biological information processing and analysis system and method based on intestinal oral axis

By constructing a multi-source bioinformatics processing system for the gut-orbital axis, the problems of single information sources and insufficient connection between analysis results and health management recommendations in existing technologies are solved. This achieves efficient integration of multi-source data and interpretability of risk assessment, and provides personalized pharmaceutical intervention recommendations.

CN122392993APending Publication Date: 2026-07-14SHENYANG PHARMA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENYANG PHARMA UNIV
Filing Date
2026-04-02
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

In existing technologies, bioinformatics-based analysis methods suffer from problems such as limited information sources, insufficient expression of complex biological relationships, and inadequate connection between analysis results and health management recommendations, making it difficult to effectively utilize the microbial ecological network structure and basic clinical characteristics of the host.

Method used

A multi-source bioinformatics processing system based on the gut-oral axis was adopted to construct a microbial interaction network by acquiring oral microbiome, host basic clinical data and host metabolome data, extract features using graph neural network, and combine it with the Cox proportional hazards model for risk assessment and pharmaceutical intervention recommendations.

Benefits of technology

It achieves efficient integration and structured expression of multi-source information, improves the utilization of microbial ecological relationships, and enhances the interpretability of health risk assessment and the effectiveness of pharmaceutical intervention recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122392993A_ABST
    Figure CN122392993A_ABST
Patent Text Reader

Abstract

The present application belongs to the field of medical artificial intelligence, and specifically relates to a multi-source biological information processing and analysis system and method based on an intestinal mouth axis, which comprises: a data acquisition module that acquires oral microbiome data, host basic clinical data and host metabolome data of a subject; a preprocessing module that performs structured processing on the data; a microbial interaction network construction module that constructs a microbial interaction network based on the oral microbiome data; a graph feature extraction module that performs feature learning on the network using a graph neural network to obtain graph representation features; a risk assessment module that inputs the graph representation features, clinical features and metabolome features into a COX proportional risk model after fusion to generate a risk score; a correlation analysis module that matches and sorts the key microbial markers and the risk score with a preset correlation rule library and outputs suggestions; and a report generation module that generates an individualized analysis report. The present application realizes efficient fusion and structured analysis of multi-source biological information and provides data support for health status evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical artificial intelligence, specifically a multi-source biological information processing and analysis system and method based on the intestinal axis. Background Technology

[0002] The development and progression of colon-related diseases are closely related to microbial imbalance, host inflammatory status, metabolic abnormalities, and basic clinical characteristics. Existing bioinformatics-based analytical methods mostly rely on clinical indicators, biochemical test results, or single-agentine data for modeling, which suffers from problems such as limited information sources, insufficient expression of complex biological relationships, and inadequate connection between analytical results and subsequent health management recommendations.

[0003] In microbial research, the oral microbiome, as an important component of the digestive tract microecology, has a potential link with the gut microecological state. Oral samples are relatively easy to obtain, making them a potential entry point for non-invasive bioinformatics analysis. However, most existing technologies only perform statistical analyses of the relative abundance or diversity of microorganisms, failing to reflect the interactions and functional hierarchical characteristics among microorganisms. This results in microbiome-based analytical models underutilizing complex ecological information.

[0004] Furthermore, traditional analytical models typically treat various input variables as independent features, failing to fully integrate the microbial ecological network structure and rarely unifying and merging host basic clinical characteristics, host metabolomic characteristics, and microbiome characteristics. Even when some technologies can output status assessment results, they often lack a clear correlation mechanism with health management recommendations, making it difficult to form a complete technical chain from multi-source data analysis and status assessment to auxiliary information output.

[0005] Therefore, it is necessary to provide a multi-source bioinformatics processing and analysis system based on the gut-oral axis to improve the ability to integrate multi-source information, enhance the ability to express microbial ecological relationships, and achieve effective connection between bioinformatics analysis results and health management auxiliary recommendations. Summary of the Invention

[0006] The purpose of this invention is to provide a multi-source bioinformatics processing and analysis system and method based on the gut-oral axis, in order to solve the problems of insufficient accuracy in risk assessment of single-source data, inadequate integration of multi-source features, insufficient utilization of microbial network information, and missing pharmaceutical decision-making chain in the existing technology.

[0007] The technical solution adopted by the present invention to achieve the above objectives is: a multi-source bioinformatics processing and analysis system based on the gut-oral axis, comprising:

[0008] The data acquisition module is used to acquire oral microbiome data, host basic clinical data, and host metabolome data of the subjects;

[0009] The data preprocessing module is used to clean, standardize, extract features, and structure oral microbiome data, host basic clinical data, and host metabolome data.

[0010] A microbial interaction network construction module is used to construct a microbial interaction network based on processed oral microbiome data. The microbial interaction network is used to characterize the ecological relationships between different microbial characteristic units in the oral microbiome.

[0011] The graph feature extraction module is used to learn features of the microbial interaction network using a graph neural network model to obtain graph representation features that characterize the ecological relationships of microorganisms.

[0012] The risk assessment module is used to fuse the graphical representation features with the processed host basic clinical features and host metabolome features, and input the fused joint features into the COX proportional hazards model to generate a colon disease risk score.

[0013] The microbiome-drug interaction inference module is used to output pharmaceutical intervention assistance suggestions based on key microbial biomarkers extracted from oral microbiome data, oral microbial functional characteristics, host metabolome characteristics, and the health risk score information.

[0014] The report generation module is used to generate a personalized analysis report, which includes basic information of the subject, results of oral microbiome composition and function analysis, key microbial biomarker information, health risk score information and risk stratification results, pharmaceutical intervention suggestions and their basis information.

[0015] The oral microbiome data includes: oral microbiome composition characteristics and oral microbiome functional characteristics; the oral microbiome composition characteristics are the relative abundance characteristics of microorganisms at different taxonomic levels; the oral microbiome functional characteristics are the abundance characteristics of functional pathways related to microbial metabolism, inflammatory response, or immune regulation.

[0016] The host's basic clinical data includes one or more of the following: age, sex, body mass index, lifestyle information, and medical history information;

[0017] The host metabolomics data are the abundance characteristics of metabolites related to intestinal inflammation, immune regulation, or gut microbiota metabolism.

[0018] The risk assessment module includes:

[0019] The feature fusion unit is used to connect the graph representation feature vector, the host basic clinical feature vector, and the host metabolome feature vector in a feature splicing manner to form a joint feature vector;

[0020] The COX proportional hazards modeling unit is used to output a colon disease risk score based on the joint feature vector and to stratify subjects by risk according to a preset threshold or risk distribution range.

[0021] The microbial-drug interaction inference module includes:

[0022] The key microbial biomarker screening unit is used to screen key microbial biomarkers based on at least one of the following: microbial abundance difference characterization information, functional pathway abnormality characterization information, network structure importance characterization information, and graphical model contribution characterization information.

[0023] The association rule matching unit is used to perform matching analysis on key microbial biomarkers, oral microbial functional characteristics, host metabolome characteristics and health risk score information based on the microbiome-drug association rule base, and to prioritize candidate pharmaceutical intervention recommendations based on the colon disease risk score.

[0024] A multi-source bioinformatics processing and analysis method based on the gut-oral axis includes the following steps:

[0025] Step S1: Obtain oral microbiome data, host basic clinical data, and host metabolome data from the subjects;

[0026] The oral microbiome data is derived from oral sample sequencing results, the host basic clinical data is derived from electronic medical records or structured information, and the host metabolome data may be derived from metabolome testing results.

[0027] Step S2: Preprocess the oral microbiome data, host basic clinical data, and host metabolome data to generate structured oral microbiome composition characteristics, oral microbiome functional characteristics, host basic clinical characteristics, and host metabolome characteristics.

[0028] Step S3: Construct a microbial interaction network based on the compositional and functional characteristics of the oral microbiome;

[0029] The microbial interaction network uses microbial feature units as nodes and ecological relationships between microorganisms as edges to characterize the ecological relationships between different microbial feature units in the oral microbiome.

[0030] Step S4: Use a graph neural network model to learn features of the microbial interaction network to obtain graph representation features that characterize the ecological relationships of microorganisms;

[0031] Step S5: Multimodal fusion of the graphical representation features with the host's basic clinical features and the host's metabolome features to form a joint feature vector. The joint feature vector is then input into the Cox proportional hazards model to generate intermediate result information characterizing the degree of colon health risk. The intermediate result information is health risk score information.

[0032] Step S6: Based on key microbial biomarkers, oral microbial functional characteristics, host metabolome characteristics, and intermediate result information, perform matching analysis with a preset microbial-drug association rule base, prioritize candidate pharmaceutical intervention auxiliary suggestions according to the intermediate result information, and output association analysis results for constructing pharmaceutical intervention suggestions;

[0033] Step S7: Generate an individualized analysis report that integrates basic subject information, oral microbiome analysis results, key microbial biomarker information, intermediate results information, and correlation analysis results.

[0034] Step S2 specifically includes:

[0035] Step S2-1: For oral microbiome data, perform quality control, noise removal, species annotation, relative abundance calculation and functional pathway annotation in sequence to generate relative abundance features of microorganisms at different taxonomic levels as oral microbiome composition features, and generate functional pathway abundance features related to microbial metabolism, inflammatory response or immune regulation as oral microbial functional features.

[0036] Step S2-2: For the host's basic clinical data, perform missing value processing, variable encoding, category unification and standardization processing in sequence to generate one or more host basic clinical characteristics including age, gender, body mass index, lifestyle information and medical history information;

[0037] Steps S2-3: For the host metabolome data, perform metabolite identification, outlier handling, normalization and abundance standardization in sequence to generate metabolite abundance features related to intestinal inflammation, immune regulation or microbiota metabolism as host metabolome features.

[0038] Step S3 specifically includes:

[0039] Step S3-1: Construct a bipartite graph containing microbial nodes and functional pathway nodes, using microbial taxonomic units in the composition features of the processed oral microbiome as the first type of node and functional pathways in the functional features of the processed oral microbiome as the second type of node; or construct a monopartite graph using microbial taxonomic units as the single type of node.

[0040] Step S3-2: Calculate the association strength between microbial characteristic units, wherein the association strength is obtained in the following manner:

[0041] (a) Calculate the co-occurrence strength among microbial taxa using the Jaccard similarity coefficient or the SparCC correlation coefficient;

[0042] (b) Calculate the correlation strength between microbial taxa, using Pearson correlation coefficient or Spearman rank correlation coefficient, and perform significance screening through permutation test;

[0043] (c) Calculate the strength of functional associations between functional pathways based on shared enzymes, metabolites or known biological pathway associations in functional pathway databases;

[0044] (d) Calculate the association strength between microbial taxa and functional pathways, based on microbial function prediction results or known microbial-functional association databases;

[0045] Step S3-3: Set a correlation strength threshold, which includes the absolute value of the correlation coefficient, the co-occurrence strength value, or the functional coupling strength value. Only node pairs with a correlation strength greater than or equal to the threshold are retained as candidate edges. For correlation strength using a significance test, the corrected p-value is further required to be less than a preset significance level.

[0046] Step S3-4: Generate a microbial interaction network based on the determined set of nodes and edge connections; the network is a weighted undirected graph or a directed graph, where nodes represent microbial feature units, edges represent associations that satisfy the conditions, and the weight of the edge is set to the corresponding association strength value.

[0047] Step S3-5: Extract network features from the generated microbial interaction network, including one or more of node features, edge features, and network topology features, as microbial interaction network features.

[0048] Step S4 specifically includes:

[0049] Step S4-1: Construct a graph neural network model, wherein the graph neural network model is selected from one or more of graph convolutional networks, graph attention networks, or graph sampling aggregation networks;

[0050] Step S4-2: Input the adjacency matrix and node feature matrix of the microbial interaction network generated in step S3 into the graph neural network model, and optionally input the edge feature matrix at the same time;

[0051] Step S4-3: Update the node representation layer by layer through multi-layer propagation of the graph neural network model, where the propagation method of each layer is as follows:

[0052] a) If a graph convolutional network is used, the normalized adjacency matrix is ​​multiplied by the current layer node representation matrix, and then updated by a linear transformation and a nonlinear activation function;

[0053] b) If a graph attention network is used, the attention coefficient between each node and its neighboring nodes is calculated, the features of the neighboring nodes are weighted and aggregated based on the attention weights, and then fused through a multi-head attention mechanism.

[0054] c) If a graph sampling aggregation network is used, a fixed number of neighboring nodes are sampled, and the information of the sampled neighbors is aggregated using an aggregation function (mean aggregation, pooling aggregation, or long short-term memory network aggregation), and then updated after being concatenated with the features of the node itself.

[0055] Step S4-4: Extract multi-level graph representation features from the output of the graph neural network model. The multi-level graph representation features include one or more of node layer representation features, sub-layer representation features, and global graph representation features.

[0056] Step S4-5: Select the global graph representation features as graph representation feature vectors to characterize microbial ecological relationships, which will be used for subsequent fusion with host basic clinical features and host metabolome features.

[0057] Step S5 specifically includes:

[0058] Step S5-1: Perform dimension unification or standardization on the generated graph representation feature vector, the generated host basic clinical feature vector, and the host metabolome feature vector to make them all within the same dimension or numerical range.

[0059] Step S5-2: Using the feature splicing method, the graph representation feature vector, the host basic clinical feature vector, and the host metabolome feature vector are connected in a preset order to form a joint feature vector. The joint feature vector is used to characterize the comprehensive feature information of the subject in terms of microbial ecological structure, host basic clinical status, and host metabolic status.

[0060] Step S5-3: Input the joint feature vector into the pre-trained Cox proportional hazards model. The Cox proportional hazards model uses each dimension of the joint feature vector as a covariate and outputs intermediate result information that characterizes the degree of colon health risk of the subject.

[0061] Step S5-4: Based on the preset threshold or risk distribution range, perform risk stratification on the intermediate result information, classify the subjects into high-risk, medium-risk or low-risk categories, and output the risk stratification results.

[0062] Step S6 specifically includes:

[0063] Step S6-1: Based on at least one of the following: microbial abundance difference characterization information, functional pathway abnormality characterization information, network structure importance characterization information, and graph model contribution characterization information, the oral microbiome data is screened to determine key microbial biomarkers; the screening indicators include one or more of the following: fold change, significance level, node degree, centrality, betweenness centrality, graph representation feature contribution value, or risk correlation indicators;

[0064] Step S6-2: Match the key microbial biomarkers, oral microbial functional characteristics, host metabolome characteristics and intermediate results information with the rule items in the microbiome-drug association rule base item by item to obtain the correspondence between the candidate intervention directions and the abnormal microbial biomarkers, abnormal functional pathways, abnormal metabolites and risk stratification.

[0065] The rule items in the microbiome-drug association rule base include at least the association rules between abnormal microbial biomarkers and intervention directions, the association rules between abnormal functional pathways and intervention directions, the association rules between abnormal metabolites and intervention directions, and the association rules between risk stratification and intervention priority.

[0066] Step S6-3: Based on the risk level represented by the intermediate result information, and combined with one or more of the abnormality levels of key microbial markers, functional pathways, and metabolites, prioritize the corresponding relationships of the candidate intervention directions.

[0067] Step S6-4: Output the association analysis results after priority sorting for use in constructing pharmaceutical intervention recommendations.

[0068] The present invention has the following beneficial effects and advantages:

[0069] 1. This invention improves the ability to integrate multi-source information by simultaneously introducing oral microbiome data, host basic clinical data, and host metabolome data, thus overcoming the problem of insufficient information dimensions in single-source data modeling.

[0070] 2. This invention constructs a microbial interaction network and extracts node features, edge features, and network topology features, thereby enabling a structured expression of the ecological relationships between microorganisms and improving the utilization of microbial ecological information.

[0071] 3. This invention uses node features, edge features, network topology features, and joint feature vectors as microbial ecological relationship representation information and risk modeling input indicators, respectively, so that complex microbial network information, host basic clinical information, and host metabolic information can be uniformly represented and quantitatively utilized, thereby improving the structuring degree and interpretability of the risk assessment process.

[0072] 4. This invention uses graph neural networks to learn features of microbial interaction networks, which can extract graph representation features that characterize the ecological relationships of microorganisms, thereby improving the expressive power of complex microbial network features.

[0073] 5. This invention integrates graph representation features, host basic clinical features, and host metabolome features by employing feature splicing, and inputs these features into the Cox proportional hazards model, thereby achieving an effective connection between complex network representation and risk modeling and improving the stability of health risk scoring information and risk stratification results.

[0074] 6. This invention constructs a technical chain from risk analysis to the output of correlation analysis results for intervention direction by using rule matching based on key microbial biomarkers, oral microbial functional characteristics, host metabolomics characteristics and risk scores, thereby enhancing the ability to output auxiliary information.

[0075] 7. This invention improves the deliverability and ease of application of analysis results by generating personalized analysis reports that include basic information, microbial community analysis results, key microbial biomarker information, risk analysis results, and the corresponding relationships of candidate intervention directions. Attached Figure Description

[0076] Figure 1 Overall structural block diagram of the system of the present invention;

[0077] Figure 2 A flowchart illustrating the multi-source bioinformatics processing and analysis method for the intestinal axis of the present invention;

[0078] Figure 3 A schematic diagram illustrating the multi-source fusion of oral microbiome data, host basic clinical data, and host metabolome data in this invention;

[0079] Figure 4 A schematic diagram of the microbial interaction network construction and graph feature extraction process of the present invention. Detailed Implementation

[0080] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0081] Example 1:

[0082] like Figure 1 As shown, this invention provides a multi-source bioinformatics processing and analysis system based on the gut-oral axis for intelligent analysis of colonic diseases and pharmaceutical decision support, including a data acquisition module, a data preprocessing module, a microbial interaction network construction module, a graph feature extraction module, a risk assessment module, a microbial-drug interaction inference module, and a report generation module.

[0083] The data acquisition module is used to acquire oral microbiome data, host basic clinical data, and host metabolome data from the subjects. Oral microbiome data may be derived from sequencing results of saliva samples, oral swab samples, or other oral samples. Host basic clinical data may include one or more of the following: age, sex, body mass index, smoking status, alcohol consumption, dietary habits, past medical history, and family history. Host metabolome data may include one or more of the following: short-chain fatty acid-related metabolites, bile acid-related metabolites, tryptophan metabolism-related metabolites, and amino acid metabolism-related metabolites.

[0084] The data preprocessing module is used to preprocess the acquired data. For oral microbiome data, sequence quality control, noise reduction, species annotation, relative abundance statistics, and functional pathway annotation can be performed sequentially; for host basic clinical data, missing value processing, variable encoding, and standardization can be performed; for host metabolome data, normalization, outlier processing, and metabolite abundance standardization can be performed.

[0085] The microbial interaction network construction module constructs a microbial interaction network based on the compositional and functional characteristics of the oral microbiome. Network nodes can represent different microbial characteristic units, and node attributes may include relative abundance, functional pathway strength, and their combined encoding information; network edges can represent co-occurrence relationships, correlation relationships, or functional association relationships. Further calculations on the network yield node features, edge features, and network topology features. These node features, edge features, and network topology features collectively constitute the characterization information of the microbial ecological structure. One or more of the following characteristics—node degree, centrality, betweenness centrality, proximity centrality, clustering coefficient, modularity, and community structure—can serve as indicators for evaluating the importance of microbial nodes, network connectivity, and the degree of modular clustering.

[0086] The graph feature extraction module employs a graph neural network model to learn features from the microbial interaction network. The graph neural network model can be one or more of graph convolutional networks, graph attention networks, and graph sampling aggregation networks. After learning from the graph model, node layer representation features, sublayer representation features, and global graph representation features can be obtained.

[0087] The risk assessment module includes a feature fusion unit and a Cox proportional hazards modeling unit. The feature fusion unit connects graphical representation feature vectors, host basic clinical feature vectors, and host metabolome feature vectors in a preset order to form a joint feature vector. The Cox proportional hazards modeling unit receives the joint feature vector and outputs a colonic disease risk score. Based on preset thresholds or stratification intervals, subjects can be further stratified into high-risk, intermediate-risk, and low-risk groups.

[0088] The microbiome-drug interaction inference module receives key microbial biomarkers, oral microbiome functional characteristics, host metabolome characteristics, and health risk score information. To improve the reproducibility of key microbial biomarker screening, the key microbial biomarkers can be extracted based on preset screening indicators. These screening indicators may include one or more of the following: abundance difference fold, significance level, network topology importance index, graph model contribution index, and correlation index with risk score. The screening of key microbial biomarkers can be completed based on at least one of the following methods:

[0089] First, based on the results of the difference analysis between the disease group and the control group, we screened for significantly different microbial characteristics;

[0090] Secondly, important nodes are screened based on node degree, centrality, betweenness centrality, or other topological indicators in the microbial interaction network;

[0091] Third, key microbial characteristics are screened based on the importance representation results that contribute significantly to risk outcomes in the graph neural network output.

[0092] The association rule matching unit calls upon the microbiome-drug association rule base to perform matching analysis on the above input features. The rule items in the rule base may include: association rules corresponding to intervention directions when a key microbial biomarker is abnormal; association rules corresponding to intervention directions when a functional pathway is abnormal; association rules corresponding to intervention directions when a metabolite is abnormal; and priority rules corresponding to risk scores falling within a specific risk stratification interval. After association rule matching, candidate intervention direction correspondences are formed, and priority is ranked according to health risk score information.

[0093] The report generation module integrates the analysis results to generate a personalized analysis report. The report includes at least the subject's basic information, oral microbiome composition and function analysis results, key microbial biomarker information, health risk score information and risk stratification results, and the corresponding relationships of candidate intervention directions and their basis.

[0094] Example 2: Implementation of Method Steps

[0095] like Figure 2 As shown, this invention provides a multi-source bioinformatics processing method based on the gut-oral axis, comprising the following steps:

[0096] Step S1: Obtain multi-source data

[0097] To obtain oral microbiome data, host basic clinical data, and host metabolome data from the subjects.

[0098] Oral microbiome data are obtained through high-throughput sequencing of saliva or oral swab samples; host basic clinical data are obtained through electronic medical record systems or structured acquisition methods, including one or more of age, sex, body mass index, lifestyle information, and medical history; host metabolome data are obtained by detecting blood, urine, or fecal samples using liquid chromatography-mass spectrometry or gas chromatography-mass spectrometry.

[0099] Step S2: Data Preprocessing

[0100] Step S2-1: Preprocessing of oral microbiome data. This involves sequentially performing quality control, noise removal, species annotation, relative abundance calculation, and functional pathway annotation. Quality control includes quality assessment and quality trimming; noise removal uses the DADA2 or Deblur algorithm to obtain amplicon sequence variant feature tables; species annotation assigns classification labels from kingdom to species based on a reference database; relative abundance calculation uses summation normalization to obtain microbial relative abundance matrices at different taxonomic levels; functional pathway annotation uses the PICRUSt2 tool to output functional pathway abundance matrices related to microbial metabolism, inflammatory responses, or immune regulation.

[0101] Step S2-2: Preprocess the host's basic clinical data. This involves sequentially performing missing value handling, variable encoding, category unification, and standardization. Missing value handling involves deletion, imputation, or labeling as an independent category based on the missing value rate. Variable encoding includes label encoding, ordinal encoding, or one-hot encoding. Category unification includes unit unification, classification mapping, and value range verification. Standardization uses Z-score normalization or Min-Max normalization to generate the host's basic clinical feature matrix.

[0102] The Z-score standardization formula is:

[0103]

[0104] in, These are the original eigenvalues. The characteristic mean, Standard deviation These are the standardized feature values.

[0105] Step S2-3: Preprocessing of host metabolome data. This involves sequentially performing metabolite identification, outlier handling, normalization, and abundance standardization. Metabolite identification is achieved through peak detection, peak alignment, and database comparison; outlier handling identifies abnormal samples at the sample level using principal component analysis and outliers at the metabolite level using box plots; normalization employs total peak area normalization, internal standard normalization, or median normalization; abundance standardization uses logarithmic transformation combined with Z-score normalization or Pareto scaling to generate a metabolome feature matrix.

[0106] Abundance standardization employs a combination of logarithmic transformation and Z-score standardization. The logarithmic transformation formula is as follows: .in, This is the original abundance value. It is a smoothing constant greater than 0. These are the abundance values ​​after logarithmic transformation. For features with a standard deviation of 0, the original value can be retained or set to a preset constant; the Z-score standardization is preferentially applied to continuous variables.

[0107] Step S3: Constructing a microbial interaction network

[0108] Step S3-1: Construct a monopart graph using microbial taxonomic units in the composition characteristics of the processed oral microbiome as nodes; or construct a bipart graph using microbial taxonomic units and functional pathways as nodes.

[0109] Step S3-2: Calculate the association strength between microbial characteristic units. This can be done using at least one of the following methods:

[0110] (a) Calculate the co-occurrence strength between microbial taxa using the Jaccard similarity coefficient or the SparCC correlation coefficient. The formula for the Jaccard similarity coefficient is:

[0111]

[0112] in, and These represent the sets of two microbial taxonomic units present in the sample. This represents the number of samples where both exist simultaneously. This indicates that at least one of the two has a sample number. This represents the co-occurrence intensity value.

[0113] (b) Calculate the correlation strength between microbial taxa using the Pearson correlation coefficient or Spearman rank correlation coefficient. The formula for the Pearson correlation coefficient is:

[0114]

[0115] in, and These are two microbial taxonomic units in the first... Relative abundance in each sample, and These are their means, The total number of samples, The correlation coefficient is defined as a value ranging from -1 to 1. Significance is determined using a permutation test.

[0116] (c) Calculate the strength of functional associations between functional pathways based on shared enzymes, metabolites or known biological pathway associations in functional pathway databases;

[0117] (d) Calculate the association strength between microbial taxa and functional pathways, based on microbial function prediction results or known microbial-functional association databases.

[0118] Step S3-3: Set an association strength threshold and retain node pairs with association strength greater than or equal to the threshold as candidate edges; for association strength using significance testing, the corrected p-value is further required to be less than the preset significance level.

[0119] Step S3-4: Generate a microbial interaction network based on the determined set of nodes and edge connections. The network is a weighted undirected graph or a directed graph, where nodes represent microbial feature units, edges represent associations that satisfy certain conditions, and the weight of the edges is set to the corresponding association strength value.

[0120] Step S3-5: Extract network features from the generated microbial interaction network, including one or more of node features, edge features, and network topology features, as microbial interaction network features. Node features include relative abundance of microorganisms, functional pathway strength, or node attribute information; edge features include correlation strength, co-occurrence strength, or functional coupling strength; network topology features include one or more of node degree, betweenness centrality, clustering coefficient, modularity, or community structure features.

[0121] Step S4: Feature Learning of Graph Neural Network

[0122] Step S4-1: Construct a graph neural network model, wherein the graph neural network model is selected from one or more of graph convolutional networks, graph attention networks, or graph sampling aggregation networks.

[0123] Step S4-2: Input the adjacency matrix and node feature matrix of the microbial interaction network generated in step S3 into the graph neural network model, and optionally input the edge feature matrix at the same time.

[0124] Step S4-3: Update the node representation layer by layer through multi-layer propagation of the graph neural network model.

[0125] If a Graph Convolutional Network (GCN) is used, the propagation method is as follows:

[0126]

[0127] in, To add a self-loop adjacency matrix, This is the original adjacency matrix. It is the identity matrix; for The degree matrix, ; For the first The node representation matrix of the layer, The input node feature matrix; For the first The learnable weight matrix of the layer; It is a non-linear activation function (such as ReLU); For the first The node representation matrix of the layer.

[0128] If a graph attention network (GAT) is used, the formula for calculating the attention coefficient is:

[0129]

[0130] in, and They are nodes and nodes eigenvectors; It is a linear transformation matrix; This is a learnable attention vector; This represents a vector concatenation operation; For nodes The set of neighboring nodes; For nodes For nodes The attention coefficients are denoted by ; LeakyReLU is the activation function.

[0131] If a graph sampling aggregation network (GraphSAGE) is used, the node update method is as follows:

[0132] ;

[0133]

[0134] in, For nodes In the The feature vector of the layer; For nodes The set of sampled neighbor nodes; This is an aggregation function (mean aggregation, pooling aggregation, or long short-term memory network aggregation) used to aggregate features of neighboring nodes; The features of the aggregated neighbors; For the first The learnable weight matrix of the layer; CONCAT is a vector concatenation operation; It is a non-linear activation function; For nodes In the The updated feature vectors after the layer.

[0135] Step S4-4: Extract multi-level graph representation features from the output of the graph neural network model, including one or more of node layer representation features, sub-layer representation features, and global graph representation features.

[0136] Step S4-5: Select the global graph representation features as graph representation feature vectors to characterize microbial ecological relationships, which will be used for subsequent fusion with host basic clinical features and host metabolome features.

[0137] Step S5: Feature Fusion and Risk Assessment

[0138] Step S5-1: Perform dimension unification or standardization on the generated graph representation feature vector, host basic clinical feature vector, and host metabolome feature vector to make them fall within the same dimension or numerical range.

[0139] Step S5-2: Using a feature splicing method, the three feature vectors are connected in a preset order to form a joint feature vector. The joint feature vector is used to characterize the comprehensive feature information of the subject in terms of microbial ecological structure, host basic clinical status and host metabolic status.

[0140] Step S5-3: Input the joint feature vector into the pre-trained Cox proportional hazards model. The Cox proportional hazards model uses each dimension of the joint feature vector as covariates and outputs intermediate result information that characterizes the degree of health risk of the subject. The intermediate result information can be represented in the form of a health risk score.

[0141] The risk function of the Cox proportional hazards model is:

[0142]

[0143] in, It is a time variable; Joint eigenvectors; Here, represents the baseline risk function, indicating the risk function when all covariates are zero. For the regression coefficient vector, It is a linear combination; This is the risk proportionality factor. The model outputs intermediate results representing the degree of health risk of the subjects, namely the risk score:

[0144]

[0145] in, The dimension of the joint feature vector. For the first The regression coefficients of each feature, For the first Each feature value.

[0146] Step S5-4: Based on the preset threshold or risk distribution range, perform risk stratification on the intermediate result information, classify the subjects into high-risk, medium-risk or low-risk categories, and output the risk stratification results.

[0147] Step S6: Association rule matching and priority sorting

[0148] Step S6-1: Screen key microbial biomarkers based on at least one of the following: microbial abundance difference characterization information, functional pathway abnormality characterization information, network structure importance characterization information, and graphical model contribution characterization information. Screening indicators include one or more of the following: fold change, significance level, nodality, centrality, betweenness centrality, graphical representation feature contribution value, or risk relevance indicators.

[0149] Step S6-2: Construct a microbiome-drug association rule base. This rule base is established through literature evidence compilation, database knowledge extraction, and clinical experience rule summarization. The rule items include association rules between abnormal microbial biomarkers and intervention directions, abnormal functional pathways and intervention directions, abnormal metabolites and intervention directions, and risk stratification and intervention priority association rules. Key microbial biomarkers, oral microbiome functional characteristics, host metabolome characteristics, and intermediate result information are matched item by item with the rule items in the rule base to obtain the correspondence between candidate intervention directions.

[0150] Step S6-3: Based on the risk level represented by the intermediate result information, and combined with one or more of the abnormality levels of key microbial markers, functional pathways, and metabolites, prioritize the candidate intervention direction correspondences to generate a ranked candidate intervention direction correspondence list.

[0151] The priority score calculation formula is as follows:

[0152]

[0153] in, Priority scores for candidate intervention recommendations; The pre-defined base weights in the rule base reflect the clinical importance of the intervention recommendation; The normalized health risk score information of the subjects is used to characterize the overall risk level; The normalized anomaly level is used to characterize the degree to which the abnormal indicator that triggers the rule deviates from the normal range. The calculation formula is the average deviation of the abnormal indicator from the normal value. , , Let be the weighting coefficient, satisfying The parameters are set according to the clinical application scenario. The basic weights, health risk score information, and abnormality degree are normalized or standardized before calculating the priority score to ensure that different indicators are within a comparable numerical range.

[0154] Step S6-4: Output the association analysis results after priority sorting. The association analysis results are used to construct auxiliary suggestion information.

[0155] Step S7: Generate a personalized analysis report

[0156] The report generation module integrates the above analysis results to generate a personalized analysis report. The report includes basic information about the subjects, results of oral microbiome composition and function analysis, key microbial biomarkers, risk scores and risk stratification results, and candidate auxiliary suggestions and their basis.

[0157] like Figure 3 As shown, Figure 3 This diagram illustrates the multi-source fusion of oral microbiome data, host basic clinical data, and host metabolome data described in this invention, used to represent the relationship between multi-source feature input, joint feature formation, risk scoring, and auxiliary suggestion output.

[0158] This invention collects oral microbiome data, host basic clinical data, and host metabolome data through a data acquisition module, and forms structured features after preprocessing. Specifically, the oral microbiome data, after microbial interaction network construction and graph neural network feature learning, generates a global graph representation feature vector; the host basic clinical data, after preprocessing, generates a clinical feature vector; and the host metabolome data, after preprocessing, generates a metabolome feature vector. These three feature vectors are concatenated by a feature fusion unit to form a joint feature vector, which simultaneously contains information on microbial ecological structure, host basic clinical status, and host metabolic status. The joint feature vector is input into a Cox proportional hazards model, outputting a risk score, which is further stratified according to preset thresholds. Simultaneously, the risk score, along with key microbial biomarkers, abnormal functional pathways, and abnormal metabolites, serves as input to an association rule matching unit. After rule base matching and priority ranking, candidate auxiliary suggestions are output. The above process fully presents the information processing chain from multi-source data input to risk score output and then to auxiliary suggestion generation.

[0159] like Figure 4 As shown, Figure 4 The diagram illustrates the process of constructing a microbial interaction network and extracting graph features according to the present invention. It is used to illustrate the process of constructing a microbial interaction network from the composition features and functional features of the oral microbiome, and extracting graph representation features through a graph neural network.

[0160] like Figure 4As shown, the process of constructing a microbial interaction network and extracting graph features includes the following steps: First, based on the preprocessed composition and functional characteristics of the oral microbiome, a microbial interaction network is constructed. Network nodes represent microbial taxonomic units or functional pathways, and edges represent co-occurrence, correlation, or functional relationships between nodes. After network construction, node features, edge features, and network topology features are extracted. These features collectively constitute the microbial interaction network features.

[0161] Subsequently, the adjacency matrix and node feature matrix of the microbial interaction network are input into the graph neural network model. The graph neural network aggregates neighbor node information layer by layer through multi-layer propagation, updating the node representation. Finally, a global graph representation feature vector is extracted from the output of the graph neural network. This vector serves as a compressed representation of microbial ecological relationships and is used for subsequent fusion with clinical features and metabolomics features. Figure 4 It clearly demonstrates the entire transformation process from raw microbiome data to structured graph representation features.

[0162] In summary, and in conjunction with the embodiments, this invention provides a multi-source bioinformatics processing and analysis system and method based on the gut-oral axis. By constructing a multi-source information fusion architecture integrating oral microbiome data, host basic clinical data, and host metabolome data, it utilizes microbial interaction networks and graph neural networks to extract microbial ecological structure features, and combines this with a Cox proportional hazards model for risk scoring and stratification. Finally, it outputs auxiliary suggestion information through association rule matching. This invention realizes a complete technical chain from multi-source heterogeneous data input to structured analysis result output, and features strong information integration capabilities, sufficient expression of microbial ecological relationships, and high interpretability of analysis results.

[0163] Those skilled in the art will understand that the above description is merely a preferred embodiment of the present invention, and the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. This is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0164] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.

Claims

1. A multi-source bioinformatics processing and analysis system based on the gut-oral axis, characterized in that, include: The data acquisition module is used to acquire oral microbiome data, host basic clinical data, and host metabolome data of the subjects; The data preprocessing module is used to clean, standardize, extract features, and structure oral microbiome data, host basic clinical data, and host metabolome data. A microbial interaction network construction module is used to construct a microbial interaction network based on processed oral microbiome data. The microbial interaction network is used to characterize the ecological relationships between different microbial characteristic units in the oral microbiome. The graph feature extraction module is used to learn features of the microbial interaction network using a graph neural network model to obtain graph representation features that characterize the ecological relationships of microorganisms. The risk assessment module is used to fuse the graphical representation features with the processed host basic clinical features and host metabolome features, and input the fused joint features into the COX proportional hazards model to generate a colon disease risk score. The microbiome-drug interaction inference module is used to output pharmaceutical intervention suggestions based on key microbial biomarkers extracted from oral microbiome data, oral microbial functional characteristics, host metabolome characteristics, and the colonic disease risk score. The report generation module is used to generate a personalized analysis report, which includes basic information of the subject, results of oral microbiome composition and function analysis, key microbial biomarker information, health risk score information and risk stratification results, pharmaceutical intervention suggestions and their basis information.

2. The multi-source bioinformatics processing and analysis system based on the gut-oral axis according to claim 1, characterized in that, The oral microbiome data includes: oral microbiome composition characteristics and oral microbiome functional characteristics; the oral microbiome composition characteristics are the relative abundance characteristics of microorganisms at different taxonomic levels; the oral microbiome functional characteristics are the abundance characteristics of functional pathways related to microbial metabolism, inflammatory response, or immune regulation. The host's basic clinical data includes one or more of the following: age, sex, body mass index, lifestyle information, and medical history information; The host metabolomics data are the abundance characteristics of metabolites related to intestinal inflammation, immune regulation, or gut microbiota metabolism.

3. The multi-source bioinformatics processing and analysis system based on the gut-osteo-axis according to claim 1, characterized in that, The risk assessment module includes: The feature fusion unit is used to connect the graph representation feature vector, the host basic clinical feature vector, and the host metabolome feature vector in a feature splicing manner to form a joint feature vector; The COX proportional hazards modeling unit is used to output a colon disease risk score based on the joint feature vector and to stratify subjects by risk according to a preset threshold or risk distribution range.

4. The multi-source bioinformatics processing and analysis system based on the gut-oral axis according to claim 1, characterized in that, The microbial-drug interaction inference module includes: The key microbial biomarker screening unit is used to screen key microbial biomarkers based on at least one of the following: microbial abundance difference characterization information, functional pathway abnormality characterization information, network structure importance characterization information, and graphical model contribution characterization information. The association rule matching unit is used to perform matching analysis on key microbial biomarkers, oral microbial functional characteristics, host metabolome characteristics, and colonic disease risk scores based on the microbiome-drug association rule base, and to prioritize candidate pharmaceutical intervention recommendations based on the colonic disease risk scores.

5. A multi-source bioinformatics processing and analysis method based on the gut-oral axis, using the system described in any one of claims 1-4, characterized in that, Includes the following steps: Step S1: Obtain oral microbiome data, host basic clinical data, and host metabolome data from the subjects; The oral microbiome data is derived from oral sample sequencing results, the host basic clinical data is derived from electronic medical records or structured information, and the host metabolome data may be derived from metabolome testing results. Step S2: Preprocess the oral microbiome data, host basic clinical data, and host metabolome data to generate structured oral microbiome composition characteristics, oral microbiome functional characteristics, host basic clinical characteristics, and host metabolome characteristics. Step S3: Construct a microbial interaction network based on the compositional and functional characteristics of the oral microbiome; The microbial interaction network uses microbial feature units as nodes and ecological relationships between microorganisms as edges to characterize the ecological relationships between different microbial feature units in the oral microbiome. Step S4: Use a graph neural network model to learn features of the microbial interaction network to obtain graph representation features that characterize the ecological relationships of microorganisms; Step S5: Multimodal fusion of the graphical representation features with the host's basic clinical features and the host's metabolome features to form a joint feature vector, and inputting the joint feature vector into the Cox proportional hazards model to generate intermediate result information characterizing the degree of colon health risk; Step S6: Based on key microbial biomarkers, oral microbial functional characteristics, host metabolome characteristics, and intermediate result information, perform matching analysis with a preset microbial-drug association rule base, prioritize candidate pharmaceutical intervention auxiliary suggestions according to the intermediate result information, and output association analysis results for constructing pharmaceutical intervention suggestions; Step S7: Generate an individualized analysis report that integrates basic subject information, oral microbiome analysis results, key microbial biomarker information, intermediate results information, and correlation analysis results.

6. The multi-source bioinformatics processing and analysis method based on the gut-oral axis according to claim 5, characterized in that, Step S2 specifically includes: Step S2-1: For oral microbiome data, perform quality control, noise removal, species annotation, relative abundance calculation and functional pathway annotation in sequence to generate relative abundance features of microorganisms at different taxonomic levels as oral microbiome composition features, and generate functional pathway abundance features related to microbial metabolism, inflammatory response or immune regulation as oral microbial functional features. Step S2-2: For the host's basic clinical data, perform missing value processing, variable encoding, category unification and standardization processing in sequence to generate one or more host basic clinical characteristics including age, gender, body mass index, lifestyle information and medical history information; Steps S2-3: For the host metabolome data, perform metabolite identification, outlier handling, normalization and abundance standardization in sequence to generate metabolite abundance features related to intestinal inflammation, immune regulation or microbiota metabolism as host metabolome features.

7. The multi-source bioinformatics processing and analysis method based on the intestinal axis according to claim 5, characterized in that, Step S3 specifically includes: Step S3-1: Construct a bipartite graph containing microbial nodes and functional pathway nodes, using microbial taxonomic units in the composition features of the processed oral microbiome as the first type of node and functional pathways in the functional features of the processed oral microbiome as the second type of node; or construct a monopartite graph using microbial taxonomic units as the single type of node. Step S3-2: Calculate the association strength between microbial characteristic units, wherein the association strength is obtained in the following manner: (a) Calculate the co-occurrence strength among microbial taxa using the Jaccard similarity coefficient or the SparCC correlation coefficient; (b) Calculate the correlation strength between microbial taxa, using Pearson correlation coefficient or Spearman rank correlation coefficient, and perform significance screening through permutation test; (c) Calculate the strength of functional associations between functional pathways based on shared enzymes, metabolites or known biological pathway associations in functional pathway databases; (d) Calculate the association strength between microbial taxa and functional pathways, based on microbial function prediction results or known microbial-functional association databases; Step S3-3: Set a correlation strength threshold, which includes the absolute value of the correlation coefficient, the co-occurrence strength value, or the functional coupling strength value. Only node pairs with a correlation strength greater than or equal to the threshold are retained as candidate edges. For correlation strength using a significance test, the corrected p-value is further required to be less than a preset significance level. Step S3-4: Generate a microbial interaction network based on the determined set of nodes and edge connections; the network is a weighted undirected graph or a directed graph, where nodes represent microbial feature units, edges represent associations that satisfy the conditions, and the weight of the edge is set to the corresponding association strength value. Step S3-5: Extract network features from the generated microbial interaction network, including one or more of node features, edge features, and network topology features, as microbial interaction network features.

8. The multi-source bioinformatics processing and analysis method based on the gut-oral axis according to claim 5, characterized in that, Step S4 specifically includes: Step S4-1: Construct a graph neural network model, wherein the graph neural network model is selected from one or more of graph convolutional networks, graph attention networks, or graph sampling aggregation networks; Step S4-2: Input the adjacency matrix and node feature matrix of the microbial interaction network generated in step S3 into the graph neural network model, and optionally input the edge feature matrix at the same time; Step S4-3: Update the node representation layer by layer through multi-layer propagation of the graph neural network model, where the propagation method of each layer is as follows: a) If a graph convolutional network is used, the normalized adjacency matrix is ​​multiplied by the current layer node representation matrix, and then updated by a linear transformation and a nonlinear activation function; b) If a graph attention network is used, the attention coefficient between each node and its neighboring nodes is calculated, the features of the neighboring nodes are weighted and aggregated based on the attention weights, and then fused through a multi-head attention mechanism. c) If a graph sampling aggregation network is used, a fixed number of neighboring nodes are sampled, and the information of the sampled neighbors is aggregated using an aggregation function (mean aggregation, pooling aggregation, or long short-term memory network aggregation), and then updated after being concatenated with the features of the node itself. Step S4-4: Extract multi-level graph representation features from the output of the graph neural network model. The multi-level graph representation features include one or more of node layer representation features, sub-layer representation features, and global graph representation features. Step S4-5: Select the global graph representation features as graph representation feature vectors to characterize microbial ecological relationships, which will be used for subsequent fusion with host basic clinical features and host metabolome features.

9. A multi-source bioinformatics processing and analysis method based on the gut-oral axis according to claim 5, characterized in that, Step S5 specifically includes: Step S5-1: Perform dimension unification or standardization on the generated graph representation feature vector, the generated host basic clinical feature vector, and the host metabolome feature vector to make them all within the same dimension or numerical range. Step S5-2: Using the feature splicing method, the graph representation feature vector, the host basic clinical feature vector, and the host metabolome feature vector are connected in a preset order to form a joint feature vector. The joint feature vector is used to characterize the comprehensive feature information of the subject in terms of microbial ecological structure, host basic clinical status, and host metabolic status. Step S5-3: Input the joint feature vector into the pre-trained Cox proportional hazards model. The Cox proportional hazards model uses each dimension of the joint feature vector as a covariate and outputs intermediate result information that characterizes the degree of colon health risk of the subject. Step S5-4: Based on the preset threshold or risk distribution range, perform risk stratification on the intermediate result information, classify the subjects into high-risk, medium-risk or low-risk categories, and output the risk stratification results.

10. A multi-source bioinformatics processing and analysis method based on the gut-oral axis according to claim 5, characterized in that, Step S6 specifically includes: Step S6-1: Screen the oral microbiome data based on at least one of the following: microbial abundance difference characterization information, functional pathway abnormality characterization information, network structure importance characterization information, and graphical model contribution characterization information, and identify key microbial biomarkers; The screening indicators include one or more of the following: fold difference, significance level, node degree, centrality, betweenness centrality, graph representation feature contribution value, or risk correlation indicators; Step S6-2: Match the key microbial biomarkers, oral microbial functional characteristics, host metabolome characteristics and intermediate results information with the rule items in the microbiome-drug association rule base item by item to obtain the correspondence between the candidate intervention directions and the abnormal microbial biomarkers, abnormal functional pathways, abnormal metabolites and risk stratification. The rule items in the microbiome-drug association rule base include at least the association rules between abnormal microbial biomarkers and intervention directions, the association rules between abnormal functional pathways and intervention directions, the association rules between abnormal metabolites and intervention directions, and the association rules between risk stratification and intervention priority. Step S6-3: Based on the risk level represented by the intermediate result information, and combined with one or more of the abnormality levels of key microbial markers, functional pathways, and metabolites, prioritize the corresponding relationships of the candidate intervention directions. Step S6-4: Output the association analysis results after priority sorting for use in constructing pharmaceutical intervention recommendations.