Disease Prediction Method and System Based on Adaptive Fusion of Microbiome Stratification Features
Through the method of hierarchical feature extraction and multi-scale feature fusion, combined with the dynamic sampling module, the problems of insufficient utilization of hierarchical information and sample imbalance in the disease prediction of microbiome data are solved, and high accuracy and stable disease prediction are achieved.
Patent Information
- Application Number
- CN202510361147.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The existing microbiome data disease prediction methods cannot fully utilize the microbial classification hierarchical information, have limited feature extraction capabilities, and are difficult to capture the structured patterns and environmentally dependent characteristics of microbiome data. In addition, the problem of sample imbalance is insufficiently considered in multiple disease classification scenarios, resulting in insufficient diagnostic accuracy and stability.
The disease prediction method based on adaptive fusion of microbial components is adopted. The feature data at five classification levels is extracted in multiple levels through the hierarchical feature extraction module. Combined with the multi-scale feature fusion module and the dynamic sampling module, the sample sampling weight is dynamically adjusted, and the multi-dimensional uncertainty evaluation strategy is used to predict disease.
It significantly improves the accuracy and stability of disease prediction of microbiome data, can effectively capture the evolution and functional connections at the microbial classification level, enhances the all-round characterization of microbiome data, solves the problem of sample imbalance, and improves the generalization ability of the model.
Smart Images

Figure CN119889701B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microbiome data analysis, and particularly to a disease prediction method and system based on adaptive fusion of microbiome hierarchical features. Background Art
[0002] With the rapid development of high-throughput sequencing technology, significant progress has been made in microbiome research in recent years. As a microecosystem inside the human body, the microbial community is closely related to human health. In particular, the "gut-brain axis" relationship between the gut microbiome and the nervous system has been confirmed by multiple studies. Microbiome data analysis shows great potential and value in the field of nervous system disease diagnosis, providing new perspectives and methods for early diagnosis and prognosis assessment.
[0003] However, different from traditional clinical data, microbiome data has significant particularities, posing severe challenges to data analysis: Firstly, microbiome data is characterized by high dimensionality, high noise, and high sparsity, often containing thousands or even tens of thousands of features, while the sample size is usually small, which easily leads to model overfitting. Secondly, microbiome data exhibits obvious hierarchical structural characteristics. The taxonomic hierarchy from phylum, class, order, family to genus reflects the evolutionary relationship and functional characteristics of microorganisms. This structural information is of great value for disease diagnosis, but traditional analysis methods often cannot effectively utilize it. In addition, microbiome data also has multi-dimensional representation requirements. The characteristics of a microbial community can be interpreted from multiple perspectives such as individual characteristics, community structure patterns, and environmental dependence. Single-dimensional feature extraction is difficult to comprehensively capture the complex characteristics of the microbiome. Finally, in the multi-disease classification scenario, the samples of different diseases usually have an imbalance in quantity. How to balance the diagnostic accuracy of various diseases while maintaining the stability of the model is another major challenge in microbiome data analysis.
[0004] Existing disease prediction methods for microbiome data mainly include: machine learning-based classification methods such as random forest, support vector machine, gradient boosting decision tree, etc.; and emerging deep learning methods such as multi-layer perceptron, convolutional neural network, recurrent neural network, etc. However, these methods generally have the following limitations: First, they cannot fully utilize the hierarchical knowledge of microbial taxonomy and ignore the information complementarity between different taxonomic levels. Second, their feature extraction ability is limited, and it is difficult to capture the structured patterns, environmental dependence features, and individual characteristics of microbiome data simultaneously. Third, they do not consider the sample imbalance problem sufficiently, and their performance is unstable in the multi-disease classification scenario.
[0005] Therefore, there is an urgent need to develop a prediction method that can fully utilize the microbial taxonomic hierarchical information, achieve multi-dimensional feature representation, and effectively address the sample imbalance problem to improve the diagnostic accuracy and stability of nervous system diseases. Summary of the Invention
[0006] The object of the present invention is to solve the problems in the prior art.
[0007] The technical solution adopted by the present invention to solve its technical problems is: to provide a disease prediction method based on adaptive fusion of microbiome hierarchical features, including the following steps:
[0008] Preprocess multi-source microbiome data samples to obtain hierarchical feature data samples containing five levels of phylum, class, order, family, and genus;
[0009] Construct a disease prediction model, train it using the hierarchical feature data samples, and dynamically adjust the sample sampling weights during the training process by combining a class balance mechanism and a multi-dimensional uncertainty evaluation strategy;
[0010] Use the trained disease prediction model to predict diseases based on microbiome data;
[0011] In the disease prediction model, the hierarchical feature extraction module performs multi-level feature extraction on the hierarchical feature data of the five classification levels, and integrates them to obtain multi-level fusion features containing hierarchical information; the multi-scale feature fusion module processes the multi-level fusion features through three parallel paths, respectively outputs features containing structured information, environment-dependent information, and individual feature information, and then integrates the features output by the three paths through a dynamic weighted fusion mechanism to obtain multi-dimensional fusion features; the classifier predicts diseases based on the multi-scale fusion features.
[0012] Preferably, the process of the hierarchical feature extraction module obtaining multi-level fusion features includes the following steps:
[0013] The hierarchical feature extraction component extracts the features of the five levels and combines them to obtain a hierarchical feature representation , expressed as:
[0014] ;
[0015] ;
[0016] ;
[0017] Among them, , , , , respectively represent the features of the five levels of phylum, class, order, family, and genus; the feature of each level is a vector composed of the aggregated abundances of all taxonomic units at that level, The data matrix representing the input disease prediction model, represents the eigenvector corresponding to the taxonomic unit ; is the set composed of all descendant taxonomic units of the taxonomic unit , calculated by recursively using the child node mapping ; is a taxonomic unit element in which records the set of all descendants; records the set of all child taxonomic units with the taxonomic unit as the parent node, constructed by reverse mapping the parent-child relationship information in the mapping ; is the known data, recording the five-level and parent node information of each microbial taxonomic unit ;
[0018] The classification level weight component dynamically adjusts the feature weights to obtain the weighted feature representation , expressed as:
[0019] ;
[0020] ;
[0021] where X represents the sample data matrix; is the base weight of the level where the taxonomic unit is located, measures the discriminative ability of the feature, represents the sparsity degree of the feature, and are the balance parameters, represents the diagonal matrix with the adjusted weight vector as the diagonal elements;
[0022] The classification relationship modeling component constructs a multi-dimensional relationship matrix based on the classification path and evolutionary distance between features, and calculates the relationship feature representation , expressed as:
[0023] ;
[0024] ;
[0025] ;
[0026] where represents the common path length of the features and , and respectively represent reaching the feature and the complete classification path length; indicating feature and the evolutionary distance between, is a parameter for controlling the influence intensity of the distance;
[0027] Multilevel fusion features are obtained through feature fusion , expressed as:
[0028] ;
[0029] is a learnable fusion weight matrix, optimized during training by gradient descent.
[0030] Preferably, the multi-scale feature fusion module takes the multilevel fusion feature as the input feature , processes the multilevel fusion feature through three parallel paths, and finally calculates the multi-dimensional fusion feature, including the following steps:
[0031] The pattern flow path processes the input feature using the correlation attention mechanism to obtain the pattern flow output feature , expressed as:
[0032] ;
[0033] ;
[0034] ;
[0035] wherein, , , are respectively learnable weight matrices for query, key, and value, , , respectively represent the learned query, key, and value projections, is a scaling factor for ensuring the stability of the gradient flow during training;
[0036] The context flow path uses the global scope processing module to establish long-range dependencies by calculating the similarity matrix between input features, and calculates the upstream and downstream output features through the feed-forward neural network of the context flow , expressed as:
[0037] ;
[0038] ;
[0039] wherein, Represents a global attention operation, and respectively represent two fully connected layer operations, is an activation function, is an operation to prevent overfitting;
[0040] The content flow path processes the input features through a similarity mapping mechanism and a channel focusing module to obtain content flow output features , which is expressed as:
[0041] ;
[0042] ;
[0043] Among them, and are learnable weight matrices, is an activation function, is an attention mapping matrix, represents a channel focusing operation;
[0044] Calculate the importance weights of each path, and then perform weighted combination on the output features of each path based on the weights to obtain multi-dimensional fusion features , which is expressed as:
[0045] ;
[0046] ;
[0047] Among them, represents the learnable fusion weights corresponding to the three paths, represents the learnable weight vectors corresponding to the three paths, represents a global pooling operation, represents the output features corresponding to the three paths; represents the pattern flow path, represents the up and down flow path, represents the content flow path.
[0048] Preferably, disease prediction is performed on the multi-dimensional fusion features through a classifier. Specifically, disease prediction is performed on the multi-scale fusion features through a random forest classifier, and the classification results and confidence levels are output.
[0049] Preferably, the classification mode can be configured as binary classification or multi-class classification according to actual prediction needs;
[0050] In the binary classification setting, the classification results are distinguished into disease states and non-disease states;
[0051] In a multi-classification setting, the classification results can be recognized and distinguished into multiple different disease types.
[0052] Preferably, the dynamic adjustment of the sample sampling weight is implemented by a dynamic sampling module, including the following steps:
[0053] Calculate the base weight based on the class balance mechanism, expressed as:
[0054] ;
[0055] ;
[0056] ;
[0057] ;
[0058] where represents the base weight of the sample ; represents the comprehensive weight of the class to which the sample belongs; is the disease level weight, representing the inverse frequency weight of the class; is the disease-to-disease weight, representing the normalized weight of the relative sample distribution between classes; is the number of samples in the disease class , is the total number of samples, is the total number of disease classes;
[0059] Calculate the composite uncertainty score based on the multi-dimensional uncertainty evaluation mechanism, expressed as:
[0060] ;
[0061] ;
[0062] ;
[0063] ;
[0064] ;
[0065] where represents the composite uncertainty score of the sample ; , and are the normalized entropy value of the sample , the confidence level difference and the inter-class difference , , and all represent weight coefficients; represents the input belonging to category prediction probability, represents the unnormalized logistic output of category ; represents the probability that the sample is correctly classified; represents the -th prediction probability after being sorted in descending order;
[0066] Combining the base weight and the composite uncertainty score to calculate the comprehensive weight, which is expressed as:
[0067] ;
[0068] ;
[0069] where represents the current iteration number, represents the maximum iteration number, is a parameter that controls the proportion of the uncertainty weight, which gradually increases during training but does not exceed the preset maximum value .
[0070] Preferably, the dynamic sampling module further includes a temperature parameter adjustment mechanism, which is expressed as:
[0071] ;
[0072] ;
[0073] where and are respectively the minimum and maximum values of the temperature parameter, is the temperature reduction rate. At the initial stage of training, a higher temperature is adopted to make the sampling more uniform, and at the later stage of training, the temperature is reduced to make the sampling more focused on high-weight samples. represents the calculation of the final sampling weight of the sample in the -th iteration.
[0074] Preferably, using the trained disease prediction model to perform disease prediction based on the microbiome data includes the following steps:
[0075] Collect the microbiome data to be detected and perform preprocessing to obtain hierarchical feature data including five levels of phylum, class, order, family, and genus;
[0076] Input the hierarchical feature data into the trained disease prediction model to obtain the classification result and confidence level.
[0077] The present invention also provides a disease prediction system based on the adaptive fusion of microbiome hierarchical features, comprising:
[0078] A preprocessing module that preprocesses multi-source microbiome data samples to obtain hierarchical feature data samples containing five levels of phylum, class, order, family, and genus;
[0079] A training module that constructs a disease prediction model and uses the hierarchical feature data samples for training. During the training process, a class balance mechanism and a multi-dimensional uncertainty evaluation strategy are combined to dynamically adjust the sample sampling weights;
[0080] A prediction module that uses the trained disease prediction model to perform disease prediction based on microbiome data;
[0081] In the disease prediction model, a hierarchical feature extraction module performs multi-level feature extraction on the hierarchical feature data of the five classification levels, and integrates them to obtain multi-level fusion features containing hierarchical information; a multi-scale feature fusion module processes the multi-level fusion features through three parallel paths, respectively outputs features containing structured information, environment-dependent information, and individual feature information, and then integrates the features output by the three paths through a dynamic weighted fusion mechanism to obtain multi-dimensional fusion features; a classifier performs disease prediction based on the multi-scale fusion features.
[0082] The present invention has the following beneficial effects:
[0083] (1) The present invention explicitly models the hierarchical relationship of microbiome taxonomy through a hierarchical feature extraction module, fully utilizes the rich biological information contained in the microbiome classification levels, effectively captures the evolutionary and functional connections between phylum, class, order, family, and genus, and improves the biological significance and interpretability of feature representation;
[0084] (2) The multi-scale feature fusion module designed by the present invention deeply interprets the microbiome data from three dimensions of structured pattern, environmental dependence, and individual characteristics through three parallel feature processing streams, realizes the all-round characterization of the microbiome data, and significantly enhances the model's ability to identify complex microbiome patterns;
[0085] (3) The dynamic sampling module of the present invention combines a class balance mechanism and a multi-dimensional uncertainty evaluation strategy, effectively solves the problem of sample imbalance in multi-disease classification scenarios, significantly improves the model's learning ability for rare diseases and difficult samples, and enhances the stability and accuracy of classification;
[0086] (4) The present invention adopts an end-to-end training method, where each module collaborates and optimizes, reducing intermediate processing steps, improving the overall performance and generalization ability of the model, and providing more robust technical support for the diagnosis of nervous system diseases based on the microbiome.
[0087] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments, but the present invention is not limited to the embodiments. Description of the Drawings
[0088] Figure 1 It is a method step diagram of an embodiment of the present invention;
[0089] Figure 2 It is a schematic flow diagram of a hierarchical feature extraction module of an embodiment of the present invention;
[0090] Figure 3 It is a schematic flow diagram of a multi-scale feature fusion module of an embodiment of the present invention;
[0091] Figure 4 It is a schematic flow diagram of a dynamic sampling module of an embodiment of the present invention;
[0092] Figure 5 It is a performance comparison diagram on various nervous system disease diagnosis tasks of an embodiment of the present invention;
[0093] Figure 6 It is a system structure diagram of an embodiment of the present invention. Specific Embodiments
[0094] The data used in the embodiments of the present invention include integrated data of the gut microbiome from 31 independent studies, covering three nervous system diseases: Alzheimer's disease (AD, 6 datasets, 284 cases in the disease group and 264 cases in the control group), Parkinson's disease (PD, 11 datasets, 1267 cases in the disease group and 942 cases in the control group), and autism spectrum disorder (ASD, 14 datasets, 708 cases in the disease group and 584 cases in the control group).
[0095] As Figure 1 shown, it is a method step diagram of the present invention, including the following steps:
[0096] S101, preprocess the multi-source microbiome data samples to obtain hierarchical feature data samples containing five levels: phylum, class, order, family, and genus;
[0097] S102, construct a disease prediction model, and use the hierarchical feature data samples for training. During the training process, combine the class balance mechanism and the multi-dimensional uncertainty evaluation strategy to dynamically adjust the sample sampling weights;
[0098] S103, use the trained disease prediction model to predict diseases based on the microbiome data.
[0099] Specifically, the disease prediction model uses a hierarchical feature extraction module to perform multi-level feature extraction on the feature data of five classification levels, and integrates them to obtain multi-level fusion features; uses a multi-scale feature fusion module to process the multi-level fusion features through three parallel channels of pattern stream, context stream, and content stream, and then integrates the features output by the three channels through a dynamic weighted fusion mechanism to achieve an all-round representation of the microbiome data; finally, a classifier is used to predict diseases based on the multi-scale fusion features.
[0100] Specifically, the preprocessing process involves multiple key steps: First, use the FastQC tool to evaluate the quality of the original sequencing data, and remove the first 10 nucleotides of each read to reduce systematic bias; perform ASV inference through the DADA2 process, which can distinguish true biological variations and sequencing errors at the single-base resolution level; use the SILVA 138.1 database for taxonomic annotation, and set a minimum bootstrap confidence threshold of 80%; exclude ASV sequences of chloroplasts or mitochondria and ASVs with a frequency lower than 1%; finally, uniformly aggregate the classification levels of all datasets to the genus level, and perform standardization processing through rarefaction sampling to eliminate the influence of differences in sequencing depth between different samples.
[0101] Specifically, the process of the disease prediction model for prediction includes the following steps:
[0102] S201, preprocess the multi-source microbiome data to be detected to obtain hierarchical feature data including five levels of phylum, class, order, family, and genus;
[0103] S202, use the hierarchical feature extraction module to perform multi-level feature extraction on the hierarchical feature data, and fuse them to obtain multi-level fusion features (i.e., multi-level feature fusion representation);
[0104] S203, use the multi-scale feature fusion module to process the multi-level feature fusion representation through three parallel channels of pattern stream, context stream, and content stream, and then integrate the features output by the three channels through a dynamic weighted fusion mechanism to obtain multi-scale fusion features (i.e., multi-dimensional feature representation);
[0105] S204, finally, use a classifier to predict diseases based on the multi-dimensional feature representation.
[0106] Specifically, S202 is as Figure 2 shown. The hierarchical feature extraction module adopts a bottom-up feature processing strategy, starting from the genus level, gradually constructing the feature representation to the phylum level, and finally obtaining. This module contains three key feature processing components:
[0107] (1) Hierarchical feature extraction component: Through the mapping function and the subset function Construct a complete classification relationship network. For each taxonomic unit , return its complete classification path (e.g., M("Lactobacillus") = {"Bacteria", "Firmicutes", "Bacilli", "Lactobacillales", "Lactobacillaceae", "Lactobacillus"}), return all its subsets (e.g., R("Lactobacillaceae") = {"Lactobacillus", "Pediococcus",...}). Hierarchical feature representation is composed of the features of each classification level:
[0108] ;
[0109] Among them, the features of each level are vectors composed of the aggregated abundances of all taxonomic units at that level:
[0110] ;
[0111] For each taxonomic unit the aggregated abundance is calculated as the sum of the abundances of all its descendant taxonomic units:
[0112] ;
[0113] Among them, is the set of all descendants (including child nodes, grandchild nodes, etc.) of the taxonomic unit , and is calculated recursively using .
[0114] (2) Classification level weight component: Introduce an adaptive weighted feature scheme, with the basic weights designed to increase from the phylum to the genus level: 0.6 at the phylum level, 0.7 at the class level, 0.8 at the order level, 0.9 at the family level, and 1.0 at the genus level. At the same time, design a dynamic weight adjustment mechanism to adjust the weights according to the sparsity and discriminability of the features:
[0115] ;
[0116] Among them, is the basic weight of the level where the taxonomic unit is located, measures the discriminative ability of the feature, determined by calculating the coefficient of variation of the feature among different disease categories, represents the sparsity degree of the feature, determined by calculating the zero-value proportion of the feature, and are balance parameters and are set to 0.3 and 0.2 respectively.
[0117] Feature representation generated by the classification level weight component Calculated as:
[0118] ;
[0119] in is the original feature matrix, represents a diagonal matrix with the adjusted weight vector as its diagonal elements.
[0120] (3) Taxonomic relationship modeling component: Construct a multidimensional relationship matrix to quantify the taxonomic distance between features by analyzing the taxonomic paths between feature pairs and determining the most recent common ancestor. The relationship strength between features is calculated using the following formula:
[0121] ;
[0122] in, Representation characteristics and The common path length, and Represents the arrival characteristics and The complete classification path length.
[0123] To further improve the accuracy of relationship modeling, the module also introduces an enhanced relationship metric that takes evolutionary distance into account:
[0124] ;
[0125] in, Representation characteristics and The evolutionary distance between Set this to 0.5 to control the strength of the distance effect.
[0126] Feature representation generated by the classification relationship modeling component Calculated as:
[0127] ;
[0128] The final feature representation This is achieved through a multi-level feature fusion strategy, expressed as:
[0129] ;
[0130] in, is the original feature weighted by the classification level weight, is the hierarchical aggregation feature generated by the hierarchical feature extraction algorithm, is the relationship feature generated based on the classification relationship matrix, is the learnable fusion weight matrix, which is optimized during the training process by the gradient descent method.
[0131] Specifically, the S203 is as follows Figure 3 As shown, the multi-scale feature fusion module performs multi-dimensional interpretation of the microbiome data through three parallel feature processing streams:
[0132] (1) Pattern Stream: Responsible for identifying and extracting the structured patterns in the microbial community. This stream first projects the input features through , , , and then processes the transformed features using the Correlation Attention mechanism, receiving the input feature . The core attention calculation formula is:
[0133] ;
[0134] ;
[0135] ;
[0136] where , , are the learnable weight matrices of the query, key, and value respectively, , , represent the learned query, key, and value projections respectively, and the scaling factor ensures the stability of the gradient flow during the training process.
[0137] (2) Context Stream: Focuses on the environment-dependent features of the microbial community. Adopts the Global Range Processing module to establish long-range dependencies by calculating the similarity matrix between features. The feed-forward neural network calculation expression of the context stream is:
[0138] ;
[0139] ;
[0140] where represents the global attention operation, and are two fully connected layers, is an activation function, is an operation to prevent overfitting.
[0141] (3) Content Stream: Focuses on processing the individual feature information of microorganisms. The features are processed through a similarity mapping mechanism and a channel focusing module. The similarity calculation process can be expressed as:
[0142] ;
[0143] ;
[0144] where and are learnable weight matrices, is an activation function, is an attention mapping matrix, represents the channel focusing operation.
[0145] These three feature streams dynamically weight and combine their contributions through a feature fusion module. The fusion process first calculates the importance weights of each stream and then combines the weighted features:
[0146] ;
[0147] ;
[0148] where , , represent learnable fusion weights determined by the input features, is a learnable weight vector, represents the global pooling operation, represents the output feature of the corresponding feature stream, is the finally fused output feature.
[0149] Specifically, in S102, the disease prediction model also dynamically adjusts the sample sampling weights based on a dynamic sampling module combined with a class balance mechanism and a multi-dimensional uncertainty evaluation strategy, strengthening the model's learning ability for rare diseases and difficult samples. As Figure 4 shown, the dynamic sampling module includes a class balance mechanism and a multi-dimensional uncertainty evaluation mechanism:
[0150] (1) Class balance mechanism: Calculates class-specific weights for disease samples according to their inverse frequency in the dataset. The disease level weight of the sample belonging to disease is defined as:
[0151] ;
[0152] Among them is the number of samples in the disease category , is the total number of samples, is the total number of disease categories.
[0153] For the imbalance between diseases, a normalization factor considering the relative sample distribution between disease categories is introduced to calculate the weights between diseases :
[0154] ;
[0155] Among them represents the number of samples in each disease category. This logarithmic formula provides a smooth adjustment to prevent over-weighting and at the same time compensates for the imbalance between diseases.
[0156] Based on this, weights corresponding to each category are assigned to each sample to form the basic sampling weights:
[0157] ;
[0158] ;
[0159] (2) Multi-dimensional uncertainty evaluation mechanism: Based on the category probability distribution of the samples by the model, the logical output of the network is converted into a probability estimate for each disease category through the Softmax transformation:
[0160] ;
[0161] Among them represents the unnormalized logical output of category , represents the input belongs to category predicted probability.
[0162] Based on this probability distribution, three complementary uncertainty evaluation approaches are implemented:
[0163] Information entropy uncertainty: Measures the dispersion of probability mass in the sample prediction distribution, defined as:
[0164] ;
[0165] Confidence uncertainty: Measures the gap between the model prediction result and the true category, defined as:
[0166] ;
[0167] Among them represents the probability that the sample is correctly classified.
[0168] Inter-class difference uncertainty: Used to identify difficult samples near the decision boundary, evaluated by calculating the difference between the probabilities of adjacent classes in the predicted probability distribution:
[0169] ;
[0170] where represents the th predicted probability after sorting in descending order.
[0171] These three uncertainty measures are combined into a composite uncertainty score:
[0172] ;
[0173] where , and are the normalized entropy value, confidence difference, and inter-class difference respectively, , and are the weight coefficients of each index, satisfying . In actual implementation, set , and .
[0174] To achieve a smooth transition from class balance to hard sample focusing, an iterative-dependent time adaptation mechanism is introduced:
[0175] ;
[0176]
[0177] where represents the current iteration number, represents the maximum iteration number, is a parameter that controls the weight ratio of uncertainty, gradually increasing during training but not exceeding the preset maximum value (set to 0.8), is set to 0.2, is the base weight, is the normalized composite uncertainty index.
[0178] To further optimize the sampling process of microbiome data, a temperature parameter adjustment mechanism is introduced. The temperature parameter controls the smoothness of the sampling weight distribution. A higher temperature is used at the beginning of training to make the sampling more uniform, ensuring that the model has a basic understanding of various microbiome patterns; the temperature is lowered at the end of training to make the sampling more focused on high-weight samples, refining the model's understanding of microbiome characteristics:
[0179] ;
[0180] Among them, and are the minimum and maximum values of the temperature parameter respectively, is the temperature reduction rate. In actual implementation, set , , .
[0181] Final sample at the th iteration, the sampling weight is calculated as:
[0182] ;
[0183] Specifically, in S204, a random forest classifier is used to predict diseases for the fused features, and the classification result and confidence are output. The classifier adopts the random forest algorithm, which realizes classification by constructing multiple decision trees and integrating their prediction results. In this embodiment, the number of decision trees is set to 500, the maximum tree depth is unlimited, the Gini impurity index is used for feature selection, and the minimum number of samples in a leaf node is 5. Each decision tree in the random forest randomly samples and selects samples from the original training set with replacement during training, and randomly selects a subset of features (usually the square root of the total number of features) when splitting nodes. This randomness makes the model have good generalization ability and robustness. In addition, the random forest can also output the probability estimates of each sample belonging to various categories, which is convenient for subsequent uncertainty analysis. During the training process, a 5-fold cross-validation method is used to evaluate the model performance, and a grid search method is used to optimize the key hyperparameters, including the number of decision trees, the maximum number of features, and the minimum number of samples for splitting.
[0184] To verify the effectiveness of the present invention, experimental evaluations were carried out using the integrated data of the gut microbiome from 31 independent studies. In terms of performance evaluation, a 5-fold cross-validation method was adopted, and multiple evaluation metrics such as accuracy, precision, recall, F1-score, and AUC-ROC were used for comprehensive evaluation. The experimental results are as Figure 5 shown. The method (MSAFI) of the embodiment of the present invention is significantly superior to the existing methods in all evaluation metrics. In terms of the overall classification accuracy, MSAFI reaches 68.96% ± 0.90%, the F1-score is 66.53% ± 0.74%, and the AUC-ROC value is as high as 90.96% ± 0.75%. It has obvious advantages compared with traditional machine learning methods such as random forest (accuracy 65.64% ± 1.95%, F1-score 63.79% ± 1.60%).
[0185] Specifically, as shown in Figure 6 , the system structure diagram of the embodiment of the present invention includes:
[0186] A preprocessing module 601 preprocesses multi-source microbiome data samples to obtain hierarchical feature data samples including five levels of phylum, class, order, family, and genus.
[0187] A training module 602 constructs a disease prediction model and trains it using the hierarchical feature data samples.
[0188] A prediction module 603 uses the trained disease prediction model to predict diseases based on microbiome data.
[0189] The present invention makes full use of microbiological classification level information to achieve multi-dimensional feature representation, effectively addresses the problem of sample imbalance, and improves the diagnostic accuracy and stability for neurological diseases.
[0190] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A disease prediction method based on adaptive fusion of hierarchical features of microbiome, characterized in that: The following steps are involved: Preprocess the multi-source microbiome data samples to obtain hierarchical feature data samples at five levels: phylum, class, order, family, and genus; Build a disease prediction model and use stratified feature data samples for training. During the training process, combine the category balance mechanism and multi-dimensional uncertainty assessment strategy to dynamically adjust the sample sampling weights. Use trained disease prediction models to predict diseases based on microbiome data; In the disease prediction model, the hierarchical feature extraction module performs multi-level feature extraction on the hierarchical feature data of the five classification levels, and integrates them to obtain multi-level fusion features containing hierarchical information; the multi-scale feature fusion module processes the multi-level fusion features through three parallel pathways, and outputs features containing structural information, environmental dependency information, and individual feature information respectively, and then integrates the features output by the three pathways through a dynamic weighted fusion mechanism to obtain multi-dimensional fusion features; The classifier predicts diseases based on multi-scale fusion features; The process of obtaining multi-level fusion features by the hierarchical feature extraction module includes the following steps: The hierarchical feature extraction component extracts features from five levels and combines them to obtain the hierarchical feature representation F. hierarchical , expressed as: F hierarchical =[H phylum ,H class ,H order ,H family ,H genus ]; H l =[F l (t1),F l (t2),...,F l (t n )]; Among them, H phylum , H class , H order , H family , H genus Represent the characteristics of the five levels of phylum, class, order, family, and genus respectively; the characteristics of each level H l is the aggregate abundance F of all taxa t at this level. l (t), l∈{phylum, class, order, family, genus}, X represents the data matrix input to the disease prediction model, X(d) represents the feature vector corresponding to the classification unit d; D(t) is the set of all descendant classification units of classification unit t, which is calculated by recursively using the child node mapping R(t), d is a classification unit element in D(t) that records the set of all descendants; R(t) records the set of all child classification units with classification unit t as the parent node, which is constructed by reverse mapping the parent-child relationship information in the mapping M(t); M(t) is known data, which records the five levels and parent node information of each microbial classification unit t; The classification level weight component dynamically adjusts the feature weights to obtain the weighted feature representation F weighted , expressed as: F weighted =X·diag(w adjusted ); w adjusted (t)=w base (t)·(1+α·discriminative(t))·(1-β·sparsity(t)); Where X represents the sample data matrix; w base (t) is the basic weight of the level where the classification unit t is located, discriminative(t) measures the discriminative ability of the feature, sparsity(t) indicates the sparsity of the feature, α and β are balance parameters, diag(w adjusted ) represents a diagonal matrix with the adjusted weight vector as the diagonal element; The classification relationship modeling component constructs a multidimensional relationship matrix based on the classification path and evolutionary distance between features, and calculates the relationship feature representation F through the multidimensional relationship matrix. relational , expressed as: F relational =X·R enhanced ; R enhanced (i,j)=R(i,j)·exp(-γ·d evolutionary (i,j)); Among them, CommonPath(i,j) represents the common path length of features i and j, Path(i) and Path(j) represent the complete classification path lengths to features i and j respectively; d evolutionary (i, j) represents the evolutionary distance between features i and j, and γ is a parameter that controls the strength of the distance effect; Through feature fusion, we can obtain multi-level fusion features F final , expressed as: F final =concat(F weighted ,F hierarchical ,F relational )·W fusion ; W fusion is a learnable fusion weight matrix that is optimized during training by gradient descent.
2. The disease prediction method based on adaptive fusion of hierarchical features of microbiome according to claim 1, characterized in that: The multi-scale feature fusion module uses the multi-level fusion feature as the input feature X1, processes the multi-level fusion feature through three parallel paths, and finally calculates the multi-dimensional fusion feature, including the following steps: The pattern flow path uses the correlation attention mechanism to process the input features and obtain the pattern flow output feature F P , expressed as: F P =LayerNorm(S P +X1); Q=X1W Q ,K=X1W K ,V=X1W V ; Among them, W Q , W K , W V are the learnable weight matrices for query, key, and value, respectively. Q, K, and V represent the learned query, key, and value projections, respectively. is a scaling factor used to ensure the stability of the gradient flow during training; The context flow path uses a global scope processing module to establish long-range dependencies by calculating the similarity matrix between input features, and calculates the upstream and downstream output features F through the feedforward neural network of the context flow C , expressed as: F C =LayerNorm(Dense2(Dropout(ReLU(Dense1(F att ))))+F att ); F att =GlobalAttention(X1); Among them, GlobalAttention represents the global attention operation, Dense1 and Dense2 represent two fully connected layer operations respectively, ReLU is the activation function, and Dropout is an operation to prevent overfitting; The content flow channel processes the input features through the similarity mapping mechanism and the channel focusing module to obtain the content flow output features F S , expressed as: F S =ChannelFocus(X1,M)=X1⊙M; M = sigmoid(W2ReLU(W1X1)); Among them, W1 and W2 are learnable weight matrices, sigmoid is the activation function, M is the attention mapping matrix, and ChannelFocus represents the channel focusing operation; Calculate the importance weight of each channel, and then perform weighted combination of the output features of each channel based on the weight to obtain the multi-dimensional fusion feature F out , expressed as: Among them, α i (F) represents the learnable fusion weights corresponding to the three pathways, w i represents the learnable weight vector corresponding to the three paths, Pool represents the global pooling operation, and F i Represents the output characteristics corresponding to the three channels; i=P represents the mode flow channel, i=C represents the up and down flow channels, and i=S represents the content flow channel.
3. The disease prediction method based on adaptive fusion of hierarchical features of microbiome according to claim 1, characterized in that: Disease prediction is performed on multi-dimensional fusion features through classifiers. Specifically, disease prediction is performed on multi-scale fusion features through random forest classifiers, and classification results and confidence levels are output.
4. The disease prediction method based on adaptive fusion of hierarchical features of microbiome according to claim 3, characterized in that: The classifier is configured as binary classification or multi-classification according to actual prediction requirements; In the binary classification setting, the classification results are distinguished into disease status and non-disease status; In a multi-classification setting, the classification results are differentiated into multiple different disease types.
5. The disease prediction method based on adaptive fusion of hierarchical features of microbiome according to claim 1, characterized in that: The dynamic adjustment of the sample sampling weight is implemented by a dynamic sampling module, including the following steps: The basic weight is calculated based on the category balancing mechanism and is expressed as: W class (y i )=w d ·w i ; Among them, W base (x i ) represents the sample x i The basic weight of W class (y i ) represents the sample x i Category i The comprehensive weight of d is the disease level weight, which represents the inverse frequency weight of the category; w i is the inter-disease weight, which represents the normalized weight of the relative sample distribution between categories; N d is the number of samples in disease category d, N is the total number of samples, and K is the total number of disease categories; The composite uncertainty score is calculated based on the multi-dimensional uncertainty assessment mechanism and is expressed as: D conf (x i )=1-P(y true |x i ); Among them, U combined (x i ) represents the sample x i The composite uncertainty score of and They are samples x i The normalized entropy value H(P(y|x i )), Confidence Difference D conf (x i ) and the inter-class difference G class (x i ), w1, w2 and w3 all represent weight coefficients; P(y=k|x) represents the predicted probability that input x belongs to category k, z k represents the unnormalized logistic output of category k; P(y true |x i ) represents the sample x i The probability of being correctly classified; represents the j-th predicted probability after being sorted in descending order; The comprehensive weight is calculated by combining the basic weight and the composite uncertainty score, which is expressed as: Among them, t represents the current number of iterations, T1 represents the maximum number of iterations, and α t It is a parameter that controls the ratio of uncertainty weights, which gradually increases as training progresses but does not exceed the preset maximum value α max .
6. The disease prediction method based on adaptive fusion of hierarchical features of microbiome according to claim 5, characterized in that: The dynamic sampling module also includes a temperature parameter adjustment mechanism, which is expressed as: Among them, τ min and τ max are the minimum and maximum values of the temperature parameter, β1 is the temperature reduction rate. In the early stage of training, a higher temperature is used to make the sampling more uniform, and in the later stage of training, the temperature is lowered to make the sampling more focused on high-weight samples. final (x i , t) represents the sample x i Final sampling weights calculation at iteration t.
7. The disease prediction method based on adaptive fusion of hierarchical features of microbiome according to claim 1, characterized in that: The method of using the trained disease prediction model to predict diseases based on microbiome data includes the following steps: Collect and preprocess the microbiome data to be tested to obtain hierarchical feature data at five levels: phylum, class, order, family, and genus; The hierarchical feature data is input into the trained disease prediction model to obtain the classification results and confidence.
8. A disease prediction system based on adaptive fusion of hierarchical features of microbiome, characterized in that: include: The preprocessing module preprocesses the multi-source microbiome data samples to obtain hierarchical feature data samples at five levels: phylum, class, order, family, and genus; The training module builds a disease prediction model and uses stratified feature data samples for training. During the training process, the category balance mechanism and multi-dimensional uncertainty assessment strategy are combined to dynamically adjust the sample sampling weights. The prediction module uses the trained disease prediction model to predict diseases based on microbiome data; In the disease prediction model, the hierarchical feature extraction module performs multi-level feature extraction on the hierarchical feature data of the five classification levels, and integrates them to obtain multi-level fusion features containing hierarchical information; the multi-scale feature fusion module processes the multi-level fusion features through three parallel pathways, and outputs features containing structural information, environmental dependency information, and individual feature information respectively, and then integrates the features output by the three pathways through a dynamic weighted fusion mechanism to obtain multi-dimensional fusion features; The classifier predicts diseases based on multi-scale fusion features; The process of obtaining multi-level fusion features by the hierarchical feature extraction module includes the following steps: The hierarchical feature extraction component extracts features from five levels and combines them to obtain the hierarchical feature representation F. hierarchical , expressed as: F hierarchical =[H phylum ,H class ,H order ,H family ,H genus ]; H l =[F l (t1),F l (t2),...,F l (t n )]; Among them, H phylum , H class , H order , H family , H genus Represent the characteristics of the five levels of phylum, class, order, family, and genus respectively; the characteristics of each level H l is the aggregate abundance F of all taxa t at this level. l (t), l∈{phylum, class, order, family, genus}, X represents the data matrix input to the disease prediction model, X(d) represents the feature vector corresponding to the classification unit d; D(t) is the set of all descendant classification units of classification unit t, which is calculated by recursively using the child node mapping R(t), d is a classification unit element in D(t) that records the set of all descendants; R(t) records the set of all child classification units with classification unit t as the parent node, which is constructed by reverse mapping the parent-child relationship information in the mapping M(t); M(t) is known data, which records the five levels and parent node information of each microbial classification unit t; The classification level weight component dynamically adjusts the feature weights to obtain the weighted feature representation F weighted , expressed as: F weighted =X·diag(w adjusted ); w adjusted (t)=w base (t)·(1+α·discriminative(t))·(1-β·sparsity(t)); Where X represents the sample data matrix; w base (t) is the basic weight of the level where the classification unit t is located, discriminative(t) measures the discriminative ability of the feature, sparsity(t) indicates the sparsity of the feature, α and β are balance parameters, diag(w adjusted ) represents a diagonal matrix with the adjusted weight vector as the diagonal element; The classification relationship modeling component constructs a multidimensional relationship matrix based on the classification path and evolutionary distance between features, and calculates the relationship feature representation F through the multidimensional relationship matrix. relational , expressed as: F relational =X·R enhanced ; R enhanced (i,)=R(i,j)·exp(-γ·d evolutionary (i,j)); Among them, CommonPath(i,j) represents the common path length of features i and j, Path(i) and Path(j) represent the complete classification path lengths to features i and j respectively; d evolutionary (i, j) represents the evolutionary distance between features i and j, and γ is a parameter that controls the strength of the distance effect; Through feature fusion, we can obtain multi-level fusion features F final , expressed as: F final =concat(F weighted ,F hierarchical ,F relational )·W fusion ; W fusion is a learnable fusion weight matrix that is optimized during training by gradient descent.
Citation Information
Patent Citations
Multi-level fusion skin disease diagnosis system based on multi-modal image data
CN115546217A
Microorganism-disease symbol correlation prediction method based on symbol information transmission
CN118448060A
Cited By
Biological characteristic intelligent classification prediction method based on adaptive weight multi-modal fusion
CN121030422A
An adaptive weight multi-modal fusion biometric intelligent classification and prediction method
CN121030422B