This invention relates to the field of
artificial intelligence technology, specifically to a method for predicting microbial diseases based on the cross-fusion of phylogenetic and abundance data. First, it collects raw metagenomic
sequencing data and corresponding species abundance information to construct an initial dataset, builds a model-specific microbial
lexicon, and transforms biological entities into symbolic representations that can be processed by
deep learning models. Then, it vectorizes the microbial data. Next, a Cross-Attention module is used to achieve cross-
modal interaction and fusion between sequence features and abundance features, providing a powerful model for
disease phenotype prediction tasks, which is then trained and optimized. Finally, through transfer learning, the discriminative knowledge learned by the model in classification tasks is transformed into topological connections in the network. This allows network centrality analysis to quantify the pivotal importance of each
microorganism from the "model decision-making perspective," thereby identifying key biomarkers that are both strongly correlated with
disease and located at the ecological core, achieving a transformation from prediction to discovery.