A Method for Mining Key Multi-omics Molecules and Pathways Based on Heterogeneous Graph Framework
By employing a multi-omics data integration method based on a heterogeneous graph framework and utilizing integrated biological knowledge graphs and deep learning techniques, the challenges in multi-omics data integration were addressed, enabling efficient identification of key disease molecules and biological functional pathways, thereby improving biological interpretability and predictive accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DALIAN INSTITUTE OF CHEMICAL PHYSICS CHINESE ACADEMY OF SCIENCES
- Filing Date
- 2024-11-26
- Publication Date
- 2026-05-26
AI Technical Summary
Existing multi-omics data integration methods struggle to effectively integrate data from different omics, especially when faced with small sample sizes, high data sparsity, and insufficient biological interpretability, making it difficult to uncover key disease molecules and important pathways with high biological significance.
We employ a heterogeneous graph framework approach to integrate multiple open-source databases to construct a comprehensive biological knowledge graph. We utilize various machine learning algorithms to transform omics data into machine learning attribute matrices, and combine graph-structured variational autoencoders and attention-based graph convolutional neural networks to train disease-specific networks. We then identify key biological modules and pathways through community discovery algorithms.
It achieves high prediction accuracy, high interpretability and high computational efficiency in the integration of multi-omics data, and can identify key disease molecules and biological functional pathways, providing new pathways for the study of disease molecular mechanisms and the discovery of therapeutic targets.
Smart Images

Figure CN122090965A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics technology, and in particular to a method for mining key multi-omics molecules and pathways based on a heterogeneous graph framework. Background Technology
[0002] With the rapid development of advanced omics technologies, researchers are now able to characterize various molecules in detail, providing a comprehensive medical perspective on biological systems or individual phenotypes. By aggregating multi-omics data, scientists can simultaneously capture multiple layers of information, thereby gaining a deeper understanding of the molecular characteristics of diseases and decoding complex biological mysteries. However, extracting valuable insights from these massive amounts of multi-omics data poses significant challenges to data analysis techniques.
[0003] Currently, methods for integrating multi-omics data are fragmented, primarily relying on complex multivariate statistics, machine learning, and omics-specific knowledge bases. Analysis of single-omics data typically follows one or several well-established workflows, but the unique characteristics of analytical methods for different omics datasets make it difficult to generalize these methods to other omics datasets. Most existing multi-omics integration methods have specific assumptions and requirements regarding data compatibility, sample size, and preprocessing, which may limit their widespread application. Currently developed multi-omics integration methods are mainly based on two strategies: data-driven (L. Ning, Y.-L. Zhou, H. Sun, Y. Zhang, C. Shen, Z. Wang, B. Xuan, Y. Zhao, Y. Ma, Y. Yan, T. Tong, X. Huang, M. Hu, X. Zhu, J. Ding, Y. Zhang, Z. Cui, J.-Y. Fang, H. Chen, J. Hong, Nat Commun 2023, 14, 7135.) and knowledge-driven (MA Reyna, M.D. M. Leserson, B.J. Raphael, Bioinformatics 2018, 34, i972.). Data-driven methods reveal key features and molecular trends across omics by analyzing shared patterns and correlations between different omics layers, and their computation does not rely on prior knowledge. However, these data-based algorithms depend on rigorously matched high-quality samples and advanced analytical techniques to identify new discoveries, and often require additional steps to interpret the discovered molecular features. While data-driven approaches are the mainstream method for multi-omics integration and have spawned various combinations, these methods often struggle to capture the nonlinear relationships and interactions between multiple omics factors. Knowledge-driven approaches, on the other hand, incorporate prior biological knowledge to reduce computational complexity and enhance model interpretability, and have been applied in various medical fields, including cancer diagnosis and radiology. However, the effectiveness of knowledge-driven methods is limited by the quality and comprehensiveness of the underlying knowledge base, especially in the absence of supporting experimental data. Therefore, combining the advantages of data-driven and knowledge-driven methods to achieve efficient integration of multi-omics data is a promising data processing approach.
[0004] Advances in Deep Learning (DL) have provided novel technical approaches for integrating multi-omics data. However, applying DL algorithms to solve biological mechanism problems in complex biological regulatory networks faces numerous challenges. Heterogeneous graphs have attracted considerable attention in various fields due to their ability to represent diverse entities and complex relationships, such as recommender systems, knowledge graph construction, social network analysis, and bioinformatics networks (X. Zhao, X. Zhao, M. Yin, Brief Bioinform 2022, 23, bbab407.). Compared to homogeneous graphs, which contain only one type of edge, heterogeneous graphs allow different types of nodes and edges to coexist, resulting in richer and more complex graph structures (S. Ji, S. Pan, E. Cambria, P. Martinen, PS Yu, IEEE Trans Neural Network Learn Syst 2022, 33, 494.). Heterogeneous graphs can more naturally represent multiple entities and their relationships in the real world, and can better characterize molecular mechanisms and cellular heterogeneity in complex biological networks. However, existing models based on heterogeneous graph neural networks still face many limitations in integrating multi-omics data. First, biomedical research typically only obtains a limited sample size, and these samples contain high-dimensional molecular data. Furthermore, identifying a large number of biomarkers from small sample data usually requires extensive exploratory and functional validation experiments, which are not only time-consuming but also costly. Most importantly, the varying numbers of molecules and connections result in different local sparsity in the network, increasing the difficulty of model training. For example, genomics, relying on well-developed microarray detection technology, can obtain information on tens of thousands of genes, while metabolomics, limited by detection coverage and molecular annotation capabilities, can only detect a maximum of a few thousand molecules (A. Ghavasieh, M. De Domenico, Nat. Phys. 2024, 20, 512.). Finally, there are limitations in the mapping from the molecular level to pathways. Due to database updates and limitations, the same significant pathways may be frequently identified in multiple diseases, and in current functional exploration studies, most newly identified significant molecules often fail to be successfully mapped to known pathways, leading to the loss of a significant amount of information during molecular functional analysis. These issues significantly limit the discovery of novel and meaningful disease insights. Therefore, while traditional methods are relatively mature in processing single-omics data, they still have significant limitations in integrating multi-omics data and elucidating biological functions, especially in identifying key disease molecules and pathways with high biological significance. Summary of the Invention
[0005] To address the limitations of existing technologies and overcome the aforementioned challenges, this invention proposes a method for mining key multi-omics molecules and pathways based on a heterogeneous graph framework. This method relies on prior knowledge graphs and multi-omics experimental data to discover key disease molecules and pathways involved in disorder mechanisms. It overcomes the problems currently faced in multi-omics data integration, such as small sample size, high data sparsity, insufficient biological interpretability, and information noise between different omics data. This method offers advantages such as high prediction accuracy, high interpretability, and high computational efficiency, facilitating the full utilization of omics data to elucidate the biological mechanisms of complex diseases.
[0006] The technical solution adopted by the present invention to achieve the above objectives is as follows:
[0007] A method for mining key multi-omics molecules and pathways based on a heterogeneous graph framework includes the following steps:
[0008] 1) By integrating multiple open-source databases, a comprehensive biological knowledge graph composed of information on different omics molecules, diseases, and biological pathways is constructed;
[0009] 2) Employ various machine learning methods to transform the raw omics data into a multi-dimensional machine learning attribute matrix;
[0010] 3) Based on the input molecules, target molecules, and metapaths, extract disease-specific networks from the comprehensive biological knowledge graph;
[0011] 4) Develop a deep learning framework based on heterogeneous graph structure to integrate multi-omics data, and train disease-specific networks based on machine learning attribute matrices;
[0012] 5) Using a community detection algorithm, key biological modules are identified in the trained disease-specific network, and important biological pathways and their related molecules are determined.
[0013] Step 1) includes the following steps:
[0014] 1.1) Statistically classify the types of molecules and intermolecular interactions contained in different databases;
[0015] 1.2) Construct a comprehensive node information knowledge base, which contains molecular types involved in different databases, summarizes different naming rules for different molecules, and standardizes molecular naming in different databases;
[0016] 1.3) Extract intermolecular types from different databases and unify their naming;
[0017] 1.4) After merging information from different databases, redundant information and information with confidence levels below the threshold are removed to obtain a comprehensive biological knowledge graph.
[0018] Step 2) includes the following steps:
[0019] 2.1) Employ multiple machine learning algorithms, using the features of omics data as explanatory variables and the classification labels of samples as dependent variables, and obtain the discriminative ability or coefficient of features as the importance score of molecules, i.e. the contribution of each feature to the predictive ability in different energy models.
[0020] 2.2) The importance scores of various omics molecules are integrated by splicing to obtain the molecular machine learning attribute matrix corresponding to the original omics data.
[0021] The machine learning algorithm includes:
[0022] Nonparametric tests / one-way ANOVA, linear discriminant analysis, random forest, support vector machine, logistic regression, decision tree, Gaussian mixture model, K-nearest neighbor, partial least squares discriminant analysis, extreme gradient boosting, Boruta, and recursive feature elimination.
[0023] Step 3) includes the following steps:
[0024] 3.1) Extract subnetworks from the comprehensive biological knowledge graph based on target molecules and metapaths;
[0025] 3.2) Extract subnetworks from the subnetwork based on the input numerator and the set neighbor order;
[0026] 3.3)) After removing the custom high centrality nodes in the subnetwork, a disease-specific network is obtained for subsequent training.
[0027] The multi-omics data integration model includes a knowledge completion module and a feature extraction module, wherein:
[0028] The knowledge completion module processes attribute matrices generated by various machine learning methods and disease-specific networks, and uses graph-structured variational autoencoders to generate latent variable learning data features to achieve structural attribute completion of disease-specific networks.
[0029] The feature extraction module processes the completed network structure and uses a graph convolutional neural network with the attention mechanism of the purifier to extract key information; it obtains molecules and their corresponding scores, where the score represents the strength of the association between the molecule and the disease.
[0030] Step 5) includes the following steps:
[0031] 5.1) Extract biological network modules from structurally complete disease-specific networks using the overlapping community algorithm;
[0032] 5.2) Quantitatively evaluate each biological network module based on molecular scores, calculate the molecular scores involved in each biological network module, and sort them according to the scores to identify the core biological modules.
[0033] The present invention has the following beneficial effects and advantages:
[0034] 1. This invention integrates multiple open-source databases to establish a comprehensive prior biological knowledge graph, which includes different types of molecules and connections, making the subsequent training of the model highly flexible and generalizable.
[0035] 2. This invention provides a method for identifying key disease molecules and biological functional pathways based on a heterogeneous graph neural network framework and multi-omics data. This method utilizes the ability of heterogeneous graphs to represent different types of molecules and their connections, allowing the model training process to closely resemble complex biological regulatory networks and improving model interpretability. Furthermore, a graph variational structure autoencoder is used to address the high data sparsity problem in multi-layered biological networks, combined with Markov random fields for knowledge learning and derivation. To combat interference from malicious neighbors, a graph convolutional neural network with a purifier-equipped attention mechanism is introduced to accurately capture key information. Therefore, this invention offers advantages such as high prediction accuracy, high interpretability, and high computational efficiency, enabling the integration of multi-omics data to mine key disease molecules and analyze disease-related disordered pathways.
[0036] 3. The method provided in this invention uses a disease-specific multilayer network and a machine learning attribute matrix generated from multi-omics data as input. It transforms the abundance data of different omics into comparable machine learning attribute data, effectively balancing the impact of noise between different omics data on data integration. Furthermore, the introduction of richer machine learning algorithms effectively improves the stability of network training and the robustness of the results.
[0037] 4. The method provided by this invention can identify key molecules and biological functional pathways in diseases, with the advantages of high accuracy and high biological interpretability. This makes the method not only technically innovative but also of great value in practical applications, providing a new approach for the study of the molecular mechanisms of diseases and the discovery of therapeutic targets. Attached Figure Description
[0038] Figure 1 A schematic diagram of the process of this invention;
[0039] Figure 2 A schematic diagram of a multi-layered knowledge network specific to prostate cancer.
[0040] Figure 3 MODAPro interpretable diagram of capturing key disease molecules;
[0041] Among them, (A) is a pie chart of the proportion of key molecule diversity; (B) is the key hidden molecule database verification information captured by MODAPro.
[0042] Figure 4 Schematic diagram of the biological pathways of MODAPro hidden molecules related to specific disease disorders. Detailed Implementation
[0043] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.
[0044] This invention provides a method for identifying key disease molecules and biological functional pathways based on a heterogeneous graph deep learning framework and multi-omics data. The specific steps are as follows:
[0045] S1: By integrating multiple open-source databases, a comprehensive biological knowledge graph is constructed, which covers information on different omics molecules, diseases, and biological pathways;
[0046] S2: Employs multiple machine learning algorithms to transform raw omics data into a multi-dimensional machine learning attribute matrix;
[0047] S3: Extract disease-specific networks from a comprehensive biological knowledge graph based on input molecules, target molecules, and metapaths;
[0048] S4: Based on heterogeneous graph deep learning technology, predictive models for key molecules and biological pathways are established using graph structure variational autoencoders and optimized attention mechanisms. The input is a multi-dimensional machine learning attribute matrix, and the output is a molecular importance score.
[0049] S5: Apply community discovery algorithms to identify key biological modules in the trained disease-specific network, and ultimately determine important biological pathways and their related molecules.
[0050] Through these steps, the present invention can effectively identify key molecules and biological functional pathways closely related to the occurrence and development of diseases from multi-omics data, providing a new analytical method for the study of the molecular mechanisms of diseases and the discovery of therapeutic targets.
[0051] In step S1, the comprehensive biological knowledge graph is constructed by integrating multiple open-source databases, including but not limited to Blood Exposome Database, CTD, HMDB, HumanTFDB, LIPIDMAPS, SMPDB, Pubchem, Drugbank, BATMAN-TCM, BIOGRID, BRENDA, Ensembl, GO, HGNC, hTFtarget, HumanNet, HURi, Reactome, and STRING database. A custom-constructed background biological knowledge graph is used for subsequent model training and analysis.
[0052] Preferably, the types of molecules and intermolecular interactions contained in different databases are statistically analyzed and classified;
[0053] Preferably, a comprehensive node information knowledge base is constructed, which includes the types of molecules involved in different databases, summarizes the different naming rules such as naming and abbreviation of different molecules, and standardizes the naming of molecules in different databases;
[0054] Preferably, intermolecular types are extracted from different databases and named uniformly;
[0055] Preferably, after merging information from different databases, redundant and low-confidence information is removed to ultimately form a comprehensive biological knowledge graph.
[0056] In step S2, an attribute matrix is constructed using multiple machine learning algorithms. The machine learning training uses the classification labels of the samples as the dependent variable and the omics data features as the explanatory variables. The machine learning algorithms include, but are not limited to, nonparametric tests / one-way ANOVA, linear discriminant analysis, random forest, support vector machine, logistic regression, decision tree, Gaussian mixture model, K-nearest neighbors, partial least squares discriminant analysis, extreme gradient boosting, Boruta, and recursive feature elimination. Two or more machine learning algorithms are selected to complete the subsequent training process.
[0057] Prioritize the use of selected machine learning algorithms to process various omics data matrices and derive importance scores for each omics molecule based on sample labels;
[0058] Prioritize the integration of importance scores for various omics molecules to convert the original omics spectrum vectors into molecular machine learning attribute matrices.
[0059] In step S3, the user selects target molecules and metapaths based on the research objectives and background, and then extracts disease-specific networks from the comprehensive biological knowledge graph;
[0060] Preferably, subnetworks are first extracted from the comprehensive biological knowledge graph based on the target molecules and metapaths;
[0061] Preferably, the subnetwork is extracted based on the input molecule and the set neighbor order;
[0062] Preferably, after removing high-centrality nodes from the network, the resulting subnetwork will be used as the final disease-specific multilayer network for subsequent training.
[0063] In step S4, attribute matrices generated by various machine learning methods and disease-specific multilayer networks are input into the knowledge completion module. The graph structure variational autoencoder is used to generate latent variable learning data features to complete the network structure attributes and complete the information derivation based on experimental data.
[0064] In step S4, the complete network structure obtained by the knowledge completion module is delivered to the feature extraction module, and the graph convolutional neural network with the attention mechanism of the purifier is used to extract key information; the final output is the molecule and its corresponding score, and the score represents the strength of the association between the molecule and the disease.
[0065] In step S5, the overlapping community detection algorithm is used to analyze the disease-specific network after attribute completion to determine the core biological modules.
[0066] Preferably, the overlapping community algorithm is used to extract biological network modules;
[0067] Preferably, molecular scores obtained based on a deep learning framework are used to quantitatively evaluate these biological network modules and rank them according to the scores, thereby identifying the core biological modules.
[0068] The deep learning framework consists of two modules, which have the following characteristics: Module 1 is an autoencoder based on graph variational structure, which combines Markov random fields to complete knowledge generation and derivation; Module 2 is a graph convolutional neural network based on an attention mechanism equipped with a purifier, which extracts accurate information after pruning malicious neighbors.
[0069] Example 1
[0070] To verify the effectiveness and practicality of the present invention, a set of multi-omics data containing prostate cancer (PRAD) tissue and adjacent normal tissue samples was used as an example. These data covered metabolomics and transcriptomics information, and the method of the present invention was used to capture key molecules and biological functional pathways of PRAD.
[0071] See Figure 1 This document illustrates a flowchart of a method for identifying key disease molecules and biological functional pathways based on a heterogeneous graph deep learning framework and multi-omics data, provided by the present invention, in one embodiment. The method includes the following steps:
[0072] Step S1: Integrate multiple open-source databases to construct a comprehensive biological knowledge graph that contains information on different omics molecules, diseases, and biological pathways.
[0073] S11: Downloaded databases including Blood Exposome Database, CTD, HMDB, HumanTFDB, LIPIDMAPS, SMPDB, Pubchem, Drugbank, BATMAN-TCM, BIOGRID, BRENDA, Ensembl, GO, HGNC, hTFtarget, HumanNet, HURi, Reactome, and STRING. Integrated these databases based on molecular classification and interaction relationships, and removed duplicate information.
[0074] S12: During the database integration process, standard names are used for different types of molecules to unify the naming conventions across different databases. For example, compounds are uniformly named using PubChem ID numbers, and gene / protein names are uniformly named using NCBI ID numbers.
[0075] S13: After merging the results from different databases, nodes without molecular information are removed. Then, a comprehensive biological knowledge graph is constructed, ensuring the correct mapping relationship exists between node files and edge files. Ultimately, the comprehensive biological knowledge graph contains 327,973 nodes and 16,710,307 edges.
[0076] Step S2: Use various machine learning methods to transform the raw omics data of prostate cancer (PRAD) into a multi-dimensional machine learning attribute matrix.
[0077] S21: Using the PRAD stage as the dependent variable, perform 12 machine learning algorithms, including one-way ANOVA, linear discriminant analysis, random forest, support vector machine, logistic regression, decision tree, Gaussian mixture model, K nearest neighbors, partial least squares discriminant analysis, extreme gradient boosting, Boruta, and recursive feature elimination, to obtain a machine learning attribute matrix containing N (number of molecules) multiplied by 12.
[0078] Step S3: Extract disease-specific networks from the integrated biological knowledge graph based on input molecules, target molecules, and metapaths.
[0079] S31: Select transcription factors, genes / proteins, enzymes, and metabolites as target molecules, and extract sub-networks from the comprehensive knowledge graph based on meta-pathways such as transcription factor-gene / protein, gene / gene-protein / protein, gene / protein-enzyme, enzyme-metabolite, and metabolite-metabolite.
[0080] S32: Multi-omics molecules of prostate cancer (PRAD) are mapped into subnetworks, and the subnetworks, including input molecules and their second-order neighbors, are extracted. Then, after removing molecules with a centrality greater than 100 from the network, a PRAD-specific network is constructed. This PRAD-specific network contains 72.16% attributed nodes and 27.84% attributeless molecules. Figure 2 ).
[0081] Step S4: Construct predictive models for key molecules and biological pathways using heterogeneous graph deep learning technology. The input is a multi-dimensional machine learning attribute matrix, and the output is a molecular importance score. The deep learning model is trained based on a disease-specific network.
[0082] S41: The constructed prostate cancer (PRAD) omics molecular machine learning attribute matrix is mapped to a disease-specific network and input into a deep learning framework. Attribute nodes are randomly divided into training and validation sets at an 80%:20% ratio, with hidden nodes used as the test set for training.
[0083] S42: Input the prepared data into Module 1 and execute the attribute completion program. Set the intermediate hidden layer dimension to 512, the learning rate to 0.0005, the weight decay to 0.0005, and the dropout to 0.4. Use the molecular scores of the training set and validation set as labels for iteration. Set the number of iterations to 1000 and set the adjustment parameters. Stop training when the loss function no longer decreases.
[0084] S43: The program automatically inputs the completed knowledge network and attribute matrix into module 2, and uses a graph convolutional neural network with a configured purifier attention mechanism to capture key molecules. Training continues, using molecular scores as labels to calculate the loss function, with 1000 iterations and a setter parameter; training stops when the loss function no longer decreases. The runtime is 1285.506 seconds. MODAPro identified 815 key molecules and their scores, including 1.72% transcription factors, 30.31% metabolites, and 67.98% proteins. This demonstrates that MODAPro can fully utilize network retrieval to uncover hidden molecules and improve the richness of key disease molecules identified (see [link to relevant documentation]). Figure 3 A). Searching multiple disease-related molecular databases (OncoKB, NCG, Malacards, DigSee, and CTD) revealed 491 and 19 hidden molecules, respectively, that were confirmed to be associated with prostate cancer progression. The confirmation of these hidden molecules' association with PRAD progression in these databases indirectly demonstrates MODAPro's superior performance in uncovering biologically interpretable hidden information (see [link]). Figure 3 B).
[0085] Step S5: The trained disease-specific network was analyzed using a community detection algorithm to identify key biological modules, and 15 prostate cancer (PRAD)-specific disease modules were successfully identified.
[0086] S51: By applying three pathway enrichment analysis methods, we conducted an in-depth exploration of the biological pathways involved in these 15 PRAD-specific disease modules. Compared with the disordered biological pathways revealed by other methods, the hidden molecules identified by MODAPro revealed eight additional new and significant key biological pathways. This finding demonstrates MODAPro's superior performance in uncovering biologically interpretable hidden information. Figure 4 ).
[0087] The above results demonstrate that the deep learning model constructed in this invention effectively identifies key molecules of diseases and their associated disordered pathways, and these results have considerable biological reliability and accuracy.
Claims
1. A method for mining key multi-omics molecules and pathways based on a heterogeneous graph framework, characterized in that, Includes the following steps: 1) By integrating multiple open-source databases, a comprehensive biological knowledge graph composed of information on different omics molecules, diseases, and biological pathways is constructed; 2) Employ various machine learning methods to transform the raw omics data into a multi-dimensional machine learning attribute matrix; 3) Based on the input molecules, target molecules, and metapaths, extract disease-specific networks from the comprehensive biological knowledge graph; 4) Develop a deep learning framework based on heterogeneous graph structure to integrate multi-omics data, and train disease-specific networks based on machine learning attribute matrices; 5) Using a community detection algorithm, key biological modules are identified in the trained disease-specific network, and important biological pathways and their related molecules are determined.
2. The method for mining key multi-omics molecules and pathways based on a heterogeneous graph framework according to claim 1, characterized in that, Step 1) includes the following steps: 1.1) Statistically classify the types of molecules and intermolecular interactions contained in different databases; 1.2) Construct a comprehensive node information knowledge base, which contains molecular types involved in different databases, summarizes different naming rules for different molecules, and standardizes molecular naming in different databases; 1.3) Extract intermolecular types from different databases and unify their naming; 1.4) After merging information from different databases, redundant information and information with confidence levels below the threshold are removed to obtain a comprehensive biological knowledge graph.
3. The method for mining key multi-omics molecules and pathways based on a heterogeneous graph framework according to claim 1, characterized in that, Step 2) includes the following steps: 2.1) Employ multiple machine learning algorithms, using the features of omics data as explanatory variables and the classification labels of samples as dependent variables, and obtain the discriminative ability or coefficient of features as the importance score of molecules, i.e. the contribution of each feature to the predictive ability in different energy models. 2.2) The importance scores of various omics molecules are integrated by splicing to obtain the molecular machine learning attribute matrix corresponding to the original omics data.
4. The method for mining key multi-omics molecules and pathways based on a heterogeneous graph framework according to claim 3, characterized in that, The machine learning algorithm includes: Nonparametric tests / one-way ANOVA, linear discriminant analysis, random forest, support vector machine, logistic regression, decision tree, Gaussian mixture model, K-nearest neighbor, partial least squares discriminant analysis, extreme gradient boosting, Boruta, and recursive feature elimination.
5. The method for mining key multi-omics molecules and pathways based on a heterogeneous graph framework according to claim 1, characterized in that, Step 3) includes the following steps: 3.1) Extract subnetworks from the comprehensive biological knowledge graph based on target molecules and metapaths; 3.2) Extract subnetworks from the subnetwork based on the input numerator and the set neighbor order; 3.3)) After removing the custom high centrality nodes in the subnetwork, a disease-specific network is obtained for subsequent training.
6. The method for mining key multi-omics molecules and pathways based on a heterogeneous graph framework according to claim 1, characterized in that, The multi-omics data integration model includes a knowledge completion module and a feature extraction module, wherein: The knowledge completion module processes attribute matrices generated by various machine learning methods and disease-specific networks, and uses graph-structured variational autoencoders to generate latent variable learning data features to achieve structural attribute completion of disease-specific networks. The feature extraction module processes the completed network structure and uses a graph convolutional neural network with the attention mechanism of the purifier to extract key information; it obtains molecules and their corresponding scores, where the score represents the strength of the association between the molecule and the disease.
7. The method for mining key multi-omics molecules and pathways based on a heterogeneous graph framework according to claim 1, characterized in that, Step 5) includes the following steps: 5.1) Extract biological network modules from structurally complete disease-specific networks using the overlapping community algorithm; 5.2) Quantitatively evaluate each biological network module based on molecular scores, calculate the molecular scores involved in each biological network module, and sort them according to the scores to identify the core biological modules.