Metabolite-disease association prediction method based on hierarchical comparative learning and stepwise orthogonal multi-modal fusion
Through hierarchical comparison learning and gradual orthogonal multimodal fusion methods, the problems of information loss and insufficient feature extraction in the prior art are solved, and the high accuracy and efficiency of metabolite-disease association prediction are achieved.
Patent Information
- Application Number
- CN202510291674.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-11
AI Technical Summary
The existing metabolites and disease association prediction methods have information loss and insufficient feature extraction when information fusion is integrated, and cannot effectively capture high-order neighbor information, limiting prediction performance and accuracy.
Using hierarchical comparison learning and stepwise orthogonal multimodal fusion methods, we use multiple similarity networks to build hypergraphs and perform feature extraction, and use hypergraph convolution and multi-layer perceptron for training and prediction, and introduce orthogonal loss function to control feature fusion.
It improves the accuracy of metabolite-disease association prediction, can capture advanced neighbor information, extract diversified and complementary characteristics, and improves the accuracy of prediction.
Smart Images

Figure CN120299749A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of metabolite-disease association prediction, and particularly to a metabolite-disease association prediction method based on hierarchical contrast learning and stepwise orthogonal multimodal fusion. Background Art
[0002] The occurrence and development of diseases are closely related to the content of metabolites in the body. Traditional biological experimental methods rely on a large amount of manpower for repeated experiments, which are time-consuming and laborious. In recent years, with the development of deep learning technology, researchers have begun to explore using deep learning methods to predict the association between metabolites and diseases. These methods extract the features of metabolites and diseases and use machine learning models for prediction, thereby improving the prediction efficiency and accuracy.
[0003] Currently, a latest solution is to propose a deep learning method called MDA-AENMF to predict the association between metabolites and diseases. This prediction model uses AutoEncoder, Graph AutoEncoder (GAE), and Non-Negative Matrix Factorization (NMF) to extract the features of metabolites and diseases, and uses a Multi-Layer Perceptron (MLP) classifier for prediction.
[0004] Existing metabolite-disease association prediction methods usually first fuse different similarity networks in various ways, then extract features through a Graph Convolutional Network (GCN) module, and finally perform classification.
[0005] However, when fusing similarities in the prior art, the fusion method will selectively choose information from each similarity network with some bias and fuse this information, inevitably resulting in information loss. In addition, when extracting features through GCN, due to the limitation of the number of layers of GCN, the model cannot capture the information of the high-order neighbors of the current node. These defects limit the prediction performance and accuracy of the existing methods. Summary of the Invention
[0006] In view of this, the present invention provides a metabolite-disease association prediction method based on hierarchical contrast learning and stepwise orthogonal multimodal fusion, aiming to solve the problems of information loss and insufficient feature extraction in the prior art and improve the accuracy and efficiency of prediction.
[0007] To this end, the present invention provides the following technical solutions:
[0008] The present invention provides a metabolite-disease association prediction method based on hierarchical contrast learning and stepwise orthogonal multimodal fusion, including:
[0009] Calculate the semantic similarity, Gaussian kernel similarity, and information entropy-based similarity of diseases respectively; and splice the metabolite-disease adjacency matrix with the semantic similarity, Gaussian kernel similarity, and information entropy-based similarity of diseases to form three initial features of diseases.
[0010] Calculate the structural similarity, Gaussian kernel similarity, and information entropy-based similarity of metabolites respectively; and splice the metabolite-disease adjacency matrix with the structural similarity, Gaussian kernel similarity, and information entropy-based similarity of metabolites to form three initial features of metabolites.
[0011] Construct a hypergraph and perform hierarchical contrast learning to process the three initial features of diseases and the three initial features of metabolites, extract unimodal latent features, perform step-by-step orthogonal multimodal fusion on the unimodal latent features to obtain latent features, and send the extracted latent features into a multi-layer perceptron MLP for training and prediction.
[0012] Furthermore, calculate the semantic similarity, Gaussian kernel similarity, and information entropy-based similarity of diseases respectively, including:
[0013] Construct a directed acyclic graph with medical subject headings as disease descriptors, and calculate the semantic similarity of diseases based on the acyclic graph.
[0014] Input the disease-related vector into the Gaussian kernel function to obtain the disease Gaussian kernel similarity.
[0015] Utilize information entropy and the common information between diseases and metabolites to obtain the information entropy-based similarity of diseases.
[0016] Calculate the structural similarity, Gaussian kernel similarity, and information entropy-based similarity of metabolites respectively, including:
[0017] Input the SMELLS of metabolites into the PaDEL-Descriptor software to obtain vectors describing the chemical properties of each metabolite, and then calculate the structural similarity of metabolites based on these vectors.
[0018] Input the metabolite-related vector into the Gaussian kernel function to obtain the metabolite Gaussian kernel similarity.
[0019] Utilize information entropy and the common information between diseases and metabolites to obtain the information entropy-based similarity of metabolites.
[0020] Furthermore, before constructing a hypergraph and performing hierarchical contrast learning to process the three initial features of diseases and the three initial features of metabolites, it also includes:
[0021] Due to the imbalance between positive and negative samples, adopt a 1:1 sampling method to randomly select negative samples with the same number as the positive samples.
[0022] Further, construct a hypergraph, including: constructing a hypergraph based on the K-nearest neighbor algorithm;
[0023] Construct a hypergraph based on K-means.
[0024] Further, constructing a hypergraph based on the K-nearest neighbor algorithm includes:
[0025] When using the K-nearest neighbor algorithm to construct a hypergraph, calculate the Euclidean distance between nodes;
[0026] Select the nearest k among all distances and connect them with a hyperedge.
[0027] Further, constructing a hypergraph based on K-means includes:
[0028] When using K-means to construct a hypergraph, randomly select the cluster centers and update the clustering according to the Euclidean distance between each node and each cluster center. When the cluster centers no longer change, connect each class of nodes with a hyperedge.
[0029] Further, perform hierarchical contrast learning to process three initial features of diseases and three initial features of metabolites, including:
[0030] Perform hypergraph convolution operations on two hypergraphs for each initial feature to obtain KN features and KM features;
[0031] For the KN features and KM features obtained by hypergraph convolution, classify them as features in two views and perform intra-modal contrast learning respectively;
[0032] The intra-modal contrast learning includes: for the same node, use its two features in two views as positive samples, and use its features different from other nodes in this view and features different from other nodes in the other view as negative samples for training, aiming to seek the consistency of the same node and the differences of different nodes in different views; concatenate the features of the two views.
[0033] Further, perform step-by-step orthogonal multi-modal fusion on the single-modal latent features to obtain latent features, including:
[0034] Concatenate the corresponding row vectors of the two features, and the dimension changes from the original X dimension to 2X dimension, and then map it to the original dimension X through a fully connected layer;
[0035] Fuse the best-performing similarity network and the features after the first fusion in the same way to obtain the latent features of metabolites and diseases.
[0036] Further, the orthogonality of the vectors is controlled by the loss function, and a regularization term is introduced into the loss function to promote orthogonality between the two modalities being fused. The loss function is as follows:
[0037]
[0038] where Loss all is the loss for model contrastive learning and classification. The summation of the regularization terms promotes the orthogonality of the fused embeddings at each layer. i represents the number of fusions, and j represents the disease or metabolite.
[0039] Advantages and positive effects of the present invention: The metabolite-disease association prediction method provided by the present invention first obtains various similarities between metabolites and diseases, then extracts the features of metabolites and diseases through hypergraph convolutional contrastive learning and step-by-step orthogonal multi-modal fusion strategies, and finally inputs the extracted latent features into the MLP for metabolite-disease association prediction. Compared with common deep learning methods, the present invention introduces hypergraphs, contrastive learning, and step-by-step orthogonal multi-modal fusion strategies, which can capture the high-level neighbor information of nodes and can extract diverse and complementary features. The method of the present invention helps to improve the accuracy of metabolite-disease association prediction and has certain value for actual disease discovery, judging the disease development stage, and subsequent treatment of diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 is a flowchart of the metabolite-disease association prediction method in an embodiment of the present invention;
[0042] Figure 2 is a flowchart of hypergraph convolutional contrastive learning in an embodiment of the present invention;
[0043] Figure 3 is a schematic diagram of internal contrastive learning in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0045] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0046] As Figure 1 shown, the metabolite-disease association prediction method based on hierarchical contrast learning and step-by-step orthogonal multimodal fusion mainly includes three links: data preparation, model construction, and model testing.
[0047] (1) Data preparation:
[0048] Step 11: Calculate the similarity network.
[0049] Medical Subject Headings (MeSH) are used as disease descriptors to construct a directed acyclic graph (DAG), from which the semantic similarity of diseases is calculated;
[0050] The SMELLS (Simplified Molecular Input Line Entry System, a linear text encoding for describing the structure of organic molecules) of metabolites are input into the PaDEL-Descriptor software to obtain vectors describing the chemical properties of each metabolite, and then the structural similarity of metabolites is calculated based on these vectors. Among them, PaDEL-Descriptor is a chemical molecular descriptor calculation tool that can calculate a series of numerical or text data for describing the characteristics of atoms, molecules, and materials.
[0051] According to the Gaussian kernel function, the disease and metabolite-related vectors are respectively input into it to obtain the disease Gaussian kernel similarity and the metabolite Gaussian kernel similarity.
[0052] By using information entropy and the mutual information between diseases and metabolites, the similarity of diseases based on information entropy and the similarity of metabolites based on information entropy can be obtained.
[0053] Step 12: Form initial features.
[0054] For the three similarity networks of metabolites, the metabolite-disease adjacency matrix is concatenated with them to form three initial features; for the three similarity networks of diseases, the transpose of the metabolite-disease adjacency matrix is concatenated with them to form three initial features.
[0055] (II) Model construction:
[0056] Step 21: Negative sample extraction.
[0057] Due to the imbalance between positive and negative samples, a sampling method of 1:1 is adopted to randomly extract negative samples with the same number as the positive samples.
[0058] Step 22: Hypergraph convolutional contrast learning.
[0059] Hypergraph convolution: Taking the disease semantic similarity as an example, each row in the initial features constructed by the disease semantic similarity is used as the initial feature of each disease. When constructing a hypergraph using KNN, the Euclidean distance between nodes is first calculated, then the k nearest ones are selected and connected by a hyperedge; when constructing a hypergraph using K-means, the clustering centers are randomly selected and the clustering is updated according to the Euclidean distance between each node and each clustering center. When the clustering centers no longer change, each class of nodes is connected by a hyperedge. After constructing the hypergraph, hypergraph convolution operations are performed. For the two features obtained by hypergraph convolution, they are classified as features in two views, and intra-modal contrast learning is performed respectively. The specific process is as Figure 2 shown. Other similarities can be processed in the same way.
[0060] Intra-modal contrast learning: For the same node, its two features in two views are used as positive samples, and its features with different nodes in this view and the features with different nodes in the other view are used as negative samples for training to achieve the goal of seeking the consistency of the same node and the difference of different nodes in different views. The intra-modal contrast learning is as Figure 3 shown, and the other two views are processed through the above process. Then the features of the two views are concatenated.
[0061] After the above steps, the features corresponding to each similarity of metabolites and diseases can be obtained. Then, contrast learning between similarities is performed. Taking diseases as an example. Since the Gaussian kernel similarity has the best effect, taking the disease Gaussian kernel similarity as the benchmark, contrast learning is performed on the disease semantic similarity and the disease similarity based on information entropy.
[0062] Step 23: Gradual orthogonal multi-modal fusion.
[0063] To extract diverse and complementary features from the above features, a gradual orthogonal fusion method is used for multi-modal information fusion. To obtain the fusion order, experiments are conducted using single similarity, and their experimental effects are sorted. The two with the worst effects are fused first. The specific fusion process is to concatenate the corresponding row vectors of the two features, and the dimension changes from the original X dimension to 2X dimensions. Then, it is mapped to the original dimension X through a fully connected layer. Then, the similarity network with the best effect and the features after the first fusion are fused in the same way. Finally, the potential features of metabolites and diseases are obtained.
[0064] The orthogonality of the vectors is controlled by the loss function, and a regularization term is introduced in the loss function to promote orthogonality between the two fused modalities. Specifically as follows:
[0065]
[0066] Where Lossall is the loss of the model's contrastive learning and classification. The subsequent sum of the regularization terms promotes the orthogonality of the embeddings fused in each layer. i represents the number of fusion times, and j represents diseases or metabolites.
[0067] Step 24: Feed the extracted potential features into a multi-layer perceptron (MLP) for training and prediction.
[0068] (III) Model testing:
[0069] Step 31: Use various similarity calculation methods to obtain metabolite-disease similarities. Process each similarity by constructing a hypergraph and performing hierarchical contrastive learning, extract single-modal potential features, perform gradual orthogonal multi-modal fusion on the single-modal potential features to obtain potential features, and finally feed the extracted potential features into the MLP for training and prediction.
[0070] Step 32: Test various parameters and important modules that affect the model performance. Observe the influence of the hypergraph convolutional contrastive learning and the gradual orthogonal multi-modal fusion module in the model on the model test results. Develop model variants that only contain one hypergraph construction method, remove the contrastive learning within one similarity, remove the contrastive learning between similarities, change the multi-modal information fusion to simple addition and concatenation, etc., and test the importance of different modules to the overall model results.
[0071] Step 33: Test the model's ability to identify potential metabolite-disease associations. For several common diseases in real society, predict and observe the model's recognition of related metabolites.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A metabolite-disease association prediction method based on hierarchical contrastive learning and stepwise orthogonal multimodal fusion, characterized in that Including: Calculate the semantic similarity, Gaussian kernel similarity, and information entropy-based similarity of diseases respectively; And splice the metabolite-disease adjacency matrix with the semantic similarity, Gaussian kernel similarity, and information entropy-based similarity of diseases to form three initial features of diseases; Calculate the structural similarity, Gaussian kernel similarity, and information entropy-based similarity of metabolites respectively; and splice the metabolite-disease adjacency matrix with the structural similarity, Gaussian kernel similarity, and information entropy-based similarity of metabolites to form three initial features of metabolites; Construct a hypergraph and perform hierarchical contrast learning to process the three initial features of diseases and the three initial features of metabolites, extract unimodal latent features, perform step-by-step orthogonal multimodal fusion on the unimodal latent features to obtain latent features, and send the extracted latent features into a multi-layer perceptron MLP for training and prediction.
2. The metabolite-disease association prediction method based on hierarchical contrastive learning and progressive orthogonal multimodal fusion according to claim 1, wherein, Calculate the semantic similarity, Gaussian kernel similarity, and information entropy-based similarity of diseases respectively, including: Construct a directed acyclic graph using Medical Subject Headings as disease descriptors, and calculate the semantic similarity of diseases based on the acyclic graph; Input the disease-related vector into the Gaussian kernel function to obtain the disease Gaussian kernel similarity; Utilize information entropy and the common information between diseases and metabolites to obtain the information entropy-based similarity of diseases; Calculate the structural similarity, Gaussian kernel similarity, and information entropy-based similarity of metabolites respectively, including: Input the SMELLS of metabolites into the PaDEL-Descriptor software to obtain vectors describing the chemical properties of each metabolite, and then calculate the structural similarity of metabolites based on these vectors; Input the metabolite-related vector into the Gaussian kernel function to obtain the metabolite Gaussian kernel similarity; Utilize information entropy and the common information between diseases and metabolites to obtain the information entropy-based similarity of metabolites.
3. A metabolite-disease association prediction method based on hierarchical contrast learning and progressive orthogonal multimodal fusion according to claim 1, characterized in that Before constructing a hypergraph and performing hierarchical contrast learning to process the three initial features of diseases and the three initial features of metabolites, it also includes: Due to the imbalance between positive and negative samples, adopt a 1:1 sampling method to randomly select negative samples with the same number as the positive samples.
4. A metabolite-disease association prediction method based on hierarchical contrast learning and step-by-step orthogonal multimodal fusion according to claim 1, characterized in that Construct a hypergraph, including: constructing a hypergraph based on the K-nearest neighbor algorithm; Constructing a hypergraph based on K-means.
5. A metabolite-disease association prediction method based on hierarchical contrast learning and stepwise orthogonal multimodal fusion according to claim 4, characterized in that, Constructing a hypergraph based on the K-nearest neighbor algorithm, including: When using the K-nearest neighbor algorithm to construct a hypergraph, calculate the Euclidean distance between nodes; Select the nearest k among all distances and connect them with a hyperedge.
6. The metabolite-disease association prediction method based on hierarchical contrast learning and stepwise orthogonal multimodal fusion according to claim 4, wherein Constructing a hypergraph based on K-means, including: When using K-means to construct a hypergraph, randomly select cluster centers and update the clustering according to the Euclidean distance between each node and each cluster center. When the cluster centers no longer change, connect each class of nodes with a hyperedge.
7. A metabolite-disease association prediction method based on hierarchical contrast learning and stepwise orthogonal multimodal fusion according to claim 4, characterized in that Perform hierarchical contrast learning to process the three initial features of diseases and the three initial features of metabolites, including: Perform hypergraph convolution operations on the two hypergraphs for each initial feature to obtain KN features and KM features; For the KN features and KM features obtained by hypergraph convolution, classify them as features in two views and perform intra-modal contrast learning respectively; The modal internal contrast learning includes: for the same node, using its two features in two views as positive samples, and using the features of different nodes from this view and the features of different nodes in the other view as negative samples for training, to achieve the goal of seeking the consistency of the same node and the difference of different nodes in different views; splicing the features of the two views together.
8. A metabolite-disease association prediction method based on hierarchical contrast learning and step-by-step orthogonal multimodal fusion according to claim 1, characterized in that Gradually performing orthogonal multi-modal fusion on the single-modal latent features to obtain latent features, including: Splicing the corresponding row vectors of the two features together, changing the dimension from the original X dimension to 2X dimensions, and then mapping it to the original dimension X through a fully connected layer; Fusing the similarity network with the best effect and the features after the first fusion in the same way to obtain the latent features of metabolites and diseases.
9. The metabolite-disease association prediction method based on hierarchical contrast learning and step-by-step orthogonal multi-modal fusion according to claim 8, characterized in that The orthogonality of the vectors is controlled by the loss function, and a regularization term is introduced in the loss function to promote orthogonality between the two fused modalities. The loss function is: Among them, Loss all is the loss for the model's contrastive learning and classification. The summation of the regularization terms encourages the embeddings fused at each layer to be orthogonal. Here, i represents the number of fusion times, and j represents the disease or metabolite.