Metabolite-disease association prediction method

By constructing a metabolite-disease-gene multi-source fusion heterogeneous network and local-global feature extraction model LGDFE, combined with the deep learning algorithm CNN, the problems of low accuracy and high cost of metabolite-disease association prediction are solved, and efficient and accurate metabolite-disease association prediction are achieved.

CN120299729APending Publication Date: 2025-07-11QIQIHAR UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510358344.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has low accuracy and high cost in predicting metabolite-disease associations, traditional methods are time-consuming and labor-material investments are huge.

Method used

A metabolite-disease-gene multi-source fusion heterogeneous network was constructed, combined with the local-global feature extraction model LGDFE and the deep learning algorithm CNN, and the deep learning algorithm CNN were extracted through the metabolite Gaussian interaction nuclear spectrum similarity, disease semantic similarity and gene association matrix, deep-level relationships in the multi-source heterogeneous network were extracted, and different levels of characteristics were automatically fused to predict the association relationship between metabolites and diseases.

Benefits of technology

It significantly improves the accuracy and efficiency of metabolite-disease association prediction, reduces experimental costs, and can automatically fuse features at different levels, solving the problem that traditional methods are difficult to explore deep-level connections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299729A_ABST
    Figure CN120299729A_ABST
Patent Text Reader

Abstract

The invention discloses a metabolite-disease association prediction method, and belongs to the technical field of bioinformatics. According to the method, a metabolite-disease-gene multi-source fusion heterogeneous network is constructed by using metabolite omics information, not only is the relationship among different biological substances considered, but also the problem of matrix sparsity is solved, and a local-global feature extraction model LGDFE is proposed, so that the method has the advantages of being high in robustness and high in robustness. The deep relationship in the multi-source fusion heterogeneous network is extracted from global and local angles, the problem that the deep relationship between metabolites and diseases is difficult to fully excavate in a traditional method is effectively solved, the deep learning algorithm CNN is used for predicting data, features of different levels can be automatically fused, the parameter complexity of a common classifier is reduced, and the classification efficiency is improved. The unknown correlation between the metabolite and the disease can be predicted, the prediction efficiency is improved, the correlation between the metabolite and the disease does not need to be obtained through a large number of experiments, and the experiment cost is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for predicting metabolite-disease associations, belonging to the technical field of bioinformatics. Background Art

[0002] Metabolites, also known as intermediate metabolites, refer to organic molecules or inorganic ions involved in the metabolic processes within living organisms. They are chemical substances within cells and organisms, and the content and types of metabolites are affected by the physiological state, disease state, and environmental factors of the organism. In recent years, metabolomics has emerged. The study of metabolomics aims to reveal the biochemical processes within living organisms, the interactions between different metabolites, and the associations between metabolites and diseases by large-scale determination and analysis of the quantity and types of metabolites.

[0003] With the continuous progress of medicine, numerous scientists have revealed the reasons for abnormal changes in metabolite content in the bodies of certain patients and used advanced means of metabolomics to explore disease treatment strategies. Research has shown that analyzing the body fluid metabolites of Parkinson's disease patients and Parkinson's high-risk patients can enable early diagnosis of the disease or monitoring of disease progression. There are also many rare human genetic diseases, such as mental illnesses, which show disorders of single endogenous small molecule metabolites. Some studies have shown that disorders or changes in endogenous central nervous system metabolites have a greater impact on certain mental illnesses. In addition, by studying the urine microbiota and serum metabolites of patients with diabetic kidney disease (DKD), researchers have discovered the effects of arginine metabolites on DKD patients. All of the above studies indicate that the levels of metabolites in the human body have a non-negligible impact on diseases. Therefore, studying the associations between metabolites and diseases is beneficial for disease prevention, early diagnosis, and treatment, and is crucial for disease physiology research and clinical treatment of diseases.

[0004] Previous studies mostly relied on traditional methods, which were not only time-consuming but also required huge investments in human and material resources. In contrast, using an efficient computational model to predict the associations between metabolites and diseases can significantly reduce the biological experimental steps and greatly improve the research efficiency. Summary of the Invention

[0005] The present invention solves the problems of low accuracy and high cost in predicting metabolite-disease associations in the prior art, and proposes a method for predicting metabolite-disease associations.

[0006] A method for predicting metabolite-disease associations, characterized by comprising the following steps:

[0007] Step 1: Construct a metabolite-disease association matrix A1 based on the metabolite-disease association dataset, a disease-gene association matrix A2 based on the disease-gene association dataset, and a metabolite-gene association matrix A3 based on the metabolite-gene association dataset;

[0008] Step 2: Obtain a metabolite Gaussian interaction kernel spectral similarity matrix M K and a disease Gaussian interaction spectral kernel similarity matrix D K ;

[0009] Step 3: Input the chemical formula of the metabolite into SMILES (Simplified Molecular Input Line Entry System) to obtain a metabolite molecular fingerprint similarity matrix M S , obtain medical subject term descriptors from the MeSH (Medical Subject Headings) database, visualize the topology of each disease as a directed acyclic graph DAG according to the medical subject term descriptors, calculate the semantic contribution degree and semantic value of each disease, calculate the semantic similarity between different diseases according to the semantic contribution degree and semantic value of each disease, and obtain a disease semantic similarity matrix D S ;

[0010] Step 4: Fuse the metabolite Gaussian interaction kernel spectral similarity matrix M K obtained in Step 2 with the metabolite molecular fingerprint similarity matrix M S obtained in Step 3 to obtain a metabolite fusion similarity matrix M F , fuse the disease Gaussian interaction spectral kernel similarity matrix D K obtained in Step 2 with the disease semantic similarity matrix D S obtained in Step 3 to obtain a disease fusion similarity matrix D F ;

[0011] Step 5: Combine the metabolite-disease association matrix A1, the disease-gene association matrix A2, and the metabolite-gene association matrix A3 obtained in Step 1 with the metabolite fusion similarity matrix M F and the disease fusion similarity matrix D F obtained in Step 4 to form a metabolite-disease-gene multi-source heterogeneous network G;

[0012] Step 6: Construct a Local-Global Feature Extraction LGDFE model. The LGDFE model includes a local feature extraction module and a global feature extraction module. The local feature extraction module performs a random walk on the multi-source heterogeneous network G in Step 5 based on a predetermined meta-path, extracts the context information of the nodes through the skip-gram model to obtain node local features, and trains the skip-gram model to obtain a metabolite local feature representation matrix M l , a disease local feature representation matrix D l ;

[0013] Use the multi-source heterogeneous network G to train the global feature extraction module of the LGDFE model to obtain a trained global feature extraction module. Input the local feature representation into the trained global feature extraction module to obtain the metabolite local-global feature matrix M lg and the disease local-global feature matrix D lg ;

[0014] Step seven: Concatenate the metabolite local-global feature matrix M lg and the disease local-global feature matrix D lg to obtain the final embedded feature T. Use the metabolite-disease association matrix A1 in step one as the label set, and the final embedded feature T as the training set. Perform positive and negative sampling on the label set and the training set to make the number of positive and negative samples the same. Use the sampled label set and training set to train the CNN model until the specified number of training times is reached to obtain a trained CNN model;

[0015] Step eight: Predict the association relationship between metabolites and diseases: Obtain the metabolite local feature matrix and global feature matrix of the association relationship to be predicted, and the disease local feature matrix and global feature matrix of the association relationship to be predicted. Associate the metabolite-disease association matrix A1, the metabolite local-global feature matrix of the association relationship to be predicted, and the disease local-global matrix of the association relationship to be predicted to obtain the metabolite-disease final embedded feature of the association relationship to be predicted. Input the metabolite-disease final embedded feature of the association relationship to be predicted into the trained CNN model to obtain the association relationship between the metabolite and disease of the association relationship to be predicted.

[0016] The metabolite-disease association matrix A1, disease-gene association matrix A2, and metabolite-gene association matrix A3 in step one are as follows:

[0017]

[0018] A1(i,j) is the element in the i-th row and j-th column of matrix A1, N m is the number of metabolite species, N d is the number of disease species, d i is the i-th metabolite, m j is the j-th disease, g k is the k-th gene.

[0019] The process of obtaining the metabolite Gaussian interaction kernel spectral similarity matrix M K and the disease Gaussian interaction spectral kernel similarity matrix D K from the metabolite-disease association matrix A1 in step two is as follows:

[0020]

[0021] M k (m i ,m i' ) is the Gaussian interaction spectrum kernel similarity between the i-th metabolite and the i'-th metabolite, D k (d j ,d j' ) is the Gaussian interaction spectrum kernel similarity between the j-th disease and the j'-th disease, η m is the normalized metabolite kernel bandwidth, η d is the normalized disease kernel bandwidth, η′ m is the original metabolite kernel bandwidth, η′ d is the original disease kernel bandwidth, is the i-th row vector of matrix A1, is the i'-th row vector of matrix A1, is the j-th column vector of matrix A1, is the j'-th column vector of matrix A1.

[0022] In step three, the metabolite molecular fingerprint similarity matrix M S is:

[0023]

[0024] Define the semantic value of disease d i as the sum of the contributions of all disease pairs in DAG(d i ) to the semantics of disease d i . In DAG(d i ), the closer a disease is to disease d i , the greater its semantic contribution to it. The semantic contribution degree of disease d i is as:

[0025]

[0026] t is the entry in the MeSH database that describes the disease, i.e., the disease node of the directed acyclic graph, t′ is the sub-entry, ω e is the semantic contribution factor of the edge connecting entry t and sub-entry t′, is the set of edges in the directed acyclic graph;

[0027] The semantic value SV(d i ) of disease d i is:

[0028]

[0029] is the entry set of disease d i ;

[0030] For a given disease d i and d j , the semantic similarity between the two is:

[0031]

[0032] In step four, the metabolite fusion similarity matrix M F and the disease fusion similarity matrix D F are:

[0033]

[0034] M F (m i , m i' ) is the fusion similarity between metabolite m i and metabolite m i' , and D F (d j , d j' ) is the fusion similarity between disease d j and disease d j' .

[0035] In step five, the metabolite-disease-gene multi-source heterogeneous network G is:

[0036]

[0037] In step six, the local feature extraction module randomly walks on the multi-source heterogeneous network G described in step five based on a predetermined meta-path, extracts the context information of the nodes through the skip-gram model to obtain the node local features, and trains the skip-gram model to obtain the metabolite local feature representation matrix M l and the disease local feature representation matrix D l . The process is as follows: Randomly walk on the multi-source heterogeneous network G using the predetermined meta-path to obtain a set of node sequences V = {v0, v1,..., v t} containing metabolites, diseases, and genes. The probability of random walk is P. Then, in the random walk at step t, the transition probability from node v t-1 to node v t is:

[0038]

[0039] N(v t-1 ) is the set of neighbor nodes of node v t-1 , represents the similarity or association strength between node v t-1 and node v t , Denote node v t-1 The similarity or association strength with node v k , and the transition probability P(v t |v t-1 ) is used to generate the node sequence;

[0040] Use one - hot variables to encode the node sequence, and map the node sequence to a low - dimensional embedding space with a fully - connected layer to obtain metabolite, disease, and gene feature representations y m , y d and y g :

[0041] y m =W M x m +b m

[0042] y d =W D x d +b d

[0043] y g =W G x g +b g

[0044] W M 、W D and W G are the weight matrices of the fully - connected layer, b m 、b d and b g are the bias vectors of the fully - connected layer, x m 、x d and x g are the node vector representations of metabolites, diseases, and genes after one - hot processing;

[0045] Use the skip - gram model to extract node context information to obtain node local features:

[0046]

[0047] V is the number of nodes, y i is the embedding representation of the given central node v i ;

[0048] To maximize the central node v i , train the skip - gram model, and set the objective function of the training model as:

[0049]

[0050] Let \(D\) be the set of all central nodes and context nodes in the training dataset, and \(W\) represent the weight matrix. Since calculating the normalization term of softmax is costly, skip-gram uses negative sampling to approximate this objective, and the objective function is:

[0051]

[0052] \(\sigma(x)\) is the sigmoid function, \(K\) is the number of negative samples, and \(P\) n (v) is the negative sample distribution.

[0053] In step six, the global feature extraction module includes an encoder unit and a decoder unit. Both units incorporate the multi-head self-attention mechanism to capture metabolite and disease features:

[0054]

[0055] \(Z = [Z_1, Z_2, \ldots, Z\) h is the output of the multi-head self-attention layer, \(h\) is the number of heads, \(Q\) (i) , \(K\) (i) and \(V\) (i) respectively represent the query, key, and value of the \(i\)-th head, \(W\) Q , \(W\) K and \(W\) V are learnable weight matrices, \(d\) k is the dimension of the key matrix, \(H=\{h_1, h_2, \ldots, h\) n \} is the node feature representation, that is, the metabolite local feature representation matrix \(M\) l and the disease local feature representation matrix \(D\) l ;

[0056] Both the encoder unit and the decoder unit consist of a multi-head self-attention layer and a linear layer. The multi-head self-attention layer of the encoder unit is:

[0057] \(E = ReLU(Z\cdot W\) o +b o )

[0058] ReLU is the activation function, \(Z\) is the encoder input feature matrix, \(W\) o is the encoder initial weight matrix, \(b\) o is the encoder bias vector;

[0059] The multi-head self-attention layer of the decoder unit is:

[0060] \(D = ReLU(Z'\cdot W'+b')\) o +b' o )

[0061] ReLU is the activation function, Z' is the decoder input feature matrix, and W o ' is the initial weight matrix of the decoder, and b o ' is the decoder bias vector;

[0062] The loss function used to train the global feature extraction module of the LGDFE model in step six is:

[0063]

[0064] h i represents the initial sample features, represents the sample features after auto - encoding reconstruction, and n represents the number of samples.

[0065] The final embedded feature T obtained in step seven is:

[0066]

[0067] is the i - th row disease feature of M lg , is the j - th row disease feature of D lg , is and 's concatenated combination, and T i' is the i'-th row of matrix T;

[0068] The process of training the CNN model is as follows:

[0069] Divide the data sets x train and x test and the label sets y train and y test . Input the training set x train into the CNN model. First, it passes through the first convolutional layer and then through the first pooling layer:

[0070] A (1) = MaxPooling(Conv1D(X test , W (1) , b (1) ))

[0071] Then the output of the first pooling layer is connected to the second convolutional layer, and the output of the second convolutional layer is connected to the second pooling layer:

[0072] A (2) = MaxPooling(Conv1D(A (1) , W (2) , b (2) ))

[0073] Flatten the data output from the second pooling layer, input it into the linear layer, then input it into the dropout layer for regularization to prevent overfitting of the model, and then input it into the second linear layer. Finally, pass through the output layer, and the activation function of the output layer is set to Sigmoid:

[0074]

[0075] Z′ = Dropout(Z)

[0076]

[0077] The loss function used for training the CNN model is:

[0078]

[0079] where N is the number of known associations between metabolites and diseases, y i represents the label of the positive example, represents the probability of being predicted as a positive example, 1 - y i represents the label of the negative example, represents the probability of being predicted as a negative example.

[0080] The beneficial effects of the present invention are:

[0081] The present invention constructs a multi-source fusion heterogeneous network of metabolite-disease-gene by using metabolomics information, which not only considers the relationships between different biological substances, but also solves the problem of matrix sparsity, and proposes a local-global feature extraction model LGDFE to extract deep relationships in the multi-source fusion heterogeneous network from both global and local perspectives, effectively solving the problem that traditional methods are difficult to fully mine the deep connections between metabolites and diseases. Moreover, the deep learning algorithm CNN is used to predict data, which can automatically fuse features at different levels and reduce the parameter complexity of ordinary classifiers, facilitating the prediction of unknown associations between metabolites and diseases, improving the prediction efficiency. And the present invention does not need to obtain the association relationship between metabolites and diseases through a large number of experiments, significantly reducing the experimental cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 is a flowchart of the metabolite-disease association prediction method;

[0083] Figure 2 is a detailed process diagram of the metabolite-disease association prediction method;

[0084] Figure 3 is a schematic diagram of the metabolite-disease-gene multi-source heterogeneous network;

[0085] Figure 4 is a structural diagram of the global feature extraction module of the LGDFE model;

[0086] Figure 5 It is the structure diagram of the CNN model;

[0087] Figure 6 They are the ROC graphs obtained by the LGDFE model under 2-fold, 5-fold, and 10-fold cross-validation. Specific implementation manner

[0088] This implementation manner is a metabolite-disease association prediction method, including the following steps:

[0089] Step 1: Construct a metabolite-disease association matrix A1 based on the metabolite-disease association dataset, construct a disease-gene association matrix A2 based on the disease-gene association dataset, and construct a metabolite-gene association matrix A3 based on the metabolite-gene association dataset:

[0090]

[0091] A1(i,j) is the element in the i-th row and j-th column of matrix A1, N m is the number of metabolite species, N d is the number of disease species, d i is the i-th metabolite, m j is the j-th disease, g k is the k-th gene.

[0092] Step 2: Obtain the metabolite Gaussian interaction kernel spectral similarity matrix M K and the disease Gaussian interaction spectral kernel similarity matrix D K :

[0093]

[0094] M k (m i ,m i' ) is the Gaussian interaction spectral kernel similarity between the i-th metabolite and the i'-th metabolite, D k (d j ,d j' ) is the Gaussian interaction spectral kernel similarity between the j-th disease and the j'-th disease, η m is the normalized metabolite kernel bandwidth, η d is the normalized disease kernel bandwidth, η′ m is the original metabolite kernel bandwidth, η′ d is the original disease kernel bandwidth, is the i-th row vector of matrix A1, is the i'-th row vector of matrix A1, is the j-th column vector of matrix A1, is the j'-th column vector of matrix A1.

[0095] Step 3: Input the chemical formula of the metabolite into SMILES (Simplified Molecular Input Line Entry Specification) to obtain the metabolite molecular fingerprint similarity matrix M S :

[0096]

[0097] M s (m i , m i' ) is the molecular fingerprint similarity between metabolite m i and metabolite m i' ;

[0098] Obtain MeSH (Medical Subject Headings) database medical subject heading descriptors, and visualize the topology of each disease as a directed acyclic graph DAG. Define the semantic value of each disease d i as the sum of the semantic contributions of all diseases to this disease d i in DAG(d i ). Calculate the semantic contribution degree S of each disease accordingly di :

[0099]

[0100] t is the entry in the MeSH database that describes the disease, i.e., the disease node of the directed acyclic graph, t' is the sub-entry, and ω e is the semantic contribution factor of the edge connecting entry t and sub-entry t'; is the set of edges in the directed acyclic graph;

[0101] The semantic value SV(d i ) of disease d i is:

[0102]

[0103] is the entry set of disease d i ;

[0104] For the given diseases d i and d j , the semantic similarity between the two is:

[0105]

[0106] Obtain the disease semantic similarity matrix D S .

[0107] Step 4: Combine the metabolite Gaussian interaction kernel spectral similarity matrix M K from Step 2 and the metabolite molecular fingerprint similarity matrix M S from Step 3, as shown in Figure 2 A, to obtain the metabolite combined similarity matrix M F :

[0108]

[0109] Combine the disease Gaussian interaction spectral kernel similarity matrix D K from Step 2 and the disease semantic similarity matrix D S from Step 3, as shown in Figure 2 A, to obtain the disease combined similarity matrix D F :

[0110]

[0111] Step 5: Combine the metabolite-disease association matrix A1, disease-gene association matrix A2, and metabolite-gene association matrix A3 from Step 1 with the metabolite combined similarity matrix M F and the disease combined similarity matrix D F to form a metabolite-disease-gene multi-source heterogeneous network G, as shown in Figure 3 : Specifically:

[0112]

[0113] Step 6: Construct a Local-Global Feature Extraction (LGDFE) model, which includes a local feature extraction module and a global feature extraction module;

[0114] First, randomly walk on the multi-source heterogeneous network G using a predefined meta-path, as shown in Figure 2 B(1), to obtain a set of node sequences V = {v0, v1,..., v t}, where the probability of random walk is P. Then, in the random walk at step t, the transition probability from v t-1 to v t is:

[0115]

[0116] N(v t-1 ) is the set of neighbor nodes of node v t-1 , represents the similarity or association strength between node v t-1 and node v t , represents the similarity or association strength between node v t-1 and node v kThe similarity or association strength, the transition probability P(v t |v t-1 ) is used to generate the node sequence;

[0117] Then, the node sequence is encoded using one-hot variables, and a fully connected layer is used to map the node sequence into a low-dimensional embedding space to obtain metabolite, disease, and gene feature representations y m 、y d and y g :

[0118] y m =W M x m +b m

[0119] y d =W D x d +b d

[0120] y g =W G x g +b g

[0121] W M 、W D and W G are the weight matrices of the fully connected layer, b m 、b d and b g are the bias vectors of the fully connected layer, x m 、x d and x g are the node vector representations of metabolites, diseases, and genes after one-hot processing;

[0122] Then, the skip-gram model is used to extract the node context information to obtain the node local features:

[0123]

[0124] V is the number of nodes, y i is the embedding representation of the given central node v i ;

[0125] To maximize the central node v i , the skip-gram model is trained, and the objective function of the training model is set as:

[0126]

[0127] Let \(D\) be the set of all central nodes and context nodes in the training dataset, and \(W\) represent the weight matrix. Since calculating the normalization term of softmax is costly, skip-gram uses negative sampling to approximate this objective, and the objective function is:

[0128]

[0129] \(\sigma(x)\) is the sigmoid function, \(K\) is the number of negative samples, and \(P\) n (v) is the negative sample distribution.

[0130] The node local features are obtained through the skip-gram model, and the skip-gram model is trained to obtain the metabolite local feature representation matrix \(M\) l and the disease local feature representation matrix \(D\) l .

[0131] Then the global features are extracted. The global feature extraction module includes an encoder unit and a decoder unit, as Figure 2 shown in B(3). Both units incorporate the multi-head self-attention mechanism to capture metabolite and disease features:

[0132]

[0133] \(Z = [Z_1, Z_2,..., Z\) h is the output of the multi-head self-attention layer, \(h\) is the number of heads, \(Q\) (i) , \(K\) (i) and \(V\) (i) represent the query, key, and value of the \(i\)-th head respectively, \(W\) Q , \(W\) K and \(W\) V are learnable weight matrices, \(d\) k is the dimension of the key matrix, and \(H=\{h_1, h_2,..., h\) n \} is the node feature representation, that is, the metabolite local feature representation matrix \(M\) l and the disease local feature representation matrix \(D\) l ;

[0134] Both the encoder unit and the decoder unit are composed of a multi-head self-attention layer and a linear layer. The multi-head self-attention layer of the encoder unit is:

[0135] \(E = ReLU(Z\cdot W\) o + b o )

[0136] ReLU is the activation function, \(Z\) is the encoder input feature matrix, \(W\) o is the encoder initial weight matrix, and \(b\) o is the encoder bias vector;

[0137] The multi-head self-attention layer of the decoder unit is as follows:

[0138] D = ReLU(Z′·W′ o + b′ o )

[0139] ReLU is the activation function, Z' is the decoder input feature matrix, W o ' is the decoder initial weight matrix, b o ' is the decoder bias vector;

[0140] Use the multi-source heterogeneous network G to train the global feature extraction module of the LGDFE model to obtain the trained global feature extraction module. The loss function of the training model is:

[0141]

[0142] h i represents the initial sample feature, represents the sample feature after auto-encoding reconstruction, and n represents the number of samples;

[0143] Input the local feature representation into the trained global feature extraction module to obtain the metabolite local-global feature matrix M lg and the disease local-global feature matrix D lg .

[0144] Step 7: Concatenate the metabolite local-global feature matrix M lg and the disease local-global feature matrix D lg to obtain the final embedded feature T:

[0145]

[0146] is the disease feature of the i-th row of M lg , is the disease feature of the j-th row of D lg , is the concatenation combination of and , and T i' is the i'-th row of the matrix T;

[0147] Use the metabolite-disease association matrix A1 described in Step 1 as the label set, and the final embedded feature T as the training set, and perform positive and negative sampling on the label set and the training set to make the number of positive and negative samples the same:

[0148]

[0149] X is the training set, y ijLabel set, \(P = \{(i, j) | A 1ij = 1\}\) represents selecting samples where \(A 1ij = 1\) from the association matrix \(A1\) as the positive sample set, \(N'=\{(i, j) | A 1ij = 0\}\) selects samples where \(A 1ij = 0\) from the association matrix \(A1\) as the negative sample set, and randomly samples \(|P|\) negative samples to make the number of positive and negative samples equal. Specifically:

[0150]

[0151] Use the sampled label set and training set to train the CNN model until the set number of training times is reached to obtain the trained CNN model. The process of training the CNN model is as follows:

[0152] Divide the data set \(x train and \(x test and the label set \(y train and \(y test . Input the training set \(x train into the CNN model. As shown in Figure 2 C, it first passes through the first convolutional layer and then through the first pooling layer:

[0153] A (1) = MaxPooling(Conv1D(X test , W (1) , b (1) ))

[0154] Then the output of the first pooling layer is connected to the second convolutional layer, and the output of the second convolutional layer is connected to the second pooling layer:

[0155] A (2) = MaxPooling(Conv1D(A (1) , W (2) , b (2) ))

[0156] Flatten the data output by the second pooling layer, input it into the linear layer, then input it into the dropout layer for regularization to prevent overfitting of the model, then input it into the second linear layer, and finally pass through the output layer. The activation function of the output layer is set to Sigmoid:

[0157]

[0158] Z' = Dropout(Z)

[0159]

[0160] The loss function used for training the CNN model is:

[0161]

[0162] N is the number of known associations between metabolites and diseases, y i represents the label of the positive example, represents the probability of being predicted as a positive example, 1 - y i represents the label of the negative example, represents the probability of being predicted as a negative example.

[0163] Step Eight: Predict the association between metabolites and diseases: Obtain the metabolite local feature matrix and global feature matrix of the association to be predicted, the disease local feature matrix and global feature matrix of the association to be predicted, associate the metabolite and disease association matrix A1, the metabolite local - global feature matrix of the association to be predicted, and the disease local - global matrix of the association to be predicted to obtain the final embedded feature of the metabolite - disease of the association to be predicted, and input the final embedded feature of the metabolite - disease of the association to be predicted into the trained CNN model to obtain the association between the metabolite and disease of the association to be predicted.

[0164] Example:

[0165] To verify the beneficial effects of the present invention, the following experiments were conducted:

[0166] In this example, 2 - fold, 5 - fold, and 10 - fold cross - validations were used to evaluate the performance of the evaluation model LGDFE proposed by the present invention. The ROC images obtained based on 2 - CV, 5 - CV, and 10 - CV are as Figure 6 shown, and the comparison of AUC values between 2 - CV, 5 - CV, 10 - CV and other models is shown in Table 1.

[0167] Table 1 Comparison of AUC values of the LGDFE model and other models under 2 - fold, 5 - fold, and 10 - fold cross - validations on the same dataset

[0168]

[0169] In the prediction results, the present invention proved the top 10 metabolites related to Alzheimer's disease and the top 10 metabolites related to type 2 diabetes through relevant literature. The verification results are shown in Tables 2 and 3.

[0170] Table 2 Comparison results of the top ten metabolites related to Alzheimer's disease

[0171]

[0172] Table 3 Comparison results of the top ten metabolites related to type 2 diabetes

[0173]

[0174]

[0175] Through performance evaluation, the AUC values of the LGDFE model under 2-, 5-, and 10-fold cross-validation all reached above 0.96, indicating that the present invention has application value.

Claims

1. A metabolite-disease association prediction method, characterized in that, It includes the following steps: Step 1: Construct a metabolite-disease association matrix A1 based on the metabolite-disease association dataset, construct a disease-gene association matrix A2 based on the disease-gene association dataset, and construct a metabolite-gene association matrix A3 based on the metabolite-gene association dataset; Step 2: Obtain the metabolite Gaussian interaction kernel spectral similarity matrix M according to the metabolite-disease association matrix A1 K and the disease Gaussian interaction spectral kernel similarity matrix D K ; Step 3: Input the chemical formula of the metabolite into the SMILES system to obtain the metabolite molecular fingerprint similarity matrix M S , obtain MeSH descriptors from the MeSH database, visualize the topology of each disease as a directed acyclic graph (DAG) according to the MeSH descriptors, calculate the semantic contribution degree and semantic value of each disease, and calculate the semantic similarity between different diseases based on the semantic contribution degree and semantic value of each disease to obtain the disease semantic similarity matrix D S ; Step 4: Combine the metabolite Gaussian interaction kernel spectral similarity matrix M K from Step 2 with the metabolite molecular fingerprint similarity matrix M S from Step 3 to obtain the metabolite fusion similarity matrix M F . Combine the disease Gaussian interaction spectral kernel similarity matrix D K from Step 2 with the disease semantic similarity matrix D S from Step 3 to obtain the disease fusion similarity matrix D F ; Step 5. Combine the metabolite-disease association matrix A1, disease-gene association matrix A2, and metabolite-gene association matrix A3 in Step 1 with the metabolite fusion similarity matrix M F , disease fusion similarity matrix D F to form a metabolite-disease-gene multi-source heterogeneous network G; Step 6: Construct a Local-Global Feature Extraction (LGDFE) model. The LGDFE model includes a local feature extraction module and a global feature extraction module. The local feature extraction module randomly walks on the multi-source heterogeneous network G in Step 5 based on a predetermined meta-path, extracts the context information of nodes through the skip-gram model to obtain node local features, trains the skip-gram model, and obtains the metabolite local feature representation matrix M l , disease local feature representation matrix D l ; Use the multi-source heterogeneous network G to train the global feature extraction module of the LGDFE model to obtain a trained global feature extraction module, input the local feature representation into the trained global feature extraction module, and obtain the metabolite local-global feature matrix M lg and the disease local-global feature matrix D lg ; Step 7: Concatenate the metabolite local-global feature matrix M lg and the disease local-global feature matrix D lg to obtain the final embedded feature T. Use the metabolite-disease association matrix A1 in Step 1 as the label set and the final embedded feature T as the training set. Perform positive and negative sampling on the label set and the training set to make the number of positive and negative samples the same. Use the sampled label set and training set to train the CNN model until the specified number of training times is reached to obtain the trained CNN model; Step 8: Predict the association relationship between metabolites and diseases: Obtain the local and global feature matrices of metabolites and the local and global feature matrices of diseases for which the association relationship is to be predicted. Associate the metabolite-disease association matrix A1, the local-global feature matrix of metabolites for which the association relationship is to be predicted, and the local-global matrix of diseases for which the association relationship is to be predicted to obtain the final embedded features of metabolites and diseases for which the association relationship is to be predicted. Input the final embedded features of metabolites and diseases for which the association relationship is to be predicted into the trained CNN model to obtain the association relationship between metabolites and diseases for which the association relationship is to be predicted.

2. The metabolite-disease association prediction method according to claim 1, wherein The metabolite-disease association matrix A1, disease-gene association matrix A2, and metabolite-gene association matrix A3 in Step 1 are: A1(i,j) is the element in the i-th row and j-th column of matrix A1, N m is the number of metabolite species, N d is the number of disease species, d i is the i-th metabolite, m j is the j-th disease, g k is the k-th gene.

3. The metabolite-disease association prediction method according to claim 1, wherein In the second step, a metabolite Gaussian interaction kernel spectral similarity matrix M is obtained based on the metabolite-disease association matrix A1 K and a disease Gaussian interaction spectral kernel similarity matrix D K The process is as follows: M k (m i ,m i' ) is the Gaussian interaction spectrum kernel similarity between the i-th metabolite and the i'-th metabolite, D k (d j ,d j' ) is the Gaussian interaction spectrum kernel similarity between the j-th disease and the j'-th disease, η m is the normalized metabolite kernel bandwidth, η d is the normalized disease kernel bandwidth, η' m is the original metabolite kernel bandwidth, η' d is the original disease kernel bandwidth, is the i-th row vector of matrix A1, is the i'-th row vector of matrix A1, is the j-th column vector of matrix A1, is the j'-th column vector of matrix A1.

4. The metabolite-disease association prediction method according to claim 1, wherein The metabolite molecular fingerprint similarity matrix M in step three S is as follows: M s (m i ,m i′ ) is the molecular fingerprint similarity between metabolite m i and metabolite m i′ ; Define the semantic value of disease d i as the sum of the contributions of all diseases to the semantics of disease d i in DAG(d i ). The semantic contribution degree of disease d i is as follows: ​ t is an entry in the MeSH database that describes a disease, i.e., a disease node in a directed acyclic graph, t′ is a sub-entry, and ω e is an edge is the semantic contribution factor connecting the entry t and the sub-entry t′, is the set of edges in the directed acyclic graph; Disease d i The semantic value SV(d i ) is as follows: The entry set for disease d i ; For a given disease d i and d j , the semantic similarity between the two is:

5. A metabolite-disease association prediction method according to claim 1, characterized in that The metabolite fusion similarity matrix M in step 4 F and the disease fusion similarity matrix D F are as follows: M F (m i ,m i′ ) is the fusion similarity of metabolite m i with metabolite m i′ . D F (d j ,d j′ ) is the fusion similarity of disease d j with disease d j′ .

6. The metabolite-disease association prediction method according to claim 1, characterized in that, The metabolite-disease-gene multi-source heterogeneous network G in Step 5 is:

7. A metabolite-disease association prediction method according to claim 1, characterized in that In Step 6, the local feature extraction module performs random walks on the multi-source heterogeneous network G described in Step 5 based on a predetermined meta-path, extracts the context information of nodes through the skip-gram model to obtain node local features, trains the skip-gram model, and obtains the metabolite local feature representation matrix M l , the disease local feature representation matrix D l The process is as follows: Randomly walk on the multi-source heterogeneous network G using a predefined meta-path to obtain a set of node sequences V = {v0, v1,..., v t}, where the probability of random walk is P. Then, in the random walk of t steps, the transition probability from node v t-1 to node v t is as follows: N(v t-1 ) is the set of neighbor nodes of node v t-1 , represents the similarity or association strength between node v t-1 and node v t ; represents the similarity or association strength between node v t-1 and node v k . The transition probability P(v t |v t-1 ) is used to generate a node sequence; Encode the node sequences using one-hot variables, and map the node sequences into a low-dimensional embedding space with a fully-connected layer to obtain metabolite, disease, and gene feature representations y of the same dimension m , y d , and y g : y m = W M x m + b m y d = W D x d + b d y g = W G x g + b g W M 、W D and W G are the weight matrices of the fully connected layers, b m 、b d and b g are the bias vectors of the fully connected layers, x m 、x d and x g are the node vector representations of metabolites, diseases, and genes after one-hot processing; Use the skip-gram model to extract the node context information to obtain the node local features: Let \(V\) be the number of nodes, and \(y\) i be the embedding representation of the given central node \(v\) i ; To maximize the central node v i , the skip-gram model is trained, and the objective function of the training model is set as: D is the set of all central nodes and context nodes in the training dataset, and W represents the weight matrix. For simplified calculation, the objective function is approximated as: σ(x) is the sigmoid function, K is the number of negative samples, and P n (v) is the negative sample distribution.

8. A metabolite-disease association prediction method according to claim 1, characterized in that, The global feature extraction module in Step 6 includes an encoder unit and a decoder unit, and both units incorporate the multi-head self-attention mechanism to capture metabolite and disease features: Z = [Z1, Z2,..., Z h is the output of the multi-head self-attention layer, h is the number of heads, Q (i) , K (i) and V (i) respectively represent the query, key, and value of the i-th head, W Q , W K and W V are learnable weight matrices, d k is the dimension of the key matrix, H = {h1, h2,..., h n} is the node feature representation, that is, the metabolite local feature representation matrix M l and the disease local feature representation matrix D l ; Both the encoder unit and the decoder unit are composed of a multi-head self-attention layer and a linear layer. The multi-head self-attention layer of the encoder unit is: E = ReLU(Z·W o + b o ) ReLU is the activation function, Z is the encoder input feature matrix, and W o is the initial weight matrix of the encoder, and b o is the bias vector of the encoder; The multi-head self-attention layer of the decoder unit is: D = ReLU(Z′·W′ o + b′ o ) ReLU is the activation function, Z' is the decoder input feature matrix, W o ' is the initial weight matrix of the decoder, b o ' is the bias vector of the decoder.

9. The metabolite-disease association prediction method according to claim 1, wherein The loss function for training the global feature extraction module of the LGDFE model in Step 6 is: h i represents the initial sample features represents the sample features after auto - encoding reconstruction, and n represents the number of samples.

10. A metabolite-disease association prediction method according to claim 1, characterized in that, The final embedded feature T obtained in Step 7 is: is the disease feature of the i-th row of M lg and is the disease feature of the j-th row of D lg ; is the concatenated combination of and i' ; T is the i'-th row of the matrix T The process of training the CNN model is: Divide the dataset x train and x test and the label set y train and y test , input the training set x train into the CNN model, first through the first convolutional layer, and then through the first pooling layer: A (1) = MaxPooling(Conv1D(X test , W (1) , b (1) )) Then the output of the first pooling layer is connected to the second convolutional layer, and the output of the second convolutional layer is connected to the second pooling layer: A (2) = MaxPooling(Conv1D(A (1) , W (2) , b (2) )) Flatten the data output by the second pooling layer, input it into the linear layer, then input it into the dropout layer for regularization to prevent the model from overfitting, then input it into the second linear layer, and finally pass through the output layer. The activation function of the output layer is set to Sigmoid: Z′ = Dropout(Z) The loss function used for training the CNN model is: where N is the number of known associations between metabolites and diseases, and y i represents the label of a positive example, represents the probability of being predicted as a positive example, and 1 - y i represents the label of a negative example, represents the probability of being predicted as a negative example.