Correlation prediction method for microorganisms and diseases based on complementary fusion of double correlation graphs

By adopting the complementary fusion method of pun-correlation graphs in the microorganism and disease association prediction, the problems of incomplete feature extraction and insufficient information fusion in the prior art are solved, and higher prediction accuracy and robustness are achieved.

CN120108756AActive Publication Date: 2025-06-06BEIJING JIAMEI KANGLIAN LIFE SCIENCE TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510246324.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-06
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The existing microbial and disease association prediction technology has the problem of incomplete feature extraction, insufficient topological attribute acquisition, and lack of interaction and fusion mechanism for correlation graph information, resulting in insufficient prediction accuracy and robustness.

Method used

Using a method based on complementary fusion of pun-related graphs, a restart random walk is built with graph enhancement feature fusion extraction module and topological enhancement, combining the complementary attention mechanism of pun-related graphs and the soft label KL divergence loss function, a node coordinated representation matrix is ​​constructed to improve information fusion effect and prediction accuracy.

Benefits of technology

It significantly improves the accuracy and robustness of microorganisms and disease association prediction, can capture complex relationships more accurately, and enhances the stability and interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108756A_ABST
    Figure CN120108756A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of biomedicine, and particularly relates to a microorganism and disease association prediction method based on double association graph complementary fusion, which comprises the following specific steps: acquiring microorganism and disease association information, constructing a first microorganism-disease heterogeneous association graph, and representing the first microorganism-disease heterogeneous association graph as a matrix B1; acquiring correlation information of microorganisms and drugs and correlation information of drugs and diseases, and constructing a second microorganism-disease heterogeneous correlation graph which is expressed as a matrix B2; according to the method, multi-dimensional feature information is effectively fused through the graph enhancement feature fusion extraction module, and the accuracy and comprehensiveness of feature representation are improved; besides, on the basis of the microorganism-disease double-correlation graph fusion matrix, the double-correlation graph complementary fusion model is combined with a double-correlation graph complementary attention mechanism and a soft label KL divergence loss function, so that the problem that information fusion of different correlation graphs is insufficient is effectively solved, and the limitation that complementarity exertion between the correlation graphs is limited is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical technology, and in particular to a method for predicting the association between microorganisms and diseases based on the complementary fusion of dual association graphs. Background Art

[0002] In recent years, the prediction of the association between microorganisms and diseases has become a research hotspot in the biomedical field. Changes in microbial communities are believed to be closely related to the occurrence of a variety of diseases, such as the relationship between intestinal microbiota and obesity, diabetes, inflammatory bowel disease and other diseases. Therefore, how to effectively reveal the association between microorganisms and diseases has become an important topic in disease research and precision medicine.

[0003] In order to improve the accuracy, interpretability and scalability of the prediction of the association between microorganisms and diseases, graph structure models have gradually become an effective tool and are widely used in the study of the relationship between microorganisms and diseases. Feature extraction based on graph structure is a key step in the prediction of the association between microorganisms and diseases. In the prior art, feature extraction mainly relies on two methods: matrix decomposition and graph convolutional network (GCN), but both have certain limitations. First, the matrix decomposition method tends to ignore the global structure and local details of the data during the dimensionality reduction process, resulting in a large amount of original information loss and the information cannot be fully expressed. On the other hand, although GCN can extract high-order features in the graph structure through the message passing mechanism, as the number of layers increases, the features of the nodes tend to be the same, resulting in insufficient expression of the differences between the nodes. In addition, the feature aggregation method of GCN often cannot fully integrate multi-dimensional features.

[0004] With the deepening of research, association graph fusion technology has gradually become a new research direction. Association graph fusion technology improves the prediction accuracy of the association between microorganisms and diseases by fusing information from different association graphs. At present, the association graph fusion method mainly uses a deep learning framework to integrate and propagate information in the modeling of graph structure data. In this way, the model can capture the complex relationship between nodes in multiple association graphs, realize the propagation of information and the learning of features. However, this type of method still has the following disadvantages:

[0005] 1. In the feature extraction stage of existing technologies for predicting the association between microorganisms and diseases, there is a problem of incomplete integration of multi-dimensional features, which makes the feature set lack integrity and cannot fully reflect the complex internal relationship between microorganisms and diseases;

[0006] 2. In the process of obtaining topological properties, the existing restart random walk method has the problems of insufficient information propagation and incomplete capture of long-distance relationships;

[0007] 3. In the existing association graph fusion methods for predicting the association between microorganisms and diseases, simple splicing or weighted summation is usually used, lacking an effective mechanism for interactive fusion of association graph information; this results in the inability to fully optimize the interactions between different association graphs, and the inability to effectively fuse information from different association graphs, which ultimately affects the overall performance of the model and the accuracy of prediction;

[0008] 4. The existing association graph fusion methods for predicting the association between microorganisms and diseases often fail to fully explore the complementary information between the association graphs while maintaining the uniqueness of each association graph; this leads to the difficulty in effectively distinguishing redundant information from unique information during the fusion process, thus causing information conflicts and redundant propagation, which ultimately affects the overall performance of the model and reduces the accuracy and robustness of the prediction.

[0009] Therefore, a method for predicting the association between microorganisms and diseases based on the complementary fusion of dual association graphs is invented. Summary of the invention

[0010] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:

[0011] A method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs includes the following specific steps:

[0012] S1: Obtain the association information between microorganisms and diseases, and construct the first microorganism-disease heterogeneous association graph, represented as matrix B 1 ; Obtain the association information between microorganisms and drugs, drugs and diseases, and construct the second microorganism-disease heterogeneous association graph, represented as matrix B 2 ;

[0013] S2: Apply the graph enhancement feature fusion extraction module GEFFE to obtain the same 1 The corresponding feature matrix Z 1 , and with B 2 The corresponding feature matrix Z 2 ;

[0014] S3: Using B 1 , B 2 , Z 1 , Z 2 , construct the dual-association graph fusion matrix B, apply the dual-association graph complementary fusion model DAG-CFM to update the node information, and obtain the node collaborative representation matrix H corresponding to B;

[0015] S4: Using H, we predict the association between microorganisms and diseases through MLP.

[0016] As a preferred solution of the method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs described in the present invention, the specific steps of S1 are as follows:

[0017] S11: Obtain the association information between microorganisms and diseases, and collect the association information between microorganisms and drugs, and drugs and diseases, and delete the duplicate entries and non-human disease entries in the collected information;

[0018] S12: Based on the collected information on the association between known microorganisms and diseases, construct a microorganism-disease adjacency matrix A 1 ; Based on the collected association information between known microorganisms and drugs, and drugs and diseases, construct the microorganism-disease adjacency matrix A 2 ;

[0019] S13: Based on A 1 Constructing the microbial comprehensive similarity matrix ZM 1 , based on A 2 Constructing the microbial comprehensive similarity matrix ZM 2 ; Due to ZM 1 The structural process and ZM 2 Similar, so ZM is used v ,v={1,2} to represent ZM 1 or ZM 2 , when v = 1, ZM v Indicates ZM 1 ; When v = 2, ZM v Indicates ZM 2 ;

[0020] S14: Based on A 1 Constructing disease comprehensive similarity matrix ZD 1 , based on A 2 Constructing disease comprehensive similarity matrix ZD 2 ; Due to ZD 1 The construction process and ZD 2 Similar, so ZD is used v ,v={1,2} to represent ZD 1 or ZD 2 , when v = 1, ZD v Indicates ZD 1 ; When v = 2, ZD v Indicates ZD 2 ;

[0021] S15: Based on matrix A 1 , ZM 1 , ZD 1 , construct the first microorganism-disease heterogeneous association graph, represented as a matrix Based on the matrix A2 , ZM 2 , ZD 2 , construct the second microorganism-disease heterogeneous association graph, represented as a matrix The specific construction process is as follows:

[0022]

[0023] As a preferred solution of the method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs described in the present invention, the specific steps of S12 are as follows:

[0024] S121: Based on the collected information on the association between known microorganisms and diseases, let nm and nd represent the number of microorganisms and diseases respectively, and obtain the microorganism-disease adjacency matrix The specific construction method is as follows: For any given microorganism m i and diseases j , if there is a known relationship between them, then A 1 (i,j)=1, otherwise A 1 (i,j)=0;

[0025] S122: Based on the collected association information between known microorganisms and drugs, and drugs and diseases, construct a microorganism-disease adjacency matrix The specific construction method is as follows: For any given microorganism m i , disease j , if there is a known m i With some drug k The relationship between k With d j There is also a known relationship between them, then A 2 (i,j)=1, otherwise A 2 (i,j)=0.

[0026] As a preferred solution of the method for predicting the association between microorganisms and diseases based on the complementary fusion of dual association graphs described in the present invention, wherein: ZM in S13 v The construction process is as follows:

[0027] S131: Based on A 1 Construction of microbial GIP core similarity matrix GM 1 , based on A 2 Construction of microbial GIP core similarity matrix GM 2 ; Due to GM 1 The construction process and GM 2 Similar, so GM is used v ,v={1,2} to represent GM 1 or GM2 , when v = 1, GM v Indicates GM 1 ; When v = 2, GM v Indicates GM 2 ; Using Gaussian kernel interaction spectra to construct the microbial GIP kernel similarity matrix GM v , the specific calculation formula is as follows:

[0028]

[0029]

[0030] Among them, λ 1 represents the normalized kernel bandwidth; |||| F represents the Frobenius norm; in the following, A is used to facilitate the description of the general calculation process. v ,v={1,2} means A 1 or A 2 , that is, when v = 1, A v Indicates A 1 ; When v = 2, A v Indicates A 2 ; IP v (m i ) and IP v (m j ) are based on A v Constructed microorganisms i and microorganisms j A binary vector, recording m i and m j Interactions with all diseases;

[0031] S132: Based on A 1 Constructing microbial cosine similarity matrix CM 1 , based on A 2 Constructing microbial cosine similarity matrix CM 2 ; Due to CM 1 The construction process and CM 2 Similar, so CM is used v ,v={1,2} to represent CM 1 and CM 2 , when v = 1, CM v Indicates CM 1 ; When v = 2, CM v Indicates CM 2 ; Using m i and m j The cosine similarity between them is used to construct the microbial cosine similarity matrix CM v , the specific calculation formula is as follows:

[0032]

[0033] Among them, A v (i,:) represents A v The i-th row of A v (j,:) represents A v The jth row of

[0034] S133: Based on GM v and CM v , perform microbial comprehensive similarity calculation and obtain the microbial comprehensive similarity matrix ZM 1 and ZM 2 ;m i and m j The comprehensive similarity calculation formula between them is as follows:

[0035]

[0036] As a preferred solution of the method for predicting the association between microorganisms and diseases based on the complementary fusion of dual association graphs described in the present invention, wherein: ZD in S14 v The construction process is as follows:

[0037] S141: Based on A 1 Construct disease GIP core similarity matrix GD 1 , based on A 2 Construct disease GIP core similarity matrix GD 2 ; Due to GD 1 The construction process and GD 2 Similar, so GD is used v ,v={1,2} to represent GD 1 or GD 2 , when v = 1, GD v Indicates GD 1 ; When v = 2, GD v Indicates GD 2 ; Use Gaussian kernel interaction spectrum GIP to construct disease GIP kernel similarity matrix GD v , the specific calculation formula is as follows:

[0038]

[0039]

[0040] Among them, λ 2 is the normalized kernel bandwidth; IP v (d i ) and IP v (d j) are based on A v Constructed Disease i and diseases j Binary vectors representing d i and d j Interactions with all microorganisms;

[0041] S142: Based on A 1 Construct disease cosine similarity matrix CD 1 , based on A 2 Construct disease cosine similarity matrix CD 2 ; Due to CD 1 The construction process and CD 2 Similar, so CD v ,v={1,2} to represent CD 1 and CD 2 , when v = 1, CD v Indicates CD 1 ; When v = 2, CD v Indicates CD 2 ; Using d i and d j The cosine similarity between the two constructs the disease cosine similarity matrix CD v , the specific calculation formula is as follows:

[0042]

[0043] Among them, A v (:,i) represents A v The i-th column of A v (:,j) represents A v The jth column of

[0044] S143: Based on GD v and CD v Calculate the comprehensive similarity of diseases and obtain the comprehensive similarity matrix ZD 1 and ZD 2 ;d i and d j The comprehensive similarity calculation formula between them is as follows:

[0045]

[0046] As a preferred solution of the method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs described in the present invention, the specific steps of S2 are as follows:

[0047] S21: Based on protein functional association, calculate the m i and m jFunctional similarity between them, constructing a microbial functional similarity matrix Based on gene interaction information, calculate for any given d i and d j Functional similarity between them, construct disease function similarity matrix

[0048] S22: For the matrix ZM 1 , ZM 2 , ZD 1 , ZD 2 Perform the topologically enhanced restarted random walk operation TERRW to obtain the first microbial topological attribute matrix in turn Second microbial topological attribute matrix First disease topological attribute matrix Second disease topological attribute matrix The definition of the TERRW operation is as follows:

[0049]

[0050] in, represents the walking probability distribution of the i-th node in the l-th step; c i is the dynamic adjustment probability calculated by node i based on its local topology information; is the extended transition probability matrix; ∈ i is the initial probability vector of node i;

[0051] S23: To retain more original features, the functional similarities, topological properties, and adjacency relationships of microorganisms and diseases are spliced ​​together to construct B 1 The corresponding feature matrix and B 2 The corresponding feature matrix Its Z 1 and Z 2 The definition is as follows:

[0052]

[0053] As a preferred solution of the method for predicting the association between microorganisms and diseases based on the complementary fusion of dual association graphs described in the present invention, wherein: in S22, due to the MM 2 ,DD 1 ,DD 2 The construction process and MM 1 Similar, its MM 1 The construction process is as follows:

[0054] S221: Set the initial probability vector ∈ of node i i , the specific construction formula is as follows:

[0055]

[0056] S222: Based on ZM 1 , construct the transition probability matrix M by normalizing each row of elements. The specific construction formula is as follows:

[0057]

[0058] Construct two-hop transfer probability matrix M based on M 2 , used to capture the two-hop neighbor information of the node. The specific construction formula is as follows:

[0059] M 2 =M×M;

[0060] S223: Through (M+M 2 ) is normalized to obtain the extended transition probability matrix

[0061] S224: For each node i, calculate the dynamic adjustment probability c according to its local topology structure i , the specific calculation process is as follows:

[0062]

[0063] where deg(i) is the degree of node i, represents the neighbor set of node i;

[0064] S225: Use the TERRW formula to iteratively update the probability distribution of the node until the probability distribution converges; the specific calculation formula is as follows:

[0065]

[0066] S226: Integrate the probability vectors of all nodes into MM 1 The specific construction process is as follows:

[0067] MM 1 =[p 1 ,p 2 ,…,p nm ]

[0068] Among them, p i is the probability distribution of node i after TERRW converges.

[0069] As a preferred solution of the method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs described in the present invention, the specific steps of S3 are as follows:

[0070] S31: Based on B 1 and B2 , construct the dual correlation graph fusion matrix B;

[0071] S32: Based on B, the dual association graph complementary attention mechanism DAG-CA is applied to calculate the attention weight between each node and the neighboring nodes in the same association graph and the attention weight between the neighboring nodes across the association graph;

[0072] S33: Combined with Z 1 , Z 2 And the neighbor node weights obtained by DAG-CA are used to fuse the dual-association graph information and construct the node collaborative representation matrix H corresponding to B;

[0073] S34: The model uses the Adam optimizer to minimize the soft label KL divergence loss L KL to train collaborative representation learning; soft label KL divergence loss L KL It is used to measure the difference between the target distribution and the soft label distribution, which is defined as follows:

[0074]

[0075] Where C is the number of clusters; t ij is the soft label assignment probability, indicating the probability that i belongs to cluster j; q ij is the target distribution, obtained by normalizing the square of the soft label assignment probability;

[0076] t ij The calculation process is as follows:

[0077]

[0078] Among them, μ j is the center of a randomly initialized cluster j, and each cluster center j is initialized to a random node in H; ||*|| represents the Euclidean distance;

[0079] q ij The calculation process is as follows:

[0080]

[0081] As a preferred solution of the method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs described in the present invention, the specific steps of S31 are as follows:

[0082] S311: Constructing a cross-association graph connection matrix B cross The value calculation formula for any element in is as follows:

[0083]

[0084] in, Indicates B cross The element corresponding to the i-th row and j-th column;

[0085] S312: Based on B 1 , B 2 and B cross , construct the dual correlation graph fusion matrix The specific construction process is as follows:

[0086]

[0087] The DAG-CA execution process in S32 is as follows:

[0088] S321: Z is transformed by linear transformation 1 Converted to query matrix Q 1 , key matrix K 1 , through linear transformation Z 2 Converted to query matrix Q 2 , key matrix K 2 ; The specific conversion process is as follows:

[0089]

[0090] in, is the learned weight matrix parameter, and the scale is

[0091] S322: Based on B, calculate B 1 The attention weight α between the middle node i and the neighbor node j in the same association graph ij ; j represents B in B 1 The node corresponding to the non-zero position of the i-th row; α ij The specific calculation formula is as follows:

[0092]

[0093] Among them, d k yes and The dimension of represents a scaling factor; Indicates Q 1 The query vector corresponding to the i-th row in the matrix, Indicates that K 1 The key vector corresponding to the jth row in the matrix, N 1 (i) indicates that node i is in B 1 The set of neighbors in ;

[0094] S323: Based on B, calculate B 2The attention weight γ between the middle node i and the neighbor node j in the same association graph ij , j represents B in B 2 The node corresponding to the non-zero position of the i-th row; γ ij The specific calculation formula is as follows:

[0095]

[0096] in, Indicates Q 2 The query vector corresponding to the i-th row in the matrix, Indicates that K 2 The key vector corresponding to the jth row in the matrix, N 2 (i) indicates that node i is in B 2 The set of neighbors in ;

[0097] S324: Based on B, calculate B 1 The attention weight β between the middle node i and the neighbor node n across the association graph in ; The cross-association graph neighbor nodes of node i are connected through B in B cross , B 2 The specific method of obtaining is as follows: If Then B 2 The node corresponding to the non-zero position in the tth column is the neighbor node of i across the association graph; in The specific calculation formula is as follows:

[0098]

[0099] Among them, Comp-Softmax is an activation function that focuses on irrelevant supplementary information across association graphs, which will enhance irrelevant information and weaken relevant information. The formula is as follows:

[0100] Comp-softmax(R)=softmax(-R)

[0101] Among them, R represents the strength of the relationship between nodes;

[0102] S325: Based on B, calculate B 2 The attention weight δ between the middle node i and the neighbor node n across the association graph in ; The cross-association graph neighbor nodes of node i are connected through B in B cross T and B 1 The specific method of obtaining is as follows: If Then B 1 The node corresponding to the non-zero position in the tth column is the neighbor node of i across the association graph; in The specific calculation formula is as follows:

[0103]

[0104] The specific steps of S33 are as follows:

[0105] S331: Based on B 1 The nodes in the double-correlation graph are fused to obtain the matrix H 1 , the information fusion process corresponding to node i is as follows:

[0106]

[0107] in, B is the information fusion of the dual correlation graph 1 The node representation vector of node i; N 1 (i) N 2 (i) indicates the node i in B 1 and B 2 The set of neighbors in |N 1 (i) | and |N 2 (i)| indicates that node i is in B 1 and B 2 The number of neighbors in W 1 , W 2 , W 3 is the learned weight matrix, and Respectively represent Z 1 i-th row, Z 1 The jth row and Z 2 The initial feature vector corresponding to the nth row, σ is the activation function;

[0108] S332: Based on B 2 The nodes in the double-correlation graph are fused to obtain the matrix H 2 , the information fusion process corresponding to node i is as follows:

[0109]

[0110] in, B is the information fusion of the dual correlation graph 2 The node i in the middle node represents a vector; W 4 , W 5 , W 6 is the learned weight matrix;

[0111] S333: Construct the node collaboration representation matrix H corresponding to B, and its calculation formula is as follows:

[0112]

[0113] As a preferred solution of the method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs described in the present invention, the specific steps of S4 are as follows:

[0114] S41: Using H, define each microorganism synergy representation vector H i With each disease co-representation vector H j The relationship between MDA ij , MDA ij The calculation is as follows:

[0115] MDA ij =-||H i -H j ||

[0116] Among them, H i -H j Represents the difference between microorganisms and diseases. The smaller the difference, the higher the possibility that the microorganism-disease pairing is associated.

[0117] S42: Utilizing MDA ij , using MLP to determine the predicted association probability of microorganisms and diseases ij ; The specific method is as follows:

[0118] ass ij =MLP(MDA ij ).

[0119] Compared with existing technologies:

[0120] The present invention effectively integrates multi-dimensional feature information through a graph-enhanced feature fusion extraction module, thereby improving the accuracy and comprehensiveness of feature representation; at the same time, combined with the topologically enhanced restarted random walk TERRW, the global perception ability of node representation is enhanced, and the complex relationship between nodes can be captured more accurately; in addition, the dual-correlation graph complementary fusion model, based on the microorganism-disease dual-correlation graph fusion matrix, combines the dual-correlation graph complementary attention mechanism and the soft label KL divergence loss function, effectively solves the problem of insufficient information fusion of different correlation graphs, breaks through the limitation of limited complementarity between correlation graphs, and thus significantly improves the information fusion effect and prediction accuracy; at the same time, the model can capture the potential nonlinearity and cross-domain interactions between microorganisms and diseases, improves the robustness, stability and interpretability of the model, and is particularly suitable for processing complex heterogeneous network data. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0122] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0123] The present invention provides a method for predicting the association between microorganisms and diseases based on the complementary fusion of dual association graphs. Figure 1 , including the following specific steps:

[0124] S1: Obtain the association information between microorganisms and diseases, and construct the first microorganism-disease heterogeneous association graph, represented as matrix B 1 ; Obtain the association information between microorganisms and drugs, drugs and diseases, and construct the second microorganism-disease heterogeneous association graph, represented as matrix B 2 ;

[0125] Among them, the specific steps of S1 are as follows:

[0126] S11: Obtain the association information between microorganisms and diseases, and collect the association information between microorganisms and drugs, and drugs and diseases, and delete the duplicate entries and non-human disease entries in the collected information;

[0127] S11 includes but is not limited to the following embodiments:

[0128] We downloaded known associations between microorganisms and diseases from the public database MDAD (http: / / www.cheng roup.cumt.edu.cn / MDAD / ). Among the 73 microorganisms and 109 human diseases collected, 502 microorganisms and diseases had literature evidence of association. We downloaded known associations between microorganisms and drugs, and drugs and diseases from the DrugBank database (http: / / www.drugbank.ca / ). We collected 1470 associations between the corresponding 73 microorganisms and 686 drugs, and 1121 associations between the 233 drugs and the corresponding 109 diseases.

[0129] S12: Based on the collected information on the association between known microorganisms and diseases, construct a microorganism-disease adjacency matrix A 1 ; Based on the collected association information between known microorganisms and drugs, and drugs and diseases, construct the microorganism-disease adjacency matrix A 2 ;

[0130] Among them, the specific steps of S12 are as follows:

[0131] S121: Based on the collected information on the association between known microorganisms and diseases, let nm and nd represent the number of microorganisms and diseases respectively, and obtain the microorganism-disease adjacency matrix The specific construction method is as follows: For any given microorganism m i and diseases j, if there is a known relationship between them, then A 1 (i,j0=1, otherwise A 1 (i,j0=0;

[0132] S122: Based on the collected association information between known microorganisms and drugs, and drugs and diseases, construct a microorganism-disease adjacency matrix The specific construction method is as follows: For any given microorganism m i , disease j , if there is a known m i With some drug k The relationship between k With d j There is also a known relationship between them, then A 2 (i,j0=1, otherwise A 2 (i,j)=0;

[0133] S13: Based on A 1 Constructing the microbial comprehensive similarity matrix ZM 1 , based on A 2 Constructing the microbial comprehensive similarity matrix ZM 2 ; Due to ZM 1 The structural process and ZM 2 Similar, so ZM is used v ,v={1,2} to represent ZM 1 or ZM 2 , when v = 1, ZM v Indicates ZM 1 ; When v = 2, ZM v Indicates ZM 2 ;

[0134] Among them, ZM in S13 v The construction process is as follows:

[0135] S131: Based on A 1 Construction of microbial GIP core similarity matrix GM 1 , based on A 2 Construction of microbial GIP core similarity matrix GM 2 ; Due to GM 1 The construction process and GM 2 Similar, so GM is used v ,v={1,2} to represent GM 1 or GM 2 , when v = 1, GM v Indicates GM 1 ; When v = 2, GM v Indicates GM 2; Use Gaussian Interaction Profile (GIP) to construct the microbial GIP nuclear similarity matrix GM v , the specific calculation formula is as follows:

[0136]

[0137] Among them, λ 1 represents the normalized kernel bandwidth; |||| F represents the Frobenius norm; in the following, A is used to facilitate the description of the general calculation process. v ,v={1,2} means A 1 or A 2 , that is, when v = 1, A v Indicates A 1 ; When v = 2, A v Indicates A 2 ; IP v (m i ) and IP v (m j ) are based on A v Constructed microorganisms i and microorganisms j A binary vector, recording m i and m j Interactions with all diseases;

[0138] Among them, Gaussian Interaction Profile (GIP) is a publicly available method for representing the similarity between species; the GIP kernel similarity matrix can quantify the similarity between different microorganisms and the similarity between different diseases through the Gaussian kernel function; the Frobenius norm is an important concept in matrix analysis, which is often used to measure the "size" or "length" of the matrix. It is the square root of the sum of the squares of the matrix elements;

[0139] S132: Based on A 1 Constructing microbial cosine similarity matrix CM 1 , based on A 2 Constructing microbial cosine similarity matrix CM 2 ; Due to CM 1 The construction process and CM 2 Similar, so CM is used v ,v={1,2} to represent CM 1 and CM 2 , when v = 1, CM v Indicates CM 1 ; When v = 2, CM vIndicates CM 2 ; Using m i and m j The cosine similarity between them is used to construct the microbial cosine similarity matrix CM v , the specific calculation formula is as follows:

[0140]

[0141] Among them, A v (i,:) represents A v The i-th row of A v (j,:) represents A v The jth row of

[0142] Among them, cosine similarity is an indicator that measures the similarity between two vectors; it measures their similarity by calculating the cosine value of the angle between the two vectors, focusing mainly on the direction of the vector rather than the size;

[0143] S133: Based on GM v and CM v , perform microbial comprehensive similarity calculation and obtain the microbial comprehensive similarity matrix ZM 1 and ZM 2 ;m i and m j The comprehensive similarity calculation formula between them is as follows:

[0144]

[0145] Among them, comprehensive similarity is a concept used to measure the similarity between two or more objects. It does not rely on a single similarity to make a judgment, but instead uses a variety of similarity calculation methods to more comprehensively and accurately evaluate the similarity relationship between objects.

[0146] S14: Based on A 1 Constructing disease comprehensive similarity matrix ZD 1 , based on A 2 Constructing disease comprehensive similarity matrix ZD 2 ; Due to ZD 1 The construction process and ZD 2 Similar, so ZD is used v ,v={1,2} to represent ZD 1 or ZD 2 , when v = 1, ZD v Indicates ZD 1 ; When v = 2, ZD v Indicates ZD 2 ;

[0147] Among them, S14 ZDv The construction process is as follows:

[0148] S141: Based on A 1 Construct disease GIP core similarity matrix GD 1 , based on A 2 Constructing the disease GIP core similarity matrix GD 2 ; Due to GD 1 The construction process and GD 2 Similar, so GD is used v ,v={1,2} to represent GD 1 or GD 2 , when v = 1, GD v Indicates GD 1 ; When v = 2, GD v Indicates GD 2 ; Use Gaussian kernel interaction spectrum GIP to construct disease GIP kernel similarity matrix GD v , the specific calculation formula is as follows:

[0149]

[0150] Among them, λ 2 is the normalized kernel bandwidth; IP v (d i ) and IP v (d j ) are based on A v Constructed Disease i and diseases j Binary vectors representing d i and d j Interactions with all microorganisms;

[0151] S142: Based on A 1 Construct disease cosine similarity matrix CD 1 , based on A 2 Construct disease cosine similarity matrix CD 2 ; Due to CD 1 The construction process and CD 2 Similar, so CD v ,v={1,2} to represent CD 1 and CD 2 , when v = 1, CD v Indicates CD 1 ; When v = 2, CD v Indicates CD 2 ; Using d i and d j The cosine similarity between the two constructs the disease cosine similarity matrix CD v, the specific calculation formula is as follows:

[0152]

[0153] Among them, A v (:,i) represents A v The i-th column of A v (:,j) represents A v The jth column of

[0154] S143: Based on GD v and CD v Calculate the comprehensive similarity of diseases and obtain the comprehensive similarity matrix ZD 1 and ZD 2 ;d i and d j The comprehensive similarity calculation formula between them is as follows:

[0155]

[0156] S15: Based on matrix A 1 , ZM 1 , ZD 1 , construct the first microorganism-disease heterogeneous association graph, represented as a matrix Based on the matrix A 2 , ZM 2 , ZD 2 , construct the second microorganism-disease heterogeneous association graph, represented as a matrix The specific construction process is as follows:

[0157]

[0158]

[0159] S2: Apply the graph-enhanced fusion feature extraction module GEFFE (Graph-Enhanced Fusion Feature Extraction) to obtain 1 The corresponding feature matrix Z 1 , and with B 2 The corresponding feature matrix Z 2 ;

[0160] Among them, the specific steps of S2 are as follows:

[0161] S21: Based on protein functional association, calculate the m i and m j Functional similarity between them, constructing a microbial functional similarity matrix Based on gene interaction information, calculate for any given d i and dj Functional similarity between them, construct disease function similarity matrix

[0162] Among them, the method for calculating the functional similarity between microorganisms based on protein functional association is an existing public method and is not the content of the present invention; this method calculates the functional similarity between microorganisms (i.e., the microbial functional association index, MFI) by constructing a protein-protein functional association network between microorganisms; firstly, the functional association information of proteins in the microbial genome is extracted, and the nodes (gene families) in the network are marked according to the presence or absence of gene families; then, the connections (edges) between different types of gene families are counted, and finally the functional similarity between two microorganisms is calculated by a formula; this method can be used to generate a functional similarity matrix between microorganisms to reflect the degree of functional association of different microorganisms;

[0163] The method for calculating the functional similarity between diseases based on gene interaction information is an existing public method and is not the content of the present invention; first, a related gene set is derived for each disease; then the similarity between diseases is measured according to the functional similarity between genes (assessed by log-likelihood score); specifically, the similarity between disease pairs is calculated by calculating the maximum functional similarity between related gene sets and comprehensively considering the functional associations of all genes, and finally a functional similarity matrix of diseases is obtained;

[0164] S22: For the matrix ZM 1 , ZM 2 , ZD 1 , ZD 2 Perform the topology-enhanced restarted random walk operation TERRW (Topology-Enhanced Restarted Random Walk) to obtain the first microbial topology attribute matrix in turn Second microbial topological attribute matrix First disease topological attribute matrix Second disease topological attribute matrix The definition of the TERRW operation is as follows:

[0165]

[0166] in, represents the walking probability distribution of the i-th node in the l-th step; c i is the dynamic adjustment probability calculated by node i based on its local topology information; is the extended transition probability matrix; ∈ i is the initial probability vector of node i;

[0167] Among them, due to MM in S22 2,DD 1 ,DD 2 The construction process and MM 1 Similar, its MM 1 The construction process is as follows:

[0168] S221: Set the initial probability vector ∈ of node i i , the specific construction formula is as follows:

[0169]

[0170] S222: Based on ZM 1 , construct the transition probability matrix M by normalizing each row of elements. The specific construction formula is as follows:

[0171]

[0172] Construct two-hop transfer probability matrix M based on M 2 , used to capture the two-hop neighbor information of the node. The specific construction formula is as follows:

[0173] M 2 =M×M;

[0174] S223: Through (M+M 2 ) is normalized to obtain the extended transition probability matrix

[0175] S224: For each node i, calculate the dynamic adjustment probability c according to its local topology structure i , the specific calculation process is as follows:

[0176]

[0177] where deg(i) is the degree of node i, represents the neighbor set of node i;

[0178] S225: Use the TERRW formula to iteratively update the probability distribution of the node until the probability distribution converges; the specific calculation formula is as follows:

[0179]

[0180] S226: Integrate the probability vectors of all nodes into MM 1 The specific construction process is as follows:

[0181] MM 1 =[p 1 ,p 2 ,…,p nm ]

[0182] Among them, pi is the probability distribution of node i after TERRW convergence;

[0183] S23: To retain more original features, the functional similarities, topological properties, and adjacency relationships of microorganisms and diseases are spliced ​​together to construct B 1 The corresponding feature matrix and B 2 The corresponding feature matrix Its Z 1 and Z 2 The definition is as follows:

[0184]

[0185] S3: Using B 1 , B 2 , Z 1 , Z 2 , construct the dual-associated graph fusion matrix B, apply the dual-associated graph complementary fusion model DAG-CFM (Dual Associated-Graph Complementary Fusion Model) to update the node information, and obtain the node collaborative representation matrix H corresponding to B;

[0186] The specific steps of S3 are as follows:

[0187] S31: Based on B 1 and B 2 , construct the dual correlation graph fusion matrix B;

[0188] Among them, the specific steps of S31 are as follows:

[0189] S311: Constructing a cross-association graph connection matrix B cross The value calculation formula for any element in is as follows:

[0190]

[0191] in, Indicates B cross The element corresponding to the i-th row and j-th column;

[0192] S312: Based on B 1 , B 2 and B cross , construct the dual correlation graph fusion matrix The specific construction process is as follows:

[0193]

[0194] S32: Based on B, the dual associated-graph complementary attention mechanism DAG-CA (Dual Associated-Graph Complementary Attention) is applied to calculate the attention weight between each node and its neighboring nodes in the same associated graph and the attention weight between neighboring nodes across the associated graph;

[0195] The DAG-CA execution process in S32 is as follows:

[0196] S321: Z is transformed by linear transformation 1 Converted to query matrix q 1 , key matrix K 1 , through linear transformation Z 2 Converted to query matrix Q 2 , key matrix K 2 ; The specific conversion process is as follows:

[0197]

[0198] in, is the learned weight matrix parameter, and the scale is

[0199] S322: Based on B, calculate B 1 The attention weight α between the middle node i and the neighbor node j in the same association graph ij ; j represents B in B 1 The node corresponding to the non-zero position of the i-th row; α ij The specific calculation formula is as follows:

[0200]

[0201] Among them, d k yes and The dimension of represents a scaling factor; Indicates Q 1 The query vector corresponding to the i-th row in the matrix, Indicates that K 1 The key vector corresponding to the jth row in the matrix, N 1 (i) indicates that node i is in B 1 The set of neighbors in ;

[0202] S323: Based on B, calculate B 2 The attention weight γ between the middle node i and the neighbor node j in the same association graph ij , j represents B in B 2 The node corresponding to the non-zero position of the i-th row; γ ij The specific calculation formula is as follows:

[0203]

[0204] in, Indicates Q 2 The query vector corresponding to the i-th row in the matrix, Indicates that K 2 The key vector corresponding to the jth row in the matrix, N 2 (i) indicates that node i is in B 2 The set of neighbors in ;

[0205] S324: Based on B, calculate B 1 The attention weight β between the middle node i and the neighbor node n across the association graph in ; The cross-association graph neighbor nodes of node i are connected through B in B cross , B 2 The specific method of obtaining is as follows: If Then B 2 The node corresponding to the non-zero position in the tth column is the neighbor node of i across the association graph; in The specific calculation formula is as follows:

[0206]

[0207] Among them, Comp-Softmax is an activation function that focuses on irrelevant supplementary information across association graphs, which will enhance irrelevant information and weaken relevant information. The formula is as follows:

[0208] Comp-softmax(R)=softmax(-R)

[0209] Among them, R represents the strength of the relationship between nodes;

[0210] S325: Based on B, calculate B 2 The attention weight δ between the middle node i and the neighbor node n across the association graph in ; The cross-association graph neighbor nodes of node i are connected through B in B cross T and B 1 The specific method of obtaining is as follows: If Then B 1 The node corresponding to the non-zero position in the tth column is the neighbor node of i across the association graph; in The specific calculation formula is as follows:

[0211]

[0212] S33: Combined with Z 1 , Z 2And the neighbor node weights obtained by DAG-CA are used to fuse the dual-association graph information and construct the node collaborative representation matrix H corresponding to B;

[0213] Among them, the specific steps of S33 are as follows:

[0214] S331: Based on B 1 The nodes in the double-correlation graph are fused to obtain the matrix H 1 , the information fusion process corresponding to node i is as follows:

[0215]

[0216] in, B is the information fusion of the dual correlation graph 1 The node representation vector of node i; N 1 (i) N 2 (i) indicates the node i in B 1 and B 2 The set of neighbors in |N 1 (i) | and |N 2 (i)| indicates that node i is in B 1 and B 2 The number of neighbors in W 1 , W 2 , W 3 is the learned weight matrix, and Respectively represent Z 1 i-th row, Z 1 The jth row and Z 2 The initial feature vector corresponding to the nth row, σ is the activation function;

[0217] S332: Based on B 2 The nodes in the double-correlation graph are fused to obtain the matrix H 2 , the information fusion process corresponding to node i is as follows:

[0218]

[0219] in, B is the information fusion of the dual correlation graph 2 The node i in the middle node represents a vector; W 4 , W 5 , W 6 is the learned weight matrix;

[0220] S333: Construct the node collaboration representation matrix H corresponding to B, and its calculation formula is as follows:

[0221]

[0222] S34: The model uses the Adam optimizer to minimize the soft label KL divergence loss L KL to train collaborative representation learning; soft label KL divergence loss L KL It is used to measure the difference between the target distribution and the soft label distribution, which is defined as follows:

[0223]

[0224] Where C is the number of clusters; t ij is the soft label assignment probability, indicating the probability that i belongs to cluster j; q ij is the target distribution, obtained by normalizing the square of the soft label assignment probability;

[0225] t ij The calculation process is as follows:

[0226]

[0227] Among them, μ j is the center of a randomly initialized cluster j, and each cluster center j is initialized to a random node in H; ||*|| represents the Euclidean distance;

[0228] q ij The calculation process is as follows:

[0229]

[0230] Among them, Adam is a publicly available method, an optimization algorithm that combines the advantages of momentum optimization and RMSProp optimization, and can dynamically adjust the learning rate of each parameter;

[0231] Euclidean distance is a commonly used distance measurement method used to measure the straight-line distance between two points in Euclidean space. It is a measure of the shortest path distance between two points defined in geometry.

[0232] S4: Using H, we predict the association between microorganisms and diseases through MLP;

[0233] Among them, the specific steps of S4 are as follows:

[0234] S41: Using H, define each microorganism synergy representation vector H i With each disease co-representation vector H j The relationship between MDA ij , MDA ij The calculation is as follows:

[0235] MDA ij =-||H i -H j ||

[0236] Among them, H i -H j Represents the difference between microorganisms and diseases. The smaller the difference, the higher the possibility that the microorganism-disease pairing is associated.

[0237] S42: Utilizing MDA ij , using MLP to determine the predicted association probability of microorganisms and diseases ij ; The specific method is as follows:

[0238] ass ij =MLP(MDA ij );

[0239] Among them, MLP is a multi-layer perceptron, which is a simple publicly available neural network model consisting of two layers of fully connected neural networks. It is used to process data and learn the complex relationship between input and output.

[0240] Based on the above, the present invention significantly enhances the diversity and depth of feature representation by integrating multiple information such as functional similarity, topological properties and adjacency relationship between microorganisms and diseases; the module first calculates the functional similarity between microorganisms and diseases, and extracts the topological properties of nodes in combination with the topologically enhanced restarted random walk TERRW method, so as to better capture the complex global and local relationships in the graph; finally, the module splices these multidimensional features, which not only retains the original feature information but also enhances its expression ability, so that the model can more comprehensively reflect the complex relationship between microorganisms and diseases, and improves the performance and accuracy of subsequent prediction tasks; in addition, the present invention aims to integrate the information from The information of two association graphs is used to enhance the representation ability of nodes; specifically, based on the dual association graph fusion matrix, the model designs a dual association graph complementary attention mechanism DAG-CA, which assigns different weights to neighbor nodes in the same association graph and across association graphs, ensuring that the complementarity between association graphs can be fully utilized in the process of information fusion; in addition, during the training process, the soft label KL divergence loss optimization strategy is adopted to ensure the accuracy of soft label assignment, thereby enhancing the collaborative representation ability of nodes; the model can fully integrate the dual association graph information, improve the stability, efficiency and interpretability of the model, and is particularly suitable for processing complex heterogeneous network data such as microorganisms and diseases, showing significant performance advantages.

[0241] Although the present invention has been described above with reference to the embodiments, various modifications may be made thereto and parts thereof may be replaced by equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the embodiments disclosed in the present invention may be used in combination with each other in any manner, and the fact that these combinations are not exhaustively described in this specification is only for the sake of omitting space and saving resources. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs, characterized in that: The specific steps are as follows: S1: Obtain the association information between microorganisms and diseases, construct the first microorganism-disease heterogeneous association graph, represented as matrix B1; obtain the association information between microorganisms and drugs, drugs and diseases, and construct the second microorganism-disease heterogeneous association graph, represented as matrix B2; S2: Apply the graph enhancement feature fusion extraction module GEFFE to obtain the feature matrix Z1 corresponding to B1 and the feature matrix Z2 corresponding to B2; S3: Use B1, B2, Z1, and Z2 to construct a dual-association graph fusion matrix B, apply the dual-association graph complementary fusion model DAG-CFM to update the node information, and obtain the node collaborative representation matrix H corresponding to B; S4: Using H, we predict the association between microorganisms and diseases through MLP.

2. The method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs according to claim 1, characterized in that: The specific steps of S1 are as follows: S11: Obtain the association information between microorganisms and diseases, and collect the association information between microorganisms and drugs, and drugs and diseases, and delete the duplicate entries and non-human disease entries in the collected information; S12: Based on the collected association information between known microorganisms and diseases, a microorganism-disease adjacency matrix A1 is constructed; based on the collected association information between known microorganisms and drugs, and drugs and diseases, a microorganism-disease adjacency matrix A2 is constructed; S13: Based on A1, a microbial comprehensive similarity matrix ZM1 was constructed, and based on A2, a microbial comprehensive similarity matrix ZM2 was constructed. Since the construction process of ZM1 is similar to that of ZM2, ZM v ,v={1,2} to represent ZM1 or ZM2. When v=1, ZM v Indicates ZM1; when v = 2, ZM v Indicates ZM2; S14: Construct the disease comprehensive similarity matrix ZD1 based on A1, and construct the disease comprehensive similarity matrix ZD2 based on A2; since the construction process of ZD1 is similar to that of ZD2, ZD2 is used. v ,v={1,2} to represent ZD1 or ZD2. When v=1, ZD v Indicates ZD1; when v = 2, ZD v Indicates ZD2; S15: Based on matrices A1, ZM1, and ZD1, a first microorganism-disease heterogeneous association graph is constructed, represented as a matrix Based on matrices A2, ZM2, and ZD2, a second microorganism-disease heterogeneous association graph is constructed, represented as a matrix The specific construction process is as follows:

3. The method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs according to claim 2, characterized in that: The specific steps of S12 are as follows: S121: Based on the collected information on the association between known microorganisms and diseases, let nm and nd represent the number of microorganisms and diseases respectively, and obtain the microorganism-disease adjacency matrix The specific construction method is as follows: For any given microorganism m i and diseases j , if there is a known association between them, then A1(i,j)=1, otherwise A1(i,j)=0; S122: Based on the collected association information between known microorganisms and drugs, and drugs and diseases, construct a microorganism-disease adjacency matrix The specific construction method is as follows: For any given microorganism m i , disease j , if there is a known m i With some drug k The relationship between k With d j If there is a known association between them, then A2(i,j)=1, otherwise A2(i,jv=0.

4. The method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs according to claim 2, characterized in that: The S13 ZM v The construction process is as follows: S131: Construct the microbial GIP core similarity matrix GM1 based on A1, and construct the microbial GIP core similarity matrix GM2 based on A2; since the construction process of GM1 is similar to that of GM2, GM v ,v={1,2} to represent GM1 or GM2. When v=1, GM v Indicates GM1; when v = 2, GM v Represents GM2; constructs the microbial GIP nuclear similarity matrix GM using Gaussian nuclear interaction spectrum v , the specific calculation formula is as follows: where λ1 represents the normalized kernel bandwidth; || || F represents the Frobenius norm; in the following, A is used to facilitate the description of the general calculation process. v ,v={1,2} means A1 or A2, that is, when v=1, A v represents A1; when v=2, A v Indicates A2; IP v (m i ) and IP v (m j ) are based on A v Constructed microorganisms i and microorganisms j A binary vector, recording m i and m j Interactions with all diseases; S132: Construct the microbial cosine similarity matrix CM1 based on A1, and construct the microbial cosine similarity matrix CM2 based on A2; since the construction process of CM1 is similar to that of CM2, CM v ,v={1,2} to represent CM1 and CM2. When v=1, CM v Indicates CM1; when v = 2, CM v Indicates CM2; using m i and m j The cosine similarity between them is used to construct the microbial cosine similarity matrix CM v , the specific calculation formula is as follows: Among them, A v (i,:) represents A v The i-th row of A v (j,:) represents A v The jth row of S133: Based on GM v and CM v , perform microbial comprehensive similarity calculation to obtain microbial comprehensive similarity matrices ZM1 and ZM2; m i and m j The comprehensive similarity calculation formula between them is as follows:

5. The method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs according to claim 2, characterized in that: The S14 ZD v The construction process is as follows: S141: Construct the disease GIP core similarity matrix GD1 based on A1, and construct the disease GIP core similarity matrix GD2 based on A2; since the construction process of GD1 is similar to that of GD2, GD v ,v={1,2} to represent GD1 or GD2. When v=1, GD v Indicates GD1; when v = 2, GD v Represents GD2; constructs the disease GIP kernel similarity matrix GD using the Gaussian kernel interaction spectrum GIP v , the specific calculation formula is as follows: Where λ2 is the normalized kernel bandwidth; IP v (d i ) and IP v (d j ) are based on A v Constructed Disease i and diseases j Binary vectors representing d i and d j Interactions with all microorganisms; S142: Construct the disease cosine similarity matrix CD1 based on A1, and construct the disease cosine similarity matrix CD2 based on A2; since the construction process of CD1 is similar to that of CD2, CD v ,v={1,2} to represent CD1 and CD2. When v=1, CD v Indicates CD1; when v = 2, CD v Indicates CD2; using d i and d j The cosine similarity between the two constructs the disease cosine similarity matrix CD v , the specific calculation formula is as follows: Among them, A v (:,i) represents A v The i-th column of A v (:,j) represents A v The jth column of S143: Based on GD v and CD v Calculate the comprehensive disease similarity and obtain the comprehensive disease similarity matrices ZD1 and ZD2; d i and d j The comprehensive similarity calculation formula between them is as follows:

6. The method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs according to claim 1, characterized in that: The specific steps of S2 are as follows: S21: Based on protein functional association, calculate the m i and m j Functional similarity between them, constructing a microbial functional similarity matrix Based on gene interaction information, calculate for any given d i and d j Functional similarity between them, construct disease function similarity matrix S22: Perform topologically enhanced restarted random walk operations TERRW on matrices ZM1, ZM2, ZD1, and ZD2 respectively to obtain the first microbial topological attribute matrix in turn. Second microbial topological attribute matrix First disease topological attribute matrix Second disease topological attribute matrix The definition of the TERRW operation is as follows: in, represents the walking probability distribution of the i-th node in the l-th step; c i is the dynamic adjustment probability calculated by node i based on its local topology information; is the extended transition probability matrix; ∈ i is the initial probability vector of node i; S23: To retain more original features, the functional similarities, topological properties, and adjacency relationships of microorganisms and diseases are spliced ​​together to construct the feature matrix corresponding to B1 The characteristic matrix corresponding to B2 Its Z1 and Z2 are defined as follows:

7. The method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs according to claim 6, characterized in that: In S22, since the construction process of MM2, DD1, and DD2 is similar to that of MM1, the construction process of MM1 is as follows: S221: Set the initial probability vector ∈ of node i i , the specific construction formula is as follows: S222: Based on ZM1, construct a transition probability matrix M by normalizing the elements in each row. The specific construction formula is as follows: Construct two-hop transfer probability matrix M based on M 2 , used to capture the two-hop neighbor information of the node. The specific construction formula is as follows: M 2 =M×M; S223: Through (M+M 2 ) is normalized to obtain the extended transition probability matrix S224: For each node i, calculate the dynamic adjustment probability c according to its local topology structure i , the specific calculation process is as follows: where deg(i) is the degree of node i, represents the neighbor set of node i; S225: Use the TERRW formula to iteratively update the probability distribution of the node until the probability distribution converges; the specific calculation formula is as follows: S226: Integrate the probability vectors of all nodes into MM1. The specific construction process is as follows: MM1=[p1,[2,…,p nm ] Among them, p i is the probability distribution of node i after TERRW converges.

8. The method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs according to claim 1, characterized in that: The specific steps of S3 are as follows: S31: Based on B1 and B2, construct a dual correlation graph fusion matrix B; S32: Based on B, the dual association graph complementary attention mechanism DAG-CA is applied to calculate the attention weight between each node and the neighboring nodes in the same association graph and the attention weight between the neighboring nodes across the association graph; S33: Combine Z1, Z2 and the neighbor node weights obtained by DAG-CA to perform dual association graph information fusion and construct the node collaborative representation matrix H corresponding to B; S34: The model uses the Adam optimizer to minimize the soft label KL divergence loss L KL to train collaborative representation learning; soft label KL divergence loss L KL It is used to measure the difference between the target distribution and the soft label distribution, which is defined as follows: Where C is the number of clusters; t ij is the soft label assignment probability, indicating the probability that i belongs to cluster j; q ij is the target distribution, obtained by normalizing the square of the soft label assignment probability; t ij The calculation process is as follows: Among them, μ j is the center of a randomly initialized cluster j, and each cluster center j is initialized to a random node in H; ||*|| represents the Euclidean distance; q ij The calculation process is as follows:

9. The method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs according to claim 8, characterized in that: The specific steps of S31 are as follows: S311: Constructing a cross-association graph connection matrix B cross The value calculation formula for any element in is as follows: in, Indicates B cross The element corresponding to the i-th row and j-th column; S312: Based on B1, B2 and B cross , construct the dual correlation graph fusion matrix The specific construction process is as follows: The DAG-CA execution process in S32 is as follows: S321: Convert Z1 into a query matrix Q1 and a key matrix K1 respectively through linear transformation, and convert Z2 into a query matrix Q2 and a key matrix K2 respectively through linear transformation; the specific conversion process is as follows: in, is the learned weight matrix parameter, and the scale is S322: Based on B, calculate the attention weight α between node i in B1 and its neighbor node j in the same association graph ij ; j represents the node corresponding to the non-zero position of the i-th row of B1 in B; α ij The specific calculation formula is as follows: Among them, d k yes and The dimension of represents a scaling factor; represents the query vector corresponding to the i-th row in the Q1 matrix, represents the key vector corresponding to the jth row in the K1 matrix, and N1(i) represents the neighbor set of node i in B1; S323: Based on B, calculate the attention weight γ between node i in B2 and its neighbor node j in the same association graph ij , j represents the node corresponding to the non-zero position of the i-th row of B2 in B; γ ij The specific calculation formula is as follows: in, represents the query vector corresponding to the i-th row in the Q2 matrix, represents the key vector corresponding to the jth row in the K2 matrix, and N2(i) represents the neighbor set of node i in B2; S324: Based on B, calculate the attention weight β between node i in B1 and the neighbor node n under the cross-association graph in ; The cross-association graph neighbor nodes of node i are connected through B in B cross , B2 is obtained, and the specific method of obtaining is as follows: Then the node corresponding to the non-zero position in the tth column of B2 is the neighbor node of i across the association graph; β in The specific calculation formula is as follows: Among them, Comp-Softmax is an activation function that focuses on irrelevant supplementary information across association graphs, which will enhance irrelevant information and weaken relevant information. The formula is as follows: Comp-softmax(R)=softmax(-R) Among them, R represents the strength of the relationship between nodes; S325: Based on B, calculate the attention weight δ between node i in B2 and the neighbor node n under the cross-association graph in ; The cross-association graph neighbor nodes of node i are connected through B in B cross T and B1, the specific way of obtaining is as follows: if Then the node corresponding to the non-zero position in the tth column of B1 is the neighbor node of i across the association graph; δ in The specific calculation formula is as follows: The specific steps of S33 are as follows: S331: Based on the nodes in B1, the information fusion of the dual association graph is performed to obtain the matrix H1. The information fusion process corresponding to the node i is as follows: in, is the node representation vector of node i in B1 after the dual association graph information is fused; N1(i) and N2(i) represent the neighbor sets of node i in B1 and B2 respectively; |N1(i)| and |N2(i)| represent the number of neighbors of node i in B1 and B2; W1, W2, and W3 are the weight matrices obtained by learning, and They represent the initial feature vectors corresponding to the i-th row of Z1, the j-th row of Z1, and the n-th row of Z2, respectively, and σ is the activation function; S332: Based on the nodes in B2, the information fusion of the dual association graph is performed to obtain the matrix H2. The information fusion process corresponding to the node i is as follows: in, is the node representation vector of node i in B2 after the information of the dual association graph is fused; W4, W5, and W6 are the weight matrices obtained through learning; S333: Construct the node collaboration representation matrix H corresponding to B, and its calculation formula is as follows:

10. The method for predicting the association between microorganisms and diseases based on complementary fusion of dual association graphs according to claim 1, characterized in that: The specific steps of S4 are as follows: S41: Using H, define each microorganism synergy representation vector H i With each disease co-representation vector H j The relationship between MDA ij , MDA ij The calculation is as follows: MDA ij =-||H i -H j || Among them, H i -H j Represents the difference between microorganisms and diseases. The smaller the difference, the higher the possibility that the microorganism-disease pairing is associated. S42: Utilizing MDA ij , using MLP to determine the predicted association probability of microorganisms and diseases ij ; The specific method is as follows: ass ij =MLP(MDA ij )。

Citation Information

Patent Citations

  • Microorganism-disease relevance prediction method based on multi-order similarity fusion learning

    CN117219173A

Cited By

  • Microorganism-disease association prediction method and device based on multi-level perceptual aggregation and deep hierarchical graph mixed learner

    CN121075433A