Method and system for predicting circRNA-disease association based on graph isomorphic Transformer
By constructing a CDA knowledge graph based on graph isomorphic Transformer and combining a dual-flow neural predictor, the problem of processing data sparseness and nonlinear association in the prediction of circRNA-disease association in the prior art is solved, achieving a more accurate and robust prediction effect.
Patent Information
- Application Number
- CN202411793935.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art fails to adequately process data sparsity and capture complex nonlinear associations between biological data when predicting circRNA and disease associations.
The CDA knowledge graph is constructed by a graph isomorphic Transformer method. Through the graph isomorphic Transformer model and a dual-stream neural predictor, local and global correlation information are fully utilized to capture nonlinear correlations, thereby achieving effective processing and prediction of CDA sparse data.
It significantly improves the accuracy and efficiency of circRNA-disease association prediction, ensures that the prediction is more comprehensive and robust, and can effectively handle complex biological interactions in multi-source heterogeneous datasets.
Smart Images

Figure CN119943335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of circRNA-disease association, and in particular to a method and system for predicting circRNA-disease association based on graph isomorphism Transformer. Background Art
[0002] Circular RNA (circRNA) is an emerging non-coding RNA with a unique circular stable structure. In recent years, circRNA has attracted extensive attention due to its important significance in the study of human disease mechanisms such as cancer, neurological diseases and cardiovascular diseases. The ubiquity and different roles of circRNA make it a key regulatory factor in the gene expression regulatory network. Therefore, the discovery and study of circRNA provides a new perspective for people to understand the RNA world.
[0003] Comprehensive research on the association between circRNA and diseases will not only bring new insights and strategies for clarifying the pathogenesis of diseases, but will also have a profound impact on the treatment and prevention of human diseases. However, traditional biological methods are inefficient and require sophisticated experimental equipment, which hinders in-depth research on CDA. To meet this challenge, researchers have proposed computational methods for predicting CDA, which are generally divided into information propagation-based methods, traditional machine learning methods, and deep learning methods.
[0004] Methods based on information propagation usually construct heterogeneous networks by using disease semantic similarity, functional similarity, and circRNA-target gene interactions between CDAs. Although such methods have low computational resource requirements, they are highly dependent on known biological networks (such as gene regulatory networks), and prediction results may be seriously affected when network information is incomplete or inaccurate.
[0005] Traditional machine learning-based methods have developed rapidly in the field of CDA prediction. Such methods can usually provide relatively accurate prediction results, but they require manual design and selection of features, and may require retraining of the model when dealing with new circRNA-disease relationships in complex biological networks.
[0006] Deep learning-based methods can automatically obtain high-order features and explore potential correlation information, and have achieved great success in predicting CDA tasks. Although such methods have made good progress in feature extraction, they have limitations in distinguishing differences between features and handling complex feature interactions and nonlinear relationships.
[0007] In summary, although current methods have achieved remarkable success in CDA prediction, they still have two limitations. First, existing methods cannot fully handle the sparsity of CDA data, resulting in suboptimal CDA prediction models. Second, existing models are not sufficient to explore various high-order information and cannot capture the complex nonlinear associations between various biological data. Summary of the invention
[0008] In order to at least partially solve the problem that existing methods cannot fully handle the sparsity of CDA data and existing models cannot capture the complex nonlinear associations between various biological data, the present invention provides a method and system for predicting circRNA-disease associations based on graph isomorphism transformers, and constructs a comprehensive knowledge graph containing multiple similarity associations between diseases, circRNAs, lncRNAs and miRNAs, so as to fully understand CDA information. Subsequently, a graph isomorphism transformer model is proposed to fully utilize the local and global association information of CDA, promote the diversified expression of CDA knowledge, and achieve full processing of CDA sparse data. Finally, a two-stream neural predictor is introduced to more comprehensively predict CDA affinity scores, which can capture the complex nonlinear associations between various biological data, thereby achieving robust and comprehensive CDA prediction.
[0009] In order to achieve the above object, the technical solution of the present invention is:
[0010] The first aspect of the present invention proposes a method for predicting circRNA-disease association based on graph isomorphism Transformer, comprising:
[0011] Step 1: Use multi-source heterogeneous data sets to build a CDA knowledge graph, which is convenient for alleviating the data sparsity problem in CDA prediction tasks and is conducive to extracting local deep information from the knowledge graph;
[0012] Step 2: Establish a fusion similarity network based on the CDA knowledge graph to facilitate feature extraction;
[0013] Step 3: Input the output of the fused similarity network into the graph isomorphism Transformer model to obtain the embedding features, so as to obtain high-quality embedding representation;
[0014] Step 4: Input the embedded features into the two-stream neural predictor to complete the identification of potential circRNA-disease associations.
[0015] Furthermore, the CDA knowledge graph is expressed by the following formula:
[0016] G=(E,R)
[0017] Among them, G is the CDA knowledge graph representation, E is the entity set, E = {miRNA, circRNA, disease, lncRNA}; R is the relationship set, R = {r1, r2, r3, r4, r5}; among them, r1 is circ-miRNA0, r2 is miRNA-disease 1, r3 is miRNA-lncRNA2, r4 is lncRNA-disease 3, and r5 is circ-disease4;
[0018] The internal form of the CDA knowledge graph is represented by the following triples:
[0019] (h, r, t)
[0020] Among them, h,t∈E, E represents the entity set (miRNA, circRNA, disease and lncRNA), h is the head entity node, t is the tail entity node, r is the relationship between entity nodes, r∈R.
[0021] Furthermore, the graph isomorphism Transformer model includes a Transformer encoder and a graph isomorphism layer connected in sequence;
[0022] The Transformer encoder includes a multi-head attention mechanism, a feedforward neural network and a relation matrix multiplication connected in sequence; the multi-head attention mechanism is used to embed the head entity, the relation and the tail entity of the fusion similarity network to obtain the embedded rich embedding representation of the head entity, the relation and the tail entity; the feedforward neural network is used to enhance the rich embedding representation to obtain the enhanced embedded representation of the head entity, the relation and the tail entity; the relation matrix multiplication is used to further enhance the enhanced embedded representation of the head entity, the relation and the tail entity to obtain the final embedded representation of the head entity, the relation and the tail entity;
[0023] The graph isomorphism layer includes an information propagation unit and an information aggregation unit connected in sequence; the information propagation unit is used to propagate the final embedded representations of the head entity, relationship and tail entity to obtain the relationship between the entities; the information aggregation unit is used to aggregate the relationship between the entities to obtain the relationship between the aggregated entities.
[0024] Furthermore, the multi-head attention mechanism is used to process the embedding of the head entity, relationship and tail entity in the fusion similarity network, thereby obtaining a rich embedding representation. The specific process is expressed by the following formula:
[0025] Multihead(Q,K,V)=[A1,A2,...,A i ]W o
[0026]
[0027] Among them, Multihead (Q, K, V) is a rich embedding representation, A i is the attention score of the i-th head, W o is the projection weight matrix, softmax is the normalized exponential function, Q i is the query of the i-th head, V i is the value of the i-th head, K i The transpose K i , K i is the key of the i-th head, d is the embedding dimension of the head entity, relation, and tail entity embedding in the fused similarity network, and [] is concatenation.
[0028] Furthermore, the feedforward neural network is used to enhance the nonlinear ability of the output of the multi-head attention mechanism to obtain the enhanced embedding representation of the head entity, relationship and tail entity. The specific process is expressed by the following formula:
[0029] FFN(Multihead(Q,K,V))=max(Multihead(Q,K,V)W1+b1,0)W2+b2
[0030] Among them, FFN(x) is the embedded representation of the enhanced head entity, relation and tail entity, W1 is the hidden layer weight, b1 is the hidden layer bias, W2 is the output layer weight, and b2 is the output layer bias.
[0031] Furthermore, the information propagation unit is used to perform propagation processing on the relationship matrix multiplication output to obtain the relationship between entities. The specific process is expressed by the following formula:
[0032]
[0033] in, is the egocentric network of node h, θ(h,r,t) is the normalized information propagation factor, e t is the embedding of the tail entity, W r is the transformation matrix of the relationship, e h is the embedding of the head entity, e r is the embedding of the relation, tanh is the hyperbolic tangent function, is the information propagation factor before normalization.
[0034] Furthermore, the information aggregation unit is used to aggregate the relationship between entities obtained by the information dissemination unit to obtain the relationship between the aggregated entities. The specific process is expressed by the following formula:
[0035]
[0036] in, for and Aggregation, is the head entity of layer l-1, is the egocentric network of node h in layer l-1, LeakyReLU is the activation function, ∈ is a learnable parameter, MLP is a multi-layer perception mechanism, is the cumulative realization of neighbor nodes, and l is the number of layers.
[0037] Furthermore, the two-stream neural predictor is used to process the embedded features obtained in the graph isomorphism Transformer model to complete the prediction of circRNA-disease association. The specific process is expressed by the following formula:
[0038]
[0039] in, is the prediction result, σ l , W l and b l are the activation function, learnable parameter matrix and bias term of the lth layer respectively, δ is the fusion feature, MLP is the multi-layer perception mechanism, is the embedding of the first layer of circRNA, is the embedding of the disease at the lth layer, and || is the concatenation.
[0040] Furthermore, when constructing the graph isomorphic Transformer model and the two-stream neural predictor, the graph isomorphic Transformer model and the two-stream neural predictor are trained by a total loss function, and the total loss function is expressed by the following formula:
[0041] L=L1+L2+λ||α|| 2
[0042]
[0043] Where λ is the regularization parameter, α is the model parameter set, L1 is the BPR loss function, O is the training set, O = {(h,r,t,h′,t′)|(h,r,t)∈R + ,(h′,r,t′)∈R -}, R + For positive interaction, R - is the negative interaction set, σ is the sigmoid function, is the prediction score of the model for the positive sample (h, r, t), is the prediction score of the model for negative samples (h′, r, t′), and P is the set of all unobserved positive and negative triplet samples in equal proportion.
[0044] The second aspect of the present invention proposes a system for predicting circRNA-disease association based on graph isomorphism Transformer, comprising:
[0045] The knowledge graph module is used to build a CDA knowledge graph using multi-source heterogeneous data sets, which is convenient for alleviating the data sparsity problem in CDA prediction tasks and is conducive to extracting local deep information from the knowledge graph;
[0046] The fusion similarity network module is used to establish a fusion similarity network based on the CDA knowledge graph to facilitate feature extraction;
[0047] The feature module is used to input the output of the fusion similarity network into the graph isomorphism Transformer model to obtain embedded features, so as to obtain high-quality embedded representation;
[0048] The recognition module is used to input the embedded features into the two-stream neural predictor to complete the identification of potential circRNA-disease associations.
[0049] Beneficial effects of the present invention:
[0050] (1) This paper proposes an efficient knowledge representation learning method called graph isomorphism Transformer model, which has powerful diverse information processing capabilities and can deeply explore local and global interactions in knowledge graphs, thereby achieving more robust knowledge representation.
[0051] (2) The present invention designs a two-stream neural predictor that can effectively learn the nonlinear associations and interactions between biological data, and predict the CDA affinity score by learning biological data features from different perspectives, thereby significantly improving the accuracy and efficiency of the prediction and ensuring a more comprehensive and robust prediction.
[0052] (3) Extensive experimental studies have verified the superiority of the present invention in CDA prediction research and demonstrated its great potential in the field of non-coding RNA (ncRNA) and protein relationship prediction.
[0053] (4) This paper constructs a heterogeneous knowledge graph that integrates multi-source datasets to capture the diverse association information between diseases, circRNAs, lncRNAs, and miRNAs. This comprehensive graph can provide a comprehensive understanding of the complex biological interactions involved in CDA. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1A flowchart of a method for predicting circRNA-disease association based on graph isomorphism Transformer provided in an embodiment of the present invention.
[0055] Figure 2 A schematic diagram of the distribution of the number of association relationships in a data set provided in an embodiment of the present invention, including five types of relationships.
[0056] Figure 3 A schematic diagram of CircRNA-disease association analysis provided in an embodiment of the present invention.
[0057] Figure 4 A schematic diagram of a knowledge graph overview provided for an embodiment of the present invention.
[0058] Figure 5 The second flowchart of the method for predicting circRNA-disease association based on graph isomorphism Transformer provided in an embodiment of the present invention.
[0059] Figure 6 A schematic diagram of the expressive power ranking of Sum, Mean and Max on a multiset provided by an embodiment of the present invention.
[0060] Figure 7 A schematic diagram of an injective-based graph isomorphic network model provided for an embodiment of the present invention.
[0061] Figure 8 A schematic diagram of the AUC and AUPR performance comparison on the data set 1 provided in an embodiment of the present invention, wherein Figure A is AUC and Figure B is AUPR.
[0062] Fig. 9 A schematic diagram of the AUC and AUPR performance comparison on the dataset 2 provided by the embodiment of the present invention, wherein Figure A is AUC and Figure B is AUPR.
[0063] Fig.10 A schematic diagram of comparing the average association numbers of the dataset 1 and the dataset 2 with accurate recognition provided by the embodiment of the present invention, wherein A is the dataset 1 and B is the dataset 2.
[0064] Fig.11 A schematic diagram of a violin plot comparing the performance of different aggregation methods and prediction methods provided in an embodiment of the present invention, wherein A is the AUC value of data set 1, and B is the AUC value of data set 2.
[0065] Fig.12 A schematic diagram of the performance comparison of different neighbor feature calculation methods provided by the embodiment of the present invention. A and B represent the AUC value and AUPR value on data set 1, respectively; C and D represent the AUC value and AUPR value on data set 2, respectively.
[0066] Fig.13 A schematic diagram showing the performance comparison of different attention heads and encoder layers provided in an embodiment of the present invention.
[0067] Among them, A represents the AUC heat map on dataset 1; B represents the AUC heat map on dataset 2.
[0068] Fig.14 An architecture diagram of a system for predicting circRNA-disease associations based on graph isomorphism Transformer provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0070] Example 1
[0071] like Figure 1 As shown in the figure, the method for predicting circRNA-disease association based on graph isomorphism Transformer includes:
[0072] Step 1: Build a CDA knowledge graph using multi-source heterogeneous data sets.
[0073] Specifically, the multi-source heterogeneous data set includes data set 1 and data set 2. Data set 1 includes non-cancer related data, including 79 diseases, 330 circRNAs, 404 lncRNAs, and 265 miRNAs. In addition, 1,327 associations were extracted from it. Data set 2 includes cancer-related data to further verify the robustness of the present invention. After deduplication, 62 diseases, 514 circRNAs, 666 lncRNAs, and 573 miRNAs, as well as a total of 3509 associations, were obtained. The summary information of data set 1 and data set 2 is shown in Table 1.
[0074] Table 1. Correlation information between dataset 1 and dataset 2
[0075]
[0076] In order to describe the data distribution more intuitively, the data set is visualized. Figure 2The quantitative distribution of five types of association relationships in datasets 1 and 2 (lnc-dis(lncRNA-disease), lnc-mi(lncRNA-miRNA), mi-dis(miRNA-disease), circ-mi(circRNA-miRNA) and circ-dis(circRNA-disease)) are shown, and the quantitative distribution differences between each type of relationship are obvious.
[0077] In addition, the relationship between circRNA and disease was analyzed and visualized. Figure 3 As shown in Figure 2, approximately 85% of diseases are associated with no more than 10 circRNAs. Therefore, the circRNA-disease association in Datasets 1 and 2 is very sparse.
[0078] Use the previously constructed dataset to build a multi-source heterogeneous knowledge graph. Figure 4 An overview of the knowledge graph is provided, covering four entities, miRNA, circRNA, lncRNA and disease, and five different types of relationships, namely miRNA-circRNA0, circRNA-disease 1, disease-lncRNA 2, disease-miRNA 3 and miRNA-lncRNA 4. The ncRNA-disease knowledge graph can be expressed as "G = (E, R)", where E represents the entity set (miRNA, circRNA, disease and lncRNA) and R represents the relationship set (r1, r2, r3, r4 and r5). The internal form of the knowledge graph can be expressed as a (h, r, t) triple, where h, t∈E, h represents the head entity, r represents the tail entity node, and r∈R represents the relationship between entity nodes.
[0079] For example, the extracted triple data such as (1, 0, 655) indicates that the relationship between the entity with serial number 1 and the entity with serial number 655 is 0 (i.e., miRNA-circRNA).
[0080] Due to the characteristics of the knowledge graph, when constructing the knowledge graph, not only the knowledge of the original triple (h1, r, h2) is obtained, but also the knowledge of the corresponding reverse relationship (h1, r′, h2) and the self-relation (h1, r″, h2), where r′ represents the reverse relationship of relationship r, and r″ represents the self-relation of relationship r. This construction method alleviates the data sparsity problem in the CDA prediction task to a certain extent, and is conducive to the model to extract local deep information from the knowledge graph.
[0081] Step 2: Establish a fusion similarity network based on the CDA knowledge graph.
[0082] Specifically, a fusion similarity network is established based on the relationship matrix according to the existing technical means, such as the paper (KGETCDA: an efficient representation learning framework based on knowledge graphencoder from transformer for predicting circRNA-disease associations). The relationship matrix is constructed according to the data association of the CDA knowledge graph. If there is an association between two entity nodes, the value in the relationship matrix is 1, and if there is no association, it is 0. Then DSS, CFS and CGS are calculated according to the relationship matrix. Among them, DSS is disease semantic similarity, CFS is CircRNA functional similarity, and CGS is CircRNA Gaussian interaction profile kernel similarity. Finally, a fusion similarity network is constructed based on the above information. The fusion similarity network not only contains the initial association relationship, but also includes these three similarity relationship information.
[0083] Step 3: Input the output of the fused similarity network into the graph isomorphism Transformer model to obtain embedded features.
[0084] Step 4: Input the embedded features into the two-stream neural predictor to complete the identification of potential circRNA-disease associations.
[0085] like Figure 5As shown, the present invention proposes a method for predicting circRNA-disease associations based on graph isomorphism transformers to alleviate the data sparsity problem and explore high-order circRNA-disease interaction information. First, a heterogeneous knowledge graph integrating multi-source datasets is constructed to capture the diverse association information between diseases, circRNAs, lncRNAs, and miRNAs. Through this comprehensive graph, a comprehensive understanding of the complex biological interactions involved in disease associations (CDA) can be achieved. Next, a graph isomorphism transformer model is proposed to capture high-quality biological information and obtain a diverse knowledge representation of CDA. Finally, a dual-stream neural predictor is designed to predict CDA affinity scores by learning biological data features from different perspectives, thereby ensuring a more comprehensive and robust prediction and completing the identification of potential circRNA-disease associations.
[0086] Example 2
[0087] Based on the above embodiments, the present invention provides a method for constructing a graph isomorphism Transformer model, which specifically includes:
[0088] Knowledge Representation Learning (KRL) has gained wide recognition for its effectiveness in solving complex relational problems in knowledge graphs. KRL is able to model complex relational graphs and generate reliable entity and relation representations to support subsequent tasks. This paper proposes a graph isomorphism Transformer model to encode entities and relations to obtain high-quality embedding representations. Specifically, we define h 、e r and e t They represent the embeddings of the head entity, relation, and tail entity respectively, and it is assumed that the embedding dimensions of the three are the same, which is d. Then, a multi-head attention mechanism is used to obtain a richer embedding representation. The specific calculation process is as follows:
[0089] First, the query (Q), key (K), and value (V) of each head are calculated through linear transformation:
[0090]
[0091] Among them, Q i is the query of the i-th head, K i is the key of the ith head, V i is the value of the i-th head, is the projection weight of the query of the i-th head, is the projection weight of the key of the i-th head, is the projection weight of the value of the i-th head.
[0092] Then the scaled dot product attention mechanism is used to calculate the attention score for each head, which is expressed as follows:
[0093]
[0094] Among them, A i is the attention score of the i-th head, softmax is the normalized exponential function, and d is the embedding dimension of the head entity, relation, and tail entity embedding in the fusion similarity network.
[0095] The final output is expressed as follows:
[0096] Multihead(Q,K,V)=[A1,A2,...,A i ]W o
[0097] Among them, Multihead (Q, K, V) is a rich embedding representation, W o is the projection weight matrix, and [] is concatenation.
[0098] In addition, in order to fully capture the complex node relationships in graph structure data and enhance the nonlinear representation ability of the model, a feedforward neural network is further used to enhance the modeling ability of the model:
[0099] FFN(Multihead(Q,K,V))=max(Multihead(Q,K,V)W1+b1,0)W2+b2
[0100] Among them, FFN(x) is the embedded representation of the enhanced head entity, relation and tail entity, W1 is the hidden layer weight, b1 is the hidden layer bias, W2 is the output layer weight, and b2 is the output layer bias.
[0101] Through the above steps, the embedding representation of entities and relations is obtained. When calculating the rationality of the triple (h, r, t), previous work uses a simple dot product operation as the scoring function, but this cannot capture the complex associations between entities and relations. Therefore, a relation matrix multiplication is designed to enhance the expressive power of the model:
[0102]
[0103] Among them, f(h,r,t) is the output of relational matrix multiplication, w r is the relationship transformation matrix, Update the transpose of the embedding for the encoded relation, e t is the embedding of the tail entity.
[0104] Then, the BPR loss function is applied to train the knowledge representation learning part. Compared with the traditional pairwise ranking-based loss function, the BPR loss calculation method can increase the reliability of the model and better learn the relationship between entities in the knowledge graph:
[0105]
[0106] Among them, L1 is the BPR loss function, O is the training set, O = {(h,r,t,h′,t′)|(h,r,t)∈R + ,(h′,r,t′)∈R -}, R + For positive interaction, R - is the negative interaction set, σ is the sigmoid function, is the prediction score of the model for the positive sample (h, r, t), is the prediction score of the model for negative samples (h′, r, t′).
[0107] In addition, in order to reduce the noise in negative samples and improve the model's ability to capture real relationships, negative samples are generated based on the similarity negative sampling method. In summary, the model has the ability to capture local and global information from complex graph structure data, which can provide rich and accurate input for downstream tasks.
[0108] In the knowledge graph structure, the traditional information propagation method regards all neighbor nodes as homogeneous, which is not friendly to the CDA prediction task. In order to better distinguish the differences between different nodes, the attention information propagation architecture is improved based on previous work, and a graph isomorphism layer is constructed, which consists of two parts: information propagation unit and information aggregation unit.
[0109] The key to ensuring high-quality expression is to make full use of the information of entities (such as circRNA and diseases) in the information propagation unit. In the knowledge graph, the relationship between entities can be obtained not only from the directly connected edges, but also indirectly through transitive means. In addition, it is necessary to recognize that the neighbor information of an entity has different degrees of influence on the importance of the entity.
[0110] Specifically, for a node h, the set of all its neighboring nodes and their relationship triples can be expressed as N h ={(h,r,,t)|h,t∈E,r∈R}, to represent the entity, the ego-network of node h is defined as follows:
[0111]
[0112] in, is the egocentric network of node h, θ(h,r,t) is the normalized information propagation factor, which determines the amount of information propagated from neighbor t to node h, and e t The tail entity.
[0113] Since the importance of different neighbor nodes varies, knowledge-aware attention is used to calculate the propagation factor of each neighbor node. The calculation formula is as follows:
[0114]
[0115] in, is the information propagation factor, W r is the transformation matrix of the relation, which projects the entity into the relation space, e h is the embedding of the head entity, e r is the embedding of the relationship, and tanh is the hyperbolic tangent function.
[0116] In addition, in order to better understand the process of information propagation in the network and strengthen the focus on important nodes, it is also necessary to normalize the above propagation factors:
[0117]
[0118] Among them, θ(h,r,t) is the normalized information propagation factor.
[0119] After the above processing, the feature representations between different nodes can be well distinguished. At the same time, in the process of attention propagation, this normalization processing also helps to reduce the impact of noise, making the focus more focused on important nodes.
[0120] Information Aggregation Unit (CIN): In graph neural networks, single injection (one-to-one mapping) is a key indicator for evaluating the ability of aggregation expression. However, the conventional average method tends to learn the data distribution, while taking the maximum value may ignore repeated information. In contrast, the summation method can retain the complete information of neighboring nodes, which meets the definition of single injection. Figure 6 The information expression of the above three strategies is shown, where the Input part represents the aggregated neighbor network, Sum learns the entire multiset, Mean learns the overall distribution, and Max ignores repeated information.
[0121] The information aggregation unit aims to improve the discriminative ability of the graph neural network to capture complex graph structure patterns. The distinguishability of neighborhood aggregation is achieved by adding weighted self-edges to each node. As a result, the network can more accurately understand the relationship between nodes. This process is shown in Figure 2. Figure 7As shown in the figure, (A) is adding a weighted self-edge to a node. (B) is the graph structure state of neighbor nodes that exchange weighted self-edges. (C) is the aggregation of the corresponding part (B).
[0122] In addition, through multi-layer neighborhood aggregation and updating, the information aggregation unit can gradually capture high-order features in the graph and reflect the global graph structure through continuous updating of node features. Among them, the summation method adopted by the information aggregation unit convolution layer (GINConv) not only ensures the learning of the entire network structure, but also improves the model's ability to express graph structure networks. Considering the excellent performance of the information aggregation unit in global network learning, the GINConv layer is used to aggregate entity h and its ego-network to better capture complex relationship patterns. The aggregation definition of entity h and its ego-network is as follows:
[0123]
[0124] in, for and Aggregation, is the head entity of layer l-1, is the egocentric network of node h in layer l-1, LeakyReLU is the activation function, ∈ is a learnable parameter, MLP is a multi-layer perception mechanism, is the cumulative realization of neighbor nodes, and l is the number of layers.
[0125] Preferably, in order to evaluate the effectiveness of the GIN aggregation strategy, the GraphSage aggregator, the Bi-Interaction aggregator, and the currently best performing GCN aggregator are also set up for comparison.
[0126]
[0127] in, The node update function for the GraphSage aggregator at level l, is the node update function of the Bi-Interaction aggregator at layer l, is the node update function of the GCN aggregator at layer l, || is the concatenation, and ⊙ is the element-wise product.
[0128] Therefore, the formula can be further extended to describe the embedding of node h in layer l. Specifically, the embedding of node h in layer l can be obtained by Perform iterative updates:
[0129]
[0130] in, is an injective function, and f operates on multiset.
[0131] The information propagated by node h in the l-ego network is defined as:
[0132]
[0133] As the number of stacking layers increases, more reliable and rich encoding information can be obtained, thereby better describing the multi-hop neighbor interaction information of circRNA-disease. Finally, the number of layers l is set to 4, because the fourth-order information is sufficient to fully capture the potential features and interaction relationships in the knowledge graph, and there will be no problems such as transition smoothness, thereby achieving more accurate characterization.
[0134] Example 3
[0135] Based on the above embodiments, the present invention provides a method for constructing a dual-stream neural predictor, which specifically includes:
[0136] After multiple layers of information propagation and aggregation, we obtain high-quality embedding representations of each entity at each level. Next, we connect the information of each layer in series to enrich the final node representation:
[0137]
[0138] Among them, || is a series connection, is the embedding of the first layer of circRNA, is the embedding of the l-th layer disease.
[0139] According to the above formula, we obtain rich feature representations of different nodes. In order to more accurately capture the complex relationship between nodes, we define is the embedding of CDA. Then a two-stream network predictor is designed for calculation to improve the nonlinear prediction ability of the model. The specific calculation process is as follows:
[0140]
[0141] in, is the prediction result, σ l , W l and b l are the activation function, learnable parameter matrix and bias term of the lth layer respectively, δ is the fusion feature, and MLP is the multi-layer perception mechanism.
[0142] Next, define the BPR loss function to optimize the above results:
[0143]
[0144] Among them, P is the set of all unobserved positive and negative triple samples with equal proportions.
[0145] Finally, to achieve end-to-end optimization, the overall loss function is defined as follows:
[0146] L=L1+L2+λ||α|| 2
[0147] Among them, λ is the regularization parameter and α is the model parameter set.
[0148] Example 4
[0149] Based on the above embodiments, the present invention provides an evaluation method for predicting circRNA-disease association based on graph isomorphism transformer (GIT-DSP), which specifically includes:
[0150] A five-fold cross-validation method was used to evaluate the method of predicting circRNA-disease associations based on graph isomorphism Transformer and compare it with other SOTA models. In the five-fold cross-validation, the dataset will be divided into five equal parts, four of which will be selected for training each time, and the rest will be used as the test set. Each part will be used as a validation set in turn so that the model can fully learn the data. Then, the average of the five evaluation results is taken as the final result. Finally, the performance is evaluated by accuracy (Acc), recall (Rec), precision (Pre), and F1 score (F1). The formula for the evaluation index is expressed as follows:
[0151]
[0152] Among them, TP and TN represent the number of correctly identified positive samples and negative samples, respectively, FP and FN represent the number of incorrectly identified positive samples and negative samples, respectively, Acc is the accuracy, Rec is the recall rate, Pre is the precision, and F1 is the F1 score.
[0153] In addition to the above evaluation indicators, the area under the precision-recall (PR) curve (AUPR) and the area under the receiver operating characteristic (ROC) curve (AUC) are also used to comprehensively evaluate the performance of the model. The PR curve is used to show the trade-off between the precision and recall rate of the model at different thresholds, while the ROC curve shows the trade-off between the true positive rate (TPR) and the false positive rate (FPR). The mathematical expressions of TPR and FPR are as follows:
[0154]
[0155]
[0156] Among them, TPR is the true positive rate and FPR is the false positive rate.
[0157] To further measure the prediction ability of the present invention, the number of correctly identified CDAs is also used as one of the evaluation criteria of the present invention. Top-K is used to express the number of associations between the top K similar RNAs predicted by the present invention and the corresponding diseases.
[0158] The present invention is implemented on NVIDIA V100 GPU using Pytorch framework. The parameter settings are based on extensive comparative experiments.
[0159] In the training phase, in order to fully train the model, the epoch was set to 100 and the learning rate was set to 1e-4. In order to accelerate convergence and improve model stability, dynamic learning rate decay was set according to the model performance schedule, and the Adam optimizer was used for optimization. In addition, in order to fully capture high-order interaction information, a four-layer propagation structure [512-256-128-64] was selected, the hidden layer was set to 256, and the embedding size of entities and relations was set to 2048. In order to reduce overfitting, a dropout of 0.1 was applied after each propagation layer, the number of attention heads was set to 16 in the knowledge representation learning phase, and a dropout of 0.2 was applied to each KRL module. In the selection of the aggregation strategy of GIN, the sum method with the best performance was selected.
[0160] In the prediction stage, the epochs were set to 32, the learning rate was 5e-4, and a weight decay of 1e-7 was used to reduce the risk of overfitting. The MLP-based Stream 1 structure adopted a gradually decreasing strategy [6016, 3008, 1504, 752], which helps the network learn data at different levels of abstraction. The MLP-based Stream 2 structure is designed to be [6016, 6016, 3008], which aims to capture data features from multiple angles and identify different pattern associations. In addition, when the positive-negative sample ratio is set to 1:10 and the number of the closest circRNAs is set to 7, the present invention shows the best performance.
[0161] To demonstrate the superiority of the present invention, the present invention is also compared with 9 SOTAs, as shown in Table 2.
[0162] Table 2. Overview of 9 comparison models
[0163]
[0164] The comparison model covers three categories and a total of 9 methods, as follows:
[0165] KATZHCDA: Using the information in the heterogeneous graph, the interaction between circRNA and disease is predicted through the KATZ algorithm.
[0166] RWR: The restart random walk method is used to simulate the random walk process on the network, and the restart probability is introduced to adjust the direction of the walk to predict the potential circRNA-disease association.
[0167] CD-LNLP: It is a linear neighborhood label propagation method that implements label propagation based on known circRNA associations and disease association maps to predict new circRNA-disease associations.
[0168] RWR-KNN: This model combines the RWR algorithm and the K nearest neighbor algorithm. The random walk algorithm is used to evaluate node similarity, while the K nearest neighbor algorithm is used to enhance the node classification accuracy and improve the CDA prediction effect.
[0169] ICIRCDA: This model preliminarily calculates circRNA-disease associations based on diverse biological information and corrects false negative associations through local interaction spectra. Matrix decomposition is used to calculate the final circRNA-disease association score.
[0170] RNMFLP: This method uses a robust non-negative matrix factorization method to capture potential circRNA-disease association pairs, and then applies the LP algorithm to predict more accurate CDA from candidate association pairs.
[0171] DMFCDA: Based on deep matrix factorization, nonlinear relationships are modeled through multi-layer neural networks to automatically learn the potential representation of circRNA-disease.
[0172] GMNN2CD: Graph Markov neural network is used to obtain deep features from low-dimensional feature representation, and then labels are propagated based on graph autoencoders.
[0173] KGETCDA: Transformer-based knowledge representation learning followed by multi-layer perceptron prediction of circRNA-disease association scores based on embeddings.
[0174] Next, we conduct specific experiments. All models are compared using five-fold cross validation on the same device, and the parameter settings of the comparison models are the best parameter settings in each model.
[0175] For data set 1, Figure 8As shown, it can be seen that the average AUC value of the present invention reaches 0.9381, which is 3.08% higher than the optimal benchmark model, and the average AUPR value reaches 0.0388, which is 63.03% higher than the optimal benchmark model.
[0176] For dataset 2, Fig. 9 As shown, it can be found that the average AUC value of the present invention reaches 0.8728, which is 2.89% higher than the optimal benchmark model, and the average AUPR value reaches 0.0566, which is 11.20% higher than the optimal benchmark model, which once again proves the powerful prediction and generalization capabilities of the present invention.
[0177] In addition, a bar chart of the number of CDA relationships correctly predicted by the present invention is also drawn, such as Fig.10 As shown, it can be clearly seen that the present invention can identify more associations and has a stronger recognition ability. (The average number of associations accurately identified in the top 10 to top 40 of data set 1 is 30.4, 40.0, 49.8 and 59.0 respectively, and the average number of associations accurately identified in the top 10 to top 40 of data set 2 is 51.0, 65.4, 74.6 and 80.8 respectively).
[0178] Further analysis shows that the performance of other methods is usually limited when predicting associations when faced with complex multi-relation scenarios and data sparsity. This is because they often fail to make full use of the data or fail to effectively capture the high-order relationships hidden in the data. However, the present invention applies knowledge graph technology and combines it with the GIN aggregator for updating and aggregating CDA information and the two-stream neural predictor for information extraction and prediction. This integrated approach not only obtains high-quality embeddings and rich knowledge representations, but also provides excellent prediction results. Therefore, when processing multi-relation data or sparse data, GIT-DSP can still demonstrate excellent prediction capabilities and generate reliable prediction results. Compared with all baseline models, the present invention performs better.
[0179] The present invention also conducts a series of ablation experiments to verify the effectiveness of the present invention. It should be pointed out that in order to more comprehensively evaluate the performance and stability of the model, all ablation experiments of the present invention are carried out under five-fold cross validation, and only one module is replaced at a time. With other settings unchanged, the effects of different aggregation modules (GIN (Ours), GCN, GraphSage, and Bi-Interaction) and different prediction modules (DualSNP (Ours), SDualSNP, MLP) in KRL on model performance are first considered, where SDualSNP represents two MLPs with the same structure superimposed. Then, the effects of different methods (maximum value, average value, and SUM) for calculating neighborhood features in GINConv are studied. Secondly, in order to be more rigorous, the effects of different attention heads and layers in KRL on the performance of GIT-DSP are also studied.
[0180] In this section, the aggregation module and the prediction module are arranged and combined to fully demonstrate the effectiveness of the present invention. Specifically, the prediction module was first designed based on GIN, and different MLP configurations were tried. Then, taking the traditional single-layer MLP (GIN+M) as the benchmark, two MLP stacks with the same structure (GIN+S) and two MLP stacks with different structures (GIN+D(Ours)) were used respectively. Similarly, the methods based on GCN, Bi-Interaction, and GraphSage also made the above settings to fully compare the impact of different aggregation modules and different prediction modules on model performance. The results of the five-fold cross validation on datasets 1 and 2 are shown in Figure 2. Fig.11 shown.
[0181] For Dataset 1, from the perspective of aggregation strategy, the GIN aggregation method used by the model significantly outperforms the other three aggregation methods in terms of AUC indicators, which strongly proves the effectiveness of the GIN aggregation method. Furthermore, from the perspective of a single combination method, among the different prediction methods, the GIN+D method stands out among all combinations. This proves the superiority of the GIN aggregation method in CDA prediction and also highlights the performance advantage of the two-stream neural predictor strategy.
[0182] For Dataset 2, the GIN aggregation method is slightly better than the GCN-based aggregation method, and significantly better than the Bi-Interaction-based and GraphSage-based aggregation strategies. From the comparison of single combination methods, the GIN+D strategy is still better than all other comparison methods, which once again proves the excellent generalization ability and strong prediction ability of GIT-DSP.
[0183] The present invention also verifies the impact of different neighbor feature calculation methods, specifically, the impact of the three aggregation strategies of Max, Mean and Sum on the performance of the model. The AUC and AUPR values obtained by replacing the GIN neighborhood aggregation strategy on the two data sets are shown in Figure 2. Fig.12 shown.
[0184] It can be clearly seen from the bar graph that the AUC and AUPR results of the Sum method are higher than those of the Mean and Max methods in both dataset 1 and dataset 2. This result shows that rich neighborhood information can promote the learning of high-order information, thus providing a solid foundation for the CDA prediction task. This further proves the importance of feature interaction in the information aggregation stage, and effective feature interaction can improve the performance of the method.
[0185] In order to further explore the learning effect of knowledge representation based on Transformer, ablation experiments with different numbers of attention heads and encoder layers were designed. It should be emphasized that with the increase of the number of attention heads and encoder layers, the time and space complexity of the model will also increase accordingly. The number of attention heads is set to 8, 16, 32 and 64, and the number of encoder layers is set to 1, 2 and 3 layers. Through these configurations, the impact of different numbers of attention heads and encoder layers on model performance can be systematically evaluated. The specific experimental results are shown in the figure. Fig.13 shown.
[0186] According to the above experimental results, the present invention achieves the best AUC value when the number of attention heads is set to 16 and the number of encoder layers is set to 1. This shows that the present invention has superior performance in high-order interactive information exploration and nonlinear association relationship prediction. Further analysis found that when the number of attention heads is greater than 16 or the number of encoder layers is greater than 1, the AUC value decreases to varying degrees. This may be because the complex network structure introduces irrelevant noise information, thereby interfering with the prediction results.
[0187] Example 5
[0188] Based on the above embodiments, the embodiment of the present invention provides a case study process, which specifically includes:
[0189] To further evaluate the performance of GIT-DSP, case studies were conducted on Dataset 1 and Dataset 2. GIT-DSP was used to learn known association information, predict association probabilities and rank them in descending order. Finally, the predicted rankings were combined with public datasets and related literature for verification one by one.
[0190] For Dataset 1, a case study of acute myeloid leukemia (AML) was conducted. According to statistics, the 5-year relative survival rate of AML patients in the United States is only 31.9%, accounting for 1.3% of new cancer cases in the United States. In response to the challenges of AML treatment, more and more studies have begun to focus on circRNA. Effective intervention of circRNA can provide new targets and strategies for the treatment of AML. The prediction results of AML-related circRNA are shown in Table 3.
[0191] Table 3. Prediction results of the top ten candidate circRNAs related to acute myeloid leukemia
[0192]
[0193] According to Table 3, among the top 10 circRNAs with potential associations with acute myeloid leukemia, 8 (hsa_circ_100290, circ-anapc7, circpan3, hsa_circ_0035381, hsa_circ_0001187, hsa_circ_0004277, circ_aff2, and hsa_circ_0000254) have been confirmed. For example, circ-anapc7 is involved in the pathogenesis of AML by acting as a sponge for the miR-181 family. In addition, circ-anapc7 is also correlated with peripheral blood leukocyte counts and the percentage of primitive cells in the bone marrow. Circpan3 promotes drug resistance in AML by regulating autophagy. Hsa_circ_0004277 inhibits the activity of AML cells by adsorbing miR-134-5p or overexpressing it.
[0194] For Dataset 2, a case study of lung cancer is conducted. According to the latest statistics, lung cancer is the most common cancer in the world, accounting for 12.4% of global cancer cases, posing a huge threat to human health. Studying the expression pattern and function of lung cancer-related circRNA has important theoretical and clinical significance for the diagnosis, treatment and prognosis of lung cancer. The prediction results of lung cancer-related circRNA are shown in Table 4.
[0195] Table 4. Prediction results of the top ten candidate circRNAs related to lung cancer
[0196]
[0197] As shown in Table 4, among the top 10 lung cancer-related circRNAs with the highest prediction scores, 7 (hsa_circ_0007059, circMTO1, circ-PRKCI, hsa_circ_0046264, circPUM1, CDR1as, circ-ERBB2) have been supported by relevant literature. For example, hsa_circ_0007059, which ranks first, can indirectly inactivate the Wnt / β-catenin and ERK1 / 2 pathways, thereby inhibiting the proliferation of lung cancer cells. CircMTO1, which ranks third, can inactivate the Notch signaling pathway by increasing the expression of QKI-5, thereby inhibiting the growth of lung adenocarcinoma (LUAD). The latest study found that circ-PRKCI, which ranks fourth, promotes the proliferation of LUAD cells, while inhibiting the expression of circ-PRKCI will cause the LUAD cell cycle to stagnate in the G1 phase, thereby inhibiting the growth and spread of lung cancer cells.
[0198] In summary, through the case analysis of acute myeloid leukemia and lung cancer, the accuracy and reliability of GIT-DSP in circRNA-disease association prediction were verified. At the same time, this study also demonstrated its important value in disease research in the biological field, and provided new ideas and methods for in-depth exploration and understanding of the molecular mechanisms of diseases. It should be pointed out that the unverified circRNAs in the prediction results deserve further exploration and verification by biologists.
[0199] Example 6
[0200] Corresponding to the above method, such as Fig.14 As shown in the figure, the system for predicting circRNA-disease association based on graph isomorphism Transformer includes:
[0201] The knowledge graph module is used to build a CDA knowledge graph using multi-source heterogeneous data sets.
[0202] The fusion similarity network module is used to establish a fusion similarity network based on the CDA knowledge graph.
[0203] The feature module is used to input the output of the fusion similarity network into the graph isomorphism Transformer model to obtain embedded features.
[0204] The recognition module is used to input the embedded features into the two-stream neural predictor to complete the identification of potential circRNA-disease associations.
[0205] It should be noted that the system for predicting circRNA-disease association based on graph isomorphism Transformer provided in the embodiment of the present invention is to implement the method for predicting circRNA-disease association based on graph isomorphism Transformer in the above embodiment. Its specific functions can be referred to the above method embodiments, which will not be repeated here.
[0206] Example 7
[0207] Based on the above embodiments, an embodiment of the present invention provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for predicting circRNA-disease association based on graph isomorphism Transformer as described in the above embodiments is implemented.
[0208] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute the method for predicting circRNA-disease association based on graph isomorphism Transformer as described in the above embodiment.
[0209] In summary, the present invention proposes an efficient knowledge representation learning method called graph isomorphism Transformer model, which has powerful diverse information processing capabilities and can deeply explore local and global interactions in knowledge graphs, thereby achieving more robust knowledge representation. The present invention designs a two-stream neural predictor that can effectively learn nonlinear associations and interactions between biological data, and predict CDA affinity scores by learning biological data features from different angles, thereby significantly improving the accuracy and efficiency of predictions and ensuring more comprehensive and robust predictions. Extensive experimental studies have verified the superiority of the present invention in CDA prediction research and demonstrated its great potential in the field of non-coding RNA (ncRNA) and protein relationship prediction. The present invention constructs a heterogeneous knowledge graph that integrates multi-source data sets to capture diverse association information between diseases, circRNAs, lncRNAs, and miRNAs. Through this comprehensive graph, a comprehensive understanding of the complex biological interactions involved in CDA can be achieved.
[0210] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting circRNA-disease association based on graph isomorphism Transformer, characterized in that: include: Step 1: Build a CDA knowledge graph using multi-source heterogeneous data sets; Step 2: Establish a fusion similarity network based on the CDA knowledge graph; Step 3: Input the output of the fusion similarity network into the graph isomorphism Transformer model to obtain the embedded features; Step 4: Input the embedded features into the two-stream neural predictor to complete the identification of potential circRNA-disease associations.
2. The method for predicting circRNA-disease association based on graph isomorphism Transformer according to claim 1, characterized in that: The CDA knowledge graph is expressed as follows: G=(E,R) Among them, G is the CDA knowledge graph representation, E is the entity set, E = {miRNA, circRNA, disease, lncRNA}; R is the relationship set, R = {r1, r2, r3, r4, r5}; among them, r1 is circ-miRNA 0, r2 is miRNA-disease 1, r3 is miRNA-lncRNA 2, r4 is lncRNA-disease 3, and r5 is circ-disease4; The internal form of the CDA knowledge graph is represented by the following triples: (h, r, t) Among them, h,t∈E, E represents the entity set (miRNA, circRNA, disease and lncRNA), h is the head entity node, t is the tail entity node, r is the relationship between entity nodes, r∈R.
3. The method for predicting circRNA-disease association based on graph isomorphism Transformer according to claim 1, characterized in that: The graph isomorphism Transformer model includes a Transformer encoder and a graph isomorphism layer connected in sequence; The Transformer encoder includes a multi-head attention mechanism, a feedforward neural network, and a relation matrix multiplication connected in sequence; the multi-head attention mechanism is used to embed the head entity, the relation, and the tail entity of the fusion similarity network to obtain a rich embedding representation of the embedding of the head entity, the relation, and the tail entity; The feedforward neural network is used to enhance the rich embedding representation to obtain enhanced embedding representations of the head entity, the relationship, and the tail entity; The relationship matrix multiplication is used to further enhance the embedded representations of the enhanced head entity, relationship, and tail entity to obtain the final embedded representations of the head entity, relationship, and tail entity; The graph isomorphism layer includes an information propagation unit and an information aggregation unit connected in sequence; the information propagation unit is used to propagate the final embedded representation of the head entity, the relationship and the tail entity to obtain the relationship between the entities; The information aggregation unit is used to aggregate the relationships between entities to obtain the aggregated relationships between the entities.
4. The method for predicting circRNA-disease association based on graph isomorphism Transformer according to claim 3, characterized in that: The multi-head attention mechanism is used to process the embedding of the head entity, relationship and tail entity in the fusion similarity network, and then obtain a rich embedding representation. The specific process is expressed by the following formula: Multihead(Q,K,V)=[A1,A2,...,A i ]W o Among them, Multihead (Q, K, V) is a rich embedding representation, A i is the attention score of the i-th head, W o is the projection weight matrix, softmax is the normalized exponential function, Q i is the query of the i-th head, V i is the value of the i-th head, K i The transpose K i , K i is the key of the i-th head, d is the embedding dimension of the head entity, relation, and tail entity embedding in the fused similarity network, and [] is concatenation.
5. The method for predicting circRNA-disease association based on graph isomorphism Transformer according to claim 4, characterized in that: The feedforward neural network is used to enhance the nonlinear ability of the output of the multi-head attention mechanism to obtain the enhanced embedding representation of the head entity, relationship and tail entity. The specific process is expressed by the following formula: FFN(Multihead(Q,K,V))=max(Multihead(Q,K,V)W1+b1,0)W2+b2 Among them, FFN(x) is the embedded representation of the enhanced head entity, relation and tail entity, W1 is the hidden layer weight, b1 is the hidden layer bias, W2 is the output layer weight, and b2 is the output layer bias.
6. The method for predicting circRNA-disease association based on graph isomorphism Transformer according to claim 3, characterized in that: The information propagation unit is used to propagate the output of the relationship matrix multiplication to obtain the relationship between entities. The specific process is expressed by the following formula: in, is the egocentric network of node h, θ(h,r,t) is the normalized information propagation factor, e t is the embedding of the tail entity, W r is the transformation matrix of the relationship, e h is the embedding of the head entity, e r is the embedding of the relation, tanh is the hyperbolic tangent function, is the information propagation factor before normalization.
7. The method for predicting circRNA-disease association based on graph isomorphism Transformer according to claim 3, characterized in that: The information aggregation unit is used to aggregate the relationship between entities obtained by the information dissemination unit to obtain the relationship between the aggregated entities. The specific process is expressed by the following formula: in, for and Aggregation, is the head entity of layer l-1, is the egocentric network of node h in layer l-1, LeakyReLU is the activation function, ∈ is the learnable parameter, MLP is the multi-layer perception mechanism, is the cumulative realization of neighbor nodes, and l is the number of layers.
8. The method for predicting circRNA-disease association based on graph isomorphism Transformer according to claim 1, characterized in that: The two-stream neural predictor is used to process the embedded features obtained in the graph isomorphic Transformer model to complete the prediction of circRNA-disease association. The specific process is expressed by the following formula: in, is the prediction result, σ l , W l and b l are the activation function, learnable parameter matrix and bias term of the lth layer respectively, δ is the fusion feature, MLP is the multi-layer perception mechanism, is the embedding of the first layer of circRNA, is the embedding of the disease at the lth layer, and || is the concatenation.
9. The method for predicting circRNA-disease association based on graph isomorphism Transformer according to claim 1, characterized in that: When constructing the graph isomorphic Transformer model and the two-stream neural predictor, the graph isomorphic Transformer model and the two-stream neural predictor are also trained through the total loss function, and the total loss function is expressed as follows: L=L1+L2+λ||α|| 2 Where λ is the regularization parameter, α is the model parameter set, L1 is the BPR loss function, O is the training set, O = {(h,r,t,h′,t′)|(h,r,t)∈R + ,(h′,r,t′)∈R - }, R + For positive interaction, R - is the negative interaction set, σ is the sigmoid function, is the prediction score of the model for the positive sample (h, r, t), is the prediction score of the model for negative samples (h′, r, t′), and P is the set of all unobserved positive and negative triplet samples in equal proportion.
10. A system for predicting circRNA-disease associations based on graph isomorphism Transformer, characterized in that: include: The knowledge graph module is used to build a CDA knowledge graph using multi-source heterogeneous data sets; The fusion similarity network module is used to build a fusion similarity network based on the CDA knowledge graph; The feature module is used to input the output of the fusion similarity network into the graph isomorphism Transformer model to obtain embedded features; The recognition module is used to input the embedded features into the two-stream neural predictor to complete the identification of potential circRNA-disease associations.