Relationship Prediction Method Based on Multi-Layer Collaborative Attention Map Collaborative Filtering

By applying a multi-layer synergistic attention map collaborative filtering method in the relationship network of circRNA and disease, the problem of insufficient extraction of circRNA-disease collaboration signals in the prior art is solved, and more accurate relationship prediction and feature extraction effects are achieved.

CN115985387BActive Publication Date: 2025-06-24JIANGNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310025368.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-06-24
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

When existing methods use circRNA and disease relationship networks to mine potential relationships, they lack effective display and encoding of collaborative signals between key circRNA-diseases, resulting in insufficient feature extraction.

Method used

Using a collaborative filtering method based on multi-layer synergistic attention map, a multi-layer characteristic characterization of circRNA and disease is constructed, combined with a deep autoencoder and a multi-layer synergistic attention mechanism, the interaction information between circRNA and disease is extracted and fused, and relationship prediction is performed.

Benefits of technology

The deep interactive relationship between circRNA and disease was effectively explored, the defect of insufficient embedding construction in the feature extraction stage was made up for, and the accuracy of relationship prediction and the nonlinear modeling ability of the model were improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115985387B_ABST
    Figure CN115985387B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of intelligent cell biometric recognition, and specifically relates to a relationship prediction method based on multi-layer collaborative attention graph collaborative filtering. A relationship prediction method based on multi-layer collaborative attention graph collaborative filtering, the method includes four stages: construction of circRNA and disease features, multi-feature fusion of circRNA and disease, multi-layer collaborative attention representation learning, and model training and relationship prediction based on collaborative filtering. Based on the existing feature descriptors, this method constructs a propagation mechanism on the central network to deeply mine the interaction between circRNA and disease, so that the circRNA-disease relationship network can be fully utilized. In the process of feature extraction and construction, this method globally captures the key collaboration signals hidden in the circRNA-disease relationship network, and uses the interaction information to make up for the defect of insufficient embedding construction in the feature extraction stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent cell biometric identification, and particularly relates to a relationship prediction method based on multi-layer collaborative attention graph collaborative filtering. Background Art

[0002] circRNA, full name circular ribonucleic acid, is a non-coding RNA that widely participates in the regulation of transcriptional and post-transcriptional gene expression. For a long time, circRNA has been considered a by-product of gene recombination due to its low abundance and lack of known functions. With the development of high-throughput RNA sequencing and bioinformatics analysis, more and more circRNAs have been discovered and identified. Currently, more than 30,000 circRNAs have been discovered, and their unique structure is more likely to maintain stability than linear RNA. Multiple studies have found that circRNA can participate in the regulatory biological processes of various malignant tumors, including cell cycle, tumorigenesis, invasion, metastasis, apoptosis, angiogenesis, etc., by sponging RNA, binding to RNA-binding proteins (RBPs), regulating transcription or affecting translation.

[0003] Currently, circRNA has become an important research hotspot in the biomedical field. Studies have found that circRNA can bind to miRNA as an RNA sponge and increase downstream gene expression by regulating miRNA activity, thus promoting tumor progression. For example, the earliest circRNA CiRS-7 contains more than 70 miR-7 binding sites and, as a miR-7 sponge, reduces the effect of miR-7 on target mRNA. In addition, circRNA also participates in the processes of transcription, translation, splicing, and binding to RNA-binding proteins (RBPs). And circular RNA can interact with other RNA molecules such as mRNA and lncRNA, and even directly interact with DNA to promote or inhibit the transcription process. Therefore, further studying the interaction between circRNA and tumors and developing new circRNAs as molecular markers or potential targets will have broad application prospects in the early diagnosis, treatment evaluation, prognosis prediction, and even tumor gene therapy of tumors and diseases.

[0004] For the research on the relationship between circRNA and diseases, some relevant databases have been established. There are also many methods to mine potential circRNA-disease relationships from the circRNA-disease relationship network by using machine learning models. The focus is mainly on directly predicting unknown circRNA-disease relationships by using the sequence features of the original circRNA sequences or the circRNA-disease relationship network. Most methods lack the explicit encoding of the collaborative signals between key circRNA-diseases, and the key collaborative signals are generally hidden in the circRNA-disease relationship network. Therefore, how to design feature extraction methods to make up for the deficiencies of embedding construction remains an important challenge. Summary of the Invention

[0005] Most of the research on the relationship between circRNA and diseases is based on single database use cases. In the present invention, the recently established experimentally verified circR2Disease, circ2Disease, and circRNADisease are used as a unified dataset for the association between circRNA and diseases to measure the performance of the model. The present invention involves 650 known circRNA-disease relationships between 585 circRNAs and 88 diseases in the circR2Disease database, 270 circRNA-disease relationships between 249 circRNAs and 60 diseases included in the circ2Disease database, and 350 circRNA-disease relationships between 330 circRNAs and 48 diseases included in the circRNADisease database. At the same time, the similarity information between circRNA and diseases is jointly constructed by combining the PubMed database and its Mesh database below.

[0006] The technical solution of the present invention is as follows:

[0007] A relationship prediction method based on multi-layer collaborative attention graph collaborative filtering, which includes four stages: feature construction of circRNA and diseases, multi-feature fusion of circRNA and diseases, multi-layer collaborative attention representation learning, and model training and relationship prediction based on collaborative filtering, as shown below:

[0008] The first stage: the feature construction stage of circRNA and diseases. This stage includes four steps, namely the construction of the first initial feature of diseases, the construction of the second initial feature of diseases, the construction of the first initial feature of circRNA, and the construction of the second initial feature of circRNA. The specific steps are as follows:

[0009] In the Mesh database, diseases are saved in the form of a directed acyclic graph. In Mesh, nodes are represented as diseases, and edges are represented as the relationships between them. Suppose there is a disease d, which can be described as a DAG d =(d, A d , E d ), where A d represents all the ancestor nodes of d including d, and E d is the set of edges corresponding to the relationships between these diseases. If a disease e is in the DAG d , then its contribution value and semantic value to disease d can be calculated. Suppose the higher the coincidence degree of the ancestor diseases shared by two diseases in the DAG, the greater the semantic similarity between the two diseases. Thus, the first semantic similarity model SD1(d(i), d(j)) is obtained. To give due attention to diseases with fewer occurrences, a second semantic similarity model SD2(d(i), d(j)) can be introduced, that is, the semantic similarity feature of rare diseases, to increase the contribution value of diseases with fewer occurrences. Then, the two semantic similarity models of diseases are fused to obtain the semantic similarity model of diseases:

[0010]

[0011] Based on the assumption that diseases with similar functions may be affected by similar circRNAs, the present invention uses the Gaussian Interaction Profile Kernel (GIP) similarity matrix to represent the similarity between diseases and represents it as DGS(d(i), d(j)).

[0012] Similarly, based on the assumption that there may be similar circRNAs with similar functions that affect similar diseases, a GIP similarity matrix is generated to represent the similarity between circRNAs and is represented as CGS(c(i), c(j)).

[0013] The functional similarity between two circRNAs is usually based on the assumption that circRNAs corresponding to diseases with similar semantics are also similar in function. Using this method, the functional similarity of pairs of circRNAs is generated and represented as FC(c(i), c(j)).

[0014] The specific steps of this stage are as follows:

[0015] Step 1: Use the original circRNA and disease data to construct the semantic similarity information SD of diseases.

[0016] Step 2: Use the original circRNA and disease data to construct the Gaussian Interaction Profile similarity information DGS of diseases.

[0017] Step 3: Use the original circRNA and disease data to construct the Gaussian interaction profile similarity information CGS of circRNA.

[0018] Step 4: Use the original circRNA and disease data to construct the functional similarity information FC of circRNA.

[0019] Phase II: In order to facilitate the input of the multi-layer collaborative attention representation learning model, it is an effective method to fuse the features of circRNA and disease to form a complete feature. Feature fusion can not only reveal the relationship between circRNA and disease, but also express the internal connection between circRNA and disease. It is helpful to explore more potential connections between circRNA and disease. For diseases, semantic similarity features of diseases and GIP similarity features of diseases have been constructed. In order to more fully express the Mesh database. If there is a semantic similarity association between diseases d(i) and d(j), the combined disease feature DSim(d(i), d(j)) is the semantic similarity between the two diseases. Otherwise, it is the Gaussian interaction profile similarity of the disease. For circRNA, functional similarity and GIP similarity of circRNA have been constructed. Similarly, the GIP similarity of circRNA is used as the total feature CSim(c(i), c(j)) of circRNA in the circRNA feature fusion stage. If there is a functional similarity feature between circRNAc(i) and circRNAc(j), GIP similarity is replaced by functional similarity.

[0020] Since autoencoders can be used to extract unique features and identify potential biological patterns, this paper proposes a deep autoencoder to generate a unified vectorized representation for circRNA and disease after feature fusion. The encoding layer and decoding layer of the autoencoder have the same number of neurons, and the reconstruction of data features is achieved by changing the size of the hidden layer. Its encoding operation can be expressed as After obtaining the compressed representation feature Ds, the decoding operation of the autoencoder is constructed using a similar method. This process is iterated continuously until the decoded DSim′≈DSim. Then the encoder data compression result Ds∈R M×k As a new disease similarity feature matrix, k represents the length of the feature vector (set to 128 in this paper). Similarly, the compressed feature matrix Cs∈R of circRNA is obtained. N×k .

[0021] The specific steps of this stage are as follows:

[0022] Step 1: Use feature fusion to generate fused circRNA similarity information CSim and disease similarity information DSim.

[0023] Step 2: Use deep autoencoders to reconstruct circRNA and disease features to obtain a new circRNA feature matrix Cs and a new disease feature matrix Ds.

[0024] The third stage: The multi-layer collaborative attention representation learning of the present invention extracts deep features and key information from the fused circRNA and disease feature matrix. For the reconstructed circRNA and disease feature matrix, a large central network of circRNA and disease is established, and a multi-layer collaborative attention message propagation and message aggregation mechanism is established based on the network, iterating layer by layer to mine key signals on the central network.

[0025] After reconstruction by the autoencoder, the initial feature matrices Cs and Ds related to circRNA and disease are generated. Suppose vector e c ∈R k , e d ∈R k And e c is the characteristic of the circRNA represented by a row in Cs, e d is the disease feature represented by a row in Ds. Therefore, a matrix can be constructed as an embedded query table And it is used as the initial feature of circRNA and disease. Where N represents the number of circRNAs and M represents the number of diseases.

[0026] Multi-layer collaborative attention representation learning mainly consists of two modules: message construction and message aggregation. In the message construction stage, a message propagation mechanism from c to d is established based on a pair of connected circRNA-disease relationships (c, d) in the circRNA-disease relationship graph, that is, propagation on the edge (c, d). In order to explore high-order connectivity information, more message propagation layers are stacked and this high-order connectivity signal is used to score the correlation between circRNA and disease:

[0027]

[0028] in, is the trainable transformation matrix, k l is the size after transformation. Denote the features of circRNA after message propagation in the previous (l - 1) steps. The obtained aggregated information of the previous (l - 1) layers of circRNA can further represent the features of circRNA in the l-th layer. Since the messages propagated along the circRNA-to-disease path are obtained in the message construction phase. In the message aggregation phase, the message aggregation mechanism aggregates the messages passed from the neighbor nodes of disease d in the relationship graph and redefines the features of disease d:

[0029]

[0030] Since in the process of message aggregation, the message weights of nodes in the same layer are the same, which is controlled uniformly by p dc This brings limitations to fully learning the contribution values of different nodes in the same layer. Therefore, the present invention is inspired by the GAT model to learn the weights of different nodes within the same layer. However, GAT still has some limitations. For example, different attention heads are independent of each other, which fails to consider the dependencies between different attention heads. In response to this, the present invention proposes a new solution: multi-layer collaborative attention heads, which distribute different attention heads in different message layers to establish connections between attention heads. First, calculate the attention scores between circRNA and disease on the basis of the original model and perform a normalization operation:

[0031]

[0032] where denotes the attention score between the circRNA propagated in the l-th layer and the disease, and is denoted as f represents a single-layer feed-forward neural network, and W is the weight matrix of this network. The sub-network composed of itself and neighbor circRNAs (diseases) is the central network. Therefore denotes the circRNA neighbor nodes associated with disease d on the central network in the circRNA-disease relationship graph propagated in the l-th layer. denotes the magnitude of the contribution value of circRNA c to disease d during the message propagation process. In order to generate the overall structure of the disease (circRNA) established by the central network, the features of disease d are updated here using the linear combination of the weights of its central network, and the specific process is as follows:

[0033]

[0034] After obtaining the features weighted by multi-layer collaborative attention, the overall propagation process of message propagation on the circRNA-disease network is realized by proposing a matrix operation form of hierarchical propagation. The above process can be expressed as:

[0035]

[0036] in It is the characteristic representation of circRNA and disease after l transmissions. The initial state of message transmission E (0) The initial value of is E, where and I is the identity matrix, is the Laplace matrix of the circRNA-disease relationship graph. By implementing the propagation rules in matrix form, the characteristics of circRNA and disease can be updated more effectively at the same time. This saves the sampling step of the nodes, making the generalization ability of each propagation and update stronger.

[0037] The specific steps of this stage are as follows:

[0038] Step 1: Generate an embedded query table of circRNA and disease using circRNA features Cs and disease features Ds

[0039] Step 2: Using the embedded query table, a multi-layer collaborative attention graph message propagation mechanism is implemented. After multi-layer message construction and message aggregation, the final circRNA and disease feature representation graph E is generated. (l) .

[0040] Phase 4: The present invention uses a collaborative filtering-based model for model training and relationship prediction. This method integrates matrix factorization (MF) and multi-layer perceptron (MLP), combines the linear characteristics of MF with the nonlinear characteristics of MLP, and models the potential structure of circRNA and disease. The circRNA-disease relationship in the circRNA-disease relationship network is predicted using the circRNA and disease feature representation graph after a multi-layer collaborative attention representation learning network.

[0041] Generalized matrix factorization is a popular representation learning method and is widely used in many literatures. Therefore, by reproducing this model, it can be used to simulate most decomposition models. Usually, the input of this model is a one-hot encoded embedding representation, so only one fully connected mapping layer is needed to represent the dense vector of circRNA or disease. However, the embedding representation of circRNA and disease has been modeled with high-order connectivity expression in the relationship graph between circRNA and disease, and the collaborative signal is effectively injected into the embedding in an explicit way, so the fully connected mapping layer is omitted here. The first mapping layer of the generalized matrix factorization obtained is defined as follows:

[0042]

[0043] To take into account the non - linear relationship between circRNA and diseases, while performing generalized matrix factorization, the present invention simultaneously introduces a multi - layer perceptron to improve the non - linear modeling ability of the model. Here, a standard multi - layer perceptron model is used to understand the interaction between the potential features of circRNA and diseases. Its definition in neural collaborative filtering is as follows:

[0044]

[0045] Among them, \(W\) i , \(a\) i , \(b\) i (\(i\in1,2,\cdots,L\)) respectively represent the weight matrix, activation function, and bias term of the \(i\) - th layer perceptron.

[0046] To enable the prediction model to have both linear learning ability and non - linear learning ability, let matrix factorization and the multi - layer perceptron first learn separate hidden layers respectively, and connect the embeddings they finally learn and merge them to generate the prediction result. Its process is defined as follows:

[0047]

[0048] Among them, \(E\) gmf and \(E\) mlp respectively represent and the results after matrix factorization and the multi - layer perceptron. \(h\) represents the connection weight of the processing results of matrix factorization and the multi - layer perceptron. Different from ordinary collaborative filtering, in this paper, the sum of vector elements is used and denoted as sum, rather than using an activation function for mapping.

[0049] The specific steps of this stage are as follows:

[0050] The first step: Use \(E\) (l) to input into the generalized matrix factorization model for learning and training, and generate the feature \(E\) gmf after being processed by the generalized matrix factorization. At the same time, use \(E\) (l) to input into the multi - layer perceptron model for learning and training, and generate the feature \(E\) mlp after being processed by the multi - layer perceptron model.

[0051] The second step: Connect the processing results \(E\) gmf and \(E\) mlp of the two different sub - models in collaborative filtering, and merge them through a separate hidden layer to generate the prediction result \(Y'\) based on the entire graph.

[0052] The beneficial effects of the present invention:

[0053] (1) Most existing methods extract and construct features from the perspectives of circRNA and diseases respectively, and use a large number of descriptive features (such as sequence and semantic relationships) to construct an embedding function. Based on the existing feature descriptors, this method constructs a propagation mechanism on the central network to deeply mine the interaction between circRNA and diseases, making full use of the circRNA-disease relationship network.

[0054] (2) In the process of feature extraction and construction, the existing matrix completion and collaborative filtering methods based on graph relationships lack the explicit encoding of the key cooperation signals between key circRNA and diseases. In this method, by globally capturing the key cooperation signals hidden in the circRNA-disease relationship network, the defect of insufficient embedding construction in the feature extraction stage is made up for by using the interaction information. Brief Description of the Drawings

[0055] Figure 1 is the algorithm method framework diagram of the present invention;

[0056] Figure 2 is the framework diagram for extracting various similarity information of the present invention;

[0057] Figure 3 is the detailed flowchart of the three-layer embedding propagation and aggregation of the present invention;

[0058] Figure 4 is the implementation flowchart of the l-th attention head in the multi-layer collaborative attention mechanism of the present invention;

[0059] Figure 5 is the framework diagram of the collaborative filtering decision model of the present invention;

[0060] Figure 6(a) is the ROC curve of the experimental results of the model for five-fold cross-validation on the circR2Disease database;

[0061] Figure 6(b) is the recall curve of the experimental results of the model for five-fold cross-validation on the circR2Disease database.

[0062] Figure 7(a) is the ROC curve obtained by five-fold cross-validation of the model on the Circ2Disease database.

[0063] Figure 7(b) is the ROC curve obtained by five-fold cross-validation of the model on the CircRNADisease database. Detailed Embodiment

[0064] The present invention will be described in detail below with reference to the drawings and embodiments.

[0065] As Figures 1 to 5As shown, the present invention realizes a collaborative filtering model based on a multi-layer collaborative attention graph, and its architecture is as shown in Figure 1 As shown. First, the model calculates various similarity information between diseases and circRNAs, fuses the similarity information of circRNAs and diseases respectively, and generates initial embeddings of circRNAs and diseases through a deep autoencoder. Secondly, the model uses a neural graph propagation model based on multi-layer collaborative attention to construct and aggregate embeddings, effectively injecting collaborative signals between different layers into the embeddings. Its process can be divided into message propagation and message aggregation. During the message propagation process, by designing a multi-layer collaborative attention mechanism, it is possible to effectively obtain the contribution values of different nodes in the central network to the central node to optimize graph neural collaborative filtering. Finally, relationship prediction is performed by using the embeddings of circRNAs and diseases and a prediction model based on collaborative filtering to fit the initial circRNA-disease adjacency matrix.

[0066] This method extracts the functional similarity information of circRNAs, Gaussian interaction profile similarity information, semantic similarity information of diseases, and Gaussian interaction profile similarity information of diseases from the initial data respectively. Figure 2 The process of constructing similarity information is plotted.

[0067] Example 1

[0068] The performance of this method is evaluated using 5-fold cross-validation, and the final results are generated in an average manner. The ROC curve obtained in each fold of the experiment is shown in Fig. 6(a), and the precision-recall curve is shown in Fig. 6(b). The other indicators are shown in Table 1.

[0069] Table 1: Results of each experimental parameter of the model for five-fold cross-validation on the circR2Disease database

[0070]

[0071] As can be seen from Table 1, the present invention has achieved good results after five-fold cross-validation. Among them, the key indicator AUC is as high as 98.54%, and AUPR also reaches 72.49%. However, it can also be seen that under the five-fold division of the full sample set, the F1 value and AUPR still fluctuate. This may be due to the insufficient existing data, but the model pays more attention to the central graph established on the existing data. Therefore, the robustness of the model in this paper still needs to be further improved.

[0072] In addition, this method was also applied to the circ2Disease and circRNADisease databases and obtained experimental results of five-fold cross-validation. The ROC curves are shown in Figures 7(a) and 7(b) respectively. The detailed information of the AUROC and AUPR metrics is shown in Table 2. From the experimental results on the circ2Disease and circRNADisease databases, it can be seen that this model has the same good performance as that on the circR2Disease database in terms of the metrics of the two databases, which also confirms the general value of this model on multiple databases.

[0073] Table 2. Experimental results of the method in this paper for five-fold cross-validation on the Circ2Disease and CircRNADisease databases

[0074]

[0075] Example 2

[0076] To more intuitively display the performance of the model, based on the circR2Disease database, the present invention is compared with existing representative methods using AUC as the performance metric. These methods include: AE-RF, NCPCDA, iCircDA-MF, GCNCDA, PWCDA and Wang's method. In addition, the same dataset and experimental settings were used for the existing methods in the experiment. However, since the evaluation metrics used by different methods are not exactly the same, for the sake of reasonable comparison, only the common metrics they used, that is, the results obtained on the AUC metric, are given here. It should be noted that although these comparison methods are all based on the circRNA-disease relationship dataset in the circR2Disease database, the specific datasets used are not exactly the same. For example, the three methods of AE-RF, NCPCDA, and iCircDA-MF only used human data, PWCDA used data of humans and mice, while Wang's Method and GCNCDA not only used data of humans and mice, but also used data of other species. Although there are certain differences in the settings of various methods, the results in Table 3 can still reflect the superiority of this method to a certain extent.

[0077] Table 3: Comparison of the results of the method in this paper with the results of the models in relevant literatures

[0078]

[0079] Example 3

[0080] To confirm the predictive performance of the present invention, the unknown relationships between circRNAs and diseases were complemented herein and sorted according to the degree of association. Using the existing medical library index PubMed database, the complemented unknown circRNA-disease relationships were verified herein. Based on the complementation results of circRNA-disease by the model, the top ten circRNAs corresponding to two diseases, breast cancer (BC) and hepatocellular carcinoma (HCC), were found by way of example and the circRNAs were sorted in descending order according to the association scores. The results are shown in Tables 4 and 5. Further, we verified the accuracy of the prediction results based on the relationships between the two diseases and different circRNAs retrieved on PubMed. In PubMed, more than 32 million biomedical literature citations from MEDLINE, life science journals, and online books are integrated, which can provide necessary proof support for the prediction results of this model. For example, 26 related circRNAs for early breast cancer identified in cancer diagnosis, such as hsa_circ_0001946(39) (the first row in Table 3). Another example is that the identification of the epithelial-mesenchymal transition-related circRNA-miRNA-mRNA ceRNA regulatory network in breast cancer found that the high expression of the central gene acted on by hsa_circRNA_400031(40) was significantly associated with poor prognosis of breast cancer patients.

[0081] Meanwhile, this paper shows the pairs of relationships between the top ten circRNAs ranked by the association scores of the model prediction results and diseases, and the circRNAs are sorted in descending order according to the association scores. The results are shown in Table 6.

[0082] It should be noted here that although the relationships of some of the predicted circRNA-disease pairs are not supported by the literature, the possibility of their association cannot be denied, which still requires further biological experiments for verification.

[0083] Table 4. Top ten circRNAs related to breast cancer in the prediction results

[0084]

[0085] Table 5. Top ten circRNAs related to hepatocellular carcinoma in the prediction results

[0086]

[0087] Table 6. Top ten pairs of circRNAs and diseases in the prediction results

[0088]

[0089]

Claims

1. A relationship prediction method based on multi - layer collaborative attention graph collaborative filtering, characterized by the following steps: The first step: Use a semantic similarity calculation model to calculate the semantic similarity information SD between diseases in the disease relationship graph as the first initial feature of the diseases. The second step: Use the Gaussian interaction profile kernel function to calculate the Gaussian interaction profile similarity information DGS between diseases in the circRNA - disease relationship graph as the second initial feature of the diseases. The third step: Use the Gaussian interaction profile kernel function to calculate the Gaussian interaction profile similarity information CGS between circRNAs in the circRNA - disease relationship graph as the first initial feature of the circRNAs. The fourth step: Based on the disease semantic similarity obtained in the first step, use a function similarity calculation model to calculate the function similarity information FC between circRNAs in the circRNA - disease relationship graph as the second initial feature of the circRNAs. The fifth step: Use a feature fusion algorithm to merge the two features of circRNAs and form a new circRNA feature CSim, and merge the two features of diseases to form a new disease feature DSim. The sixth step: Use a deep auto - encoder DAE to reconstruct the circRNA feature and the disease feature to form a new circRNA feature Cs and a new disease feature Ds. The seventh step: Use Cs and Ds to complete the initial embedding construction and generate an embedding query matrix E of circRNAs and diseases. Step 8: Based on the generated embedding matrix E, use the multi-layer collaborative attention graph message propagation mechanism to obtain the aggregated information E of circRNA and disease features on the entire relationship network (l) , and use it as the final feature representation of circRNA and disease; The ninth step: Use the training data in the circRNA - disease relationship graph to train the collaborative filtering prediction model and calculate the prediction result Y'. The multi-layer collaborative attention graph message propagation mechanism in the eighth step uses a message propagation mechanism that includes a message construction phase and a message aggregation phase, and establishes a multi-layer collaborative attention mechanism based on the circRNA central network and the disease central network. Therefore, the feature learning process on the disease central network can be expressed as: where is the magnitude of the contribution of each circRNA to the central disease on the central network of disease d during the l-th propagation process and is expressed as: Feature learning is performed on the entire circRNA-disease network, and its matrix operation form of hierarchical propagation is expressed as Among them is the embedded feature representation after the circRNA and the disease have been propagated l times; the initial state E of message propagation (0) has an initial value of E, is the Laplacian matrix of the circRNA-disease relationship graph.

2. The relationship prediction method based on multi-layer collaborative attention graph collaborative filtering according to claim 1, characterized in that: The deep auto - encoder architecture used in the circRNA and disease feature reconstruction stage in the sixth step includes 1 encoding layer and 1 decoding layer; for the DAE of circRNAs, its encoding layer and decoding layer each consist of 3 fully - connected layers; each fully - connected layer reduces the 585 - dimensional input feature to 350 dimensions, 150 dimensions, and 64 dimensions respectively; for the DAE of diseases, its encoding layer and decoding layer each consist of 1 fully - connected layer; each fully - connected layer reduces the 88 - dimensional input feature to 64 dimensions and unifies the activation function of each layer as ReLU.

3. The relationship prediction method based on multi-layer collaborative attention graph collaborative filtering according to claim 1 or 2, characterized in that: In the initial embedding construction in the seventh step, E is the initial embedding matrix and is expressed as where the vector e c ∈R k , e d ∈R k , and e c is the feature of a certain circRNA in Cs, and e d is the feature of a certain disease in Ds.

4. The relationship prediction method based on multi-layer collaborative attention graph collaborative filtering according to claim 1 or 2, characterized in that: In the collaborative filtering model of the ninth step, it includes 1 generalized matrix factorization model and 1 multi-layer perceptron model; the generalized matrix factorization model contains 1 operation of multiplying the circRNA feature matrix by the disease feature matrix and 1 normalization layer, taking the circRNA feature and the disease feature as inputs and finally obtaining a 256-dimensional output result; the multi-layer perceptron model contains 3 convolutional layers and 2 pooling layers. The first layer concatenates the input training or test circRNA features with the corresponding disease features to form an output with a dimension of 512, and generates a 256-dimensional output result after passing through the multi-layer perceptron; the weighted sum of the processing results of the generalized matrix factorization and the multi-layer perceptron model is used to obtain the prediction result Y′, where the weight of the generalized matrix factorization result E gmf is 0.9, and the weight of the multi-layer perceptron result E mlp is 0.

1.

5. The relationship prediction method based on multi-layer collaborative attention graph collaborative filtering according to claim 3, characterized in that: In the collaborative filtering model of the ninth step, there is 1 generalized matrix factorization model and 1 multi-layer perceptron model; the generalized matrix factorization model includes 1 operation of multiplying the circRNA feature matrix by the disease feature matrix and 1 normalization layer, taking the circRNA feature and the disease feature as inputs and finally obtaining a 256-dimensional output result; the multi-layer perceptron model includes 3 convolutional layers and 2 pooling layers. The first layer concatenates the input training or test circRNA features with the corresponding disease features to form an output of dimension 512, and generates a 256-dimensional output result after passing through the multi-layer perceptron; the results of the generalized matrix factorization and the multi-layer perceptron model are weighted and summed to obtain the prediction result Y′, where the weight of the generalized matrix factorization result E gmf is 0.9, and the weight of the multi-layer perceptron result E mlp is 0.1.