A miRNA target prediction method, system and application based on a multi-layer heterogeneous graph

By constructing multi-layer heterogeneous maps and using feature aggregation and multi-head self-attention mechanisms, the real-time and accuracy problems of existing miRNA target prediction methods are solved, and efficient and accurate prediction of miRNA targets is achieved.

CN116631496BActive Publication Date: 2025-07-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310496854.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-05
Publication Date
2025-07-18
Estimated Expiration
2043-05-05

AI Technical Summary

Technical Problem

The existing miRNA target prediction methods lack real-time prediction capabilities for different targets, with large prediction errors and insufficient data volume, resulting in low prediction accuracy.

Method used

A multi-layer heterogeneous pattern consisting of seven RNA networks was constructed, and the multi-head self-attention mechanism and projection method were used to calculate and predict the target of miRNA through the average pooling layer polymerization characteristics.

Benefits of technology

It improves the accuracy and real-time prediction of miRNA targets, can better reflect the complex interaction relationship between miRNA and target, and improves the accuracy of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116631496B_ABST
    Figure CN116631496B_ABST
Patent Text Reader

Abstract

The present invention relates to a miRNA target prediction method, system and application based on a multi-layer heterogeneous graph. First, in terms of node representation, the present solution decouples the node representation into edge embedding and basic embedding, and each layer separately maintains the edge embedding of all nodes on that layer. Second, in terms of graph propagation, since shallow GCNs cannot propagate features over a large range and deep GCNs are prone to over-smoothing, we choose sampling average aggregation to solve this problem, and extract fixed k node embeddings from the node neighborhood for averaging to represent the central node. Third, in terms of the attention mechanism, the present solution makes a slight innovation on the basis of predecessors. For the multi-head attention mechanism, instead of simply concatenating vectors, a pooling layer and a fully connected layer are used, and the overall implementation is more logical and the parameter adjustment is simpler during experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to computer application technology, and relates to a method for using a computer to identify biological information, and specifically relates to a miRNA target prediction method, system and application based on a multi-layer heterogeneous graph. Background Art

[0002] Non-coding RNA (ncRNA) refers to RNA molecules that are not translated into proteins in cells. They play a variety of biological functions in cells, such as gene expression regulation, post-transcriptional modification, histone modification, RNA splicing, RNA degradation, etc. The two major categories of ncRNA include small RNA (siRNA, miRNA, etc.) and non-small RNA (lncRNA, circRNA, etc.). The dysregulation of these two types of RNA is closely related to diseases including cancer, plays an important role in cell regulation, has high clinical and scientific relevance, and will have an important impact on future medicine and disease treatment.

[0003] Among them, miRNA (microRNA) is one of the most widely studied non-coding RNAs. miRNA is a small molecule RNA with a length of about 20 to 25 nucleotides. It is widely present in the cells of eukaryotic organisms and mainly participates in the regulation of gene expression at the post-transcriptional level by complementary binding to target genes. miRNA is synthesized by the synergistic action of a series of enzymes and protein complexes in the cell. Its biosynthesis process includes steps such as miRNA gene transcription, pri-miRNA shearing, pre-miRNA release and mature miRNA binding. The mechanism of action of miRNA in cells is very complex. They can regulate the expression of target genes in two ways: one is to bind to the 3' untranslated region (UTR) of mRNA and inhibit the translation process of target gene mRNA, and the other is to bind to the coding region (CDS) of mRNA and induce the degradation of target gene mRNA. miRNA plays a very important role in the regulation of various biological processes in organisms, including cell proliferation, differentiation, apoptosis, cell cycle, etc. Therefore, miRNA has become one of the hot research areas in cell molecular biology, disease pathogenesis and new drug development.

[0004] Both ncRNAs and mRNAs can serve as miRNA targets. First, they are usually located in cells and are involved in important biological processes such as signal transduction, cell activity, and regulation of cell state. In addition, they have highly conserved structural features, which makes it easier for miRNAs to recognize and bind these ncRNAs and mRNAs as their targets. Finally, the expression levels of these ncRNAs and mRNAs are regulated by miRNAs, and this causal relationship can be used to provide further information. This regulatory relationship together creates an auxiliary regulatory network around miRNAs.

[0005] The miRNA regulatory network is a complex, self-regulating biological signal transduction system that can be used to regulate and coordinate gene expression in cells. It regulates protein synthesis and cell activities by sending a series of messages. The miRNA regulatory network includes various complex interactions between miRNAs and their targets, which can control gene expression levels, promote transcriptional level changes, lead to cell phenotype regulation, and abnormal signal transduction pathways. In this complex regulatory system, miRNAs can affect the synthesis or inhibition of host genes, as well as the expression of other miRNAs and mRNAs.

[0006] Current miRNA target prediction methods are mainly data-mining methods based on machine learning, statistics, and bioinformatics technologies, used to analyze the interactions between miRNAs and mRNAs, as well as other non-coding RNAs. This method uses machine learning techniques to extract characteristic features, such as deep learning, clustering analysis, support vector machines, etc., based on various information about miRNA and target sequences and expression levels, to identify features, and predicts miRNA targets through multivariate statistical analysis. On this basis, different AI models and data mining techniques can be used to develop more powerful miRNA target prediction models to predict the complex structure of miRNA-mRNA interactions. Such models can help biologists better understand miRNA-mRNA interactions and contribute to the study of the role of miRNAs in cell signal transduction and epigenetic regulation.

[0007] According to the data type, miRNA-lncRNA target recognition algorithms can be divided into three categories, namely sequence-based, expression data-based, and graph-based. In 2018, Zhang et al. mainly proposed a sequence-derived linear domain propagation method (SLNPM) based on sequence features. They used the linear domain similarity method to calculate the similarity between lncRNAs and miRNAs, and constructed lncRNA similarity networks and miRNA similarity networks respectively. They implemented the label propagation process on the networks to score lncRNA-miRNA pairs. In 2018, Huang et al. based on many existing evidences showed that the interaction between lncRNA and miRNA is closely related to their relative expression levels. In addition to the expression profiles, they further used the lncRNA-function, miRNA-mRNA, and the sequence data of miRNAs and lncRNAs, and used PCC and Needleman-Wunsch pairwise sequence alignment to calculate the similarity matrices of miRNAs and lncRNAs respectively, and proposed a simple bipartite graph-based model EPLMI. In the same year, Huang et al. also proposed the GBCF model by integrating the Bayesian collaborative filtering algorithm based on the same assumptions and data. Zhang mainly proposed the linear neighborhood propagation algorithm SLNPM based on sequence data. However, these algorithms actually did not use graph neural networks. Huang proposed an end-to-end prediction model GCLMI based on graph convolution and autoencoders in 2019. There is no need for data preprocessing anymore, and the influence experiment of negative sampling was carried out. Zhang experimented with an integrated model based on five graph representation learning algorithms in the same year and also achieved good results. In 2019, You et al. integrated multiple RNA-related information sources to construct a heterogeneous network model LMNLMI. First, heterogeneous network fusion was performed on lncRNAs and miRNAs respectively to obtain a new similarity network. Then, LMNLMI found the best projection from the lncRNA feature space to the miRNA space, so that the distance between the projected feature vectors of lncRNAs and the feature vectors of known interacting miRNAs was close. After that, LMNLMI inferred new interactions based on the geometric proximity of the lncRNA to the projected feature vector in the projection space, and sorted its candidate targets. Finally, LMNLMI was also compared with the collaborative filtering algorithm commonly used in recommendation systems. In 2020, Fan et al. constructed a heterogeneous graph model SNFHGILMI based on sequence and link data. Assuming that miRNAs and lncRNAs follow a Gaussian distribution, they calculated the high-order features using KL divergence, and then non-linearly fused them with the similarity network calculated by the sequence. Finally, they used the heterogeneous graph inference algorithm for prediction. H. Liu proposed the model LMFNRLMI based on the logical matrix factorization algorithm, and used neighborhood regularization to optimize the matrix factorization algorithm.

[0008] The miRNA-mRNA can also be classified according to similar criteria. In 2020, Jiang proposed a prediction algorithm miRTMC based on heterogeneous networks using the matrix completion algorithm. Calculate the similarity matrix of miRNA based on the seed region through the Needleman-Wunsh global alignment algorithm, and at the same time calculate the similarity matrix of mRNA based on complementarity with the 3'-UTR through the Smith-Waterman local alignment algorithm. Use the link data verified by biological experiments to fuse the two matrices, and transform the miRNA-target prediction problem into a low-dimensional matrix completion problem. Wang proposed the model miRACLe based on the assumption that different types of RNA in the sample are free-moving particles that randomly contact and bind with different efficiencies. Wang mainly based on the sequence similarity of the miRNA seed interval, and used the existing recommendation algorithm to construct the model miRTRS for prediction. M. Mokhtaridoost et al. searched for miRNA-mRNA regulatory modules based on the data of miRNA and mRNA expression profiles through a linear multiple regression model and low-rank matrix factorization. Fu et al. generated a miRNA-mRNA interaction network during the egg maturation process of adult female mosquitoes through the CLEAR-CLIP experiment.

[0009] The current methods mainly have the following defects:

[0010] 1. Lack of prediction ability for different targets: The current miRNA target prediction methods can only statically analyze the existing data sets and cannot perform real-time prediction according to new data. Therefore, they can often only perform single prediction for lncRNA or mRNA.

[0011] 2. Large prediction error: There are many implicit factors in the miRNA target prediction methods, and these implicit factors cannot fully reflect the correlation between miRNA and mRNA, resulting in low prediction accuracy.

[0012] 3. Insufficient data volume: The data volume relied on by the miRNA target prediction methods is insufficient, and most data sets only contain a small number of miRNA-target pairs, thus limiting the accuracy of the system prediction effect. Summary of the Invention

[0013] Technical Problems to be Solved

[0014] In order to avoid the deficiencies of the prior art, the present invention proposes a miRNA target prediction method, system and application based on a multi-layer heterogeneous graph.

[0015] Technical Solutions

[0016] A miRNA target prediction method based on a multi-layer heterogeneous graph, characterized by the following steps:

[0017] Step 1: Construct a heterogeneous graph composed of seven RNA networks, where the nodes represent one of the three RNAs: miRNA, lncRNA, and mRNA, and the seven RNA networks reflect seven different edge types;

[0018] The seven edge types are:

[0019] ① The miRNA-lncRNA interaction layer represents the known and verified LMI,

[0020] ② The miRNA-miRNA sequence similarity layer measures the sequence similarity between miRNAs,

[0021] ③ The miRNA-miRNA co-expression layer measures the co-expression relationship between miRNAs,

[0022] ④ The miRNA-mRNA interaction layer represents the known mRNAs targeted by miRNAs,

[0023] ⑤ The lncRNA-lncRNA sequence similarity layer measures the sequence similarity between lncRNAs,

[0024] ⑥ The lncRNA-lncRNA co-expression layer measures the expression similarity between lncRNAs,

[0025] ⑦ The lncRNA-mRNA interaction layer represents the known mRNAs targeted by lncRNAs;

[0026] Step 2: Adopt a method based on a specific layer to aggregate features from different layers, and use an average pooling layer for aggregation to obtain the k-order features of node i in the network layer That is, the edge embedding of each node in the multi-layer heterogeneous graph in its corresponding layer is obtained, indicating that the k-order feature of node i depends on the average value of the k-1 order features of node i and its neighbors:

[0027]

[0028] where σ(·) represents the sigmoid function, W (k) is a weight matrix that needs to be learned during the training process, mean(·) represents the averaging operation, r represents the layer number, that is, the r-th network layer, N i,r is a set of nodes that includes node i and its neighbors, represents N i,r(k-1)-order features of node j in the set, where 1 ≤ k ≤ K, and K represents the maximum feature aggregation level of each network layer;

[0029] Step 3: Denote the edge embeddings in all layers of node i as matrix U i =(u i,1 ,…,u i,l ), where U i ∈R s ×l , that is, U i is an s×l matrix, s represents the edge embedding dimension of the node, and l represents the total number of layers; Use the multi-head self-attention mechanism to encode the edge embeddings of multiple layers of node v i to obtain H i,r [k], which is:

[0030]

[0031] where: k represents the number of the attention head (k ∈ [1, m], m represents the total number of attention heads), and H i,r [k] represents the k-th head representation of node v i in the r-th layer; The calculation formula of A i,r is as follows:

[0032]

[0033] where r represents the layer number, i represents the node number, softmax(·) represents the softmax function, and are learnable matrices, where m represents the total number of attention heads, s represents the edge embedding dimension of the node, and d a represents the intermediate dimension during the change process;

[0034] Step 4: Use the projection method to project the edge embeddings into the task space, then extract the features from each layer and finally integrate them together;

[0035] Specifically:

[0036] Map the representations of multiple attention heads of a single node from R s to the final task space R d through the following formula:

[0037] P i,r [k] = H i,r [k]W p

[0038] where W p ∈R s×dis the matrix parameter to be learned through training. s represents the edge embedding dimension before the node passes through the projector, d represents the edge embedding dimension after the node passes through the projector, k represents the number of attention heads, and P i,r [k] represents the k-th head representation of the node after projection;

[0039] Select the bi-pooling operation to fuse the representations of the k attention heads of the node, and obtain the final edge embedding e of node i at layer r i,r :

[0040]

[0041] where: m represents the total number of attention heads, both j and k represent the numbers of attention heads, and p i,r [k] represents the k-th head representation of node i at the r-th layer, and p i,r [j] represents the j-th head representation of node i at the r-th layer, represents the element-wise product of two vectors, and W r,pool is the matrix parameter to be learned through training;

[0042] The basic embedding of the node vi is shared across all layers. As a message-passing medium, it fuses the edge embeddings from each layer and is passed between layers;

[0043] Step 5: By randomly generating a set of values from a Gaussian distribution, the basic embedding of each node can be randomly initialized. The following formula is used to fuse the basic embedding and the edge embedding e i,r to obtain the fused embedding of order t

[0044]

[0045] represents the basic embedding of order t - 1, represents the edge embedding of order t;

[0046] The information mixing at the adjacent neighborhood aggregation level is achieved through the mixture of the fused embedding from the previous round and the edge embedding to smooth the output between different aggregation layers. By stacking more neighborhood aggregation layers, longer-distance information can be captured, and the final representation o of any node i is obtained i ;

[0047] Step 6: Use the following cosine distance formula to calculate the distance between the two target nodes i and j in the prediction space:

[0048]

[0049] where o i, o j respectively represent the final representations of nodes vi and vj;

[0050] The node represents a miRNA, or a target of miRNA, namely mRNA or lncRNA. If the distance between a target node and a miRNA node is large, it indicates that the possibility of this mRNA / lncRNA being a target of this miRNA is relatively large.

[0051] The miRNA-lncRNA interaction layer represents the known and verified LMI as: extracting unique miRNA-lncRNA interactions with at least one CLIP sequence experimental evidence from the lncRNASNP2 database. The network layer consists of multiple unique miRNAs, multiple unique lncRNAs, and multiple unique small RNA-lncRNA edges.

[0052] The miRNA-miRNA sequence similarity layer measures the sequence similarity between miRNAs as: first retrieving miRNA sequences of multiple miRNAs from the miRbase database, and then performing global alignment on each pair of miRNA sequences using the Needleman-Wunsch algorithm implemented in the Biostring software package; the gap opening penalty is set to 0.5, and the gap opening extension penalty is set to 0.1; if the identity score of two miRNAs is greater than or equal to 40, they will be connected in this layer. The resulting network layer consists of multiple miRNAs and multiple miRNA-miRNA interactions.

[0053] The miRNA-miRNA co-expression layer measures the co-expression relationship between miRNAs as: retrieving miRNA expression profiles of multiple miRNAs from mammalian microRNA expression atlases, which are collected from major organs and cell types of multiple human subjects; the Pearson correlation coefficient PCC is used to measure the co-expression similarity between miRNAs, and two miRNAs with a PCC greater than or equal to 0.3 will be connected in this layer. The resulting network layer consists of multiple miRNA co-expression miRNA pairs among multiple miRNAs.

[0054] The miRNA-mRNA interaction layer represents the known mRNAs targeted by miRNAs as: downloading experimentally verified miRNA-mRNA interactions from miRTarBase. After removing weak miRNA-mRNA interactions, only one or more evidences from qRT-PCR, luciferase reporter assay, Western blot, microarray, immunohistochemistry, and in situ hybridization, etc. retain strong interactions.

[0055] The sequence similarity of lncRNA-lncRNA layer measures the sequence similarity between lncRNAs as follows: First, download the DNA sequences of multiple lncRNAs from the NONCODE database, and calculate the lncRNA-lncRNA sequence similarity based on sequence alignment; use the local alignment algorithm Smith-Waterman to perform the task; during the alignment process, the penalty for opening a gap is set to 10, and the incremental cost generated along the gap length is set to 4; lncRNA pairs with an alignment score greater than or equal to 400 will be retained as lncRNA-lncRNA edges in this layer; the resulting network layer consists of multiple lncRNAs and multiple lncRNA-lncRNA interactions.

[0056] The co-expression layer of lncRNA-lncRNA measures the expression similarity between lncRNAs as follows: Download the expression profiles of multiple lncRNAs from the NONCODE database, select a higher PCC threshold of 0.9, and finally there are multiple lncRNA co-expression linkage relationships among the multiple lncRNAs remaining in this layer.

[0057] The lncRNA-mRNA interaction layer represents the known mRNAs targeted by miRNAs as follows: Download multiple experimentally verified edges from the RISE database and obtain the lncRNA-mRNA edge relationship from them.

[0058] A system for the miRNA target prediction method based on a multi-layer heterogeneous graph, characterized in that it includes an aggregator, an encoder, a projector, a fuser, and a predictor; the input end of the aggregator receives a heterogeneous graph composed of seven RNA networks, and the node features of each layer of the heterogeneous graph are updated by the aggregator; then, the node features of each layer are fused with the information of other layers through the encoder to obtain the node features updated again, where each node will have multiple head representations; then, the multiple head representations are fused through the projector, and the finally fused vector is used as the edge embedding of the node; the edge embedding and the base embedding are fused in the fuser to obtain the final node representation; finally, the cosine distance is calculated in the predictor for prediction.

[0059] A miRNA target prediction method based on the multi-layer heterogeneous graph and the system, characterized in that the method and the system are used for the prediction of miRNA targets.

[0060] Beneficial effects

[0061] A miRNA target prediction method, system and application based on a multi-layer heterogeneous graph. First, in terms of node representation, this solution decouples the node representation into edge embedding and basic embedding, and each layer separately maintains the edge embedding of all nodes on that layer. Second, in terms of graph propagation, since shallow GCNs cannot spread features over a large range and deep GCNs are prone to over-smoothing, we choose sampling average aggregation to solve this problem, and extract a fixed number k of node embeddings from the node neighborhood for averaging to represent the central node. Third, in terms of the attention mechanism, this solution makes a slight innovation on the basis of predecessors. For the multi-head attention mechanism, instead of simply concatenating vectors, a pooling layer and a fully connected layer are used, which is more logical in the overall implementation and easier to adjust parameters during experiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 : Schematic diagram of the system structure of the present invention

[0063] Figure 2 : Schematic diagram of the hierarchical heterogeneous graph structure constructed by the present invention

[0064] Figure 3 : Schematic diagram of the data processing process of the method of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] The present invention will be further described below in conjunction with embodiments and drawings:

[0066] In view of the deficiencies of current algorithms in data integration, we know that for sequence data, the interaction between miRNA and its target occurs not only in the 3'-UTR but also in the 5'-UTR, and many algorithms only consider the 3'-UTR; at the same time, most methods do not apply expression data perfectly, as miRNA can not only promote target expression but also inhibit it, and many algorithms only simply consider the positive promotion relationship; moreover, current methods based on sequence and expression data have certain defects in prediction accuracy, and the problem of serious imbalance between positive and negative samples in biological data has not been solved. Based on the above deficiencies, we propose a model framework for predicting miRNA targets (i.e., lncRNA-miRNA, mRNA-miRNA interactions) based on a deep graph expression technology of a multi-layer heterogeneous biological network, as shown in Figure 1 .

[0067] First, through existing data, a multi-layer heterogeneous graph can be obtained; the node features of each layer can be updated through an aggregator; then, the node features of each layer can be fused with the information of other layers through an encoder to obtain the node features after re-updating (since the encoder uses a multi-head attention mechanism, each node will have multiple head representations); then, through a projector, the representations of multiple heads are fused, and the finally fused vector is used as the edge embedding of the node; the edge embedding and the basic embedding are fused in a fuser to obtain the final node representation; finally, the cosine distance is calculated in a predictor for prediction.

[0068] Construction of the input end

[0069] Different from a homogeneous graph where all nodes and edges are of the same type, a heterogeneous graph is a graph composed of different types of nodes and edges. In a heterogeneous graph, each node and edge has its own type. For example, in a social network, the types of user nodes and friend nodes are different, and the types of follow relationships between user nodes and user nodes are also different. Heterogeneous graphs can represent various complex relationships, which enables them to better reflect the complex relationships in the real world.

[0070] To solve the problem that most methods are imperfect for expressing data applications, we made innovations at the data input end. We constructed a heterogeneous graph composed of seven RNA networks, where the nodes represent one of the three types of RNA: miRNA, lncRNA, and mRNA. Different layers of the heterogeneous network can fully consider different edge types. In this scheme, we can simultaneously utilize various types of data such as interaction relationships and co-expression relationships. Below, we will give a detailed introduction from the data source of the multi-layer heterogeneous graph and the construction method of the multi-layer heterogeneous graph:

[0071] 1. Data source

[0072] We constructed through fusing multiple types of data as Figure 1The hierarchical heterogeneous graph shown below will have its data sources introduced in detail in this section. The limited number of lncRNA-miRNA or mRNA-miRNA interactions recorded in the current database, i.e., experimentally verified samples, greatly affects the prediction ability in LMI and MMI tasks. Other experimentally verified RNA interactions related to lncRNA and miRNA will contribute to LMI and MMI prediction tasks based on the association hypothesis. To obtain better performance, we introduced six additional types of RNA associations that have been verified in wet laboratories or have a high level of computational confidence. Therefore, we constructed a heterogeneous network consisting of seven RNA networks, representing three types of RNA: miRNA, lncRNA, and mRNA. Different layers of the heterogeneous network can fully consider different edge types, specifically as Figure 2 :

[0073] ① The miRNA-lncRNA interaction layer represents known experimentally verified LMI. Experimentally verified miRNA-lncRNA interactions were extracted from the lncRNASNP2 database. After filtering duplicate records, we extracted unique miRNA-lncRNA interactions with at least one CLIP-seq experimental evidence. The final network layer consists of 276 unique miRNAs, 908 unique lncRNAs, and 7698 unique miRNA-lncRNA edges.

[0074] ② The miRNA-miRNA sequence similarity layer measures the sequence similarity between miRNAs. We first retrieved the miRNA sequences of 267 miRNAs from the miRbase database. Then, the global alignment of each pair of miRNA sequences was performed using the Needleman-Wunsch algorithm implemented in the Biostrings software package. The gap opening penalty was set to 0.5, and the gap extension penalty was set to 0.1. If the identity score of two miRNAs is greater than or equal to 40, they will be connected in this layer. The resulting network layer consists of 71 miRNAs and 58 miRNA-miRNA interactions.

[0075] ③ The miRNA-miRNA co-expression layer measures the co-expression relationship between miRNAs. We first retrieved the miRNA expression profiles of 90 miRNAs from the mammalian microRNA expression atlas, which were collected from the major organs and cell types of 172 human subjects. The Pearson correlation coefficient (PCC) was used to measure the co-expression similarity between miRNAs, and two miRNAs with a PCC greater than or equal to 0.3 will be connected in this layer. The resulting network layer consists of 268 miRNA co-expression miRNA pairs among 62 miRNAs.

[0076] ④ The miRNA-mRNA interaction layer represents the known mRNAs targeted by miRNAs. According to previous studies, miRNAs with similar targeted mRNAs are more likely to have similar targeted lncRNAs. We downloaded experimentally verified miRNA-mRNA interactions from miRTarBase. After removing weak miRNA-mRNA interactions, only one or more pieces of evidence from qRT-PCR, luciferase reporter assay, Western blot, microarray, immunohistochemistry, and in situ hybridization, etc., were retained for strong interactions. There were a total of 83,140 experimentally verified edges, among which 10,754 were verified by strong evidence, involving 13,157 mRNAs in terms of expression.

[0077]

[0078] ⑤ The lncRNA-lncRNA sequence similarity layer measures the sequence similarity between lncRNAs. Similar to the miRNA sequence similarity layer, we first downloaded the DNA sequences of 644 lncRNAs from the NONCODE database and calculated the lncRNA-lncRNA sequence similarity based on sequence alignment. The difference is that we used a local alignment algorithm, namely Smith-Waterman, to perform this task instead of using Needleman-Wunsch. This is because lncRNAs are much longer than miRNAs, and miRNAs use short seed sequences (6-8 bases long) to bind to miRNA response elements (MREs) on target RNAs. During the alignment process, the penalty for opening a gap was set to 10, and the incremental cost generated along the gap length was set to 4. LncRNA pairs with an alignment score greater than or equal to 400 would be retained as lncRNA-lncRNA edges in this layer. The resulting network layer consisted of 92 lncRNAs and 52 lncRNA-lncRNA interactions.

[0079] ⑥ The lncRNA-lncRNA co-expression layer measures the expression similarity between lncRNAs. Similar to the miRNA co-expression layer, we downloaded the expression profiles of 548 lncRNAs from the NONCODE database, which consists of multiple samples from 24 tissues. Since the lncRNA co-expression correlation in our experiment was much stronger than that of miRNAs, we selected a higher PCC threshold of 0.9, and finally there were 1,673 lncRNA co-expression linkage relationships among the 303 lncRNAs remaining in this layer.

[0080] ⑦ The lncRNA-mRNA interaction layer represents the known mRNAs targeted by miRNAs. We downloaded 138,684 experimentally verified edges from the RISE database and screened out the lncRNA-mRNA edge relationships from them, as shown in Table 1 specifically.

[0081] Data processing process

[0082] As Figure 1 shown, our network framework can be split into five main components, namely the aggregator, encoder, projector, fusor, and predictor. For the input of a multi-layer heterogeneous graph, we can obtain the correlation between any two nodes through this network. The following is a detailed introduction to each component:

[0083] (1) Aggregator:

[0084] Most existing graph representation learning methods are based on the message passing neural network (MPNN) framework, including ChebyNet, GCN, GAT, and GIN, etc. They learn the representation of each node by aggregating the feature information of neighbors. We also build our learning model on multi-layer heterogeneous graphs based on this mechanism. In a multi-layer heterogeneous graph, the feature distribution of nodes is in these hierarchical spaces, and each space reflects different functions. Nodes have different feature spaces in different networks. Therefore, we designed a method based on specific layers to aggregate features from different layers.

[0085] We can abstract the network as G = (V, E, R), where V represents the set of nodes, E represents the set of edges, and R represents the set of different network layers. Each edge e in E ij,r corresponds to a pair of vertices v i and v j in the r-th network layer. We design the node embedding as a combination of two parts: the base embedding and the edge embedding. The edge embedding of node v i is specific to each RNA network layer, while the base embedding of the node is shared among different network layers.

[0086] For the edge embedding process of each layer, we follow the general GNN idea, that is, the edge embedding of node v i is aggregated from the local neighborhood of the node, as shown in the equation:

[0087]

[0088] where r represents the r-th network layer, represents node v in this network layeri The k - th order feature, N i,r is a set that contains node v i and the neighbors of node v i . The set, denotes the (k - 1)-th order feature of the nodes in the N i,r set, where 1 ≤ k ≤ K, K represents the maximum feature aggregation level of each network layer, aggregator(·) represents different aggregator functions. In this experiment, we use the average pooling layer as the aggregator, then the above formula can be rewritten as follows:

[0089]

[0090] where σ(·) represents the sigmoid function, W (k) is a weight matrix that needs to be learned during the training process, mean(·) represents the averaging operation. At the beginning, we randomly initialize for each node in each layer. To maintain computational efficiency, we uniformly sample a fixed number of neighbors instead of using all the neighbors of node v i in each iteration. The advantage of doing this is to prevent the explosion of the receptive field and reduce the consumption of computing power.

[0091] (2) Encoder

[0092] After passing through the aggregator, we can obtain the edge embeddings of each node in all layers in the multi - layer heterogeneous graph. We denote all the edge embeddings of node v i as matrix U i =(u i,1 ,…,u i,l ), where U i ∈R s×l , that is, U i is an s×l matrix, s represents the dimension of the node's edge embedding, and l represents the total number of layers. Then, we use the multi - head self - attention mechanism to further encode the edge embeddings of multiple layers of node v i . This can enable the node to capture the hidden dependency features in other layers, integrate the topological features and node feature information from all layers, and at the same time reduce the impact caused by the density and sample imbalance of different layers. The following is its formula:

[0093]

[0094] where r represents the layer number, i represents the node number, softmax(·) represents the softmax function, and is a learnable matrix, where s represents the edge embedding dimension before the node passes through the encoder, and d a represents the intermediate dimension during the change process, and m represents that there are m attention heads in total. The reason for dividing by a is that when the dimensions of W 2,r U i and W 3,r U i,r are relatively high, the elements in the calculated attention matrix are either too large or too small. This results in an uneven distribution of the values obtained after the softmax non-linear mapping, that is, the variance is too large. The elements with larger values are very close to 1, and the smaller ones are very close to 0. The values obtained after mapping tend to the two boundaries 0 and 1. This will cause the gradient to be unstable, making it difficult for the model to converge. Therefore, we need to divide by a to mitigate the attention matrix obtained after the non-linear mapping.

[0095]

[0096] Then, through the above formula, we can obtain the final features of the node passing through the encoder. Here, k represents the number of the attention head (k ∈ [1, m]), and H i,r [k] represents the k-th head representation of the node v i in the r-th layer.

[0097] (3) Projector

[0098] After obtaining H i,r [k], we connect the projector to the model and use the projection method to map the features to the required task space. The separation of the task space and the feature space can make the features extracted by the model more accurate and robust. In this work, the projector can project the edge embedding into the task space and can well extract the features from each layer and finally integrate them together. In the previous step, we obtained the multi-head representation of the node after encoding. At this time, the node representation vector is actually in R m×s dimensional space (m represents the total number of attention heads, and s represents the dimension of the edge embedding). So next, we need to integrate the representations from multiple attention heads. We first map the multi-attention head representation of a single node from R s to the final task space R d using the following formula:

[0099] P i,r [k] = H i,r [k]W p

[0100] where W p ∈ R s×dis the matrix parameter to be learned through training. s represents the edge embedding dimension before the node passes through the projector, d represents the edge embedding dimension after the node passes through the projector, k represents the number of attention heads, and P i,r [k] represents the k-th head representation of the node after projection.

[0101] Next, we selected a pooling operation to fuse the representations of the node from k attention heads. Specifically, we used the Bi-pool (bilinear interaction) pooling, and the formula is as follows:

[0102]

[0103] where e i,r represents the final edge embedding of node vi at layer r, m represents the total number of attention heads, j and k both represent the numbers of attention heads, and p i,r [k] represents the k-th head representation of node v i at layer r, and p i,r [j] represents the j-th head representation of node v i at layer r, represents the element-wise product of two vectors, and W r,pool is the matrix parameter to be learned through training. Generally speaking, through the projection operation, the node can capture more features from different attention heads. Multiple attention heads can not only make the importance ranking strategy of the layer more diverse, but also prevent overfitting during the training process, and finally obtain the edge embedding of the node.

[0104] (4) Fusor

[0105] As mentioned in (1) above, we designed the node embedding as a combination of two parts: the base embedding and the edge embedding. After (1)-(3), we can obtain the edge embedding of the node at a certain layer. We regard the base embedding as a randomly initialized bias vector, and the base embedding of the node in each layer is the same, which makes the model more robust. Then, we use the fusor to fuse the edge embedding and the base embedding.

[0106]

[0107] represents the fusion embedding of order t, represents the base vector of order t-1, represents the edge embedding of order t. It can be seen that the base embedding of node vi is shared across all layers. As a message passing medium, it fuses the edge embeddings from each layer and is passed between layers. The fusion operation actually realizes the adjacent neighborhood aggregation level through the mixture of the fusion embedding of the previous round and the edge embedding The information mixing on it can smooth the outputs between different aggregation layers, enabling this paper to stack more neighborhood aggregation layers, capture longer-distance information, and obtain the final representation \(o\) of any node \(i\). i .

[0108] (5) Predictor

[0109] After obtaining the final representation of the nodes, we use the cosine distance formula shown below to calculate the distance between two nodes \(i\) and \(j\) in the prediction space:

[0110]

[0111] where \(o\) i and \(o\) j represent the final representations of nodes \(v_i\) and \(v_j\) respectively.

[0112] For nodes \(v_i\), \(v_j\), \(v_k\) (assuming without loss of generality that \(v_i\) and \(v_j\) are linked, and \(v_i\) and \(v_k\) are not linked), after steps (1)-(4), they have a final representation \(o\) i , \(o\) j , \(o\) k , and in the predictor, the \(o\) i and \(o\) j with links are closer in distance, while the \(o\) i and \(o\) k without links are farther apart, so as to achieve the purpose of target prediction.

Claims

1. A miRNA target prediction method based on a multi-layer heterogeneous graph, characterized in that The steps are as follows: Step 1: Construct a heterogeneous graph composed of seven RNA networks, where the nodes represent one of the three types of RNA: miRNA, lncRNA, and mRNA, and the seven RNA networks reflect seven different edge types; The seven edge types are: ① The miRNA-lncRNA interaction layer represents known and validated LMIs, ② The miRNA-miRNA sequence similarity layer measures the sequence similarity between miRNAs, ③ The miRNA-miRNA co-expression layer measures the co-expression relationship between miRNAs, ④ The miRNA-mRNA interaction layer represents the known mRNAs targeted by miRNAs, ⑤ The lncRNA-lncRNA sequence similarity layer measures the sequence similarity between lncRNAs, ⑥ The lncRNA-lncRNA co-expression layer measures the expression similarity between lncRNAs, ⑦ The lncRNA-mRNA interaction layer represents the known mRNAs targeted by lncRNAs; Step 2: Use a layer-specific method to aggregate features from different layers, and use an average pooling layer to aggregate to obtain the k-order feature of node i in the network layer. That is, the edge embedding of each node in the layer of the multi-layer heterogeneous graph is obtained, which shows that the k-th order feature of node i depends on the average value of the k-1-th order features of node i and its neighbors: where σ(·) represents the sigmoid function, and W (k) is a weight matrix to be learned during the training process, mean(·) represents the averaging operation, r represents the layer number, i.e., the r-th network layer, and N i,r is a set of nodes that includes node i and its neighbors, denotes i,r the (k - 1)-th order feature of node j in the set N, where 1 ≤ k ≤ K, and K represents the maximum feature aggregation level of each network layer; Step 3: Denote the edge embeddings in all layers of node \(i\) as matrix \(U\) i =(u i,1 ,…,u i,l ), where \(U\) i ∈ ℝ s×l , that is, \(U\) i is an \(s×l\) matrix, \(s\) represents the edge embedding dimension of the node, and \(l\) represents the total number of layers; use the multi-head self-attention mechanism to encode the edge embeddings of multiple layers of node \(v\) i to obtain \(H\) i,r [k]\), as follows: where: k represents the number of attention heads, k ∈ [1, m], and m represents the total number of attention heads; H i,r [k] represents the k-th head representation of the node v in the r-th layer i ; The calculation formula of A i,r is as follows: where r represents the layer number, i represents the node number, and softmax(·) represents the softmax function, and are learnable matrices, where m represents the total number of attention heads, s represents the edge embedding dimension of the nodes, and d a represents the intermediate dimension during the change process; Step 4: Use the projection method to project the edge embeddings into the task space, then extract the features from each layer and finally integrate them together; Specifically: Map the representation of the multi-attention heads of a single node from $\mathbb{R}^{ s}$ to the final task space $\mathbb{R}^{ d}$ through the following formula: s d as follows:​ P i,r [k] = H i,r [k]W p where W p ∈R s×d is the matrix parameter to be learned through training, s represents the edge embedding dimension before the node passes through the projector, d represents the edge embedding dimension after the node passes through the projector, k represents the number of the attention head, and P i,r [k] represents the k-th head representation of the node after projection; Select the bilinear interaction Bi-pool for pooling operation to fuse the representations of k attention heads of the node, and obtain the final edge embedding e of node i at layer r i,r : Where: m represents the total number of attention heads, and j and k both represent the numbers of attention heads, p i,r [k] represents the k-th head representation of node i in the r-th layer, p i,r [j] represents the j-th head representation of node i in the r-th layer, represents the element-wise product of two vectors, W r,pool is a matrix parameter to be learned through training; The basic embedding of the node vi is shared across all layers and serves as a message-passing medium to fuse the edge embeddings from each layer and pass them between layers; Step 5: By randomly generating a set of values from a Gaussian distribution, the basic embedding of each node can be randomly initialized. The following formula is used to fuse the basic embedding and the edge embedding e i,r to obtain the fused embedding of order t represents the base embedding of order t-1, represents the edge embedding of order t; Implement adjacent neighborhood aggregation hierarchy through the hybrid of the previous round of fusion embedding and edge embedding Mix the information on it to smooth the outputs between different aggregation layers. Stack more neighborhood aggregation layers to capture longer-distance information and obtain the final representation o of any node i i ; Step 6: Use the cosine distance formula shown below to calculate the distance between two nodes i and j of the target in the prediction space: where o i and o j represent the final representations of nodes vi and vj respectively; The nodes represent miRNAs, or the targets of miRNAs, namely mRNAs or lncRNAs. If the distance between a target node and a miRNA node is large, it indicates that the mRNA / lncRNA is more likely to be a target of the miRNA.

2. The miRNA target prediction method based on a multi-layer heterogeneous graph according to claim 1, wherein: The miRNA-lncRNA interaction layer represents the known and validated LMI as follows: Extract unique miRNA-lncRNA interactions with at least one CLIP-seq experimental evidence from the lncRNASNP2 database. The network layer consists of multiple unique miRNAs, multiple unique lncRNAs, and multiple unique miRNA-lncRNA edges.

3. The miRNA target prediction method based on a multi-layer heterogeneous graph according to claim 1, characterized in that: The miRNA-miRNA sequence similarity layer measures the sequence similarity between miRNAs as follows: First, retrieve the miRNA sequences of multiple miRNAs from the miRbase database, and then use the Needleman-Wunsch algorithm implemented in the Biostrings software package to perform global alignment on each pair of miRNA sequences; the gap opening penalty is set to 0.5, and the gap extension penalty is set to 0.1; if the identity score of two miRNAs is greater than or equal to 40, they will be connected in this layer. The resulting network layer consists of multiple miRNAs and multiple miRNA-miRNA interactions.

4. The miRNA target prediction method based on a multi-layer heterogeneous graph according to claim 1, wherein: The co-expression relationship between miRNAs measured by the miRNA-miRNA co-expression layer is as follows: miRNA expression profiles of multiple miRNAs were retrieved from mammalian microRNA expression atlases, which were collected from major organs and cell types of multiple human subjects; the Pearson correlation coefficient (PCC) was used to measure the co-expression similarity between miRNAs, and two miRNAs with PCC greater than or equal to 0.3 would be linked in this layer, and the resulting network layer consists of multiple miRNA co-expression miRNA pairs among multiple miRNAs.

5. The miRNA target prediction method based on a multi-layer heterogeneous graph according to claim 1, wherein: The miRNA-mRNA interaction layer representing the known mRNAs targeted by miRNAs is as follows: Experimentally verified miRNA-mRNA interactions were downloaded from miRTarBase. After removing weak miRNA-mRNA interactions, only one or more pieces of evidence from qRT-PCR, luciferase reporter assay, Western blot, microarray, immunohistochemistry, and in situ hybridization retained strong interactions.

6. The miRNA target prediction method based on a multi-layer heterogeneous graph according to claim 1, wherein: The sequence similarity between lncRNAs measured by the lncRNA-lncRNA sequence similarity layer is as follows: First, DNA sequences of multiple lncRNAs were downloaded from the NONCODE database, and the lncRNA-lncRNA sequence similarity was calculated based on sequence alignment; The Smith-Waterman local alignment algorithm was used to perform the task; During the alignment process, the penalty for opening a gap was set to 10, and the incremental cost generated along the gap length was set to 4; lncRNA pairs with an alignment score greater than or equal to 400 would be retained as lncRNA-lncRNA edges in this layer; the resulting network layer consists of multiple lncRNAs and multiple lncRNA-lncRNA interactions.

7. The miRNA target prediction method based on a multi-layer heterogeneous graph according to claim 1, wherein: The expression similarity between lncRNAs measured by the lncRNA-lncRNA co-expression layer is as follows: Expression profiles of multiple lncRNAs were downloaded from the NONCODE database, and a higher PCC threshold of 0.9 was selected. Finally, there are multiple lncRNA co-expression link relationships among the multiple lncRNAs remaining in this layer.

8. The miRNA target prediction method based on a multi-layer heterogeneous graph according to claim 1, wherein: The lncRNA-mRNA interaction layer representing the known mRNAs targeted by miRNAs is as follows: Multiple experimentally verified edges were downloaded from the RISE database, and the lncRNA-mRNA edge relationship was obtained from them.

9. A system for the miRNA target prediction method based on a multi-layer heterogeneous graph according to any one of claims 1 to 8, characterized in that: It includes an aggregator, an encoder, a projector, a fuser, and a predictor; the input end of the aggregator receives a heterogeneous graph composed of seven RNA networks, and the node features of each layer of the heterogeneous graph are updated by the aggregator; then, the node features of each layer are fused with the information of other layers by the encoder to obtain the node features updated again, where each node has multiple head representations; Then, the representations of multiple heads are fused by the projector, and the finally fused vector is used as the edge embedding of the node; the edge embedding and the base embedding are fused in the fuser to obtain the final node representation; finally, the cosine distance is calculated in the predictor for prediction.

Citation Information

Patent Citations

  • QTL locus for controlling ramie skin thickness character and obtaining method and application thereof

    CN113388693A

  • Heterogeneous graph neural network traditional Chinese medicine target prediction method based on attention mechanism

    CN114121181A