MiRNA enhanced gene regulatory network inference method based on knowledge graph embedding and multi-channel heterogeneous graph
By constructing multi-channel heterogeneous graphs and GraphSAGE networks, we deeply mined gene sequence information, solved the problem of predicting the regulatory relationship of high-throughput miRNAs, and improved the prediction accuracy and comprehensiveness of gene regulatory networks.
Patent Information
- Application Number
- CN202510711921.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-05-29
AI Technical Summary
In the existing technology, high-throughput miRNA gene sequence feature extraction and deep learning of complex regulatory relationships have not been effectively solved, resulting in insufficient research on the synergistic multi-level regulatory effects between miRNA and other regulatory factors.
A multi-channel heterogeneous graph is constructed, and knowledge graph embedding and multi-channel GraphSAGE network are used to deeply mine gene sequence information through sequence-to-vector technology, capture the characteristics of nodes under different regulatory relationships, and determine the connection relationship between nodes based on the connection probability value.
It achieves efficient and accurate prediction of the regulatory relationships of high-throughput miRNAs, and improves the prediction accuracy and comprehensiveness of gene regulatory networks.
Smart Images

Figure CN120690288A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of bioinformatics, and in particular to a miRNA-enhanced gene regulatory network inference method based on knowledge graph embedding and multi-channel heterogeneous graphs. Background Art
[0002] Gene Regulatory Networks (GRNs) are complex networks composed of genes and their products interacting with each other, playing a key role in regulating gene expression and function. Studying GRNs not only helps reveal the molecular mechanisms of disease but also provides new research directions and ideas for personalized treatment. In this process of expression regulation, miRNAs (also known as microRNAs), as important regulatory factors, act as "molecular fine-tuners" within GRNs. They influence cellular function through multi-target, dynamic, and hierarchical regulation, thereby exerting a negative influence on gene expression regulation.
[0003] In the current research on gene regulatory networks, there is a lack of research on the synergistic multi-level regulatory effects between miRNA and other regulatory factors. A related literature (Luo J, Ouyang W, Shen C, et al. Multi-Relation Graph Embedding for Predicting miRNA-Target Gene Interactions by Integrating Gene Sequence Information[J]. IEEE Journal of Biomedical and Health Informatics, 2022, 26: 4345-4353.) proposed a model based on multi-relation graph embedding to predict the interaction between miRNA and target genes. By combining multi-relation graph convolution and bidirectional long short-term memory network (Bidirectional Long Short-Term Memory, Bi-LSTM), the prediction accuracy of miRNA-target gene interaction was improved.
[0004] However, although the rise of knowledge graphs and deep learning algorithms has opened up new avenues for integrating diverse biological data, the extraction of high-throughput miRNA gene sequence features and deep learning of complex regulatory relationships remain urgent challenges. To this end, this application proposes a new gene regulatory network inference method to achieve efficient and accurate regulatory relationship prediction. Summary of the Invention
[0005] The main purpose of this application is to provide a miRNA-enhanced gene regulatory network inference method based on knowledge graph embedding and multi-channel heterogeneous graph, aiming to solve the problem of how to efficiently and accurately predict the regulatory relationship of high-throughput miRNAs.
[0006] To achieve the above objectives, the present application provides a method for inferring miRNA-enhanced gene regulatory networks based on knowledge graph embedding and multi-channel heterogeneous graphs, the method comprising:
[0007] A multi-channel heterogeneous graph is constructed with transcription factors, target genes and miRNAs as nodes and the regulatory relationships between nodes as edges;
[0008] Divide the gene sequence information of transcription factors, target genes, and miRNAs into DNA fragments of length k, and pre-train the DNA fragments through a preset learning model so that the preset learning model learns the knowledge graph embedding associated with the DNA fragments;
[0009] Using GraphSAGE in each channel to sample adjacent nodes of any selected target node in the multi-channel heterogeneous graph, thereby capturing the characteristics of the target node under one or more regulatory relationships;
[0010] After traversing all nodes, determining the aggregated embedded representation of the target node's neighbor nodes under each of the control relationships;
[0011] According to the embedding representation, a connection probability value between any two nodes is calculated, and two nodes with a connection probability value greater than a preset threshold are determined as a target node pair with a connection relationship.
[0012] Optionally, the regulatory relationship includes at least one of the following:
[0013] The regulatory relationship between transcription factors and target genes;
[0014] Regulatory relationships between transcription factors;
[0015] The regulatory relationship between transcription factors and miRNAs;
[0016] The regulatory relationship between miRNAs and transcription factors;
[0017] Regulatory relationship between miRNA and target genes.
[0018] Optionally, the training objective of the preset learning model is to minimize negative log-likelihood loss.
[0019] Optionally, the expression of the embedded representation is:
[0020]
[0021] Where, It represents the embedded representation of the neighbor nodes of node v after aggregation under the relationship r; is the embedding representation associated with the neighbor node u in the kth layer; MaxPool is the maximum pooling operation, and Represent the pooling weight and bias term associated with relation r, is the neighbor node, represents the set of neighbor nodes under each relationship, is the ReLU activation function.
[0022] Optionally, the calculation expression of the connection probability value is:
[0023]
[0024] Where, Representation node and The connection probability value between For nodes The transpose of the associated embedding representation, For nodes Embedding representation of .
[0025] In addition, to achieve the above objectives, the present application also provides a miRNA-enhanced gene regulatory network inference model based on knowledge graph embedding and multi-channel heterogeneous graph, the miRNA-enhanced gene regulatory network inference model comprising:
[0026] A multi-channel heterogeneous graph construction module is used to construct a multi-channel heterogeneous graph using transcription factors, target genes, and miRNAs as nodes and the regulatory relationships between nodes as edges;
[0027] A pre-trained node feature embedding module is used to segment the gene sequence information of transcription factors, target genes, and miRNAs into DNA fragments of length k, and pre-train the DNA fragments through a preset learning model so that the preset learning model learns the knowledge graph embedding associated with the DNA fragments;
[0028] A multi-channel heterogeneous GraphSAGE network module is used to sample the adjacent nodes of any selected target node using GraphSAGE in each channel, thereby capturing the characteristics of the target node under one or more regulatory relationships; after traversing all nodes, the module determines the aggregated embedded representation of the neighboring nodes of the target node under each of the regulatory relationships; based on the embedded representation, the module calculates the connection probability value between any two nodes; and determines two nodes with a connection probability value greater than a preset threshold as a target node pair with a connection relationship.
[0029] Optionally, the number of GraphSAGE layers in the multi-channel heterogeneous GraphSAGE network module is 2.
[0030] Optionally, the embedding dimension is set to 256 dimensions.
[0031] In addition, to achieve the above-mentioned purpose, the present application also provides a computer system, which includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor. When the computer program is executed by the processor, the steps of the miRNA-enhanced gene regulatory network inference method based on knowledge graph embedding and multi-channel heterogeneous graph as described in any one of the above items are implemented.
[0032] In addition, to achieve the above-mentioned purpose, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the miRNA-enhanced gene regulatory network inference method based on knowledge graph embedding and multi-channel heterogeneous graph as described in any of the above items are implemented.
[0033] This application has at least the following beneficial effects:
[0034] 1. Using transcription factors, target genes, and miRNAs as nodes and the regulatory relationships between nodes as edges, we construct a multi-channel heterogeneous graph, and use a gene regulatory knowledge graph covering transcription factors, miRNAs, and target genes to improve the gene regulatory network.
[0035] 2. During the pre-training phase, we use sequence-to-vector technology to deeply mine and extract semantic information from nucleic acid sequences. We then decompose the multi-channel heterogeneous graph into multiple subgraphs based on different types of regulatory relationships, and use multi-channel GraphSAGE to learn node features unique to each regulatory type.
[0036] 3. Determine the connection relationship prediction results between nodes based on the connection probability values between the embedded features of each node, and realize high-throughput miRNA regulatory relationship prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a flowchart of the first embodiment of the miRNA-enhanced gene regulatory network inference method based on knowledge graph embedding and multi-channel heterogeneous graphs of this application;
[0038] Figure 2 The multi-channel heterogeneous graph involved in the embodiment of this application is ''
[0039] Figure 3 Schematic diagram of the architecture of the miRNA-enhanced gene regulatory network inference model involved in the embodiments of the present application;
[0040] Figure 4 The AUROC graph of the KGE-HGRN model involved in the embodiment of the present application in five-fold cross validation;
[0041] Figure 5 This is a schematic diagram of the graph structure analysis results of different types of edges involved in the embodiments of this application;
[0042] Figure 6 Schematic diagram of the architecture of the hardware operating environment of the computer system involved in the embodiment of the present application.
[0043] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0044] To better understand the above technical solutions, exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0045] First embodiment
[0046] In gene regulatory networks, genes are represented as nodes, and regulatory relationships are represented by directed edges. Inference of gene regulatory networks can be transformed into a link prediction task, aiming to identify possible regulatory relationships between gene nodes.
[0047] Reference Figure 1 In this embodiment, a miRNA-enhanced gene regulatory network inference method based on knowledge graph embedding and multi-channel heterogeneous graph is proposed, and the method includes the following steps:
[0048] Step S10, constructing a multi-channel heterogeneous graph with transcription factors, target genes, and miRNAs as nodes and regulatory relationships between nodes as edges;
[0049] In this step, a large-scale gene regulation knowledge graph containing multiple channels, namely a multi-channel heterogeneous graph, is constructed by integrating data from TRRUSTv2, RegNetwork and DoRothEA databases.
[0050] In an alternative embodiment, see Figure 2 The multi-channel heterogeneous graph shown in Figure 2 is represented as a directed graph. , where V, E, T, and R represent the node set, edge set, and their corresponding types, respectively. Each node v corresponds to a specific type t, and each edge e is associated with a specific edge type r. The model divides the nodes into transcription factors , target gene and miRNAs , the edge types E are classified according to their relationships: TF-TG TF-TF TF-miRNA 、miRNA-TF and miRNA-TG , respectively, refer to the regulatory relationships between transcription factors and target genes, the regulatory relationships between transcription factors, the regulatory relationships between transcription factors and miRNAs, the regulatory relationships between miRNAs and transcription factors, and the regulatory relationships between miRNAs and target genes. In addition, each node in the graph is denoted as vi, where i ranges from 1 to N, and N = |V| represents the total number of nodes.
[0051] In an alternative embodiment, the constructed knowledge graph was deployed on the Neo4j platform and a path query was performed there. Taking the transcription factor E2F1 and its target gene TOP2A as an example, the path length was set to no more than three hops, and possible partial regulatory pathways between them were extracted.
[0052] Step S20, dividing the gene sequence information of the transcription factor, target gene, and miRNA into DNA segments of length k, and pre-training the DNA segments through a preset learning model so that the preset learning model learns the knowledge graph embedding associated with the DNA segments;
[0053] In this example, the pre-training phase uses sequence-to-vector technology to convert sequences into low-dimensional numerical vectors. This method effectively captures potential information and associations within the sequence, while ensuring a close connection between this information and biological characteristics. During this process, the k-mer counting method is employed, which segments the gene sequence information into DNA segments of a fixed length k, thereby improving the efficiency of sequence embedding. These DNA segments are treated as "words," while the entire gene sequence information is considered a "document."
[0054] Furthermore, in some optional embodiments, in order to process gene sequence information more efficiently, the Doc2Vec method in NLP is used in the preset learning model. It should be noted that the core idea of the Doc2Vec stage is to generate a vector representation for each k-mer segment in the sequence set. In this process, the PV-DBOW model can be used as the preset training model to optimize the vector representation of the entire sequence for training, with the goal of predicting the k-mer segments in the sequence. This method is similar to the Skip-gram model in Word2Vec. In this way, Doc2Vec can not only effectively capture the local features of the RNA sequence, but also deeply explore its global features, thereby providing more comprehensive information for subsequent analysis.
[0055] Specifically, suppose a sequence c contains a series of k-mer segments (where i ranges from 1 to N), the entire sequence is represented as a vector , each k-mer fragment Also represented as a vector The prediction goal of the preset learning model is to predict the probability of each segment based on the vector representation of the sequence. This probability is normalized by the softmax function as follows:
[0056]
[0057] Where v represents the set of all possible k-mer fragments.
[0058] In some optional embodiments, in order to maximize the alignment between the document vector and the corresponding target segment vector while reducing the similarity with the non-target segment vector, the preset training model aims to minimize the following negative log-likelihood loss:
[0059]
[0060] Where, is the target fragment, and D is the document (gene sequence information).
[0061] It should be noted that the negative log-likelihood loss function is continuously optimized by the gradient descent method. In this process, the negative log-likelihood measures the target segment The goal is to maximize the similarity between the document vector v(D) and each target segment vector v( ). At the same time, the softmax function normalizes all possible k-mer segments and converts the goal into maximizing the probability of the target segment relative to all candidate segments. The loss function adopts a logarithmic form, which can impose a greater penalty on incorrect predictions, which can encourage the model to learn the relationship between documents and their segments more effectively during training. By comparing the document vector v(D) and the segment vector v( ), perform iterative updates, where By representing all potential k-mer segments, the model is able to gradually narrow the gap between the predicted and actual values. Ultimately, this process generates effective embeddings that can capture both global information and local features in the sequence.
[0062] Step S30, using GraphSAGE in each channel to sample the adjacent nodes of the arbitrarily selected target node, thereby capturing the characteristics of the target node under one or more regulatory relationships;
[0063] Step S40, after traversing all nodes, determining the embedded representation of the target node's neighbor nodes under each of the control relationships after aggregation;
[0064] Step S50 , calculating the connection probability value between any two nodes based on the embedded representation, and determining the two nodes with the connection probability value greater than a preset threshold as a target node pair with a connection relationship.
[0065] In this embodiment, when performing link prediction, based on the framework established in steps S10 and S20, the features of the entire heterogeneous graph and all nodes are taken as input. The GraphSAGE module is used to extract and summarize the embedded features of each node. The module takes the embedded representations of all nodes as input and calculates the probability of a connection between any two nodes.
[0066] Since information propagation differs between different edge types in a graph, but these propagation mechanisms are interrelated, in order to handle this complex multi-relationship graph structure, this embodiment adopts a GraphSAGE network based on a multi-channel heterogeneous graph framework.
[0067] Considering the complexity of multi-relational graphs, the GraphSAGE network decomposes a multi-channel heterogeneous graph into multiple subgraphs, each containing only one specific edge type. Within each subgraph, the GraphSAGE network then randomly selects adjacent nodes of a target node from its neighborhood as learning points, thereby capturing the characteristics of the target node under one or more regulatory relationships. After traversing all nodes, the sampled neighbor node features are aggregated using an aggregation function to update the feature embedding representation of each node.
[0068] In some optional implementations, the selected aggregation function is the maximum pooling function MaxPool.
[0069] For example, let the feature representation of the current node v be , where d represents the dimension of the embedding space, for each neighbor node , whose embedding is expressed as When multiple relations r affect the same node, Represents the set of neighbor nodes under each relationship. Perform a maximum pooling operation on the neighbor nodes of each relationship to obtain the embedded representation of the target node after aggregation of the neighbor nodes under each regulatory relationship:
[0070]
[0071] Where, It represents the embedded representation of the neighbor nodes of node v after aggregation under the relationship r; is the embedding representation associated with the neighbor node u in the kth layer; MaxPool is the maximum pooling operation, and Represent the pooling weight and bias term associated with relation r, is the neighbor node, represents the set of neighbor nodes under each relationship, is the ReLU activation function.
[0072] Furthermore, and optionally, after determining the embedding representation of each layer, the embedding representations of all relations are finally aggregated by weighted summation to obtain a weighted feature representation. :
[0073]
[0074] It is important to emphasize that the model is able to capture unique neighborhood information for each relationship type and generate a global representation of each node by integrating this information.
[0075] In some optional implementations, the calculation expression of the connection probability value is:
[0076]
[0077] Where, Representation node and The connection probability value between For nodes The transpose of the associated embedding representation, For nodes Embedding representation of .
[0078] Furthermore, in some optional implementations, the preset threshold is set to 0.5, that is, if the probability If it exceeds 0.5, the node is considered and nodes Conversely, if If it is less than or equal to 0.5, it is considered that there is no connection between the two.
[0079] Furthermore, in some optional implementations, in the probability prediction phase, a binary cross entropy function is used to calculate the loss, which in turn measures the computational loss during the model training process. The formula is as follows:
[0080]
[0081] In the formula, Loss represents the cross entropy loss of the entire model, V represents the set of all nodes in the graph, is the label of the edge connecting node i and node j, and It represents the probability that the model predicts whether there is a connection between node i and node j.
[0082] The technical solution provided in this embodiment addresses the difficulties of extracting features from high-throughput gene sequences and learning complex regulatory relationships. A multi-channel heterogeneous graph is constructed using transcription factors, target genes, and miRNAs as nodes, and the regulatory relationships between nodes as edges. During the pre-training phase, sequence-to-vector technology is used to deeply mine and extract semantic information from nucleic acid sequences. The multi-channel heterogeneous graph is then decomposed into multiple subgraphs based on different types of regulatory relationships. Multi-channel GraphSAGE is then used to learn the node features unique to each regulatory type. Finally, the connection relationship prediction results between nodes are determined based on the connection probability values between the embedded features of each node, enabling high-throughput miRNA regulatory relationship prediction.
[0083] Second embodiment
[0084] Based on the method in the first embodiment, as an implementation scheme, refer to Figure 3 The schematic diagram of the architecture of the miRNA-enhanced gene regulatory network inference model is shown. In this embodiment, a miRNA-enhanced gene regulatory network inference model based on knowledge graph embedding and multi-channel heterogeneous graph is provided, referred to as the KGE-HGRN model, which includes:
[0085] A multi-channel heterogeneous graph construction module is used to construct a multi-channel heterogeneous graph using transcription factors, target genes, and miRNAs as nodes and the regulatory relationships between nodes as edges;
[0086] A pre-trained node feature embedding module is used to segment the gene sequence information of transcription factors, target genes, and miRNAs into DNA fragments of length k, and pre-train the DNA fragments through a preset learning model so that the preset learning model learns the knowledge graph embedding associated with the DNA fragments;
[0087] A multi-channel heterogeneous GraphSAGE network module is used to sample the adjacent nodes of any selected target node using GraphSAGE in each channel, thereby capturing the characteristics of the target node under one or more regulatory relationships; after traversing all nodes, the module determines the aggregated embedded representation of the neighboring nodes of the target node under each of the regulatory relationships; based on the embedded representation, the module calculates the connection probability value between any two nodes; and determines two nodes with a connection probability value greater than a preset threshold as a target node pair with a connection relationship.
[0088] Comparative Example 1
[0089] Based on the miRNA-enhanced gene regulatory network inference model proposed in the second embodiment, and considering that the effectiveness of the KGE-HGRN model is affected by multiple key factors, this embodiment experiments with different neural network layer depths, embedding vector dimensions, and selected node feature aggregation strategies. In order to evaluate the impact of different numbers of layers on the performance of the KGE-HGRN model, the embedding dimension was set to 256, and 5-fold cross-validation was performed using 1-layer, 2-layer, and 3-layer GraphSAGE models. The results are shown in Table 1:
[0090] Table 1. Performance of the KGE-HGRN model at different GraphSAGE layers
[0091] The results show that the model performs best on most evaluation metrics when it has 2 layers. However, when the number of layers increases to 3, the model performance decreases significantly, which may be caused by overfitting, thus affecting the overall effect.
[0092] Comparative Example 2
[0093] Based on the miRNA-enhanced gene regulatory network inference model proposed in the second embodiment, in order to evaluate the performance of the KGE-HGRN model, the embedding dimensions were set to 64, 128, 256, 512, and 1024, respectively. The results are shown in Table 2:
[0094] Table 2. Effect of different embedding dimensions on the performance of the KGE-HGRN model
[0095] The results show that as the embedding dimension increases, the performance of the model gradually improves, reaching the best effect at 256 dimensions.
[0096] Comparative Example 3
[0097] Based on the miRNA-enhanced gene regulatory network inference model proposed in the second example, the model further uses GraphSAGE for node feature learning. Each node collects and aggregates features from neighboring nodes in each iteration. As the number of iterations increases, nodes gradually obtain information from more distant neighbors, thereby accessing a wider range of graph structure and feature information. In this example, four aggregation strategies—MAX Pool, LSTM, mean, and GCN—are used to evaluate the performance of the framework. The results are shown in Table 3:
[0098] Table 3. Effect of different GraphSAGE aggregation strategies on model performance
[0099] The results show that when the aggregation strategy of GraphSAGE is set to MAX Pool, the overall performance of the model is the best.
[0100] Third embodiment
[0101] As a verification example, in this example, the KGE-HGRN model based on the method proposed in the first example or constructed in the second example is compared with a traditional common model.
[0102] (1) Dataset construction and experimental design
[0103] In this example, the pre-constructed knowledge graph contains 1,647 transcription factors, 18,393 target genes, 807 miRNAs, and 519,464 mutual regulatory relationships between them. The KGE-HGRN model is evaluated using several evaluation metrics commonly used in gene regulatory network inference tasks, including F1 score, accuracy, precision, recall, AUROC, and AUPRC. First, a benchmark evaluation is performed to compare the model's performance with existing GRN inference methods. Next, ablation experiments are conducted and the model parameters of the knowledge graph and framework are compared to rigorously evaluate their effectiveness. Finally, this section conducts a focused case study on E2F1 to further verify the effectiveness of the KGE-HGRN model.
[0104] In bioinformatics, positive samples typically refer to gene pairs with known regulatory relationships, while negative samples refer to gene pairs without such relationships. To construct the negative sample set, the model randomly selects gene pairs from unlabeled data. However, to avoid mistakenly selecting positive samples during the selection process, the negative samples are resampled before each model calculation to ensure accuracy and diversity.
[0105] The model used K-fold cross-validation to evaluate model performance. The dataset was divided into several subsets. In each round, the model was trained on one subset and tested on another. This process was repeated multiple times, each time using a different subset as the test set. The model used the commonly used 5-fold cross-validation method. Each fold was trained for 300 rounds to ensure that the model's variance dropped below 0.01.
[0106] (2) Experimental results and analysis
[0107] In order to test the effect of different K values in K-fold cross validation on the experimental results, the specific K values are set to 3, 5, 7 and 10 respectively. The relevant results can be seen in Table 4:
[0108] Table 4. Comparison of KGE-HGRN models using different cross-validation
[0109] The results show that the best effect is achieved when K=5, which is also consistent with the experimental settings of most current studies.
[0110] Furthermore, the AUROC graph of the KGE-HGRN model in the five-fold cross validation is as follows: Figure 4 As shown in Figure 3, the performance of KGE-HGRN is compared with six other existing methods, including KGE-TGI, MGDHGS, EdgeConv, Graph Isomorphism Network (GIN), GAT, and GATv2.
[0111] The results are shown in Table 5 below:
[0112] Table 5. Performance comparison of KGE-HGRN model and other methods
[0113] Results show that the proposed model outperforms other models in terms of metrics such as AUROC, AUPRC, accuracy, precision, and F1 score. Although EdgeConv achieved a recall of 0.95, its precision was lower, suggesting that the framework may have generated a high number of false positives. In terms of AUROC and AUPRC, the KGE-HGRN model surpassed MGDHGS and KGE-TGI by 12%, 3%, 8%, and 2%, respectively, further demonstrating the model's superiority and effectiveness.
[0114] In the technical solution provided in this embodiment, the advantage of the KGE-HGRN model is that it can use multi-channel GraphSAGE to effectively learn the unique features under different regulatory types. Thanks to this strategy, the model can more accurately capture the multi-level structural characteristics of biological networks. By utilizing multi-relationship heterogeneous knowledge graphs, the model can learn complementary features from additional regulatory information, thereby playing a key role in improving prediction accuracy.
[0115] Fourth embodiment
[0116] As a verification example, this application considers that miRNA plays a vital role in the gene regulatory network. Therefore, the regulatory relationship between miRNA, transcription factors (TFs) and target genes (TGs) is incorporated into the model. Transcription factors have a dual regulatory effect: they not only regulate the transcription of mRNA, but also affect the transcription level of miRNA. On the other hand, some miRNAs can target mRNA encoding transcription factors, thereby regulating the expression of these transcription factors. In addition, transcription factors and miRNAs can also synergistically regulate the same target gene, where transcription factors control gene expression at the transcriptional level, while miRNAs fine-tune gene expression at the post-transcriptional level.
[0117] In order to evaluate the impact of adding miRNA regulatory information on model performance, this example conducted a graph structure analysis of different types of edges:
[0118] 1. Regulatory relationship between transcription factors and target genes, i.e. TF-TG ;
[0119] 2. Regulatory relationships between transcription factors, i.e. TF-TF ;
[0120] 3. Regulatory relationship between transcription factors and miRNA, i.e. TF-miRNA ;
[0121] 4. Regulatory relationship between miRNA and transcription factors, i.e. miRNA-TF ;
[0122] 5. Regulatory relationship between miRNA and target genes, i.e. miRNA-TG .
[0123] The analysis results are shown in Figure 5In the figure, ACC is the accuracy rate, Pre is the precision rate, and F1 is the F1 Score, which comprehensively considers the model's precision and recall rate. It balances Precision and Recall through the harmonic mean, thus providing a more comprehensive evaluation:
[0124] ;
[0125] AUROC stands for the area under the receiver operating characteristic curve (ROC Curve), which is obtained by plotting the true positive rate (TPR, also known as recall) and the false positive rate (FPR, 1 - true negative rate) at different thresholds; while AUPRC is obtained by calculating the precision at different recall levels, plotting the curve in the precision-recall space, and then calculating the area under the curve.
[0126] In this example, the E2F1 gene was selected as the research object. It is an important member of the E2F transcription factor family and is involved in the regulation of DNA replication and the cellular response to DNA damage. Dysregulation of E2F1 activity is closely related to the progression of various cancers. In the experiment, all known regulatory relationships were used for training, and the remaining unknown relationships were used as a candidate test set. The KGE-HGRN model was used to identify potential regulatory relationships and rank the candidate relationships according to the predicted likelihood, and finally the top ten candidate relationships were selected. This is shown in Table 6 below:
[0127] Table 6. Genes predicted to have regulatory relationships with E2F1
[0128] The results showed that by verifying these predictions through reviewing relevant literature on PubMed, it was found that 70% (or 7 out of 10) of the predicted regulatory relationships were directly or indirectly confirmed in existing studies.
[0129] As an implementation solution, Figure 6 This is a schematic diagram of the architecture of the hardware operating environment of the computer system involved in the embodiment of the present application.
[0130] like Figure 6As shown, the computer system may include: a processor 1001, such as a CPU, a memory 1005, a user interface 1003, a network interface 1004, and a communication bus 1002. The communication bus 1002 is used to implement communication between these components. The user interface 1003 may include a display and an input unit such as a keyboard. Optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as a disk storage device. The memory 1005 may also be a storage device independent of the processor 1001.
[0131] Those skilled in the art will understand that Figure 6 The computer system architecture shown in the figure does not constitute a limitation of the computer system, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0132] like Figure 6 As shown, the memory 1005 as a storage medium may include an operating system, a network communication module, a user interface module and a computer program. Among them, the operating system is a program that manages and controls the hardware and software resources of the computer system, the operation of the computer program and other software or programs.
[0133] exist Figure 6 In the computer system shown, the user interface 1003 is mainly used to connect to the terminal and communicate data with the terminal; the network interface 1004 is mainly used to communicate data with the background server; the processor 1001 can be used to call the computer program stored in the memory 1005.
[0134] In this embodiment, the computer system includes: a memory 1005, a processor 1001, and a computer program stored in the memory and executable on the processor, wherein:
[0135] When the processor 1001 calls the computer program stored in the memory 1005, it performs the following operations:
[0136] A multi-channel heterogeneous graph is constructed with transcription factors, target genes and miRNAs as nodes and the regulatory relationships between nodes as edges;
[0137] Divide the gene sequence information of transcription factors, target genes, and miRNAs into DNA fragments of length k, and pre-train the DNA fragments through a preset learning model so that the preset learning model learns the knowledge graph embedding associated with the DNA fragments;
[0138] Using GraphSAGE in each channel to sample adjacent nodes of any selected target node in the multi-channel heterogeneous graph, thereby capturing the characteristics of the target node under one or more regulatory relationships;
[0139] After traversing all nodes, determining the aggregated embedded representation of the target node's neighbor nodes under each of the control relationships;
[0140] According to the embedding representation, a connection probability value between any two nodes is calculated, and two nodes with a connection probability value greater than a preset threshold are determined as a target node pair with a connection relationship.
[0141] Furthermore, those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is a computer-readable storage medium. The program instructions are executed by at least one processor in a computer system to implement the steps in the process of the above-described method embodiment.
[0142] Therefore, the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the various steps of the miRNA-enhanced gene regulatory network inference method based on knowledge graph embedding and multi-channel heterogeneous graph as described in the above embodiment.
[0143] The computer-readable storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0144] It should be noted that since the storage medium provided in the embodiments of this application is the storage medium used to implement the method of the embodiments of this application, based on the method described in the embodiments of this application, those skilled in the art will be able to understand the specific structure and deformation of the storage medium, and therefore will not be described in detail here. All storage media used in the method of the embodiments of this application fall within the scope of protection to be provided by this application.
[0145] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0146] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0147] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0148] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0149] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claim. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present application may be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.
[0150] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0151] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A miRNA-enhanced gene regulatory network inference method based on knowledge graph embedding and multi-channel heterogeneous graph, characterized by: The method comprises the following steps: A multi-channel heterogeneous graph is constructed with transcription factors, target genes and miRNAs as nodes and the regulatory relationships between nodes as edges; Divide the gene sequence information of transcription factors, target genes, and miRNAs into DNA fragments of length k, and pre-train the DNA fragments through a preset learning model so that the preset learning model learns the knowledge graph embedding associated with the DNA fragments; Using GraphSAGE in each channel to sample adjacent nodes of any selected target node in the multi-channel heterogeneous graph, thereby capturing the characteristics of the target node under one or more regulatory relationships; After traversing all nodes, determining the aggregated embedded representation of the target node's neighbor nodes under each of the control relationships; According to the embedding representation, a connection probability value between any two nodes is calculated, and two nodes with a connection probability value greater than a preset threshold are determined as a target node pair with a connection relationship.
2. The method according to claim 1, wherein The regulatory relationship includes at least one of the following: The regulatory relationship between transcription factors and target genes; Regulatory relationships between transcription factors; The regulatory relationship between transcription factors and miRNAs; The regulatory relationship between miRNAs and transcription factors; Regulatory relationship between miRNA and target genes.
3. The method according to claim 1, wherein The training objective of the preset learning model is to minimize the negative log-likelihood loss.
4. The method according to claim 1, wherein The expression of the embedding representation is: ; Where, It represents the embedded representation of the neighbor nodes of node v after aggregation under the relationship r; is the embedding representation associated with the neighbor node u in the kth layer; MaxPool is the maximum pooling operation, and Represent the pooling weight and bias term associated with relation r, is the neighbor node, represents the set of neighbor nodes under each relationship, is the ReLU activation function.
5. The method according to claim 1, wherein The calculation expression of the connection probability value is: ; Where, Representation node and The connection probability value between For nodes The transpose of the associated embedding representation, For nodes Embedding representation of .
6. A miRNA-enhanced gene regulatory network inference model based on knowledge graph embedding and multi-channel heterogeneous graph, characterized by: The miRNA-enhanced gene regulatory network inference model includes: A multi-channel heterogeneous graph construction module is used to construct a multi-channel heterogeneous graph using transcription factors, target genes, and miRNAs as nodes and the regulatory relationships between nodes as edges; A pre-trained node feature embedding module is used to segment the gene sequence information of transcription factors, target genes, and miRNAs into DNA fragments of length k, and pre-train the DNA fragments through a preset learning model so that the preset learning model learns the knowledge graph embedding associated with the DNA fragments; A multi-channel heterogeneous GraphSAGE network module is used to sample the adjacent nodes of any selected target node using GraphSAGE in each channel, thereby capturing the characteristics of the target node under one or more regulatory relationships; after traversing all nodes, the module determines the aggregated embedded representation of the neighboring nodes of the target node under each of the regulatory relationships; based on the embedded representation, the module calculates the connection probability value between any two nodes; and determines two nodes with a connection probability value greater than a preset threshold as a target node pair with a connection relationship.
7. The miRNA-enhanced gene regulatory network inference model according to claim 6, wherein: The number of GraphSAGE layers in the multi-channel heterogeneous GraphSAGE network module is 2.
8. The miRNA-enhanced gene regulatory network inference model according to claim 6, wherein: The embedding dimension is set to 256 dimensions.
9. A computer system, characterized in that: The computer system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the miRNA-enhanced gene regulatory network inference method based on knowledge graph embedding and multi-channel heterogeneous graph are implemented as described in any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the miRNA-enhanced gene regulatory network inference method based on knowledge graph embedding and multi-channel heterogeneous graph as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Network node correlation-based identification method and system for function module in co-regulation network
CN107679367A
Generation method and device of target gene prediction model, and storage medium
CN113838527A
Ring RNA-disease association prediction method and device based on weighted graph attention and heterogeneous graph neural network, and medium
CN115798730A
Transcription factor target gene relation prediction method, system, equipment and medium
CN116230070A
Multi-view circRNA and miRNA relation prediction method and system
CN117995280A