A method and model for predicting transcription factor-target gene interactions
By using random walk algorithms and graph embedding models in isomerographic graphs, we predict the interaction between transcription factors and target genes, and solve the problem that the connection between transcription factors and target genes cannot be directly predicted and comprehensively considered in the prior art, achieving more efficient and accurate prediction effects.
Patent Information
- Application Number
- CN202111493609.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The prior art cannot directly obtain results when predicting the interaction between transcription factors and target genes, and cannot consider the connection between transcription factors and target genes more comprehensively.
A heterogeneous graph embedding algorithm based on random walk is used to construct a heterogeneous network containing transcription factors, target genes and disease nodes. The node embedding model is used to generate node embeddings, and the prediction score matrix of transcription factors and target genes is calculated.
A more comprehensive consideration of the connection between transcription factors and target genes is achieved, improving the accuracy of predicting the interaction between transcription factors and target genes, simplifying the process and reducing costs.
Smart Images

Figure CN114420203B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and in particular to a method and a model for predicting transcription factor-target gene interaction. Background Art
[0002] Transcriptional regulation coordinates gene regulation that occurs during transcription by modulating the transcription rate, which determines cell developmental fate and cellular responses to genetic and environmental perturbations. In this process, transcription factors (TFs) bind to cis-regulatory elements of DNA and activate RNA polymerase to regulate the transcription of target genes. Given their importance to biological processes, determining the interaction patterns of TFs and their target genes is crucial for the study of biology and medicine.
[0003] Yang et al. proposed the GripDL model, which consists of a convolutional neural network (CNN) that learns embedded features from in situ hybridization (ISH) images of genes (including TFs and their target genes) to infer potential interaction relationships between them. However, gene expression image data is still very expensive for some research goals and is difficult to be widely adopted in practical research. Lin et al. used collaborative filtering technology based on three-factor decomposition and used related protein-protein interactions as training data to predict potential target genes for specific TFs. However, this method does not fully consider the connection between TFs and their target genes.
[0004] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0005] In view of the above-mentioned deficiencies in the prior art, the object of the present invention is to provide a method and model for predicting transcription factor-target gene interactions, aiming to solve the problems in the prior art that when predicting the interaction between transcription factors and target genes, the results of TF-target gene interactions cannot be directly obtained and the connection between TF and target genes cannot be considered more comprehensively.
[0006] The technical solution of the present invention is as follows:
[0007] A method for predicting transcription factor-target gene interaction, comprising the steps of:
[0008] Extract the connection between transcription factors and target genes from the gene database and the connection between genes and diseases from the disease database to construct a first heterogeneous network with three node types;
[0009] Assuming that the node between the transcription factor and the target gene is TG, the first heterogeneous network and the node TG form a second heterogeneous network;
[0010] Generate sample paths from the second heterogeneous network using a meta-path-based random walk method;
[0011] Building a graph embedding model based on the sample path, given a sample pair of nodes, using the graph embedding model to connect the embedding vectors of the sample pair of nodes as input of a fully connected layer to obtain the result of the node;
[0012] According to the result, the final embedding of the node is obtained from the second heterogeneous network, and a transcription factor embedding matrix F and a target gene embedding matrix G are formed respectively. The dot product of the two matrices is calculated to obtain a prediction score matrix R, where the value of the i-th row and the j-th column represents the prediction score of the i-th transcription factor and the j-th target gene.
[0013] In the method for predicting transcription factor-target gene interaction, the node types include transcription factors, target genes, and diseases.
[0014] In the method for predicting transcription factor-target gene interaction, the node TG is set to a vector of constant 1.
[0015] In the method for predicting transcription factor-target gene interaction, the representation of the metapath is in Represents two nodes of different types V1 and V n The connection between them.
[0016] The method for predicting transcription factor-target gene interaction calculates the probability p of the k-th step transfer according to the representation of the meta-path, and the formula is as follows:
[0017]
[0018] in Representation Node V t+1 Type of adjacent nodes.
[0019] In the method for predicting transcription factor-target gene interactions, the graph embedding model consists of a first neural network model and a second neural network model; the first neural network model includes constructing a model and obtaining node embedding through the model; the second neural network model is used to calculate the similarity between a pair of node embeddings.
[0020] The method for predicting transcription factor-target gene interactions, for the first neural network model, given a heterogeneous network G = (V, E, T), where |V|>1, the heterogeneous network learns node embeddings with different types of nodes and connections by maximizing the probability that the node has a heterogeneous set of neighbors, and the objective function is expressed as:
[0021]
[0022] Where N t(υ) Represents the neighbors of node v in a heterogeneous context of different node types;
[0023] For the second neural network model, the embedding vector of each node is extracted from the hidden layer weight matrix of the first neural network model. Given a pair of nodes, their embedding vectors are concatenated as input, and the result is obtained as the model loss. The objective function (1) becomes function (2):
[0024]
[0025] Where F represents a fully connected layer, g represents the concatenation of two node embeddings, and σ(x) is calculated as follows:
[0026]
[0027] A model for predicting transcription factor-target gene interaction, wherein the model for predicting transcription factor-target gene interaction is constructed using the method for predicting transcription factor-target gene interaction as described above.
[0028] The model for predicting transcription factor-target gene interaction, wherein the model for predicting transcription factor-target gene interaction is used to make a BioTGI tool, and the BioTGI tool is used to predict the interaction between potential TF and target gene.
[0029] Beneficial effect: The present invention provides a method and model for predicting transcription factor-target gene interaction, wherein the method for predicting transcription factor-target gene interaction is based on a random walk heterogeneous graph embedding algorithm to predict the potential interaction relationship between TF and target gene, a random walk method is used to generate node statements in a heterogeneous graph, and then a sliding window method is used to extract training samples, and the features of the nodes are generated by the heterogeneous graph embedding algorithm, thereby predicting the undiscovered interaction relationship between TF and target gene. In addition, since the present invention adds TG nodes in the heterogeneous graph, in the process of using random walk to generate sample paths, a probability is set when selecting TG or disease nodes, and TG or disease is selected with a certain probability, which can better capture the node information of the heterogeneous network while solving the cold start problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A flow chart of a method for predicting transcription factor-target gene interaction according to the present invention;
[0031] Figure 2 A schematic diagram of building a heterogeneous network according to the present invention;
[0032] Figure 3 A schematic diagram of a setting element path of the present invention;
[0033] Figure 4 A schematic diagram of a random walk method based on a meta-path of the present invention;
[0034] Figure 5 It is a schematic diagram of the heterogeneous node embedding model of the present invention;
[0035] Figure 6 Schematic diagram of the scoring matrix of the present invention. DETAILED DESCRIPTION
[0036] The present invention provides a method and model for predicting transcription factor-target gene interaction. In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0037] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "back" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first" and "second" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features.
[0038] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as generally understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as herein.
[0039] In order to detect reliable interactions between TFs and target genes, different strategies have been developed based on in vitro and in vivo experiments to obtain transcription factor binding site (TFBS) spectra. For example, Chip-X technology (chromatin immunoprecipitation technology) (such as chip-chip, chip-seq, chip-pet technology) separates the binding sites of TFs, and then sequences the in vivo immunoprecipitated DNA enriched by antibody immunoprecipitation to ultimately determine the target genes of TFs. As an alternative technology for identifying the association between TFs and targets, DamID (DNA adenine methyltransferase identification) methylates the adenine bases in the GATC motif near the TF target interaction site, and these methylated sequences are subsequently amplified and detected to identify TF-target interactions. However, these technologies consume a lot of manpower and material resources, so it is necessary to find more convenient and less time-consuming methods to infer the target genes corresponding to TFs.
[0040] As a supplement to the use of tests, existing researchers have developed computational methods to predict TFBS and thus predict the interaction between TF and target genes. One of them is deep learning technology directly applied to DNA sequences. For example: KEGRU extracts k-mer features of DNA sequences and constructs a bidirectional gated recurrent unit network to learn embedding to predict potential motifs as binding sites for TFs. However, this prediction based on TFBS as intermediate data in the computational pipeline still cannot directly derive the results of TF-target interactions. To bridge this gap, Yang et al. proposed the GripDL model, which consists of a convolutional neural network (CNN) that learns embedded features from in situ hybridization (ISH) images of genes (including TFs and their target genes) to infer potential interactions between them. However, gene expression image data is still very expensive for some research goals and is difficult to be widely adopted in actual research. Lin et al. used collaborative filtering technology based on three-factor decomposition and used related protein-protein interactions as training data to predict potential target genes for specific TFs. However, this method cannot consider the connection between TFs and their target genes more comprehensively.
[0041] Based on this, please refer to the attached Figure 1 The present invention provides a method for predicting transcription factor-target gene interaction, which specifically comprises the steps of:
[0042] Step S10: constructing a first heterogeneous network:
[0043] Extract the connection between transcription factors and target genes from the gene database and the connection between genes and diseases from the disease database to construct a first heterogeneous network with three node types;
[0044] Step S20: The first heterogeneous network and the node TG form a second heterogeneous network:
[0045] Assuming that the node between the transcription factor and the target gene is TG, the first heterogeneous network and the node TG form a second heterogeneous network;
[0046] Step S30: Generate a sample path from the second heterogeneous network:
[0047] Generate sample paths from the second heterogeneous network using a meta-path-based random walk method;
[0048] Step S40: construct a graph embedding model based on the sample path:
[0049] Building a graph embedding model based on the sample path, given a sample pair of nodes, using the graph embedding model to connect the embedding vectors of the sample pair of nodes as input of a fully connected layer to obtain the result of the node;
[0050] Step S50: predicting score:
[0051] According to the result, the final embedding of the node is obtained from the second heterogeneous network, and a transcription factor embedding matrix F and a target gene embedding matrix G are formed respectively. The dot product of the two matrices is calculated to obtain a prediction score matrix R, where the value of the i-th row and the j-th column represents the prediction score of the i-th transcription factor and the j-th target gene.
[0052] The method for predicting the interaction between transcription factor and target gene is used to predict the interaction mode of TF-target gene interaction, which can more comprehensively consider the connection between TF and its target gene, and the prediction method is simple, convenient and time-saving. Since the present invention more comprehensively considers the association between TF and its target gene and the association between them and the disease, its prediction effect is significantly improved compared with the existing ones.
[0053] Specifically, Figure 2 As shown, the connection between transcription factors and target genes is extracted from the gene database, and the connection between genes and diseases is extracted from the disease database, and finally a first heterogeneous network with three different node types is constructed; wherein the node types include transcription factors (TFs), target genes, and diseases.
[0054] It is a challenge to find a target gene corresponding to a new TF (cold start problem) without any known target gene. To solve the cold start problem in heterogeneous graphs (HG), the present invention assumes that there is always a node associated with the TF and the target gene. Therefore, when we sample the path using the random walk strategy, we can sample this new node. Based on the assumption, a node called TG is added to the sample path, and the node TG is the node between the transcription factor and the target gene, such as Figure 3In the present invention, the TG node is regarded as a universal entity between TF and target gene and shares the same information with them, so the embedding of the TG node is set to a vector of constant 1.
[0055] In some implementations, a meta-path scheme is set, and the meta-path is represented as in Represents two nodes of different types V1 and V n Given a meta-path scheme, the probability p of the k-th step transition can be calculated, such as Figure 4 As shown, the formula is as follows:
[0056]
[0057] in Representation Node V t+1 Under the guidance of the meta-path, the explicit connections can be captured, and entities without explicit connections in heterogeneous networks will also be connected through the meta-path as the random walk proceeds. Therefore, the method of predicting transcription factor-target gene interaction of the present invention can more comprehensively consider the connection between TF and its target gene.
[0058] In this embodiment, a graph embedding model is constructed based on the path samples generated by random walk, and the graph embedding model is composed of a first neural network model and a second neural network model; the first neural network model includes constructing a model and obtaining node embedding through the model; the second neural network model calculates the similarity between a pair of node embeddings in a new way, such as Figure 5 shown.
[0059] Specifically, the first neural network model consists of two parts, namely, building a model and obtaining node embedding through the model. The main data required is the parameters learned by the model, that is, the weights of the hidden layer matrix.
[0060] In this embodiment, for the first neural network model, given a heterogeneous network G = (V, E, T), where |V|>1, the heterogeneous network learns node embeddings with different types of nodes and connections by maximizing the probability that the node has a heterogeneous set of neighbors. Mathematically, the objective function is defined as follows:
[0061]
[0062] Where N t(υ) Represents the neighbors of node v in a heterogeneous context of different node types;
[0063] And p(c t |v;θ) is defined by the soft-max function:
[0064]
[0065] That is to say, what is really needed in this embodiment is the parameters learned by the first neural network model. Therefore, the present invention develops another neural network model (second neural network model) to optimize the embedding learned by the first neural network model.
[0066] Specifically, the present invention utilizes the weights of the hidden layer matrix of the first neural network model to obtain the embedding vector of each node; when a sample pair of nodes is given, their node embedding vectors can be connected as the input of the fully connected layer to obtain the result of this pair of nodes; and then the final result is regarded as the training loss of the two models.
[0067] In this implementation, in order to optimize effectively, a negative sampling method is used, so function (1) can be transformed into formula (2):
[0068]
[0069] Where F represents a fully connected layer, g represents the concatenation of two node embeddings, and σ(x) is calculated as follows:
[0070]
[0071] like Figure 6 As shown, in the last step, this embodiment obtains the final embedding of the node from the second heterogeneous network, forms the TF embedding matrix F and the target gene embedding matrix G respectively, and calculates the dot product of the two matrices to obtain a prediction score matrix R, where the value of the i-th row and the j-th column represents the prediction score of the i-th TF and the j-th target gene.
[0072] In this embodiment, the embedding of the middle layer is superimposed as the input of a fully connected layer, and the obtained value is used as the loss of the second heterogeneous network for further training. The prediction effect of the model is improved to a certain extent. The method and model for predicting transcription factor-target gene interaction of the present invention have an AUC value of 85.28% in the task of predicting TF and target gene interaction.
[0073] It should be noted that the AUC value is the area under the ROC curve, and its formula is:
[0074]
[0075] Among them, e + represents the positive sample set, e - represents the negative sample set, rank e Represents the prediction score ranking of edge e.
[0076] In addition, the present invention also provides a model for predicting transcription factor-target gene interaction, wherein the model for predicting transcription factor-target gene interaction is constructed using the method for predicting transcription factor-target gene interaction as described above.
[0077] Based on this model, the BioTGI tool is made using the model for predicting transcription factor-target gene interactions. The BioTGI tool is used to predict the interaction between potential TFs and target genes. Specifically, users can query potential target genes by providing TFs. When users provide the name of TFs, BioTGI will train the entire model and eventually obtain a score vector. Each value represents the possibility of interaction with all target genes in the data set. The larger the value, the greater the probability of interaction between TFs and corresponding target genes. Therefore, users can understand which genes are most likely to become target genes for the input TFs. Similarly, users can also query potential TFs for target genes, and the calculation process is the same as that for calculating the target genes for TFs.
[0078] In summary, the present invention provides a method and model for predicting transcription factor-target gene interactions, wherein the method for predicting transcription factor-target gene interactions is based on a random walk heterogeneous graph embedding algorithm to predict the potential interaction relationship between TF and target genes, using a random walk method to generate node statements in a heterogeneous graph, and then using a sliding window method to extract training samples, and generating node features through a heterogeneous graph embedding algorithm, thereby predicting the undiscovered interaction relationship between TF and target genes. In addition, since the present invention adds TG nodes to the heterogeneous graph, in the process of using random walks to generate sample paths, when selecting TG or disease nodes, a probability is set, and TG or disease is selected with a certain probability, while solving the cold start problem, the node information of the heterogeneous network can be better captured.
[0079] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A method for predicting transcription factor-target gene interactions, characterized in that The specific steps include: Extract the connection between transcription factors and target genes from the gene database and the connection between genes and diseases from the disease database to construct a first heterogeneous network with three node types; Assuming that the node between the transcription factor and the target gene is TG, the first heterogeneous network and the node TG form a second heterogeneous network; Generate sample paths from the second heterogeneous network using a meta-path-based random walk method; Building a graph embedding model based on the sample path, given a sample pair of nodes, using the graph embedding model to connect the embedding vectors of the sample pair of nodes as input of a fully connected layer to obtain the result of the node; Obtaining the final embedding of the node from the second heterogeneous network according to the result, forming a transcription factor embedding matrix F and a target gene embedding matrix G respectively, and calculating the dot product of the two matrices to obtain a prediction score matrix R, wherein the value of the i-th row and the j-th column represents the prediction score of the i-th transcription factor and the j-th target gene; The meta-path is represented as in Represents two nodes of different types V1 and V n The relationship between According to the representation of the meta-path, the probability p of the k-th step transfer is calculated, and the formula is as follows: in Representation Node V t+1 Type of adjacent nodes; The graph embedding model consists of a first neural network model and a second neural network model; the first neural network model includes constructing a model and obtaining node embedding through the model; the second neural network model is used to calculate the similarity between a pair of node embeddings.
2. The method for predicting transcription factor-target gene interaction according to claim 1, characterized in that The node types include transcription factors, target genes, and diseases.
3. The method for predicting transcription factor-target gene interaction according to claim 1, characterized in that: The node TG is set to a vector of constant 1.
4. The method for predicting transcription factor-target gene interaction according to claim 1, characterized in that: For the first neural network model, given a heterogeneous network G = (V, E, T), where |V|>1, the heterogeneous network learns node embeddings with different types of nodes and connections by maximizing the probability that the node has a heterogeneous set of neighbors. The objective function is expressed as: Where N t(v) Represents the neighbors of node v in a heterogeneous context of different node types; For the second neural network model, the embedding vector of each node is extracted from the hidden layer weight matrix of the first neural network model. Given a pair of nodes, their embedding vectors are concatenated as input, and the result is obtained as the model loss. The objective function (1) becomes function (2): Where F represents a fully connected layer, g represents the concatenation of two node embeddings, and σ(x) is calculated as follows:
5. A model for predicting transcription factor-target gene interactions, characterized in that: The model for predicting transcription factor-target gene interaction is constructed using the method for predicting transcription factor-target gene interaction as described in any one of claims 1-4.
6. The model for predicting transcription factor-target gene interaction according to claim 5, characterized in that The model for predicting transcription factor-target gene interaction is used to prepare a BioTGI tool, and the BioTGI tool is used to predict the interaction between potential TF and target gene.
Citation Information
Patent Citations
Built-in drug target interaction prediction method based on heterogeneous network
CN109887540A
Disease gene prediction method based on rapid network embedding
CN111540405A