Method and apparatus for identifying transcription factors of tissue genes
By constructing a fully connected directed graph and reconstructing an undirected role graph using a graph neural network model, transcription factors of tissue genes are identified, solving the problem of poor accuracy in existing technologies and achieving higher recognition accuracy and dynamic recognition capabilities.
Patent Information
- Application Number
- CN202511145227.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing technologies have poor accuracy in identifying transcription factors in tissue genes, especially in novel cell types or diseases where it is difficult to distinguish between direct and indirect regulatory relationships, and deep learning algorithms are susceptible to noise.
By obtaining the initial gene expression matrix of tissue genes, a fully connected directed graph is constructed and converted into an undirected role graph. The graph neural network model is then used to reconstruct the graph, generating a directed gene regulation network to identify transcription factors.
It improves the accuracy of transcription factor recognition, effectively distinguishes between direct and indirect regulatory relationships, and meets the dynamic recognition needs of adaptive biological tissue gene regulation.
Smart Images

Figure CN120727110B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bioinformatics, in particular to a transcription factor identification method and device for tissue genes. BACKGROUND
[0002] With the development of single-cell RNA sequencing (scRNA-seq) technology, the use of gene regulatory network (GRN) inference has become a key means to deeply understand cell fate determination, disease mechanism and biological process regulation. At present, in order to accurately infer the regulatory relationship of tissue genes, prior knowledge or high-quality labeled data is usually used for deep learning in the identification of transcription factors. However, prior knowledge cannot meet the application in new cell types or diseases, and the high sparsity of single-cell data will make the deep learning algorithm susceptible to noise, making it difficult to distinguish direct and indirect regulatory relationships, greatly reducing the accuracy of transcription factor identification, and failing to adapt to the dynamic identification needs of biological tissue gene regulation. Therefore, there is an urgent need for a transcription factor identification method for tissue genes to solve the above technical problems. SUMMARY
[0003] Therefore, the present application provides a transcription factor identification method and device for tissue genes, which mainly aims to solve the problem of poor accuracy of existing transcription factor identification for tissue genes.
[0004] According to one aspect of the present application, a transcription factor identification method for tissue genes is provided, comprising:
[0005] obtaining an initial gene expression matrix of tissue genes, and constructing a full-connection directed graph based on the initial gene expression matrix;
[0006] After converting the full-connection directed graph into an undirected role graph, reconstructing the undirected role graph and a feature matrix based on a graph neural network model that has completed model training to obtain a reconstructed target gene expression matrix, the loss value of the graph neural network model being calculated after dynamically removing edges;
[0007] generating a directed gene regulatory network based on the target gene expression matrix, and identifying transcription factors for the tissue genes based on the directed gene regulatory network.
[0008] Further, before the full-connection directed graph is converted into an undirected role graph and the undirected role graph is reconstructed based on a graph neural network model that has completed model training to obtain a reconstructed target gene expression matrix, the method further comprises:
[0009] obtaining a sample data set, the sample data set comprising an undirected role graph sample and a feature matrix sample;
[0010] A two-layer graph neural network is constructed, a first layer network of the two-layer graph neural network is used to aggregate neighbor node features, and a second layer network is used to generate a high-order feature representation;
[0011] During model training of the two-layer graph neural network based on the sample data set, all network edges are traversed, and when the loss value is determined to be the minimum when the target network edge is deleted, the model training of the two-layer graph neural network is completed, and a graph neural network model after model training is obtained.
[0012] Further, the traversal of all network edges and the completion of the model training of the two-layer graph neural network when the loss value is determined to be the minimum when the target network edge is deleted include:
[0013] An initial loss reference value is determined;
[0014] The target network edge is deleted, and the loss value of the two-layer graph neural network corresponding to the deletion of the target network edge is calculated;
[0015] When the loss value is less than the initial loss reference value, the initial loss reference value is updated based on the loss value, and the network edge is deleted;
[0016] When the loss value is greater than or equal to the initial loss reference value, the target network edge is updated, and the step of calculating the loss value of the two-layer graph neural network corresponding to the deletion of the target network edge is re-executed.
[0017] Further, after the fully connected directed graph is constructed based on the initial gene expression matrix, the method further includes:
[0018] The nodes in the fully connected directed graph are split to obtain input nodes and output nodes;
[0019] The input nodes and the output nodes are encoded based on a bidirectional encoder to obtain the undirected role graph and the feature matrix.
[0020] Further, before the initial gene expression matrix of the tissue gene is obtained, the method further includes:
[0021] The ribonucleic acid sequence data of the tissue gene is determined;
[0022] The ribonucleic acid sequence data of low-expression genes and abnormal cell genes in the ribonucleic acid sequence data is deleted based on a preset pipeline, and the ribonucleic acid sequence data after deletion is normalized and standardized to obtain gene data;
[0023] The initial gene expression matrix is constructed based on the gene data.
[0024] Further, the constructing a full connection directed graph based on the initial gene expression matrix comprises:
[0025] Filtering network edges in the full connection directed graph based on a Pearson correlation coefficient or a protein interaction coefficient to obtain a full connection directed graph to be converted into an undirected graph.
[0026] Further, the identifying a transcription factor of the tissue gene based on the directed gene regulation network comprises:
[0027] Determining an out-degree of the tissue gene according to the directed gene regulation network;
[0028] Determining a maximum out-degree in the directed gene regulation network through a horizontal comparison of the out-degree, and determining the transcription factor of the tissue gene based on a cell type corresponding to the maximum out-degree.
[0029] According to another aspect of the present application, a device for identifying a transcription factor of a tissue gene is provided, comprising:
[0030] An acquisition module configured to acquire an initial gene expression matrix of a tissue gene, and construct a full connection directed graph based on the initial gene expression matrix;
[0031] A reconstruction module configured to, after converting the full connection directed graph into an undirected role graph, reconstruct the undirected role graph and a feature matrix based on a graph neural network model that has completed model training, to obtain a reconstructed target gene expression matrix, a loss value of the graph neural network model being calculated after edges are dynamically removed;
[0032] An identification module configured to generate a directed gene regulation network based on the target gene expression matrix, and identify a transcription factor of the tissue gene based on the directed gene regulation network.
[0033] Further, the device further comprises a construction module, a training module,
[0034] The acquisition module is further configured to acquire a sample data set, the sample data set comprising an undirected role graph sample and a feature matrix sample;
[0035] The construction module is configured to construct a two-layer graph neural network, a first layer network of the two-layer graph neural network being configured to aggregate neighbor node features, and a second layer network being configured to generate high-order feature representations;
[0036] The training module is configured to, in a model training process of the two-layer graph neural network based on the sample data set, traverse all network edges, and complete the model training of the two-layer graph neural network when a loss value is determined to be the minimum after a target network edge is removed, to obtain a graph neural network model that has completed model training.
[0037] Further, the training module is specifically configured to determine an initial loss benchmark value, delete a target network edge, and calculate a loss value of the two-layer graph neural network corresponding to the deletion of the target network edge; when the loss value is less than the initial loss benchmark value, update the initial loss benchmark value based on the loss value, and delete the network edge; when the loss value is greater than or equal to the initial loss benchmark value, update the target network edge, and re-execute the step of calculating the loss value of the two-layer graph neural network corresponding to the deletion of the target network edge.
[0038] Further, the device further comprises:
[0039] The splitting module is configured to split the nodes in the full-connection directed graph to obtain input nodes and output nodes.
[0040] The encoding module is configured to encode the input nodes and the output nodes based on a bidirectional encoder to obtain the undirected role graph and the feature matrix.
[0041] Further, the device further comprises a determination module and a processing module.
[0042] The determination module is configured to determine ribonucleic acid sequence data of tissue genes.
[0043] The processing module is configured to delete, based on a preset pipeline, ribonucleic acid sequence data of low-expression genes and abnormal cell genes in the ribonucleic acid sequence data, and perform normalization and standardization processing on the deleted ribonucleic acid sequence data to obtain gene data.
[0044] The construction module is further configured to construct the initial gene expression matrix based on the gene data.
[0045] Further, the construction module is further configured to filter network edges in the full-connection directed graph based on a Pearson correlation coefficient or a protein interaction coefficient to obtain a full-connection directed graph to be converted into an undirected graph.
[0046] Further, the identification module is specifically configured to determine an out-degree of the tissue gene according to the directed gene regulation network, determine a maximum out-degree in the directed gene regulation network through a horizontal comparison of the out-degree, and determine a transcription factor of the tissue gene based on a cell type corresponding to the maximum out-degree.
[0047] According to another aspect of the present application, a storage medium is provided, and the storage medium stores at least one executable instruction, and the executable instruction causes a processor to perform operations corresponding to the above-mentioned transcription factor identification method of tissue genes.
[0048] According to still another aspect of the present application, a terminal is provided, comprising a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface accomplish communication with each other through the communication bus;
[0049] The memory is configured to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the above-mentioned transcription factor identification method of tissue genes.
[0050] By means of the above technical solution, the technical solution provided by the embodiments of the present application has at least the following advantages:
[0051] The present application provides a transcription factor identification method and device of tissue genes. Compared with the prior art, the embodiments of the present application obtain an initial gene expression matrix of tissue genes, and construct a fully connected directed graph based on the initial gene expression matrix. After converting the fully connected directed graph into an undirected role graph, the graph neural network model trained based on the undirected role graph and the feature matrix is reconstructed to obtain a reconstructed target gene expression matrix, and the loss value of the graph neural network model is calculated after the edges are dynamically removed. Based on the target gene expression matrix, a directed gene regulation network is generated, and based on the directed gene regulation network, the transcription factor of the tissue genes is identified. The network reconstruction method ensures the retention effectiveness of important edges in the regulation network, which can improve the effectiveness of identifying the complex interaction between genes, greatly improve the accuracy of distinguishing direct and indirect regulation relationships, meet the dynamic identification needs of adaptive biological tissue gene regulation, and thus improve the identification accuracy of transcription factors.
[0052] The above description is only a summary of the technical solutions of the present application. In order to enable the technical means of the present application to be more clearly understood, the following detailed description of the embodiments of the present application can be implemented according to the content of the description, and in order to enable the above and other purposes, features and advantages of the present application to be more obvious and easy to understand, the following detailed description of the embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0053] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered as limiting the present application. Moreover, the same reference numerals are used to represent the same components throughout the drawings. In the drawings:
[0054] Figure 1 A flow chart of a transcription factor identification method of tissue genes provided by the embodiments of the present application is shown;
[0055] Figure 2 A directed gene regulation network reconstruction flow chart provided by the embodiments of the present application is shown;
[0056] Figure 3 A composition block diagram of a transcription factor recognition device for an organization gene is shown;
[0057] Figure 4 A structural schematic diagram of a terminal is shown. DETAILED DESCRIPTION
[0058] Exemplary embodiments of the present disclosure will be described in detail with reference to the drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood, and the scope of the present disclosure can be accurately conveyed to those skilled in the art.
[0059] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to such a process, method, product, or apparatus.
[0060] The embodiments of the present application provide a transcription factor recognition method for an organization gene, as shown in the method comprises: Figure 1
[0061] 101, obtaining an initial gene expression matrix of an organization gene, and constructing a full-connection directed graph based on the initial gene expression matrix.
[0062] In this embodiment, the current execution entity, acting as the processing end for identifying transcription factors in tissue genes, can be a terminal device or a cloud server, etc., to obtain the initial gene expression matrix of tissue genes. The tissue genes can be the gene content of tissue cells from different species in a large-scale spatiotemporal omics dataset, obtained through different dimensional segmentation, or the gene content of tissue cells collected in real time; this embodiment does not impose specific limitations. The initial gene expression matrix is a matrix used to characterize tissue genes in both cellular and feature dimensions; it is a descriptive form of tissue genes using matrices from the mathematical field, and this embodiment does not impose specific limitations. Furthermore, the fully connected directed graph is a directed structural graph containing gene nodes and potential regulatory edges. Specifically, genes can be represented as nodes in the directed structural graph, and potential regulatory relationships can be represented as edges in the directed structural graph to construct a fully connected directed graph.
[0063] 102. After converting the fully connected directed graph into an undirected role graph, the undirected role graph and the feature matrix are reconstructed based on the graph neural network model that has been trained, to obtain the reconstructed target gene expression matrix.
[0064] In this embodiment, after obtaining the fully connected directed graph, the current execution end transforms it into an undirected role graph in order to use it as input to the neural network model. The edges in the fully connected directed graph are directional, representing potential directional regulatory relationships between different gene nodes. Conversely, the edges in the undirected role graph are non-directional, representing potential non-directional regulatory relationships between different gene nodes. This embodiment does not impose specific limitations on the transformation. Furthermore, the current execution end reconstructs the target gene expression matrix based on the undirected role graph and the feature matrix. Here, the feature matrix represents the spatially derived matrix content of the expression data of tissue genes at gene sites or cells. This matrix, combined with the adjacency matrix corresponding to the undirected role graph (obtained through transformation of the undirected role graph), serves as input to the graph neural network model for reconstruction. In this embodiment, the loss value of the graph neural network model is calculated after dynamically removing edges. That is, during model training, the loss value obtained by training the model is calculated by deleting each edge in the graph one by one. At this time, dynamic edge removal means that in each round of training, one edge is selected for deletion and the loss value of this round is calculated until the loss value meets the preset condition, thus completing the model training. The graph neural network model is a two-layer graph neural network (GNN). The graph neural network model consists of nodes (such as tissue genes) and edges (such as regulatory relationships), which can be represented by an adjacency matrix. This embodiment does not impose specific limitations.
[0065] 103. generating a directed gene regulatory network based on the target gene expression matrix, and identifying transcription factors of the tissue genes based on the directed gene regulatory network.
[0066] In the embodiments of the present application, after the current execution end obtains the reconstructed target gene expression matrix, the target gene expression matrix is generated, and then an optimized target gene expression matrix is used to generate a directed network diagram, i.e., a directed gene regulatory network GRN (Gene Regulatory Network). Further, based on the directed gene regulatory network, transcription factors of the tissue genes are identified, thereby improving the identification accuracy of transcription factors in the field of biology. When generating a directed gene regulatory network based on a target gene expression matrix, the same method as that for constructing a fully connected directed graph using an initial gene expression matrix can be used for construction, thereby obtaining a directed gene regulatory network containing nodes (genes) and edges (potential regulatory relationships).
[0067] In another embodiment of the present application, in order to further illustrate and limit, before the step of converting the fully connected directed graph into an undirected role graph, and based on the graph neural network model that has completed model training, reconstructing the undirected role graph to obtain the reconstructed target gene expression matrix, the method further comprises:
[0068] obtaining a sample data set;
[0069] constructing a two-layer graph neural network;
[0070] In the process of model training of the two-layer graph neural network based on the sample data set, all network edges are traversed, and when the loss value is determined to be the minimum after deleting the target network edge, the model training of the two-layer graph neural network is completed, thereby obtaining a graph neural network model that has completed model training.
[0071] In order to achieve the purpose of optimizing the potential regulation relationship of genes based on the graph neural network model, thereby improving the accuracy of identifying transcription factors, the current execution end first acquires a sample data set, the sample data set includes an undirected role graph sample and a feature matrix sample, at this time, the feature matrix sample can be acquired from the SRT database, the SRT database includes STOmics, SOAR, SpatialDB, CROST, 10x Genomics website (which can be obtained through 10x Genomics) and Census (which can be obtained through Cellxgene) and the like, for example, containing 96,700,729 cells / points from 7,367 tissue sections, covering 365 tissue cells, including lung, skin, brain, liver, kidney, spinal cord and embryo, and covering normal tissues and various diseases, such as pancreatic ductal adenocarcinoma, amyotrophic lateral sclerosis, non-small cell lung cancer and hepatocellular carcinoma, the embodiments of the present application are not limited. At the same time, a two-layer graph neural network is constructed, at this time, the first layer network of the two-layer graph neural network is used to aggregate neighbor node features, and the second layer network of the two-layer graph neural network is used to generate high-order feature representation, the embodiments of the present application are not limited.
[0072] It should be noted that in the process of training the two-layer graph neural network based on the sample data set, the loss function used is the minimum mean square error (MSE) loss function, which is expressed as:
[0073] ;
[0074] Where N is the number of samples, is the true value, is the output obtained by the model, so as to calculate the loss value based on the loss function. At the same time, in the training process, all network edges are traversed, and when the loss value is determined to be the smallest when the target network edge is deleted, the model training of the two-layer graph neural network is completed, and the graph neural network model after the model training is obtained.
[0075] In some embodiments, the encoding process of the graph neural network GNN includes: the aggregation of neighbor node features of the first layer GNN and the generation of high-order feature representation of the second layer GNN. Wherein, the aggregation of neighbor node features of the first layer GNN, the calculation formula is:
[0076] ; wherein, is the weight of the first layer network, which is used to linearly transform the input feature matrix to extract the preliminary feature representation, is the feature of the node obtained by the first layer. is an activation function (which can preferably be a ReLU function as an activation function), i is (without i) v is a network node, and j is a neighbor node. For the feature vector of the node, The feature vectors of the neighboring nodes, For the set of neighboring nodes, This is an aggregation function, which uses the element-wise averaging of the feature vectors of all neighboring nodes as the aggregation operation, as shown below: The second-layer GNN generates higher-order feature representations, calculated using the following formula:
[0077] ; For the weights of the second layer network, For the nodes obtained in the second layer Based on the characteristics, the final output is represented as In one embodiment of this application, the feature matrix X, as input to the graph neural network model, is first derived from the set of its neighbors for each node v. Retrieve messages from the data, and then use aggregation functions to group the nodes. The graph neural network integrates its own features with those of its neighboring nodes. In this case, the preferred activation function for the graph neural network model is the ReLU function, and the preferred loss function is the mean squared error (MSE) function (denoted as...). Then, the gradient of the loss function with respect to the model parameters is calculated using the backpropagation algorithm, and the parameters are updated using the Adam optimizer, such as setting the learning rate to 0.0001, thereby minimizing the loss value.
[0078] In another embodiment of this application, for further explanation and limitation, the step of traversing all network edges and determining that the loss value is minimized by deleting the target network edge completes the model training of the two-layer graph neural network, resulting in a graph neural network model with completed model training, including:
[0079] Determine the initial loss baseline value;
[0080] Delete the target network edge and calculate the loss value of the two-layer graph neural network corresponding to the deleted target network edge;
[0081] When the loss value is less than the initial loss benchmark value, the initial loss benchmark value is updated based on the loss value, and the network edge is deleted.
[0082] When the loss value is greater than or equal to the initial loss benchmark value, the target network edge is updated, and the steps of calculating and deleting the loss value of the two-layer graph neural network corresponding to the target network edge are re-executed.
[0083] To preserve key regulatory relationships and optimize the regulatory network structure, thereby improving network reconstruction accuracy, the current execution end first determines an initial loss baseline value during model training. This initial loss baseline value serves as the initial loss baseline. The configuration can be based on the loss value obtained from the initial model training, or it can be based on empirical values from actual testing; this embodiment does not impose specific limitations. Furthermore, any edge can be selected. As the target network edge, it is deleted to calculate the loss value of the two-layer graph neural network corresponding to the deletion of this target network edge. The model is retrained based on the new graph structure to obtain a new loss value. When the loss value Less than the initial loss benchmark value When, indicate the deletion of the edge. The subsequent network model showed improved data fit. Removing this edge made the network structure fit the real-world GRN better. Therefore, the initial loss baseline was adjusted. The value is updated to When the loss value Greater than or equal to the initial loss baseline value When, explain this side This edge makes a significant contribution to the model's performance; therefore, it is retained. The target network edges are updated to evaluate the impact of each edge on the model. Then, the updated target network edges are deleted. The steps of calculating and deleting the loss values of the two-layer graph neural network corresponding to the updated target network edges are repeated until the model optimization is completed.
[0084] It should be noted that, in order to ensure that the neural network model has the ability to remove noisy connections, the current execution end calculates the change in loss during the new round of model training after deleting edges. .like If the value is less than 0, it indicates a reduction in loss. Therefore, this edge is identified as redundant and deleted. If the value is greater than 0, it indicates an increase in loss. These edges are retained as key regulatory edges, ultimately generating a sparse network that combines prediction accuracy with biological significance. This effectively eliminates noisy connections to highlight the core regulatory relationships.
[0085] In another embodiment of this application, for further explanation and limitation, after constructing a fully connected directed graph based on the initial gene expression matrix, the method further includes:
[0086] The nodes in the fully connected directed graph are split to obtain input nodes and output nodes;
[0087] The input nodes and the output nodes are encoded based on a bidirectional encoder to obtain the undirected role graph and the feature matrix.
[0088] In order to convert the undirected graph into a graph neural network-based reconstruction, the current execution end first splits the nodes in the full connection directed graph, at this time, the node represents a gene. When splitting, the node corresponding to each gene can be split into a "source node", that is, an input node, and a "target node", that is, an output node, while retaining the regulation direction information. Specifically, the full connection directed graph is represented as , V is a node, and E is an edge. Since the nodes in the directed graph have a direction, when converted into an undirected graph, the node is converted into two nodes, that is, and , The directed edge in is mapped to the undirected edge in , and n is the total number of nodes in the graph neural network.
[0089] In some embodiments, as shown in Figure 2 , after conversion, the directed edge of in the network is represented as in the undirected graph. In addition, the obtained undirected graph can also be converted back into a directed graph, thereby maintaining the integrity and biological significance of the gene regulation network. A reverse conversion method can be used, and the embodiments of the present application are not limited in particular. For the undirected graph G2 output by the graph neural network model, the undirected edge represents the edge in the directed graph, a is a node, and b-n is a target node.
[0090] In another embodiment of the present application, in order to further illustrate and limit, before the step of obtaining the initial gene expression matrix of the tissue gene, the method further comprises:
[0091] determining the ribonucleic acid sequence data of the tissue gene;
[0092] deleting the ribonucleic acid sequence data of low expression genes and abnormal cell genes in the ribonucleic acid sequence data based on a preset pipeline, and performing normalization and standardization processing on the deleted ribonucleic acid sequence data to obtain gene data;
[0093] constructing the initial gene expression matrix based on the gene data.
[0094] In order to effectively construct a regulatory network based on a graph structure and characteristics of an organizational gene, and improve the effectiveness of transcription factor identification, the current execution end determines the ribonucleic acid sequence data of the organizational gene in advance. The ribonucleic acid sequence data of the organizational gene can be single-cell RNA sequencing data or bulk cell RNA sequencing data, and a preset pipeline is a Scanpy pipeline in the field of bioinformatics, which is used to remove low-expression genes and abnormal cells, normalize and standardize the ribonucleic acid sequence data after removal, and obtain gene data. At this time, logarithmic normalization and Z-score standardization can be used to process the gene data, for example, more than three cell expressions are finally obtained, or cells expressing more than 200 genes are retained, that is, only low-expression genes and abnormal cells are removed, and an initial gene expression matrix is constructed based on the gene data. When constructing the initial gene expression matrix, that is, using the gene data to construct a matrix of cells and feature dimensions, row i in the matrix is a cell, column j is a gene, and element v is gene expression. The embodiments of the present application are not limited.
[0095] In another embodiment of the present application, in order to further illustrate and limit, the step of constructing a fully connected directed graph based on the initial gene expression matrix includes:
[0096] Based on the Pearson correlation coefficient or the protein-protein interaction coefficient, the network edges in the fully connected directed graph are screened to obtain a fully connected directed graph to be converted into an undirected graph.
[0097] In order to make the structure graph with potential regulatory relationship more effective, the current execution end can screen the network edges in the fully connected directed graph obtained after the initial conversion into a fully connected directed graph through the Pearson correlation coefficient or the protein-protein interaction coefficient, to obtain an optimal fully connected directed graph. The Pearson correlation coefficient is an index used to measure the degree of linear correlation between two continuous variables in statistics, which can be pre-configured based on the characteristics of different organizational genes, or configured based on the scene requirements, without specific limitation. The protein-protein interaction coefficient (PPIC) is an index for quantifying the binding strength or interaction probability between proteins, which can be measured or configured based on the scene requirements in the embodiments of the present application, without specific limitation.
[0098] In another embodiment of the present application, in order to further illustrate and limit, the step of identifying the transcription factor of the organizational gene based on the directed gene regulatory network includes:
[0099] The out-degree of the organizational gene is determined according to the directed gene regulatory network.
[0100] By comparing the out-degree, the maximum out-degree in the directed gene regulatory network is determined, and the transcription factor of the tissue gene is determined based on the cell type corresponding to the maximum out-degree.
[0101] In order to realize the effectiveness of the recognition of the transcription factor, the current execution end determines the out-degree of the tissue gene according to the obtained directed gene regulatory network, at this time, the out-degree (Out-degree) is an index for describing the interaction relationship between a gene and other genes or molecules, which can be determined based on the number of genes in the directed gene regulatory network by regulating other genes or pathways. Specifically, after calculating the out-degree of one tissue gene, the out-degrees of all genes are compared horizontally, the maximum out-degree in the directed gene regulatory network is determined by comparing the out-degrees horizontally, and the transcription factor of the tissue gene is determined based on the maximum out-degree and the cell type corresponding to the tissue gene. At this time, since the gene out-degree (Out-degree) represents the number of interactions between the gene and other genes or molecules in the obtained gene regulatory network, the gene with high out-degree is usually a hub node in the gene regulatory network and plays a key role in a specific cell type. Therefore, the difference between genes is the molecular basis of cell differentiation, and then through selective transcription and translation regulation, a differential gene expression profile is formed to determine different cell types. In the embodiments of the present application, the cell type corresponding to the gene with the highest out-degree is the cell type in which the gene is most active, so the development or function maintenance of this cell type plays an important role, that is, as a cell type-specific transcription factor.
[0102] It should be noted that the method for determining the cell type can be determined by single cell sequencing technology, that is, by Marker gene (marker gene) to identify the cell type, for example, the database CellMarker, PanglaoDB and the like provide query basis for marker genes of different cell types, and the tool ScType and SingleR can also be used for query, that is, the cell type can be automatically identified by using gene expression data, and the embodiments of the present application are not limited specifically.
[0103] In a specific embodiment, as shown in (a) of Figure 2 , the current execution end constructs a fully connected directed graph from the cell gene expression as an initial gene expression matrix, that is, a correlation network, and obtains an undirected role graph and a feature matrix by a bidirectional vector encoding method. As shown in (b) of Figure 2 , (d) of Figure 2 , the gene ... can be split to obtain the "source node", that is, the input node, and the "target node", that is, the output node, which correspond to ... and ... wherein each node can be composed of an input node and an output node . Further, the directed gene regulatory network GRN is obtained by reconstructing the feature matrix and the undirected role graph until the loss value is the minimum, as shown in (c) of Figure 2 . During the training process, the loss value is calculated after deleting any edge, so as to complete the model training when the loss value meets the condition.
[0104] The embodiment of the present application provides a transcription factor recognition method for tissue genes. Compared with the prior art, the embodiment of the present application obtains an initial gene expression matrix of tissue genes, and constructs a full connection directed graph based on the initial gene expression matrix; after converting the full connection directed graph into an undirected role graph, the undirected role graph and a feature matrix are reconstructed based on a graph neural network model which has completed model training, to obtain a reconstructed target gene expression matrix, and the loss value of the graph neural network model is calculated after dynamically removing edges; a directed gene regulatory network is generated based on the target gene expression matrix, and transcription factors of the tissue genes are recognized based on the directed gene regulatory network. The network reconstruction method ensures the retention effectiveness of important edges in the regulatory network, can improve the effectiveness of recognizing the complex interaction between genes, greatly improves the accuracy of distinguishing direct and indirect regulatory relationships, meets the dynamic recognition requirement of adaptive biological tissue gene regulation, and thus improves the recognition accuracy of transcription factors.
[0105] Further, as an implementation of the method shown in the above Figure 1 , the embodiment of the present application provides a transcription factor recognition device for tissue genes, as shown in Figure 3 , which comprises:
[0106] An acquisition module 21 is configured to acquire an initial gene expression matrix of tissue genes, and construct a full connection directed graph based on the initial gene expression matrix;
[0107] A reconstruction module 22 is configured to, after converting the full connection directed graph into an undirected role graph, reconstruct the undirected role graph and a feature matrix based on a graph neural network model which has completed model training, to obtain a reconstructed target gene expression matrix, and the loss value of the graph neural network model is calculated after dynamically removing edges;
[0108] An identification module 23 is configured to generate a directed gene regulatory network based on the target gene expression matrix, and recognize transcription factors of the tissue genes based on the directed gene regulatory network.
[0109] Further, the apparatus further comprises: a constructing module, a training module,
[0110] The acquisition module is further configured to acquire a sample data set, the sample data set comprising an undirected role graph sample and a feature matrix sample.
[0111] The constructing module is configured to construct a two-layer graph neural network, a first layer network of the two-layer graph neural network being configured to aggregate neighbor node features, and a second layer network of the two-layer graph neural network being configured to generate high-order feature representations.
[0112] The training module is configured to, in a model training process of the two-layer graph neural network based on the sample data set, traverse all network edges, and complete the model training of the two-layer graph neural network when a loss value corresponding to a target network edge to be deleted is the minimum, to obtain a graph neural network model after the model training.
[0113] Further, the training module is specifically configured to determine an initial loss reference value, delete the target network edge, and calculate the loss value of the two-layer graph neural network corresponding to the deletion of the target network edge; when the loss value is less than the initial loss reference value, update the initial loss reference value based on the loss value, and delete the network edge; when the loss value is greater than or equal to the initial loss reference value, update the target network edge, and re-execute the step of calculating the loss value of the two-layer graph neural network corresponding to the deletion of the target network edge.
[0114] Further, the apparatus further comprises:
[0115] The splitting module is configured to split nodes in the fully connected directed graph to obtain input nodes and output nodes.
[0116] The encoding module is configured to encode the input nodes and the output nodes based on a bidirectional encoder to obtain the undirected role graph and the feature matrix.
[0117] Further, the apparatus further comprises: a determining module, a processing module,
[0118] The determining module is configured to determine ribonucleic acid sequence data of an organization gene.
[0119] The processing module is configured to delete, based on a preset pipeline, ribonucleic acid sequence data of low-expression genes and abnormal cell genes in the ribonucleic acid sequence data, and perform normalization and standardization processing on the ribonucleic acid sequence data after the deletion to obtain gene data.
[0120] The constructing module is further configured to construct the initial gene expression matrix based on the gene data.
[0121] Further, the construction module is further configured to filter network edges in the full connection directed graph based on a Pearson correlation coefficient or a protein interaction coefficient to obtain a full connection directed graph to be converted into an undirected graph.
[0122] Further, the identification module is specifically configured to determine an out-degree of the tissue gene according to the directed gene regulation network, determine a maximum out-degree in the directed gene regulation network through transverse comparison of the out-degree, and determine a transcription factor of the tissue gene based on a cell type corresponding to the maximum out-degree.
[0123] The embodiment of the present application provides a transcription factor identification device of a tissue gene. Compared with the prior art, the embodiment of the present application obtains an initial gene expression matrix of a tissue gene, and constructs a full connection directed graph based on the initial gene expression matrix. After the full connection directed graph is converted into an undirected role graph, the undirected role graph and a feature matrix are reconstructed based on a graph neural network model that has completed model training, to obtain a reconstructed target gene expression matrix. A loss value of the graph neural network model is calculated after edges are dynamically removed. A directed gene regulation network is generated based on the target gene expression matrix, and a transcription factor of the tissue gene is identified based on the directed gene regulation network. The network reconstruction manner ensures the retention effectiveness of important edges in the regulation network, can improve the effectiveness of identifying complex interactions between genes, greatly improves the accuracy of distinguishing direct and indirect regulation relationships, meets the dynamic identification requirement of adaptive biological tissue gene regulation, and thus improves the identification accuracy of the transcription factor.
[0124] According to an embodiment of the present application, a storage medium is provided, and the storage medium stores at least one executable instruction. The computer executable instruction can execute the transcription factor identification method of the tissue gene in any method embodiment.
[0125] Figure 4 A structural schematic diagram of a terminal according to an embodiment of the present application is shown, and the specific implementation of the terminal is not limited in the embodiments of the present application.
[0126] As shown in the Figure 4 terminal can include a processor 302, a communications interface 304, a memory 306, and a communications bus 308.
[0127] The processor 302, the communications interface 304, and the memory 306 can complete mutual communication through the communications bus 308.
[0128] The communications interface 304 is configured to communicate with network elements of other devices, such as a client or other servers.
[0129] The processor 302 is configured to execute the program 310, and in particular, execute the steps of the above-described embodiments of the transcription factor identification method for tissue genes.
[0130] In particular, the program 310 can include program codes including computer operation instructions.
[0131] The processor 302 can be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the terminal can be the same type of processor, such as one or more CPUs; or can be different types of processors, such as one or more CPUs and one or more ASICs.
[0132] The memory 306 is configured to store the program 310. The memory 306 can include a high-speed RAM memory, and can also include a non-volatile memory such as at least one disk memory.
[0133] The program 310 can be specifically configured to cause the processor 302 to perform the following operations:
[0134] Obtain an initial gene expression matrix of tissue genes, and construct a fully connected directed graph based on the initial gene expression matrix;
[0135] After converting the fully connected directed graph into an undirected role graph, reconstruct the undirected role graph and a feature matrix based on a graph neural network model that has completed model training, to obtain a reconstructed target gene expression matrix, and the loss value of the graph neural network model is calculated after edges are dynamically removed;
[0136] Generate a directed gene regulation network based on the target gene expression matrix, and identify transcription factors of the tissue genes based on the directed gene regulation network.
[0137] It should be apparent to those skilled in the art that the modules or steps of the application described above can be implemented with a general purpose computing device, which can be centralized on a single computing device or distributed on a network of multiple computing devices, and optionally implemented with program codes executable by a computing device, which can be stored in a storage device and executed by a computing device, and in some cases, the steps shown or described can be executed in an order different from that shown, or made into individual integrated circuit modules, or made into a single integrated circuit module. Thus, the present application is not limited to any particular combination of hardware and software.
[0138] The preferred embodiments of the present application described above are intended to be illustrative only and the present application is not limited to the above described preferred embodiments nor the application described therein. Various modifications made to the preferred embodiments of the application will be apparent to those skilled in the art, and this application includes all such modifications without parting from the spirit and scope of the present application.
Claims
1. A method of identifying transcription factors for genes of a tissue, characterized by, The method comprises the following steps: obtaining an initial gene expression matrix of tissue genes, and constructing a full-connection directed graph based on the initial gene expression matrix; after converting the full-connection directed graph into an undirected role graph, reconstructing the undirected role graph and a feature matrix based on a graph neural network model that has completed model training, to obtain a reconstructed target gene expression matrix, and the loss value of the graph neural network model is calculated after dynamically removing edges; generating a directed gene regulation network based on the target gene expression matrix, and identifying transcription factors of the tissue genes based on the directed gene regulation network; before reconstructing the undirected role graph based on the graph neural network model that has completed model training to obtain the reconstructed target gene expression matrix, the method further comprises the following steps: obtaining a sample data set, wherein the sample data set comprises an undirected role graph sample and a feature matrix sample; constructing a two-layer graph neural network, wherein the first layer network of the two-layer graph neural network is used to aggregate neighbor node features, and the second layer network of the two-layer graph neural network is used to generate high-order feature representation; during model training of the two-layer graph neural network based on the sample data set, all network edges are traversed, and when the loss value is determined to be the minimum after deleting a target network edge, the model training of the two-layer graph neural network is completed, to obtain a graph neural network model that has completed model training; after constructing the full-connection directed graph based on the initial gene expression matrix, the method further comprises the following steps: splitting the nodes in the full-connection directed graph to obtain input nodes and output nodes; encoding the input nodes and the output nodes based on a bidirectional encoder to obtain the undirected role graph and the feature matrix.
2. The method of claim 1, wherein, The step of traversing all network edges and completing the model training of the two-layer graph neural network to obtain the graph neural network model that has completed model training when the loss value is determined to be the minimum after deleting a target network edge comprises the following steps: determining an initial loss reference value; deleting the target network edge and calculating the loss value of the two-layer graph neural network corresponding to deleting the target network edge; when the loss value is less than the initial loss reference value, updating the initial loss reference value based on the loss value, and deleting the network edge; when the loss value is greater than or equal to the initial loss reference value, updating the target network edge, and re-executing the step of calculating the loss value of the two-layer graph neural network corresponding to deleting the target network edge.
3. The method of claim 1, wherein, Before obtaining the initial gene expression matrix of the tissue genes, the method further comprises the following steps: determining ribonucleic acid sequence data of tissue genes; based on a preset pipeline, deleting ribonucleic acid sequence data of low-expression genes and abnormal cell genes in the ribonucleic acid sequence data, and performing normalization and standardization processing on the deleted ribonucleic acid sequence data to obtain gene data; constructing the initial gene expression matrix based on the gene data.
4. The method of claim 1, wherein, The step of constructing the full-connection directed graph based on the initial gene expression matrix comprises the following steps: Screening network edges in the full connection directed graph based on a Pearson correlation coefficient or a protein interaction coefficient to obtain a full connection directed graph to be converted into an undirected graph.
5. The method according to any one of claims 1 to 4, characterized in that, The identifying the transcription factor of the tissue gene based on the directed gene regulation network comprises: Determining the out-degree of the tissue gene according to the directed gene regulation network; Determining the maximum out-degree in the directed gene regulation network by transversely comparing the out-degree, and determining the transcription factor of the tissue gene based on the cell type corresponding to the maximum out-degree.
6. An apparatus for recognizing transcription factors of genes of an organism, characterized by, Comprise: An acquisition module configured to acquire an initial gene expression matrix of a tissue gene, and construct a full connection directed graph based on the initial gene expression matrix; A reconstruction module configured to reconstruct a target gene expression matrix after converting the full connection directed graph into an undirected role graph, based on a graph neural network model that has completed model training, the loss value of the graph neural network model being calculated after edges are dynamically removed; An identification module configured to generate a directed gene regulation network based on the target gene expression matrix, and identify the transcription factor of the tissue gene based on the directed gene regulation network; The device further comprises a construction module and a training module, The acquisition module is further configured to acquire a sample data set, the sample data set comprising an undirected role graph sample and a feature matrix sample; The construction module is configured to construct a two-layer graph neural network, the first layer network of the two-layer graph neural network being configured to aggregate neighbor node features, and the second layer network being configured to generate high-order feature representation; The training module is configured to traverse all network edges during model training of the two-layer graph neural network based on the sample data set, and complete the model training of the two-layer graph neural network when the loss value is determined to be the minimum after a target network edge is deleted, to obtain a graph neural network model that has completed model training; The device further comprises: A splitting module configured to split nodes in the full connection directed graph to obtain input nodes and output nodes; An encoding module configured to encode the input nodes and the output nodes based on a bidirectional encoder to obtain the undirected role graph and the feature matrix.
7. A computer readable storage medium having stored thereon computer programs / instructions, characterized in that, The computer program / instructions are executed by the processor to implement the steps of the method of claim 1.
8. A computer device comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program comprises instructions that, when executed by the processor, cause the processor to perform the method of any one of claims 1-7. The processor executes the computer program to implement the steps of the method of claim 1.
Citation Information
Patent Citations
Methods and systems for identifying gene regulatory elements and altering gene regulation and expression
WO2024178321A1
Gene regulatory network inference method based on spatiotemporal transcriptomic data
WO2025025222A1