A method for constructing a gene regulatory network based on a generative flow network
By using sparse Transformer and generative flow networks to extract features from single-cell RNA-seq data, the method addresses the challenge of constructing accurate gene regulatory networks, improving the accuracy and reliability of gene interaction analysis.
Patent Information
- Application Number
- CN202410337864.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-03-24
AI Technical Summary
Existing deep learning-based gene regulation network construction methods cannot effectively capture complex regulatory relationships between genes from sparse single-cell RNA sequencing data, resulting in low construction accuracy.
Combining sparse Transformer and generation stream network, a gene regulation network is constructed through sparse feature extraction and iterative training, and sparse features are extracted using sparse Transformer, and edges are gradually added through generation stream network to construct the probability distribution of the gene regulation network.
It improves the accuracy and construction accuracy of gene regulation networks, can better understand the complex nonlinear relationships between genes, and provides a more accurate gene regulation network model.
Smart Images

Figure CN118212989B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a causal discovery and information fusion method for single-cell RNA sequencing data in bioinformatics and genomics, aiming to discover complex regulatory relationships between different gene pairs in genomics, deepen the understanding of the interactions between different types of genes and transcription factors, and assist in the diagnosis of cell diseases. A method for constructing a gene regulatory network based on a generative flow network is designed. Background Art
[0002] The development of sequencing technology has led to an explosive growth of biological sequencing data. The field of bioinformatics has entered a new era based on sequencing technology, with algorithm models as the means and the development of science and technology as the support. Single-cell RNA sequencing (scRNA-seq) is a high-throughput biotechnology used to measure the number of RNA molecules in a single cell, thereby revealing transcriptomic differences between different cells. This technology has made significant progress in the fields of biomedical research, developmental biology, disease understanding, etc. ScRNA-seq data provides a two-dimensional matrix of cells and genes, where each value represents the expression level of a gene in a cell.
[0003] A gene regulatory network is used to describe the regulatory relationships between genes, aiming to simplify the analysis of complex biological systems. A gene regulatory network consists of transcription factors, genes, and regulatory factors. Each node in the gene regulatory network represents a transcription factor or a gene, and the edge connecting the transcription factor and the target gene represents a regulatory relationship. Gene regulatory networks play a crucial role in cells. Accurately constructing a gene regulatory network can help better understand the signal mechanism within cells. However, due to the complexity of the regulatory relationships in gene regulatory networks, it is difficult to achieve this goal solely through biological experiments. By analyzing variables such as gene expression levels in scRNA-seq data, a gene regulatory network can be constructed in another way. Specifically, a gene regulatory network is a directed graph composed of a set of nodes and directed edges, where each node represents a transcription factor or a gene, and the edge connecting the transcription factor and the target gene represents the regulatory relationship between two genes. Therefore, constructing a gene regulatory network using scRNA-seq data has become a hot issue in current bioinformatics research.
[0004] In recent years, a large number of methods for constructing gene regulatory networks from scRNA-seq data have emerged. These methods can be classified into gene regulatory network construction methods based on traditional machine learning and those based on deep learning. Gene regulatory network construction methods based on traditional machine learning include those based on mutual information, information theory models, ordinary differential equation models, Boolean network models, Bayesian network models, and so on. The above traditional machine learning methods have the advantages of simple model structure principle, easy training, and low time complexity. However, they are difficult to capture the complex non-linear regulatory relationships between genes, resulting in poor construction accuracy. With the development of deep learning technology, in recent years, some deep learning methods have provided a new perspective for the construction of gene regulatory networks, and a large number of gene regulatory network construction methods based on deep learning have emerged, such as those based on convolutional neural networks, graph neural networks, and three-dimensional convolutional neural networks. Such methods can usually process high-dimensional data and capture complex non-linear features between genes. However, due to the high sparsity, low signal-to-noise ratio and other characteristics of scRNA-seq data, the performance of existing deep learning-based gene regulatory network methods is severely restricted. Thus, accurately learning gene regulatory networks using deep learning methods remains a challenging research issue.
[0005] The Generative Flow Network (GFlowNet) is a newly proposed novel deep generative model in recent years and has now been widely applied in many fields such as causal discovery, discrete latent variable modeling, and computational graph scheduling. Since the GFlowNet has achieved good results in various fields, but the existing deep learning-based gene regulatory network construction methods cannot effectively extract the complex sparse features in scRNA-seq data,
[0006] Therefore, how to accurately construct gene regulatory networks from sparse scRNA-seq data remains a highly challenging topic in this field. Summary of the Invention
[0007] Aiming at the problem that the existing gene regulatory network construction methods cannot fully capture the complex regulatory relationships between genes from sparse scRNA-seq data, resulting in low accuracy, the present invention proposes a gene regulatory network construction method based on a generative flow network. This method combines a sparse Transformer with a generative flow network. First, the sparse Transformer is used to perform sparse feature extraction on the input sparse scRNA-seq data. By establishing a sparse association between the data and vectors, sparse features can be effectively extracted. Then, the generative flow network is used to construct a gene regulatory network for the input sparse features by gradually adding edges, and a distribution of the regulatory relationships of the gene regulatory network is obtained through iteration. Finally, when the preset number of iterations is reached, according to the designed information fusion and extraction rules, the final gene regulatory network can be extracted through threshold setting and screening.
[0008] The main idea of implementing the present invention is as follows: The generative flow network has been regarded as one of the most promising methods in multiple fields such as causal discovery, discrete latent variable modeling, and computational graph scheduling. However, the existing gene regulatory network learning methods cannot accurately learn gene regulatory networks from sparse scRNA-seq data. Through the designed gene regulatory network construction method based on the generative flow network, firstly, the sparse Transformer can be used to effectively extract sparse features in scRNA-seq data for subsequent construction of the gene regulatory network. Secondly, through the generative flow network, a latest and promising deep learning method, complex non-linear regulatory relationships between genes can be captured, thereby further improving the accuracy of constructing the gene regulatory network.
[0009] A gene regulatory network construction method based on a generative flow network includes the following steps:
[0010] (Step 1) Data acquisition: To verify the effectiveness of the method proposed in the present invention, experiments are carried out on the sc-RNA-seq dataset of mouse hematopoietic stem cells to evaluate the performance of constructing a gene regulatory network.
[0011] (Step 2) Sparse feature extraction: Since the sc-RNA-seq dataset of mouse hematopoietic stem cells is relatively sparse, the sparse Transformer is used to perform sparse feature extraction on the scRNA-seq data for more accurate subsequent construction of the gene regulatory network.
[0012] (Step 3) Construction of gene regulatory relationships: After the sparse features are extracted, the sparse features are input into the generative flow network model for the construction of the gene regulatory network.
[0013] (Step 4) Fusion and extraction of gene regulatory network: Through multiple iterative updates of the generated flow network model, a probability distribution of gene regulatory relationships associated with the mouse hematopoietic stem cell sc-RNA-seq dataset is obtained, and the final gene regulatory network is extracted by setting a regulatory extraction threshold.
[0014] Compared with the prior art, the present invention has the following obvious advantages and beneficial effects;
[0015] (1) The generated flow network model designed by the present invention can effectively extract sparse features from sparse scRNA-seq data, effectively alleviating the problem that existing methods are difficult to capture complex gene regulatory relationships from sparse scRNA-seq data.
[0016] (2) The gene regulatory network fusion and extraction strategy designed by the present invention can optimize the construction of a more accurate gene regulatory network, improving the accuracy of algorithm learning.
[0017] (3) Experimental results on the sc-RNA-seq dataset of mouse hematopoietic stem cells show that the present invention can effectively construct a gene regulatory network and can be used as an effective tool to provide reference for medical researchers to analyze gene regulatory mechanisms. Description of the Drawings
[0018] Figure 1 : Flowchart of the algorithm involved in this method.
[0019] Figure 2 : Result graph of OverallScore of five methods on the mouse hematopoietic stem cell dataset. Detailed Embodiment
[0020] The following elaborates on the specific embodiments and detailed steps of the present invention. The specific implementation process of the present invention is as Figure 1 shown, specifically including:
[0021] (Step 1) Data acquisition;
[0022] To verify the effectiveness of the algorithm proposed in the present invention, experiments were conducted on the BEELINE dataset. The BEELINE framework is a comprehensive evaluation framework that can be used to evaluate the comprehensive performance of gene regulatory network construction algorithms. This framework provides 7 sc-RNA-seq datasets, among which 5 are datasets on mouse cells and 2 are datasets on human cells. Three sc-RNA-seq datasets of mouse hematopoietic stem cells (mHSC) were selected, namely mHSC-L, mHSC-GM, and mHSC-E. For these three datasets, genes expressed in less than 10% of the cells were removed. Then, the remaining data was normalized. Only the top 1000 standard deviation genes were retained for each cell. The sc-RNA-seq datasets can be obtained at https: / / zenodo.org / record / 3701939.
[0023] (Step 2) Sparse feature extraction;
[0024] A gene regulatory network can be represented as G = <N, Z>, where N is the set of genes (nodes), and X i ∈Z (i = 1, 2,..., n) represents the i-th gene among the n genes, and Z represents the set of regulatory relationships (directed edges) between genes. X j →X i ∈Z represents the directed edge from gene X j to gene X i . A gene regulatory network can be expressed as the joint probability distribution of the gene set N = {X1, X2, X3,... X n}:
[0025]
[0026] Here, Pa(X i ) is the set of regulatory factors (parent nodes) of gene X i in N, and P(X i |Pa(X i )) quantifies the regulatory strength between gene X i and its regulatory factor Pa(X i ).
[0027] scRNA-seq data exists in the form of a two-dimensional matrix with n rows and c columns, representing the expression values of n genes in c cells. Since the original data is relatively sparse, sparse features need to be extracted first. The present invention adopts a sparse model based on the Transformer architecture, and its key theory involves the attention mechanism, the multi-head self-attention mechanism, and the extraction of sparse features.
[0028] By introducing the attention mechanism, the Transformer model can pay more attention to different features of different genes in the input scRNA-seq data. The calculation of the attention weights is as follows:
[0029]
[0030] Here, Q, K, and V represent the matrices of Query, Key, and Value respectively. Q represents the expression pattern of genes in the scRNA-seq data, that is, query the expression level of the i-th gene in the j-th cell in the scRNA-seq data. K represents the characteristics of other genes related to Q, so that the Transformer model can understand the interactions and correlations between a certain gene and other genes. V represents the expression values of the actual genes associated between Q and K. d k represents the dimension of the gene expression pattern, that is, the number of gene features. Taking the square root can make the training gradient more stable. QK T is the dot product operation of Q and K, representing the attention score matrix regarding gene features. Finally, the attention distribution corresponding to different gene features is obtained through the softmax function.
[0031] By calculating multiple attention heads in parallel and concatenating all the attention heads together, the Transformer model can better capture different expression relationships in the scRNA-seq data. The mathematical representation is as follows:
[0032] MultiHead(Q, K, V) = Concat(head1, head2, …, head h )W O (3)
[0033] where the calculation process of each attention head is the same as that of the basic attention mechanism, and W O represents the weight matrix for connecting multiple attention heads.
[0034] To process sparse data, the present invention introduces a sparse multi-head self-attention mechanism into the Transformer model. By only focusing on the non-zero elements in the gene input sequence data, sparse features can be effectively extracted. The calculation process is as follows:
[0035]
[0036] where ⊙ represents element-wise multiplication, M is the mask matrix, which means restricting the attention to the non-0 expression values of genes in cells. The sparse features are represented as follows:
[0037] SparseFeatures = SparseMultiHeadSelfAttention(I) (5)
[0038] This representation can better capture the correlations in the scRNA-seq input sequence, especially suitable for processing sparse scRNA-seq data. Here, I is the feature vector matrix output by the upper-layer sparse self-attention. Through the mapping of sparse multi-head self-attention in the sparse Transformer model, sparse features that can better reflect gene expression relationships are obtained.
[0039] The above process details the sparse multi-head self-attention mechanism of the sparse Transformer model and the processing method for sparse data, explaining the principle of feature extraction by the sparse Transformer model.
[0040] (Step 3) Construction of gene regulatory relationships;
[0041] After the sparse Transformer model extracts sparse features, the generative flow network begins to use these features to construct gene regulatory relationships. The generative flow network model is iteratively trained to transition the generative state s until the terminal state s f is obtained to get the posterior probability distribution of the gene regulatory network for scRNA-seq data.
[0042] Each transition state s' is associated with a reward function R(s') to guide the generative flow network model to learn a better terminal state s f . Each state here represents a gene regulatory network. The transition process of the state represents the generation process in which the gene regulatory network gradually adds edges from an empty graph. The terminal state is the final gene regulatory network generated by each flow in the generative flow network model. The goal of the generative flow network model is to find a flow that satisfies the following flow matching conditions for all genes:
[0043]
[0044] Pa(s') represents the set of parent nodes of state s', that is, the set of states that generate s'. Among them, Ch(s') is the set of child nodes of s', that is, the set of subsequent states generated by state s'. F θ (s → s') is the generation coefficient representing the transition from state s to state s'. R(s') represents the score evaluation of the specific regulatory goal reached by state s', that is, the transition to the terminal state. Therefore, the flow matching condition can ensure that the regulatory inflow minus the regulatory outflow of each gene regulatory network state is equal to the specific regulatory goal.
[0045] In a gene regulatory network, the reward function R(s') can be expressed as the sum of local scores, that is, the total reward of the gene regulatory network can be decomposed into the local contributions of each gene. The decomposition process is as follows:
[0046]
[0047] where n is the number of gene nodes. The generative flow network model uses the sparse features extracted from scRNA-seq data to iteratively train the process, and then stores the gene regulatory information in the state transition process s→s' in a buffer. After the iteration ends, the generative flow network model obtains the probability distribution of t terminal states output by the generative flow network model, that is, the probability distribution of gene regulatory relationships.
[0048] (Step 4) Fusion and extraction of the gene regulatory network;
[0049] After step 3 is completed, the probability distribution of t terminal states output by the generative flow network model is obtained. Next, fusion and extraction rules need to be designed to obtain the final gene regulatory network. First, the posterior probability of t terminal states is averaged according to the following formula:
[0050]
[0051] where represents the regulatory coefficient of the i-th gene on the j-th gene in the m-th terminal state, and U ij represents the final regulatory coefficient of the i-th gene on the j-th gene after fusion. If U ij > 0.4, it is set to 1; otherwise, U ij is set to 0. Finally, the optimal gene regulatory network is obtained, and the algorithm ends.
[0052] To fully verify the superiority of this method, this method was compared with 4 existing causal learning methods, DeepSEM, DGRN, GENIE3, and PIDC. These four methods are classic methods for constructing gene regulatory networks. To evaluate the performance of the learning algorithm, three evaluation indicators, AUROC, AUPRC, and OverallScore, were used to evaluate the results. AUPRC is the area under the precision curve and recall curve, and AUROC is the area under the FPR and TPR curves. AUPRC is a composite indicator reflecting the Precision and Recall of the model, and AUROC is a composite indicator reflecting the TPR and FPR of the model. They can better show the pros and cons of the method performance. For OverallScore, it is defined as follows:
[0053]
[0054] These three evaluation metrics are widely used to evaluate the performance of gene regulatory network construction. The values of the three evaluation metrics are between 0 and 1, and the larger the value, the higher the accuracy of the gene regulatory network constructed by the algorithm.
[0055] Three sub-datasets were selected from the sc-RNA-seq datasets of three groups of mouse hematopoietic stem cells (mHSC) as the input for the above five algorithms. The algorithms were run 10 times repeatedly, and the average values of all metric results were taken, as shown in Table 1.
[0056] Table 1: Comparison of various methods on three groups of mouse hematopoietic stem cell datasets
[0057]
[0058] It can be found from Table 1 that the overall performance of DeepSEM and DGRN on AUPRC and AUROC is better than that of GENIE3 and PIDC, which indicates that deep learning algorithms can capture richer non-linear gene regulatory relationships. The AUPRC of DGRN, GENIE3 and PIDC on these three datasets is even as low as 0.07. It is considered that this is because the scRNA-seq dataset of mouse hematopoietic stem cells is too sparse, which makes it difficult for the above three algorithms to capture important sparse features, resulting in poor performance. Obviously, this method is better than all other algorithms in all metrics, and is better than all comparison algorithms in AUPRC. The AUPRC value on the mHSC-GM dataset is 0.62, far higher than other algorithms, which fully proves that the designed sparse Transformer containing sparse multi-head self-attention can effectively extract sparse features in scRNA-seq data, so that the generative flow network can construct a gene regulatory network with higher accuracy.
[0059] To more intuitively display the comprehensive performance of this method, the OverallScore metrics of the five algorithms on the three datasets were visually displayed, as Figure 2 shown. Obviously, the OverallScore of this method is higher than any other algorithm on these three datasets, which fully proves that the comprehensive performance of this method is higher and it can stably and accurately construct gene regulatory networks from sparse scRNA-seq data.
[0060] The above experiments can show that the method proposed in the present invention has higher accuracy in constructing gene regulatory networks from sparse scRNA-seq data compared with other methods, and thus has great application value and prospects in computer-aided diagnosis of gene diseases.
Claims
1. A method for constructing a gene regulatory network based on a generative flow network, characterized in that, The following steps are involved: Step 1: Data acquisition: Experiments were conducted on the sc-RNA-seq dataset of mouse hematopoietic stem cells to evaluate the performance of constructing gene regulatory networks; Step 2: Sparse feature extraction: Use sparse Transformer to extract sparse features from scRNA-seq data to construct a gene regulatory network; Step 3: Construction of gene regulatory relationships: After sparse features are extracted, the sparse features are input into the generative flow network model to construct the gene regulatory network; Step 4: Fusion and extraction of gene regulatory network: By generating multiple iterative updates of the flow network model, a probability distribution of gene regulatory relationships associated with the mouse hematopoietic stem cell sc-RNA-seq dataset is obtained, and the final gene regulatory network is extracted by setting the regulation extraction threshold; In step 1, the BEELINE dataset includes 7 sc-RNA-seq datasets, 5 of which are about mouse cells and 2 are about human cells; three groups of mouse hematopoietic stem cell mHSC sc-RNA-seq datasets are selected, namely mHSC-L, mHSC-GM and mHSC-E; for these three datasets, genes expressed in less than 10% of cells are removed; then, the remaining data are normalized; In step two, the gene regulatory network is represented as G = <N, Z>, where N is the gene set, and X i ∈ Z represents the i-th gene among n genes, i = 1, 2, …, n, and Z represents the set of regulatory relationships between genes. X j → X i ∈ Z represents the directed edge from gene X j to gene X i . The gene regulatory network is represented as the joint probability distribution of the gene set N = {X1, X2, X3, … X n}: Here, Pa(X i ) is the set of regulatory factors of gene X i in N, and P(X i |Pa(X i )) quantifies the regulatory strength between gene X i and its regulatory factor Pa(X i ); scRNA-seq data exists in the form of a two-dimensional matrix with n rows and c columns, representing the expression values of n genes in c cells; By introducing the attention mechanism, the Transformer model focuses on the different features of different genes in the input scRNA-seq data; the attention weight is calculated as follows: Q, K, and V represent the matrices of Query, Key, and Value, respectively; Q represents the expression pattern of genes in scRNA-seq data, that is, querying the expression level of the i-th gene in the j-th cell in the scRNA-seq data; K represents the characteristics of other genes related to Q; V represents the expression value of the actual gene associated between Q and K; d k represents the dimension of the gene expression pattern, that is, the number of gene features; QK T is the dot product operation of Q and K, representing the attention score matrix regarding gene features. Finally, the attention distribution corresponding to different gene features is obtained through the softmax function; The Transformer model is represented as follows: MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O (3) where the calculation process of each attention head is the same as that of the basic attention mechanism, and W O represents the weight matrix connecting multiple attention heads; The sparse multi-head self-attention mechanism is introduced into the Transformer model to extract sparse features by focusing only on the non-zero elements in the gene input sequence data; The calculation process is as follows: Among them, ⊙ represents element-by-element multiplication, M is a mask matrix, which means that attention is limited to the non-zero expression values of genes in cells, and the sparse features are represented as follows: SparseFeatures=SparseMultiHeadSelfAttention(I) (5) I is the feature vector matrix output by the upper-layer sparse self-attention. Through the mapping of sparse multi-head self-attention in the sparse Transformer model, sparse features that can reflect the relationship between gene expression are obtained.
2. The method for constructing a gene regulatory network based on a generative flow network according to claim 1, wherein In step 3, after the sparse Transformer model extracts sparse features, the generative flow network starts to use these features to construct gene regulatory relationships; the generative flow network model is iteratively trained to perform the transition of the generative state s until the terminal state s is generated. f to obtain the posterior probability distribution of the gene regulatory network for scRNA-seq data; Each transition state s' is associated with a reward function R(s') to guide the learning of the generative flow network model towards a better terminal state s f ; each state represents a gene regulatory network, and the transition process of the state represents the generation process of the gene regulatory network by gradually adding edges to the empty graph. The terminal state is the final gene regulatory network generated by each flow in the generative flow network model; The goal of the generative flow network model is to find a flow that satisfies the following flow matching conditions for all genes: Pa(s') represents the set of parent nodes of state s', i.e., the set of states that generate s'; where Ch(s') is the set of child nodes of s', i.e., the set of subsequent states generated by state s'; F θ (s → s') is the generation coefficient representing the transition from state s to state s', and R(s') represents the fractional evaluation of achieving a specific regulatory target for state s', i.e., transitioning to the terminal state; the flow matching condition ensures that the regulatory inflow minus the regulatory outflow for each gene regulatory network state equals the regulatory target; In the gene regulatory network, the reward function R(s') is expressed as the sum of local scores, that is, the total reward of the gene regulatory network is decomposed into the local contribution of each gene. The decomposition process is as follows: where n is the number of gene nodes, and the generative flow network model uses the sparse features extracted from scRNA-seq data to iteratively train the process, and then stores the gene regulatory information in the state transition process s→s' in the buffer; after the iteration ends, the generative flow network model obtains the probability distribution of the t terminal states output by the generative flow network model, and obtains the probability distribution of the gene regulatory relationship.
3. A method for constructing a gene regulatory network based on a generative flow network according to claim 2, characterized in that, In step four, after step three is completed, the probability distribution of the t terminal states output by the generative flow network model is obtained; next, it is necessary to design fusion and extraction rules to obtain the final gene regulatory network; first, the posterior probabilities of the t terminal states are averaged according to the following formula: wherein represents the regulation coefficient of the i-th gene on the j-th gene in the m-th terminal state, and U ij represents the regulation coefficient of the i-th gene on the j-th gene finally obtained after fusion; If U ij > 0.4 is set to 1, otherwise, U ij is set to 0; obtain the gene regulatory network, and end.
Citation Information
Patent Citations
Generation method and device of gene regulation relation model
CN116403634A
Gene regulatory network inference method based on multi-view layered hypergraph
CN116844645A