A gene regulatory network inference method and system based on graph representation learning

By using a graph representation learning approach, utilizing single-cell RNA sequencing data and graph convolutional neural networks, and combining attention mechanisms and graph contrastive learning regularization terms, the high cost and low accuracy of traditional methods are solved, achieving efficient and accurate inference of gene regulatory networks.

CN119763665BActive Publication Date: 2025-10-28JIANGNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411631566.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-10-28
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

Traditional methods for studying gene regulatory networks are costly and computationally inefficient. Existing machine learning and deep learning methods neglect the global structure of gene regulatory networks during feature extraction, resulting in insufficient accuracy.

Method used

A graph representation learning-based approach was adopted to construct a gene regulatory network dataset using single-cell RNA sequencing data. Low-dimensional embedding features of genes were extracted using graph convolutional neural networks, and channel and spatial attention mechanisms were introduced. The loss function was optimized by combining graph contrastive learning regularization terms, and iterative training was performed to predict the regulatory relationships of gene pairs.

Benefits of technology

It significantly reduces research costs, improves the accuracy and efficiency of gene regulatory network inference, can more accurately capture complex interactions, and enhances the model's expressive and generalization abilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119763665B_ABST
    Figure CN119763665B_ABST
Patent Text Reader

Abstract

This invention relates to a method and system for inferring gene regulatory networks based on graph representation learning. The method includes: constructing a gene regulatory network dataset using single-cell RNA sequencing data; using known gene-interacting pairs as positive samples and unknown gene-interacting pairs as negative samples; and randomly dividing the dataset into training, validation, and test sets. The prior gene regulatory network and gene expression profiles from the training set are input into a pre-constructed gene regulatory network inference model for iterative training to obtain predicted gene-pair values. A loss function containing a graph contrastive learning regularization term is constructed, and the loss between predicted and true values ​​is calculated to update model parameters. When the training iterations reach a threshold, the model training parameters are output and loaded into the model to predict the regulatory relationships of unknown gene-interacting pairs. This invention solves the problems of high cost in traditional biological experiments and low accuracy of existing computational models, improving the inference accuracy of gene regulatory networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene pair regulatory relationship inference technology, and in particular to a gene regulatory network inference method and system based on graph representation learning. Background Technology

[0002] In gene expression regulation, transcription factors produced by upstream regulatory genes through transcription and translation bind to specific regions of downstream target genes, thereby activating or inhibiting the expression intensity of those genes and regulating their expression levels. Gene expression regulation plays a crucial role in the growth and development of biological cells. Genes regulate protein synthesis through transcription and translation, thus controlling various life activities such as cell growth and development. Gene regulatory networks use graph structures to describe the regulatory interactions between genes in cells, where nodes and edges represent genes and regulatory relationships, respectively. Gene regulatory networks are helpful in studying the regulatory mechanisms of various cellular life activities, elucidating the pathogenesis of complex diseases, and providing a theoretical basis for drug development for complex diseases. Furthermore, gene regulatory networks can be applied to the modeling and optimization of metabolic systems, and the design and construction of microbial cell factories. Therefore, the rational construction of gene regulatory networks has significant theoretical and applied value.

[0003] However, traditional molecular interaction experimental methods, such as microarrays and chromatin immunoprecipitation sequencing, while useful for studying gene-gene interactions, are often costly, time-consuming, and highly dependent on experimental conditions. To overcome these limitations, researchers have attempted to use theories such as differential equations and Boolean networks to mechanistically model gene regulatory networks and construct such networks. However, when dealing with large-scale gene regulatory networks, these methods often suffer from low computational efficiency and insufficient accuracy.

[0004] In recent years, with the rapid development of single-cell high-throughput sequencing technology, researchers have obtained an explosive growth in the amount of gene expression profiling data at the cellular level. This has provided an important opportunity for the application of machine learning, deep learning, and other methods in gene regulatory network reconstruction. Machine learning and deep learning have attracted much attention in the field of bioinformatics and have achieved remarkable results in areas such as the prediction and optimization of gene regulatory elements, protein function mining and design, and metabolic network analysis and design.

[0005] Current research utilizes traditional machine learning methods to predict gene interactions in gene regulatory networks. For example, the GENIE3 model calculates the interaction scores between regulatory factors and target genes using the random forest algorithm, thereby inferring the structure of the gene regulatory network; the GRNBoost2 model further improves computational efficiency by combining the GENIE3 model with the Arboreto computational framework. However, traditional machine learning methods require manually designed features, which to some extent limits the improvement of gene regulatory network inference performance.

[0006] Building upon this foundation, deep learning methods have been extensively studied and explored in supervised feature representation learning of gene regulatory networks. These methods infer the regulatory relationships between gene pairs based on prior networks and gene expression profile data, thereby reconstructing gene regulatory networks. For example, the CNNC model uses convolutional neural networks to extract features of gene pair expression values ​​and infer interactions between genes; the DeepDRIM model uses multi-layer convolutional neural networks to extract deep features from the main image and neighbor images of gene pair expression values, further inferring gene regulatory networks. The STGRNS model borrows the structure of the large language model BERT and uses transfer learning to effectively extract long-range features of gene pair expression values, inferring potential interactions between genes.

[0007] However, the aforementioned methods primarily focus on the interactions between gene pairs during feature extraction, neglecting the global structure of the gene regulatory network. To address this, the GCNG model introduces a graph convolutional neural network to extract low-dimensional gene embedding features, more effectively inferring potential interactions within the gene regulatory network. Furthermore, the GENELINK model utilizes a graph attention neural network, optimizing the low-dimensional gene embedding extraction process through an attention mechanism, further improving the accuracy of gene regulatory network inference. However, due to the sparsity of gene regulatory networks, feature extraction using only global connectivity information remains challenging. Therefore, future research needs to further explore the implicit connectivity information within gene regulatory networks to more effectively extract gene features and potential correlations between genes, thereby promoting the in-depth development of gene regulatory network research. Summary of the Invention

[0008] To address the issues of high cost and low accuracy of existing computational models in traditional biological experimental methods, this invention provides a gene regulatory network inference method based on graph representation learning, which includes the following steps:

[0009] S1: Obtain single-cell RNA sequencing data. Based on the regulatory relationships between genes in the single-cell RNA sequencing data, construct a gene regulation network dataset. In the gene regulation network dataset, gene pairs with interactions are considered as positive samples, gene pairs with unknown interactions are considered as negative samples, and positive and negative samples are randomly divided into training set, validation set, and test set.

[0010] S2: Input the prior gene regulation network and gene expression profile generated from positive samples in the training set into the pre-constructed gene regulation network inference model for iterative training, and obtain the gene pair prediction values ​​calculated in each iteration;

[0011] S3: Construct a loss function based on the binary cross-entropy loss term and the graph contrastive learning regularization term. Use the loss function to calculate the loss between the predicted and true values ​​of gene pairs to update the model parameters and determine whether the training epochs have reached a threshold.

[0012] If not, return to step S2 and continue the iterative calculation;

[0013] If so, output the model training parameters, load the model training parameters into the gene regulation network inference model, predict the regulatory relationship of gene pairs with unknown interactions, and obtain the prediction results.

[0014] Preferably, in S2, the method for obtaining the predicted gene pair values ​​calculated in each iteration is as follows:

[0015] S21: Based on the directed graph of the prior gene regulatory network generated by positive samples in the training set, construct multiple subgraphs, each subgraph corresponding to an adjacency matrix, and process all adjacency matrices to obtain an implicit connection tensor including multiple implicit connection matrices.

[0016] S22: Based on the normalized gene expression profile, multiple implicit connection matrices are sequentially input into the graph convolutional neural network. For each implicit connection matrix, a gene feature matrix is ​​output. The gene feature matrix is ​​subjected to layer normalization. The normalized gene feature matrices are concatenated to obtain the graph embedding tensor.

[0017] S23: Based on the graph embedding tensor, the enhanced gene features are obtained. The Hadamard product of the feature vectors corresponding to any two genes in the enhanced gene features is calculated to obtain the feature fusion vector. The feature fusion vector is processed by a multilayer perceptron to output the predicted value of the gene pair.

[0018] Preferably, in S21, an implicit connection tensor comprising multiple implicit connection matrices is obtained. The method is as follows:

[0019] S211: Concatenate all adjacency matrices to obtain tensor A.s ∈{0,1} 5×N×N , will A s Simultaneously input to two parameterized layers, yielding tensors respectively. and In each parameterized layer, the following relationship is satisfied:

[0020]

[0021] Where N is the number of genes in the gene regulatory network, Q (i) (j,:,:) represents the output Q of the i-th parameterized layer. (i) The j-th matrix in the matrix, where B is the tensor Q. (i) The total number of matrices in (j,:,:). Gene connection types in gene regulatory networks A s (k,:,:) represents the input tensor A. s The sub-adjacency matrix of the k-th connection type, These are the normalized training parameters in the i-th layer;

[0022] S212: The tensor Q output from the two parameterized layers (1) and Q (2) Performing an inner product along the corresponding dimensions yields the implicitly connected tensor. This process can be represented as:

[0023] A L (j,:,:)=Q (1) (j,:,:)·Q (2) (j,:,:),j=1,2,...,B

[0024] in, Indicates the output implicit connection tensor A L The j-th implicit connection matrix in the matrix.

[0025] Preferably, in S22, the graph embedding tensor Each tensor X IE The expression for (j,:,:) is as follows:

[0026]

[0027] Where H represents the gene feature dimension after graph embedding. Indicates the output implicit connection tensor A L Let N be the j-th implicit connection matrix, N be the number of genes in the gene regulatory network, and I be the identity matrix. for The degree matrix, X FHere, W represents the normalized gene expression profile, W represents the training parameters of the graph convolutional neural network, LayerNorm(·) represents the layer normalization function, and B represents the number of tensors in the graph embedding tensor.

[0028] Preferably, in S23, the method for obtaining the enhanced gene characteristics is as follows:

[0029] S231: Graph Embedding Tensors Perform a dimensional transformation to obtain

[0030]

[0031] Where M is the square root of H, Given B feature tensors Composition, i = 1, 2, ..., B, and each tensor can be represented as a feature matrix with N channels. i=1,2,...,B,j=1,2,...,N;

[0032] S232: Yes Each feature tensor Feature matrices on all N channels Global max pooling and global average pooling are performed separately to obtain two channel features. These two channel features are then input into a parameter-shared convolutional algorithm, followed by element-wise summation and a Sigmoid activation function to obtain channel attention.

[0033]

[0034] S233: Transfer the feature tensor Through broadcasting mechanism and channel attention feature N c,i After performing the Hadamard product, the output features are obtained. This process can be represented as:

[0035]

[0036] S234: Along Two spatial features are obtained by performing global max pooling and global average pooling on the channel dimension N, respectively. These features are then concatenated on the channel dimension N and followed by convolution to obtain the spatial attention feature. This process can be represented as:

[0037]

[0038] S235: Will Through broadcasting mechanisms and spatial attention characteristics N s,i After performing the Hadamard product, the enhanced feature map is obtained.

[0039]

[0040] S236: All enhanced feature maps and original feature map Each input matrix is ​​residually connected to obtain an M×M dimensional input matrix. The input matrix is ​​then flattened into a vector, and the mean value is calculated along dimension B to obtain the enhanced gene features.

[0041]

[0042] Here, Reshape(·) represents the dimension transformation operation.

[0043] Preferably, the predicted value of gene pairs The calculation method is as follows:

[0044]

[0045] Where y is the label value of the gene pair. and is the predicted value output by the gene regulatory network inference model, and n is the number of samples in the training batch.

[0046] Preferably, in S3, the expression for the loss function is as follows:

[0047] Loss = αLoss bc +βN gc

[0048] Where Loss represents the total loss function value, Loss bc N represents the binary cross-entropy loss term. gc The graph comparison learning regularization term is represented by α and β, which are both weighting coefficients.

[0049] Preferably, the binary cross-entropy loss term Loss bc The method to obtain it is as follows:

[0050]

[0051] Where y is the tag value of the gene pair. and y represents the predicted value output by the gene regulatory network inference model, where n is the number of samples in the training batch, and y is the predicted value. s Let be the tag value of the s-th gene pair. is the predicted value for the s-th gene pair.

[0052] Preferably, the graph comparison learning regularization term N gc The method to obtain it is as follows:

[0053] S31: Randomly discard the normalized gene expression profile X F After processing a subset of values ​​from the implicit connection matrix A, the results are fed into a graph convolutional neural network to obtain the gene embeddings under explicit connections.

[0054]

[0055] in, A d Let A be the join matrix after discarding some values, and I be the identity matrix. for The degree matrix, For X F Discarding a portion of the gene expression profile;

[0056] S32: For the enhanced genetic trait X GE and gene embedding X EE The feature vector u of the m-th gene m =X GE (m,:) and v m =X EE (m,:), calculate u m With v m The first gene pair function P(u) between m ,v m ) and v m with u m The second gene pair function P(v) between m ,u m ),as follows:

[0057]

[0058] Where θ is the neural network fitting parameter, τ represents the temperature coefficient, and 1 [m≠n] This represents an indicator function, which is 1 when m≠n and 0 when m=n;

[0059] S33: Based on the first gene pair function P(u m ,v m ) and the second gene pair function P(v m ,u m ), thus obtaining the graph comparison learning regularization term N. gc :

[0060]

[0061] Where N represents the number of genes.

[0062] Based on the same inventive concept, embodiments of the present invention also provide a gene regulatory network inference system based on graph representation learning. This system is used to implement the above-mentioned gene regulatory network inference method based on graph representation learning, specifically including:

[0063] The gene regulation network dataset construction module is used to acquire single-cell RNA sequencing data. Based on the regulatory relationships between genes in the single-cell RNA sequencing data, a gene regulation network dataset is constructed. In the gene regulation network dataset, gene pairs with interactions are regarded as positive samples, gene pairs with unknown interactions are regarded as negative samples, and positive and negative samples are randomly divided into training set, validation set and test set.

[0064] The gene regulation network inference model construction module is used to input the prior gene regulation network and gene expression profile generated from positive samples in the training set into the pre-constructed gene regulation network inference model for iterative training, and obtain the gene pair prediction values ​​calculated in each iteration;

[0065] The gene pair regulatory relationship prediction module is used to construct a loss function based on a binary cross-entropy loss term and a graph contrastive learning regularization term. The loss function is used to calculate the loss value between the predicted and actual values ​​of gene pairs to update the model parameters. When the training rounds reach a threshold, the model training parameters are output and loaded into the gene regulation network inference model to predict the regulatory relationship of gene pairs with unknown interactions, and the prediction results are obtained.

[0066] As can be seen from the above technical solutions, this invention application has the following beneficial effects:

[0067] First, it addresses the high cost of traditional biological experiments: This invention primarily relies on single-cell RNA sequencing data, rather than traditional biological experimental methods such as microarrays and chromatin immunoprecipitation sequencing. This significantly reduces experimental costs, as single-cell RNA sequencing data can be obtained through publicly available datasets or high-throughput sequencing technologies, making it more economical than traditional methods. By constructing a gene regulatory network dataset, it fully utilizes the gene regulatory relationships within the single-cell RNA sequencing data, treating interacting gene pairs as positive samples and gene pairs with unknown interactions as negative samples. This method improves data utilization, enabling the acquisition of more information with limited data.

[0068] Second, address the issue of low accuracy in existing gene-pair relationship calculation models:

[0069] (1) When constructing the gene regulation network inference model, the concept of implicit connection is introduced. By constructing multiple subgraphs and implicit connection matrices, more complex interaction relationships in the gene regulation network can be captured, thereby improving the inference accuracy of the model.

[0070] (2) Applying graph convolutional neural networks (GCN) to gene regulation network inference can learn low-dimensional embedding representations of genes. These representations not only preserve the original feature information of genes, but also incorporate their contextual information in the network, thereby enhancing the expressive power of the model.

[0071] (3) Use channel attention and spatial attention mechanisms to enhance gene features, and use techniques such as global max pooling and global average pooling to extract key features. These steps help to remove redundant information and retain the most representative features, thereby improving the inference accuracy of the model.

[0072] (4) Taking into account the difference between the predicted and actual values ​​of gene pairs and the positional relationship of genes in the network, a loss function containing graph comparison learning regularization term is constructed to guide the model to learn more accurate gene regulation relationships. Attached Figure Description

[0073] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the embodiments will be briefly described below. Referring to the accompanying drawings will provide a clearer understanding of the features and advantages of the present invention. The drawings are illustrative and should not be construed as limiting the present invention in any way. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort. Wherein:

[0074] Figure 1 This is a flowchart of a gene regulatory network inference method based on graph representation learning provided in Embodiment 1 of the present invention;

[0075] Figure 2 This is a flowchart illustrating the gene regulatory network inference method based on graph representation learning provided in Embodiment 1 of the present invention.

[0076] Figure 3 This is a schematic diagram of the construction of the gene regulatory network dataset and the partitioning of the sample set provided in Embodiment 1 of the present invention;

[0077] Figure 4 In the diagram, A represents the extraction of implicit connections in the gene regulatory network structure, B represents the extraction of gene features in the gene regulatory network under implicit connections, C represents the enhancement of gene feature extraction in the gene regulatory network, and D represents the output of the predicted results of gene-regulatory relationships.

[0078] Figure 5 This is a flowchart of obtaining the enhanced feature map using the attention mechanism module;

[0079] Figure 6 This is a flowchart for obtaining graph comparison learning regularization terms;

[0080] Figure 7In the diagram, A represents the AUROC index of the gene regulatory network inference model (referred to as the "GRLGRN model") proposed in this invention in seven cell types under different scales of the gold standard network framework, and B represents the distribution of the AUROC index of GRLGRN and other comparative models on the dataset.

[0081] Figure 8 In the diagram, A represents the AUPRC index of the gene regulatory network inference model (referred to as the "GRLGRN model") proposed in this invention in seven cell types under different scales of the gold standard network framework, and B represents the distribution of the AUPRC index of GRLGRN and other comparative models on the dataset.

[0082] Figure 9 This is a block diagram of a gene regulatory network inference system based on graph representation learning provided in Example 2;

[0083] Explanation of reference numerals in the instruction manual:

[0084] 100. Gene regulation network dataset construction module; 200. Gene regulation network inference model construction module; 300. Gene pair regulatory relationship prediction module. Detailed Implementation

[0085] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0086] Example 1

[0087] like Figure 1 and Figure 2 As shown in the figure, this invention proposes a gene regulatory network inference method based on graph representation learning, which includes:

[0088] S1: Obtain single-cell RNA sequencing data. Based on the regulatory relationships between genes in the single-cell RNA sequencing data, construct a gene regulation network dataset. In the gene regulation network dataset, gene pairs with interactions are considered as positive samples, gene pairs with unknown interactions are considered as negative samples, and positive and negative samples are randomly divided into training set, validation set, and test set.

[0089] S2: Input the prior gene regulation network and gene expression profile generated from positive samples in the training set into the pre-constructed gene regulation network inference model for iterative training, and obtain the gene pair prediction values ​​calculated in each iteration;

[0090] S3: Construct a loss function based on the binary cross-entropy loss term and the graph contrastive learning regularization term. Use the loss function to calculate the loss between the predicted and true values ​​of gene pairs to update the model parameters and determine whether the training epochs have reached a threshold.

[0091] If not, return to step S2 and continue the iterative calculation;

[0092] If so, output the model training parameters, load the model training parameters into the gene regulation network inference model, predict the regulatory relationship of gene pairs with unknown interactions, and obtain the prediction results.

[0093] As can be seen from the above technical solution, this invention constructs a gene regulatory network dataset using single-cell RNA sequencing data, and then trains and predicts models based on this dataset, significantly reducing the need for direct biological experiments and thus lowering research costs. By introducing graph representation learning, the gene regulatory network is transformed into graph structure data, and the loss function is optimized using graph contrastive learning regularization terms to enhance the model's generalization ability. Simultaneously, through iterative training and loss function optimization, the model can more accurately predict the regulatory relationships between gene pairs. This method not only considers gene expression profile information but also utilizes the topological structure information of the gene regulatory network, thereby improving the model's inference accuracy.

[0094] Specifically, in step S1, single-cell RNA sequencing (scRNA-seq) data for seven cell types—human embryonic stem cells (hESC), human mature hepatocytes (hHEP), mouse dendritic cells (mDC), mouse embryonic stem cells (mESC), mouse erythroid hematopoietic stem cells (mHSC-E), granulocyte-monocyte lineage mouse hematopoietic stem cells (mHSC-GM), and lymphoid lineage mouse hematopoietic stem cells (mHSCL)—were obtained from publicly available data. Each cell type corresponds to one of three different standard networks provided by the cell-type-specific ChIP-seq, non-specific ChIP-seq, and STRING databases, resulting in a total of 21 standard networks for the seven cell types. For each standard network, based on the magnitude of changes in gene expression values, the top 500 and top 1000 genes with the most significant changes in expression values, including transcription factors, were selected. Based on these key genes and their regulatory relationships, two gene regulatory networks of different sizes (a total of 42) were constructed, along with a gene expression profile matrix describing the expression intensity of each gene over time.

[0095] Furthermore, such as Figure 3As shown, gene pairs with interactions are considered positive samples and labeled with 1, while gene pairs with unknown interactions are considered negative samples and labeled with 0. Positive and negative samples are randomly divided into training, validation, and test sets in a 3:1:1 ratio. These positive and negative samples are called hard samples, and the purpose is to accelerate the training process, reduce computational complexity, and better extract gene features.

[0096] Preferably, in S2, the method for obtaining the predicted gene pair values ​​calculated in each iteration is as follows:

[0097] S21: As Figure 4 As shown in Figure A, based on the directed graph of the prior gene regulation network generated from positive samples in the training set, five subgraphs are constructed, including a directed subgraph of "regulatory gene-target gene" and its reverse graph, a directed subgraph of "regulatory gene-regulatory gene" and its reverse graph, and a gene self-connection subgraph. Each subgraph corresponds to an adjacency matrix. Data processing is performed on all adjacency matrices to obtain an implicit connection tensor including multiple implicit connection matrices.

[0098] S22: Based on the normalized gene expression profile, multiple implicit connection matrices are sequentially input into the graph convolutional neural network. For each implicit connection matrix, a gene feature matrix is ​​output. The gene feature matrix is ​​subjected to layer normalization. The normalized gene feature matrices are concatenated to obtain the graph embedding tensor.

[0099] S23: Based on the graph embedding tensor, the enhanced gene features are obtained. The Hadamard product of the feature vectors corresponding to any two genes in the enhanced gene features is calculated to obtain the feature fusion vector. The feature fusion vector is processed by a multilayer perceptron to output the predicted value of the gene pair.

[0100] Preferably, in S21, an implicit connection tensor comprising multiple implicit connection matrices is obtained. The method is as follows:

[0101] S211: Concatenate all adjacency matrices to obtain the tensor. To consider different types of implicit connection information simultaneously, A s Simultaneously input to two independent parameterization layers, yielding tensors respectively. and In each parameterized layer, the following relationship is satisfied:

[0102]

[0103] Where N is the number of genes in the gene regulatory network, Q (i) (j,:,:) represents the output Q of the i-th parameterized layer. (i) The j-th matrix in the matrix, where B is the tensor Q.(i) The total number of matrices in (j,:,:). Gene connection types in gene regulatory networks A s (k,:,:) represents the input tensor A. s The sub-adjacency matrix of the k-th connection type, These are the normalized training parameters in the i-th layer; The normalization process satisfies:

[0104]

[0105] S212: The tensor Q output from the two parameterized layers (1) and Q (2) Performing an inner product along the corresponding dimensions yields the implicitly connected tensor. This process can be represented as:

[0106] A L (j,:,:)=Q (1) (j,:,:)·Q (2) (j,:,:),j=1,2,...,B

[0107] in, Indicates the output implicit connection tensor A L The j-th implicit connection matrix in the matrix.

[0108] Further, in step S22, as Figure 4 As shown in Figure B, the normalized gene expression profile is denoted as... Where N is the number of genes in the gene regulation network, and D is the dimension of the gene expression profile. Based on the normalized gene expression profile, B implicit connection matrices are sequentially input into the graph convolutional neural network. For each implicit connection matrix, a gene feature matrix is ​​output. The gene feature matrices are then subjected to layer normalization. Finally, the normalized gene feature matrices are concatenated to obtain the graph embedding tensor. in Each tensor X IE The expression for (j,:,:) is as follows:

[0109]

[0110] Where H represents the gene feature dimension after graph embedding. Indicates the output implicit connection tensor A L Let N be the j-th implicit connection matrix, N be the number of genes in the gene regulatory network, and I be the identity matrix. for The degree matrix, X FHere, W represents the normalized gene expression profile, W represents the training parameters of the graph convolutional neural network, LayerNorm(·) represents the layer normalization function, and B represents the number of tensors in the graph embedding tensor.

[0111] Further, in step S23, as Figure 4 As shown in C, based on graph embedding tensor Obtain enhanced genetic traits The method is as follows:

[0112] S231: Graph Embedding Tensors Perform a dimensional transformation to obtain

[0113]

[0114] Where M is the square root of H, Given B feature tensors Composition, i = 1, 2, ..., B, and each tensor can be represented as a feature matrix with N channels.

[0115] i=1,2,...,B,j=1,2,...,N;

[0116] S232: As Figure 5 As shown, The input is fed into a convolutional block attention module (CBAM) consisting of channel attention and spatial attention, for... Each feature tensor Feature matrices on all N channels Global max pooling and global average pooling are performed separately to obtain two channel features. These two channel features are then input into a parameter-shared convolutional network. The neural network consists of convolution operation Conv1, a non-linear activation function (ReLU), and convolution operation Conv2. After passing through the neural network, the two channel features are summed element-wise and then passed through a sigmoid activation function to obtain channel attention.

[0117]

[0118] The convolution operation performed by inputting these two channel features into parameter sharing satisfies the following:

[0119] Conv c (·)=Conv2(ReLU(Conv1(·)))

[0120] S233: Transfer the feature tensor Through broadcasting mechanism and channel attention feature N c,iAfter performing the Hadamard product, the output features are obtained. This process can be represented as:

[0121]

[0122] S234: Along Two spatial features are obtained by performing global max pooling and global average pooling on the channel dimension N, respectively. These features are then concatenated on the channel dimension N to obtain a spatial score of size 2×M×M. After performing the Conv3 convolution operation, the spatial attention feature is obtained using the Sigmoid activation function. This process can be represented as:

[0123]

[0124] S235: Will Through broadcasting mechanisms and spatial attention characteristics N s,i After performing the Hadamard product, the enhanced feature map is obtained.

[0125]

[0126] S236: All enhanced feature maps and original feature map Each input matrix is ​​residually connected to obtain an M×M dimensional input matrix. The input matrix is ​​then flattened into a vector, and the mean value is calculated along dimension B to obtain the enhanced gene features.

[0127]

[0128] Here, Reshape(·) represents the dimensionality transformation operation. The parameter settings for convolution operations Conv1, Conv2, and Conv3 are shown in Table 1.

[0129] Table 1

[0130] Convolution operation name Parameter settings <![CDATA[Conv1]]> Number of kernel parameters = 1 × N × N / r, stride = 1 <![CDATA[Conv2]]> Number of kernel parameters = 1 × N / r × N, stride = 1 <![CDATA[Conv3]]> Number of kernel parameters = 1 × 2 × M × M, stride = 1

[0131] Furthermore, such as Figure 4 As shown in D, calculate the enhanced gene characteristics. The feature vector X corresponding to any two genes i and j GE (i,:) and X GE The Hadamard product of (j,:) is used to obtain the feature fusion vector. This feature fusion vector is then input into the nonlinear activation function ReLU and processed by a multilayer perceptron to obtain the predicted values ​​of gene pairs. as follows:

[0132]

[0133] Where y is the tag value of the gene pair. and represents the predicted value output by the gene regulation network inference model, and n is the number of samples in the training batch. The multilayer perceptron described above consists of two fully connected layers, each containing a ReLU function and a Sigmoid function.

[0134] Furthermore, in step S3, to avoid insufficient information extraction due to the extreme sparsity of the baseline network provided by Non-Specific CHIP-seq, the following method is adopted: Figure 6 The diagram illustrates a contrastive learning framework for maximizing X. GE and X EE The consistency of the feature vectors of the same gene in the graph is used to obtain the graph contrastive learning regularization term N. gc The specific method is as follows:

[0135] S31: Randomly discard the normalized gene expression profile X F After processing a subset of values ​​from the implicit connection matrix A, the results are fed into a graph convolutional neural network to obtain the gene embeddings under explicit connections.

[0136]

[0137] in, A d Let A be the join matrix after discarding some values, and I be the identity matrix. for The degree matrix, For X F Discarding a portion of the gene expression profile;

[0138] S32: For the enhanced genetic trait X GE and gene embedding X EE The feature vector u of the m-th gene m =X GE (m,:) and v m =X EE (m,:), calculate u m With v m The first gene pair function P(u) between m ,v m ) and v m with u m The second gene pair function P(v) between m ,u m ),as follows:

[0139]

[0140] Where θ is the neural network fitting parameter, τ represents the temperature coefficient, and 1 [m≠n] The indicator function is 1 when m≠n and 0 when m=n; the model hyperparameters are shown in Table 2.

[0141] Table 2

[0142]

[0143] S33: Based on the first gene pair function P(u m ,v m ) and the second gene pair function P(v m ,u m ), thus obtaining the graph comparison learning regularization term N. gc :

[0144]

[0145] Where N represents the number of genes.

[0146] Based on the above graph comparison learning regularization term N gc The expression for the loss function Loss used by the gene regulatory network inference model during the training phase is defined as follows:

[0147] Loss = αLoss bc +βN gc

[0148] Where Loss represents the total loss function value, Loss bc N represents the binary cross-entropy loss term. gc The graph contrastive learning regularization term is represented by α and β, both of which are weighting coefficients. The binary cross-entropy loss term Loss... bc The method to obtain it is as follows:

[0149]

[0150] Where y is the tag value of the gene pair. and y represents the predicted value output by the gene regulatory network inference model, where n is the number of samples in the training batch, and y is the predicted value. s Let be the tag value of the s-th gene pair. is the predicted value for the s-th gene pair.

[0151] The Adam optimization strategy was used to iteratively update the model parameters, closely monitoring the model's performance on the validation set and fine-tuning the hyperparameter configuration accordingly to achieve optimal model efficiency. Subsequently, all models were trained and tested on the same dataset, and a comprehensive comparison was made based on the AUROC and AUPRC evaluation metrics.

[0152] The AUROC metric is derived by plotting the ROC curve using the true positive rate (TPR, also known as recall) and false positive rate (FPR) at different thresholds, and then calculating the area under the curve. AUROC is mainly used to measure the model's ability to distinguish between positive and negative samples. Its value ranges from 0 to 1. The closer the AUROC value is to 1, the better the model's performance.

[0153] The AUPRC metric is derived by plotting the precision (PR) curve at different recall rates and then calculating the area under the curve. AUPRC better reflects the model's performance on the minority class when dealing with imbalanced datasets and binary classification problems. Its value ranges from 0 to 1, with higher values ​​indicating better model performance.

[0154] Among existing model systems, GNNLINK efficiently integrates high-order connection information in gene regulatory networks through multi-layer graph convolutional neural networks; GENELINK uses graph attention neural networks to calculate attention weights between gene connections, thereby inferring gene regulatory networks; STGRNS creatively segments and splices gene expression profiles and introduces BERT's encoder architecture to infer gene regulatory networks; GENIE3, based on the random forest algorithm, ranks the interaction scores between regulatory factors and target genes; GRNBoost2 cleverly combines GENIE3 with the Arboreto computational framework, significantly accelerating the inference process of gene regulatory networks; and the GNE model utilizes multi-layer perceptrons and the global structural information of gene regulatory networks to perform in-depth cluster analysis of genes. These models, each with its own characteristics, constitute representative methods in the current field of gene regulatory network inference.

[0155] Regarding the AUROC metric, the performance of the GRLGRN model proposed in this invention is compared with other models, and the results are visualized. Figure 7 and Figure 8 This is clearly demonstrated in the text. Specifically, Figure 7 China A and Figure 8Figure A shows the AUROC and AUPRC metrics of the GRLGRN model across seven cell types under different scales of the gold standard network framework. The data shows that the GRLGRN model demonstrates superior performance compared to other models on the vast majority of datasets. In the three different types of gold standard network datasets, the GRLGRN model outperforms the second-ranked model on 85.7% (12 / 14), 57.1% (8 / 14), and 92.8% (13 / 14) of the datasets, respectively, with performance improvements of approximately 6.93%, 4.3%, and 6.43%. Regarding the AUPRC metric, the GRLGRN model outperforms the second-best model on 73.8% (31 / 42) of the datasets, with an average improvement of at least approximately 7.84% under the gold standard network framework.

[0156] Figure 7 China B and Figure 8 Table B reveals the distribution of AUROC and AUPRC scores for the GRLGRN model and other comparative models on the datasets. The benchmark datasets cover ground truth networks with TFs+500 (above) and TFs+1000 (below). Finally, Table 3 summarizes the average AUROC and AUPRC scores of the GRLGRN model and six other comparative models on 42 datasets, further validating the significant advantages of the GRLGRN model in gene regulatory network inference.

[0157] Table 3

[0158]

[0159] Example 2

[0160] Based on the same inventive concept as the method in Embodiment 1, this embodiment of the invention also provides a gene regulatory network inference system based on graph representation learning. This system is used to implement the gene regulatory network inference method based on graph representation learning in Embodiment 1, such as... Figure 9 As shown, the system specifically includes:

[0161] The gene regulation network dataset construction module 100 is used to acquire single-cell RNA sequencing data. Based on the regulatory relationship between genes in the single-cell RNA sequencing data, a gene regulation network dataset is constructed. In the gene regulation network dataset, gene pairs with interactions are regarded as positive samples, gene pairs with unknown interactions are regarded as negative samples, and positive and negative samples are randomly divided into training set, validation set and test set.

[0162] The gene regulation network inference model construction module 200 is used to input the prior gene regulation network and gene expression profile generated from positive samples in the training set into the pre-constructed gene regulation network inference model for iterative training, and obtain the gene pair prediction values ​​calculated in each iteration.

[0163] The gene pair regulatory relationship prediction module 300 is used to construct a loss function based on the binary cross-entropy loss term and the graph contrast learning regularization term. The loss function is used to calculate the loss value between the predicted value and the true value of the gene pair to update the model parameters. If the training rounds reach a threshold, the model training parameters are output and loaded into the gene regulation network inference model to predict the regulatory relationship of gene pairs with unknown interactions and obtain the prediction results.

[0164] This embodiment provides a gene regulatory network inference system based on graph representation learning, used to implement the aforementioned gene regulatory network inference method based on graph representation learning. Therefore, the specific implementation of the gene regulatory network inference system based on graph representation learning can be found in the previous embodiment section of the gene regulatory network inference method based on graph representation learning. For example, the gene regulatory network dataset construction module 100, the gene regulatory network inference model construction module 200, and the gene pair regulatory relationship prediction module 300 are used to implement steps S1, S2, and S3 in the above-mentioned gene regulatory network inference method based on graph representation learning, respectively. Therefore, the specific implementation can be referred to the description of the corresponding embodiments. To avoid redundancy, it will not be repeated here.

[0165] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0166] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0167] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0168] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A gene regulatory network inference method based on graph representation learning, characterized in that, Includes the following steps: S1: Obtain single-cell RNA sequencing data. Based on the regulatory relationships between genes in the single-cell RNA sequencing data, construct a gene regulation network dataset. In the gene regulation network dataset, gene pairs with interactions are considered as positive samples, gene pairs with unknown interactions are considered as negative samples, and positive and negative samples are randomly divided into training set, validation set, and test set. S2: Input the prior gene regulation network and gene expression profile generated from positive samples in the training set into the pre-constructed gene regulation network inference model for iterative training, and obtain the gene pair prediction values ​​calculated in each iteration; S3: Construct a loss function based on the binary cross-entropy loss term and the graph contrastive learning regularization term. Use the loss function to calculate the loss between the predicted and true values ​​of gene pairs to update the model parameters and determine whether the training epochs have reached a threshold. If not, return to step S2 and continue the iterative calculation; If so, output the model training parameters, load the model training parameters into the gene regulation network inference model, predict the regulatory relationship of gene pairs with unknown interactions, and obtain the prediction results; In S2, the method for obtaining the predicted gene pair values ​​calculated in each iteration is as follows: S21: Based on the directed graph of the prior gene regulation network generated by positive samples in the training set, construct subgraphs including directed subgraphs of regulation gene-target gene and their reverse graphs, directed subgraphs of regulation gene-regulation gene and their reverse graphs, and gene self-connection subgraphs. Each subgraph corresponds to an adjacency matrix. Perform data processing on all adjacency matrices to obtain an implicit connection tensor including multiple implicit connection matrices. S22: Based on the normalized gene expression profile, multiple implicit connection matrices are sequentially input into the graph convolutional neural network. For each implicit connection matrix, a gene feature matrix is ​​output. The gene feature matrix is ​​subjected to layer normalization. The normalized gene feature matrices are concatenated to obtain the graph embedding tensor. S23: Based on graph embedding tensors, channel attention and spatial attention mechanisms are used to enhance gene features, resulting in enhanced gene features. The Hadamard product of the feature vectors corresponding to any two genes in the enhanced gene features is calculated to obtain the feature fusion vector. The feature fusion vector is processed by a multilayer perceptron to output the predicted value of the gene pair.

2. The gene regulatory network inference method based on graph representation learning according to claim 1, characterized in that, In S21, an implicit connectivity tensor containing multiple implicit connectivity matrices is obtained. The method is as follows: S211: Concatenate all adjacency matrices to obtain tensor A. s ∈{0,1} 5×N×N , will A s Simultaneously input to two parameterized layers, yielding tensors respectively. and In each parameterized layer, the following relationship is satisfied: Where N is the number of genes in the gene regulatory network, Q (i) (j,:,:) represents the output Q of the i-th parameterized layer. (i) The j-th matrix in the matrix, where B is the tensor Q. (i) The total number of matrices in (j,:,:). Gene connection types in gene regulatory networks A s (k,:,:) represents the input tensor A. s The sub-adjacency matrix of the k-th connection type, These are the normalized training parameters in the i-th layer; S212: The tensor Q output from the two parameterized layers (1) and Q (2) Performing an inner product along the corresponding dimensions yields the implicitly connected tensor. This process can be represented as: A L (j,:,:)=Q (1) (j,:,:)·Q (2) (j,:,:),j=1,2,...,B in, Indicates the output implicit connection tensor A L The j-th implicit connection matrix in the matrix.

3. The gene regulatory network inference method based on graph representation learning according to claim 1, characterized in that, In S22, the graph embedding tensor Each tensor X IE The expression for (j,:,:) is as follows: Where H represents the gene feature dimension after graph embedding. Indicates the output implicit connection tensor A L Let N be the j-th implicit connection matrix, N be the number of genes in the gene regulatory network, and I be the identity matrix. for The degree matrix, X F Here, W represents the normalized gene expression profile, W represents the training parameters of the graph convolutional neural network, LayerNorm(·) represents the layer normalization function, and B represents the number of tensors in the graph embedding tensor.

4. The gene regulatory network inference method based on graph representation learning according to claim 1, characterized in that, In S23, the method for obtaining the enhanced gene characteristics is as follows: S231: Graph Embedding Tensors Perform a dimensional transformation to obtain Where M is the square root of H, Given B feature tensors Composition, i = 1, 2, ..., B, and each tensor can be represented as a feature matrix with N channels. S232: Yes Each feature tensor Feature matrices on all N channels Global max pooling and global average pooling are performed separately to obtain two channel features. These two channel features are then input into a parameter-shared convolutional algorithm, followed by element-wise summation and a Sigmoid activation function to obtain channel attention. S233: Transfer the feature tensor Through broadcasting mechanism and channel attention feature N c,i After performing the Hadamard product, the output features are obtained. This process can be represented as: S234: Along Two spatial features are obtained by performing global max pooling and global average pooling on the channel dimension N, respectively. These features are then concatenated on the channel dimension N and followed by convolution to obtain the spatial attention feature. This process can be represented as: S235: Will Through broadcasting mechanisms and spatial attention characteristics N s,i After performing the Hadamard product, the enhanced feature map is obtained. S236: All enhanced feature maps and original feature map Each input matrix is ​​residually connected to obtain an M×M dimensional input matrix. The input matrix is ​​then flattened into a vector, and the mean value is calculated along dimension B to obtain the enhanced gene features. Here, Reshape(·) represents the dimension transformation operation.

5. The gene regulatory network inference method based on graph representation learning according to claim 1, characterized in that, Predicted value of gene pairs The calculation method is as follows: Where y is the tag value of the gene pair. and is the predicted value output by the gene regulatory network inference model, and n is the number of samples in the training batch.

6. The gene regulatory network inference method based on graph representation learning according to claim 1, characterized in that, In S3, the expression for the loss function is as follows: Loss=αLoss bc +βN gc Where Loss represents the total loss function value, Loss bc N represents the binary cross-entropy loss term. gc The graph comparison learning regularization term is represented by α and β, which are both weighting coefficients.

7. The gene regulatory network inference method based on graph representation learning according to claim 6, characterized in that, The binary cross-entropy loss term Loss bc The method to obtain it is as follows: Where y is the tag value of the gene pair. and y represents the predicted value output by the gene regulatory network inference model, where n is the number of samples in the training batch, and y is the predicted value. s Let be the tag value of the s-th gene pair. is the predicted value for the s-th gene pair.

8. The gene regulatory network inference method based on graph representation learning according to claim 6, characterized in that, The graph comparison learning regularization term N gc The method to obtain it is as follows: S31: Randomly discard the normalized gene expression profile X F After processing a subset of values ​​from the implicit connection matrix A, the results are fed into a graph convolutional neural network to obtain the gene embeddings under explicit connections. in, A d Let A be the join matrix after discarding some values, I be the identity matrix, and D be the join matrix. The degree matrix, For X F Discarding a portion of the gene expression profile; S32: For the enhanced genetic trait X GE and gene embedding X EE The feature vector u of the m-th gene m =X GE (m,:) and v m =X EE (m,:), calculate u m With v m The first gene pair function P(u) between m ,v m ) and v m with u m The second gene pair function P(v) between m ,u m ),as follows: Where θ is the neural network fitting parameter, τ represents the temperature coefficient, and 1 [m≠n] This represents an indicator function, which is 1 when m≠n and 0 when m=n; S33: Based on the first gene pair function P(u m ,v m ) and the second gene pair function P(v m ,u m ), thus obtaining the graph comparison learning regularization term N. gc : Where N represents the number of genes.

9. A gene regulatory network inference system based on graph representation learning, characterized in that, The system is used to implement the gene regulatory network inference method based on graph representation learning as described in any one of claims 1 to 8, specifically including: The gene regulation network dataset construction module is used to acquire single-cell RNA sequencing data. Based on the regulatory relationships between genes in the single-cell RNA sequencing data, a gene regulation network dataset is constructed. In the gene regulation network dataset, gene pairs with interactions are regarded as positive samples, gene pairs with unknown interactions are regarded as negative samples, and positive and negative samples are randomly divided into training set, validation set and test set. The gene regulation network inference model construction module is used to input the prior gene regulation network and gene expression profile generated from positive samples in the training set into the pre-constructed gene regulation network inference model for iterative training, and obtain the gene pair prediction values ​​calculated in each iteration; The gene pair regulatory relationship prediction module is used to construct a loss function based on a binary cross-entropy loss term and a graph contrastive learning regularization term. The loss function is used to calculate the loss value between the predicted and actual values ​​of gene pairs to update the model parameters. When the training rounds reach a threshold, the model training parameters are output and loaded into the gene regulation network inference model to predict the regulatory relationship of gene pairs with unknown interactions, and the prediction results are obtained.