A Method for Inferring Single-Cell Multi-Omics Gene Regulatory Networks Based on Contrastive Learning
The single-cell multiomics gene regulation network is constructed through comparative learning methods, which solves the problem of complex construction and high computing power requirements in the existing technology, and realizes efficient and accurate identification of gene regulation networks and determining key nodes.
Patent Information
- Application Number
- CN202411264805.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-09-10
AI Technical Summary
The existing methods are complex and require a lot of computing power to build a gene regulation network, and lack interpretability analysis of the internal mechanisms of the network.
Using a method based on contrast learning, the structural entropy of the node is calculated to determine the key nodes by obtaining single-cell RNA sequencing data and chromatin accessibility data, preprocessing and graph structure conversion is performed, and the node2vec and Skip-gram models are used to construct the feature space matrix.
It improves the accuracy and computing efficiency of the construction of gene regulation networks, can identify unlabeled regulatory relationships, determine key nodes, and be interpretable.
Smart Images

Figure CN119207582B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics technology, and particularly relates to a method for inferring gene regulatory networks of single-cell multi-omics based on contrastive learning. Background Art
[0002] Gene expression regulation determines cell function and state. Gene regulatory networks (GRNs) constitute a complex system for controlling gene expression within cells, which reflects the regulatory interactions between transcription factors (TFs) and their target genes. In this system, genes regulate the expression of other genes through the interactions between the proteins encoded by transcription. GRN inference refers to the process of summarizing gene regulatory interactions. It involves transforming complex regulatory processes into interpretable network structures through computational methods.
[0003] Studying gene regulatory networks is an important way to reveal the mechanisms of disease treatment and key genes. Especially in cancer research, certain genes, called cancer driver genes, play a core role in the occurrence and development of tumors. Mutations in these genes can endow cancer cells with the ability of uncontrolled proliferation, escape from immune surveillance, and other defense mechanisms, thus becoming the key driving factors for cancer progression. For example, mutations in genes such as TP53, PIK3CA, KMT2C, ARID1A, KMT2D, LRP1B, PTEN, RB1, FAT4, KRAS are closely related to the formation of various cancer types. By deeply exploring gene regulatory networks, the transcriptional regulatory mechanisms of genes can be clarified, and then the therapeutic targets and pathological mechanisms of diseases can be revealed. By comparing samples extracted from the healthy state and analyzing the gene regulatory networks of samples analyzed in the diseased state, key genes in diseases can be analyzed.
[0004] Currently, the inference of gene regulatory networks has gone through three main development stages, namely, using multi-cell omics data, single-cell single-omics data, and single-cell multi-omics cell data for gene regulatory network inference. The comprehensive impact of multi-omics features on gene expression can significantly narrow the search space, making the integration of diverse data a key research direction in GRN inference.
[0005] Currently, there are mainly two data integration methods for single-cell multi-omics cell data. The first method involves inputting multi-omics data into a deep learning model through limited preprocessing to integrate the multi-omics data into the feature space. For example, GLUE is a variational autoencoder model capable of integrating multi-omics data. By embedding a prior knowledge network and adjusting the variational autoencoder model, GLUE significantly reduces the impact of data heterogeneity on the model. The second method is heterogeneous network embedding, which is a computer method that can model multiple mechanisms simultaneously. The heterogeneous network abstracts various types of data and their relationships into multiple nodes and edges, and then uses machine learning methods to process and extract useful information. GEEK is a classic heterogeneous network evaluation framework that constructs various "meta-paths" to extract more information, such as protein, gene-binding associations, and chromatin accessibility to predict gene expression levels.
[0006] However, these existing methods focus on using multi-omics data to construct a single regulatory network. The method for constructing the regulatory network is complex and requires a large amount of computing power. Moreover, the existing methods only study the mechanism of the network at the overall level and require an interpretable analysis of the internal mechanism of the network. Summary of the Invention
[0007] Aiming at the above deficiencies in the prior art, the single-cell multi-omics gene regulatory network inference method based on contrastive learning provided by the present invention solves the problems that the method for constructing the regulatory network in the prior art is complex and requires a large amount of computing power.
[0008] In order to achieve the above invention object, the technical solution adopted by the present invention is as follows:
[0009] Provide a single-cell multi-omics gene regulatory network inference method based on contrastive learning, which includes the steps of:
[0010] S1. Obtain single-cell RNA sequencing data and single-cell chromatin accessibility data, and preprocess them to obtain a prior network of multi-omics data;
[0011] S2. Load the prior regulatory network into a graph structure including transcription factor nodes and target gene nodes, and use the node2vec method to convert the graph structure into a feature space matrix including several hidden vectors;
[0012] S3. Calculate the Pearson correlation coefficient of the two nodes on each edge in the graph structure. When it is greater than a preset threshold, it is used as a node pair, and the node pair and its corresponding feature vector and hidden vector are stored in the dataset;
[0013] S4. Divide the data in the dataset into several batches and randomly input them into the contrastive learning model to obtain the node type of each node, the pairing degree of the node pair, and the node feature reconstruction vector;
[0014] S5. Use a regulatory network interpreter to extract the mapping of two nodes in the reconstructed vector of node features in the feature space, and calculate the structural entropy of the nodes based on the mapping to determine whether the nodes are key nodes.
[0015] Furthermore, the method for preprocessing the prior network of multi-omics data includes:
[0016] Cluster the single-cell RNA sequencing data and single-cell chromatin accessibility data, and use the MACS2 tool to analyze the peaks of the clustered single-cell chromatin accessibility data to generate the coordinate data of the peaks;
[0017] Obtain the sequence-specific motifs related to transcription factors in the Cis-BP database and the human gene data and their corresponding location information in the human genome database;
[0018] Use the bedtools tool to pair the coordinate information of the peaks with the sequence-specific motifs and the coordinate information of the peaks with the human genes respectively;
[0019] Connect the sequence-specific motifs to the genes according to the pairing information to obtain a multi-omics network file containing gene, motif information, and peak coordinate data;
[0020] Obtain the transcription factor and motif datasets, and delete the genes and motifs in the multi-omics network file that are not located in the transcription factor and motif datasets to obtain the prior regulatory network of the multi-omics data.
[0021] The beneficial effects of the above technical solutions are as follows: This preprocessing scheme constructs an accurate multi-omics network file by clustering the single-cell RNA sequencing data and single-cell chromatin accessibility data and combining the transcription factor motif information with the gene location information, thereby improving the accuracy of regulatory network construction, providing important information for revealing gene regulatory mechanisms, and providing strong data support for the next step of the invention.
[0022] Furthermore, the method for converting the graph structure into a feature space matrix using the node2vec method includes:
[0023] Set the number of times each node in each graph structure is visited, then randomly select a node in the graph structure and start the random walk strategy to obtain the walk paths of all nodes in the graph structure;
[0024] Input the walk paths of all nodes into the Skip-gram model to obtain the hidden vectors of the walk paths of each node, and use all the hidden vectors to form a feature space matrix.
[0025] The beneficial effects of the above technical solution are as follows: The present invention adopts the node2vec method. By setting the number of node visits and starting the random walk strategy, the walk paths of all nodes in the graph structure are obtained. Subsequently, the walk paths are input into the Skip-gram model to obtain the hidden vectors of each node, and a feature space matrix is formed. This method can not only effectively capture the local and global relationships between nodes in the graph structure, but also map various node information to the same feature space, enabling the subsequent model to better understand and predict the regulatory relationships between nodes. In addition, this method is applicable to large-scale and complex graph structures, improving the computational efficiency.
[0026] Further, the objective function of the Skip-gram model is:
[0027]
[0028]
[0029] where L is the objective function; is the probability that the next node of node k is node m ; V is the set of nodes composed of all nodes in the graph structure; is V any node in k that is not node ; is the similarity or proximity function; is the Skip-gram model;
[0030] The loss function of the Skip-gram model is:
[0031]
[0032] where is the loss function.
[0033] The beneficial effects of the above technical solution are as follows: The Skip-gram model adopted in this solution can effectively quantify and predict the pairing probability between nodes by defining the conditional probability and similarity function between nodes and combining the negative log-likelihood loss function, thereby establishing the correlation between nodes in the graph structure. This method not only improves the prediction accuracy, but also is applicable to large-scale graph data. By learning the low-dimensional representation of nodes, it reduces the data complexity and improves the computational efficiency.
[0034] Further, the feature vector represents a single value of the chromatin accessibility score or gene expression level of each node.
[0035] Furthermore, the contrastive learning model includes a siamese encoder, a multi-head attention mechanism, and a multi-task decoder;
[0036] The method for obtaining the node type, the pairing degree of node pairs, and the node feature reconstruction vector of each node using the contrastive learning model includes:
[0037] Split the feature vectors and hidden vectors of the two nodes of each node pair in the dataset u and v under the corresponding nodes, and input them into a siamese encoder respectively to calculate the hidden vectors and ;
[0038] Calculate the attention weights of the hidden vector using the multi-head attention mechanism, extract the attention matrix from the attention weights, and pass it to another multi-head attention mechanism to calculate the attention weight vector of the hidden vector ;
[0039] According to the attention weight vector, use the three output heads of the multi-task decoder to output the node type, the pairing degree of node pairs, and the node feature reconstruction vector respectively.
[0040] The beneficial effects of the above technical solution are: The contrastive learning model adopted in this solution can capture the local and global features of node pairs simultaneously through the structure of the siamese encoder, the multi-head attention mechanism, and the multi-task decoder, focus on the important information in different subspaces, extract and transfer the attention matrix, and realize the joint learning of multiple objectives. This method not only improves the prediction accuracy but also retains the unique information of each node and enhances the representation ability of the model.
[0041] Furthermore, the expressions of the attention weights and the attention weight vector are respectively:
[0042]
[0043]
[0044] where is the attention weight; is a linear neural network; Q is the query vector; K is the key vector; is the transpose; is the value vector; is the dimension of the key vector and the query vector, used to scale the attention scores to stabilize the gradients during training; is the attention matrix; is the attention weight vector;
[0045] The expression for the output head prediction node type is:
[0046] ,
[0047] where, is the output head for predicting the node type; is a mathematical function that converts any real number vector into a probability distribution; is the activation function; is the intermediate value; is the Dropout layer;
[0048] The expression for the pairing degree of the output head prediction node pair is:
[0049] ,
[0050] where, is the pairing degree of the node pair; is the intermediate parameter;
[0051] The expression for the output head to reconstruct the node feature reconstruction vector is:
[0052]
[0053] where, is the node u and v 's node feature reconstruction vectors.
[0054] The beneficial effects of the above technical solution are: The contrastive learning model adopted in this solution can effectively capture the complex relationships between node pairs, focus on important information, and quantify the correlation between node pairs through the QS attention mechanism. This method not only improves the accuracy and robustness of prediction, but also enhances the generalization ability and anti-overfitting ability of the model. In addition, the reconstruction of the node feature reconstruction vector by the output head retains the unique information of each node, improves the representation ability of the model, and then infers the gene regulatory network.
[0055] Furthermore, the loss function for the reconstruction part of the node feature reconstruction vector is:
[0056]
[0057]
[0058] where, N is the total number of all types of nodes in the graph structure; is the power function of the natural logarithm; is the cosine similarity function; is the temperature coefficient of the cosine similarity function; is the midpoint between node and ; is the feature vector and hidden vector of node u ; is the feature vector and hidden vector of node v ; is the averaging function; is the summation function.
[0059] The beneficial effects of the above technical solution are: The loss function of this solution can improve the accuracy of feature reconstruction, enhance the sensitivity and flexibility of the model to similarity measurement, and capture the common features between node pairs by adopting the design of the cosine similarity function, temperature coefficient, and midpoint "anchor". This method not only improves the representation ability of the reconstructed vector for the common characteristics of node pairs, but also optimizes the parameters of the model by minimizing the reconstruction error, improving the efficiency and accuracy of node feature reconstruction.
[0060] Furthermore, the expression for calculating the structural entropy of a node based on mapping is:
[0061] ,
[0062] where is the feature vector and hidden vector of node j ; is the Euclidean distance between the mapping j of node and ; is the structural entropy of node j ; is the degree of node j ; j is any one of the two nodes corresponding to the node feature reconstruction vector.
[0063] The beneficial effects of the present invention are: After obtaining the feature space matrix of this solution, most of the irrelevant node pairs can be removed through the Pearson correlation coefficient between any two nodes, which can reduce the magnitude of the dataset, thereby overcoming the problem that the method of constructing a regulatory network is complex and requires a large amount of computing power; due to the relatively large quantity correlation in the dataset, when performing contrast learning, it can better match the unlabeled regulatory relationships, thus ensuring the recognition accuracy of the type and the pairing degree of node pairs.
[0064] When performing regulatory network inference, this solution can associate multi-omics information through preprocessing. Then, the node2vec method can be used to obtain information on different dimensions of cells, ensuring the accuracy of subsequent inference processes. By combining the method of contrastive learning, this solution can match unlabeled regulatory relationships, further improving the accuracy of node (gene and transcription factor) types. Finally, the structural entropy obtained by the regulatory network interpreter of this solution can quickly determine key nodes. Since the selection of key nodes is based on structural entropy, it also has interpretability. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a flowchart of a single-cell multi-omics gene regulatory network inference method based on contrastive learning.
[0066] Figure 2 It is a performance comparison diagram of node2vec and metapath2vec on the GM12878 dataset, where (a) is the performance comparison diagram based on the AUPR index; (b) is the performance comparison diagram based on the F-Score index.
[0067] Figure 3 It is a t-SNE visualization comparison diagram of node2vec and metapath2vec on the GM12878 dataset; where (a) is the t-SNE visualization comparison diagram of node2vec; (b) is the t-SNE visualization comparison diagram of metapath2vec. DETAILED DESCRIPTION OF THE INVENTION
[0068] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0069] Refer to Figure 1 , Figure 1 which shows a flowchart of a single-cell multi-omics gene regulatory network inference method based on contrastive learning; as Figure 1 shown, this method S includes steps S1 to S5.
[0070] In step S1, single-cell RNA sequencing data and single-cell chromatin accessibility data are obtained and preprocessed to obtain a prior network of multi-omics data.
[0071] The single-cell RNA sequencing data and single-cell chromatin accessibility data collected in this solution are from "Integrated single-cell analysis maps the continuous regulatory landscape of human hematopoietic differentiation", Chinese translation: "Integrated single-cell analysis maps the continuous regulatory landscape of human hematopoietic differentiation", author: Buenrostro et al., journal: Cell, publication date: 2018-05.
[0072] In an embodiment of the present invention, the method for preprocessing the prior network of multi-omics data includes steps S11 to S15:
[0073] S11. Cluster the single-cell RNA sequencing data and single-cell chromatin accessibility data, and use the MACS2 tool to analyze the peaks of the single-cell chromatin accessibility data after clustering to generate coordinate data of the peaks;
[0074] During implementation, the detailed implementation process of clustering in this solution is as follows:
[0075] First, read the matrices of single-cell RNA data and single-cell chromatin accessibility data; then standardize the data using the LIGER tool; next, perform joint non-negative matrix factorization (iNMF) on the standardized data to extract common and dataset-specific genes; then cluster the cells according to the genes obtained by iNMF, and achieve clustering correspondence between datasets.
[0076] The clustering in this solution is equivalent to combining and analyzing the two types of data to obtain which cells in the two types of data belong to which class, and then respectively putting the two types of data of the cells into the corresponding clusters.
[0077] S12. Obtain the sequence-specific motifs related to transcription factors in the Cis-BP database and the human gene data and their corresponding position information in the human genome database;
[0078] S13. Use the bedtools tool to pair the coordinate information of the peaks with the sequence-specific motifs and the coordinate information of the peaks with the human genes respectively;
[0079] S14. Connect the sequence-specific motifs to the genes according to the pairing information to obtain a multi-omics network file containing gene, motif information, and coordinate data of the peaks;
[0080] S15. Obtain the transcription factor and motif datasets, and delete the genes and motifs not located in the transcription factor and motif datasets in the multi-omics network file to obtain the prior regulatory network of the multi-omics data.
[0081] In step S2, the prior regulatory network is loaded into a graph structure including transcription factor nodes and target gene nodes, and the graph structure is converted into a feature space matrix including a number of hidden vectors by using the h method; preferably, this solution uses networkx a tool to load the prior regulatory network into a graph structure including transcription factor nodes and target gene nodes.
[0082] In implementation, the method steps S21 and S22 of preferably using the node2vec method to convert the graph structure into a feature space matrix include:
[0083] S21. Set the number of times each node in each graph structure is visited, and then randomly select a node in the graph structure and start the random walk strategy to obtain the walk paths of all nodes in the graph structure; preferably, this solution starts the random walk strategy according to the probabilities p and q , where the probability p controls the probability of whether to visit duplicate nodes, and the probability q determines the probability of whether to visit an already visited node or an unvisited node.
[0084] S22. Input the walk paths of all nodes into the Skip-gram model to obtain the hidden vectors of the walk paths of each node, and use all the hidden vectors to form a feature space matrix.
[0085] Preferably, the objective function of the Skip-gram model in this solution is:
[0086]
[0087]
[0088] where L is the objective function; is the probability that the next node of node k is node m ; V is the node set composed of all nodes in the graph structure; is V any node in k that is not node ; is the similarity or proximity function; is the Skip-gram model;
[0089] The loss function of the Skip-gram model is:
[0090]
[0091] where is the loss function.
[0092] In this solution, the optimization objective and the damage function are optimized through the Skip-gram model. The maximization of covering the edges in the coverage structure diagram is used as the optimization objective. The more edges are covered, the more regulatory relationships are represented. Then this solution may find more unproposed gene regulatory relationships.
[0093] In step S3, calculate the Pearson correlation coefficient of the two nodes on each edge in the graph structure. When it is greater than the preset threshold, it is used as a node pair, and the node pair, its corresponding feature vector, and hidden vector are stored in the dataset.
[0094] In step S4, divide the data in the dataset into several batches and randomly input them into the contrast learning model to obtain the node type of each node, the pairing degree of the node pair, and the node feature reconstruction vector;
[0095] In an embodiment of the present invention, the contrast learning model includes a twin encoder, a multi-head attention mechanism, and a multi-task decoder; among them, there are two twin encoders and two multi-head attention mechanisms.
[0096] The method of using the contrast learning model to obtain the node type of each node, the pairing degree of the node pair, and the node feature reconstruction vector includes steps S41 to S43:
[0097] S41. Split the feature vectors and hidden vectors of the two nodes of each node pair in the dataset u and v under the corresponding nodes, that is, and mentioned later; then input them into a twin encoder respectively to calculate the obtained hidden vectors and .
[0098] S42. Use the multi-head attention mechanism to calculate the attention weights of the hidden vector , extract the attention matrix in the attention weights, and pass it to another multi-head attention mechanism to calculate the attention weight vector of the hidden vector .
[0099] Among them, the expressions of the attention weights and the attention weight vector are respectively:
[0100]
[0101]
[0102] Among them, is the attention weight; is the linear neural network; Qis the query vector; K is the key vector; is the transpose; is the value vector; is the dimension of the key vector and the query vector, used to scale the attention scores to stabilize the gradients during training; is the attention matrix; is the attention weight vector;
[0103] S43. According to the attention weight vector, use the three output heads of the multi-task decoder to output the node type, the pairing degree of the node pair, and the node feature reconstruction vector respectively.
[0104] In this solution, the linear transformation weights of Query and Key are shared by two self-attention mechanisms, while Value has independent linear transformation weights. This solution optimizes the attention mechanism in this way to adapt to the task of quantifying node relationships.
[0105] During implementation, the preferred expression for the output head to predict the node type in this solution is:
[0106] ,
[0107] where, is the output head for predicting the node type; is a mathematical function that converts an arbitrary real vector into a probability distribution; is the activation function; is the intermediate value; is the Dropout layer.
[0108] The expression for the output head to predict the pairing degree of the node pair is:
[0109] ,
[0110] where, is the pairing degree of the node pair; is the intermediate parameter;
[0111] The expression for the output head to reconstruct the node feature reconstruction vector is:
[0112]
[0113] where, is the node u and v 's node feature reconstruction vector.
[0114] The loss function for the reconstruction part of the node feature reconstruction vector is:
[0115]
[0116]
[0117] Among them, N is the total number of all types of nodes in the graph structure; is the power function of the natural logarithm; is the cosine similarity function; is the temperature coefficient of the cosine similarity function; is the node and the midpoint between; is the node u the feature vector and hidden vector of; is the node v the feature vector and hidden vector of; is the average function; is the summation function.
[0118] In step S5, the regulatory network interpreter is used to extract the mapping of two nodes in the node feature reconstruction vector in the feature space, and the structural entropy of the node is calculated based on the mapping to determine whether the node is a key node.
[0119] When implemented, the preferred expression for calculating the structural entropy of the node based on the mapping in this solution is:
[0120] ,
[0121] Among them, is the node j the feature vector and hidden vector of; is the node j the mapping of and the Euclidean distance between; is the node j the structural entropy of; is the node j the degree of; j is any one of the two nodes corresponding to the node feature reconstruction vector.
[0122] The regulatory network interpreter of this solution has two purposes: First, it can confirm which nodes in the network are key nodes; second, it can identify important modules in the regulatory network. When identifying key nodes, a lower structural entropy is considered to indicate that the node is more stable. This solution can set a threshold, and when the structural entropy is lower than this threshold, it can indicate that the node corresponding to the structural entropy is a key node.
[0123] The following is a detailed description of the effectiveness of the single-cell multi-omics gene regulatory network inference method of this solution in combination with specific comparative experiments:
[0124] Experiment, parameter, and dataset settings
[0125] This solution evaluates the effectiveness of inferring single-cell multi-omics gene regulatory networks through performance comparison. To evaluate the performance of the framework, this solution compares the evaluation framework with seven baseline methods: scMTNI, GNAT, INDEP, SCENIC, Lasso, AMuSR-tuned, and MRTLE.
[0126] scMTNI is a multi-task learning evaluation framework designed to infer the regulatory network of a specific cell type by optimizing hyperparameters such as marginal probability, sparse penalty, and the strength of the prior network. MRTLE is an algorithm based on probabilistic graphical models that uses the phylogenetic structure between species, multi-species transcriptome data, and sequence-specific patterns to infer genome-scale regulatory networks across species simultaneously. The GNAT algorithm utilizes the tissue hierarchy to share relevant information and infers tissue-specific gene co-expression networks, each modeled with a Gaussian Markov random field (GMRF). The AMuSR algorithm uses sparse block-sparse regression to estimate the activity of transcription factors and infers gene regulatory networks from expression datasets, decomposing the model coefficient matrix through multi-task learning to capture conservatism and dataset-specific interactions. INDEP is a single-task evaluation framework version of scMTNI that independently infers the regulatory network for each cell type, uses a dependence network model, and learns the regulatory network for each cell type through a score-based search algorithm. The LASSO method uses linear regression with L1 regularization to predict the expression profile of a target gene based on the expression profiles of candidate regulatory factors and infers the regulatory factors with non-zero coefficients as the regulatory factors of that gene. The SCENIC algorithm uses GENIE3 or GRNBoost2 as part of the Arboreto evaluation framework to infer transcription factor-target gene relationships, adopting a tree ensemble method that can be directly applied to each cell type-specific dataset.
[0127] During performance evaluation, this solution uses two metrics, AUPR (area under the precision-recall curve) and F-Score, to evaluate its efficacy in inferring regulatory networks. To ensure robustness, all experimental models start from five different random numbers, and the comparative experimental results represent the average of these five iterations. The prediction of the model uses the top-k validation method, which selects the top k edges (sorted by weight) inferred by the evaluation framework of this solution for validation and compares them with the true network containing exactly k edges.
[0128] The node2vec algorithm (a multi-omics integration algorithm, i.e., random walk) is an essential part of the evaluation framework for generating embedding vectors, and its key parameters include , , , , , and . Here, represents the dimension of the embedding space, refers to the length of the random walk, refers to the context range considered, refers to the number of walks per node, refers to the number of negative samples, is the probability of returning to the previous node during the walk, is the balance between exploration and exploitation during the random walk.
[0129] The hyperparameter configuration of the evaluation framework in this scheme is as follows: is set to 64, is set to 128, is set to 0.75, is set to 128, is set to 0.0001, is set to 128, is set to 10, is set to 5, is set to 8, is set to 1, is set to 0.25, is set to 4. All experiments were conducted on a server equipped with an NVIDIA RTX 3080 GPU. These parameter settings are the best results obtained after grid search. When using the metapath2vec method in the comparative experiment, this scheme adopted the template with the best performance in the Gene-Genome and bin-Gene templates reported in the research.
[0130] This scheme used a simulated dataset to evaluate the effectiveness of the framework in inferring regulatory networks. The dataset was simulated using the BoolODE method, representing single-cell data of 2,000 cells with a sparsity of 80%. It includes three different cell types: hematopoietic stem cells (HSCs), common myeloid progenitors (CMPs), and granulocyte-monocyte progenitors (GMPs). Each cell type is represented by a sub-dataset containing 15 transcription factors and 65 genes, and each gene is accompanied by its corresponding prior regulatory network. The evaluation framework used this dataset to infer regulatory networks and evaluated the performance of the inferred networks through AUPR and F-Score metrics.
[0131] The evaluation framework utilizes human hematopoietic differentiation data to evaluate the effectiveness of various modules in this protocol's evaluation framework. This dataset integrates single-cell accessibility (scATAC-seq) and single-cell RNA sequencing (scRNA-seq) data, covering a diverse set of immunophenotypic cells. The dataset is divided into eight different cell clusters, each accompanied by its existing regulatory network. These prior regulatory networks are based on motifs of transcription factors within a 5000bp window around their gene transcription start sites (TSSs), indicating regulatory interactions between genes and transcription factors. Thus, the regulatory network of the dataset contains two types of nodes, namely transcription factors (TFs) and genes, as well as TF-Gene interaction pairs. The evaluation framework of this protocol uses this dataset to infer the regulatory network and evaluates the performance of the inferred network through AUPR and F-Score metrics. The evaluation framework uses the GM12878 dataset. This dataset contains two types of nodes, genomic bins and genes, and three pairing relationships: genomic bin-genomic bin, gene-gene, and genomic bin-gene. It integrates information on chromatin accessibility (DNase-seq), protein-protein interactions, and gene expression levels. The prior regulatory network of this dataset believes that gene pairing is determined by protein interactions, protein epigenetic modifications establish pairings between genomic bins, and genomic bin-gene pairings depend on whether the gene exists in the genomic bin. This protocol adopts the Top-K method to test the AUPR and F-Score metrics of two embedding and two inference methods in this dataset respectively.
[0132] Comparative experiment
[0133] The performance of the evaluation framework proposed in this protocol in inferring the regulatory network was tested on three simulated datasets and compared with seven baseline methods. The comparison results are shown in Tables 1 and 2.
[0134] Table 1 Comparison between this protocol and 7 existing methods on simulated datasets (regarding AUPR metric)
[0135] AUPR CLMOGRI scMTNI GNAT INDEP Scenic Lasso AMuSR_tuned MRTLE Cell type 1 0.26 0.233273 0.210597 0.213886 0.22214 0.212391 0.215322 0.234724 Cell type 2 0.287 0.263825 0.238419 0.243309 0.251842 0.233849 0.24342 0.270785 Cell type 3 0.291 0.277375 0.260223 0.266422 0.260615 0.264009 0.284234 0.281862
[0136] Table 2 Comparison between this protocol and 7 existing methods on simulated datasets (regarding F1score metric)
[0137] F1score CLMOGRI scMTNI GNAT INDEP Scenic Lasso AMuSR_tuned MRTLE Cell type 1 0.285 0.277228 0.207921 0.212871 0.237624 0.212871 0.040346 0.262376 Cell type 2 0.324 0.290323 0.253456 0.24424 0.239631 0.24424 0.078049 0.285714 Cell type 3 0.31 0.276151 0.263598 0.280335 0.263598 0.267782 0.054393 0.280335
[0138] As can be seen from Table 1 and Table 2, the proposed solution outperforms other methods in terms of both AUPR and F-score. Compared with the multi-task model scMTNI, the proposed solution achieved significant improvements of 2.7%, 2.4%, and 1.4% on three datasets respectively, indicating that the multi-task configuration of the proposed solution is more effective in inferring regulatory networks.
[0139] Similarly, the model of the proposed solution maintained an advantage of more than 1.5% in the comparison with methods MRTLE and INDEP that focus on cell evolution-related regulatory networks, further confirming the superior performance of the evaluation framework in feature extraction. Compared with single-omics machine learning methods such as Lasso and SCENIC, the model proposed in this solution also showed significant performance improvements, indicating that the joint inference of multi-omics data can achieve better performance in regulatory network inference.
[0140] To evaluate the performance of the evaluation framework of the proposed solution in integrating multi-omics data, the GM12878 dataset was used to compare the node2vec method (random walk path strategy) and the metapath2vec method in the proposed solution. The performance of the node2vec method and the metapath2vec method was described by AUPR and F-Score metrics, and the two metrics obtained refer to Figure 2 the (a) and (b) in
[0141] As Figure 2 shown, when using the same network inference method, metapath2vec performed significantly worse than node2vec on 23 chromatin datasets. This is mainly because metapath2vec embeds feature vectors based on fixed patterns, which may miss some potential regulatory relationships and cannot fully learn the structural features of the network in complex network inference tasks.
[0142] The node2vec (random walk strategy) of the proposed solution outperformed metapath2vec in terms of both AUPR and F-score. This superior performance indicates that node2vec is better at capturing complex structural and semantic relationships in gene regulatory networks because its feature vector embedding is more flexible and detailed than the fixed pattern method of metapath2vec.
[0143] The proposed solution also performed dimensionality reduction and visualization analysis on the feature vector matrices integrated by the node2vec and metapath2vec methods, referring to the (a) and (b) in Figure 3 respectively, Figure 3 where Feature 1 in is the first coordinate axis of t-SNE; Feature 2 is the second coordinate axis of t-SNE; red dots represent genes, and blue dots represent genomic fragments. FromFigure 3 It can be seen that the feature vector matrix constructed by the node2vec method is significantly better than the metapath2vec method. This further confirms the effectiveness of this solution in integrating multiple omics data to infer regulatory networks.
[0144] Generally speaking, this solution can efficiently extract structured features from multi-omics data to facilitate the inference of regulatory networks.
Claims
1. A method for inferring single-cell multi-omics gene regulatory networks based on contrastive learning, characterized in that, Including the steps: S1. Obtain single-cell RNA sequencing data and single-cell chromatin accessibility data, perform clustering on the single-cell RNA sequencing data and the single-cell chromatin accessibility data, and use the MACS2 tool to analyze the peaks of the clustered single-cell chromatin accessibility data to generate coordinate data of the peaks; Obtain the sequence-specific motifs related to transcription factors in the Cis-BP database and the human gene data and their corresponding position information in the human genome database; Use the bedtools tool to pair the coordinate information of the peaks with the sequence-specific motifs and the coordinate information of the peaks with the human genes respectively; Connect the sequence-specific motifs to the genes according to the pairing information to obtain a multi-omics network file containing gene, motif information, and coordinate data of the peaks; Obtain the transcription factor and motif datasets, delete the genes and motifs not located in the transcription factor and motif datasets in the multi-omics network file to obtain a prior regulatory network of the multi-omics data; S2. Load the prior regulatory network into a graph structure including transcription factor nodes and target gene nodes, set the number of times each node in each graph structure is visited, and then randomly select a node in the graph structure to start the random walk strategy to obtain the walking paths of all nodes in the graph structure; Input all the node walking paths into the Skip-gram model to obtain the hidden vectors of each node walking path, and use all the hidden vectors to form a feature space matrix; S3. Calculate the Pearson correlation coefficient between two nodes on each edge in the graph structure. When it is greater than the preset threshold, it is used as a node pair, and store the node pair and its corresponding feature vector and hidden vector in the dataset; S4. Divide the data in the dataset into several batches, and split the feature vectors and hidden vectors of the two nodes of each node pair in the dataset under the corresponding nodes, and then input them into a siamese encoder respectively to calculate the hidden vectors u and v ; and ; Calculate the hidden vector using the multi-head attention mechanism of the attention weights, extract the attention matrix in the attention weights, and pass it to another multi-head attention mechanism to calculate the hidden vector of the attention weight vector; According to the attention weight vector, use the three output heads of the multi-task decoder to output the node type, the pairing degree of the node pair, and the node feature reconstruction vector respectively; S5. Use the regulatory network interpreter to extract the mapping of two nodes in the node feature reconstruction vector in the feature space, and calculate the node structure entropy based on the mapping to determine whether the node is a key node; The expression for calculating the node structure entropy based on the mapping is: , Among them, is the feature vector and hidden vector of node j ; is the mapping of node j ; is the Euclidean distance between ; is the structural entropy of node j ; is the degree of node j ; j is any one of the two nodes corresponding to the node feature reconstruction vector.
2. The single-cell multi-omics gene regulatory network inference method according to claim 1, wherein The objective function of the Skip-gram model is: Among them, L is the objective function; is the node k The probability that the next node of m is node V is the set of nodes composed of all nodes in the graph structure; is V any node in k that is not node is the similarity or proximity function; is the power function of the natural logarithm; is the Skip-gram model; The loss function of the Skip-gram model is: Among them, is the loss function.
3. The single-cell multi-omics gene regulatory network inference method according to claim 1, wherein The feature vector represents a single value of the chromatin accessibility score or gene expression level of each node.
4. The single-cell multi-omics gene regulatory network inference method according to claim 1, wherein The expressions for the attention weight and the attention weight vector are respectively: Among them, is the attention weight; is a linear neural network; Q is the query vector; K is the key vector; is the transpose; is the value vector; is the dimension of the key vector and the query vector; is the attention matrix; is the attention weight vector; The expression for the output head to predict the node type is: , Among them, is the output head for predicting the node type; is a mathematical function that converts any real number vector into a probability distribution; is the activation function; is the intermediate value; is the Dropout layer; The expression for the output head to predict the pairing degree of the node pair is: , Among them, is the pairing degree of node pairs; is an intermediate parameter; The expression for the output head to reconstruct the node feature reconstruction vector is: Among them, is the node feature reconstruction vector of u and v .
5. The single-cell multi-omics gene regulatory network inference method according to claim 4, wherein The loss function for the reconstruction part of the node feature reconstruction vector is as follows: Among them, N is the total number of all types of nodes in the graph structure; is the power function of the natural logarithm; is the cosine similarity function; is the temperature coefficient of the cosine similarity function; is the node and the midpoint between; is the node u the feature vector and the hidden vector of; is the node v the feature vector and the hidden vector of; is the mean function; is the summation function.
Citation Information
Patent Citations
Single-cell RNA (Ribonucleic Acid) sequence gene regulation and control inference method based on deep learning
CN116825204A
Gene regulatory network inference method based on multi-view layered hypergraph
CN116844645A