A Single-Cell Hi-C Data Augmentation Method, System, and Storage Medium
The VGAE model of multi-scale variational graph autoencoder combines low resolution and high resolution information of single-cell Hi-C data, solving data sparsity and noise problems, improving data quality and biological significance, and supporting more accurate genomic structure analysis.
Patent Information
- Application Number
- CN202510293942.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The prior art is difficult to effectively solve the sparseness and noise problems of single-cell Hi-C data, resulting in limited enhancement effect at high resolution and difficulty in cell clustering and multi-dimensional analysis.
A multi-scale variational graph autoencoder VGAE model is designed to effectively fuse low-resolution and high-resolution Hi-C data, feature extraction and matrix reconstruction are performed through graph neural networks, and data fusion and enhancement across scales are achieved.
It improves the integrity and quality of the single-cell Hi-C contact matrix, enhances the biological significance of the data, can provide richer genomic structural information at higher resolution, and supports more accurate positioning of the boundaries of TADs and Loops domains.
Smart Images

Figure CN119811510B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to bioinformatics technology, and specifically to a single-cell Hi-C data enhancement method, system and storage medium. Background Art
[0002] Genomic DNA does not exist in a simple linear state in the cell nucleus, but exists in the form of chromatin with a highly folded and condensed specific high-level spatial conformation, and is stored in the cell nucleus as a carrier of genetic information. Genomics is to collectively characterize and quantify all the genes of an organism, study issues such as the structure, function, evolution, and localization of the genome, and analyze the relationships between them and the impacts on organisms. A large number of studies have shown that the spatial structure of chromosomes is closely related to most biological processes in cells, including genomic regulation, stable expression of the genome, DNA replication, and chromosomal translocation; there are also studies indicating that the chromatin 3D structure is also related to diseases, and there are studies showing that the insulation characteristics of the boundaries of TADs (Topologically Associating Domains) have a significant impact on the epigenetic landscape of single cells.
[0003] Since most existing three-dimensional genome modeling methods focus on constructing the chromosomal / genomic structure of cell populations derived from bulk Hi-C (High-throughput Chromosome Conformation Capture) data (usually millions of cells). The reconstructed structure is the average structure of cell populations or subpopulations and cannot well capture the variability of individual cells. Although individual cells of the same type usually have some common structural characteristics, there are differences in the interactions between topological associated domains TADs, chromosomal loops within TADs, and the distribution of chromosomal compartments among different single cells, and these structures (such as TADs, chromosomal loops, chromosomal compartments) play important roles in gene transcription, histone methylation, cell cycle, and cell development.
[0004] The development of single-cell technologies for detecting three-dimensional genome conformation, especially the development of single-cell chromosome conformation capture technology, enables us to better understand genome function than before; however, due to the extreme sparsity and high noise of single-cell Hi-C data (such as graph data in a graph), it is still difficult to study genomic structure and function using single-cell Hi-C data. It is very important to analyze and model the 3D genome / chromosome structure using single-cell Hi-C data to study cell variability. Therefore, it becomes very necessary to improve the quality of single-cell Hi-C data, and it is crucial to develop a computational method to enhance the contact map.
[0005] In recent years, some studies have applied deep learning techniques to 3D genomics, including TADs identification, Loops (the loop interaction structure of DNA or chromatin in three-dimensional space, i.e., chromatin loops) identification, and reconstruction of contact matrices. Researchers can reveal the structural and functional connections at the single-cell level that are overlooked by Bulk Hi-C through the dynamic chromosome conformation at the single-cell level.
[0006] Based on single-cell Hi-C data, it is necessary to reveal the structural features of the single-cell 3D genome through effective computational methods. For single-cell Hi-C data, the classical analysis process is as follows: First, cell barcodes are read through sequencing, and after alignment with the reference genome, they are merged into the single-cell Hi-C contact matrix; Second, the data is preprocessed, and then dimensionality reduction, imputation, embedding, and clustering are performed through effective computational methods; Finally, downstream analysis is performed on the processed data to reveal the heterogeneity and dynamics of the three-dimensional genome. Existing technologies can embed single-cell Hi-C data into a low-dimensional space, thereby capturing biologically significant changes in the 3D chromatin structure, but it is difficult to perform cell clustering.
[0007] Although existing methods have made certain progress in improving the resolution and analysis performance of Hi-C data, there are still the following limitations:
[0008] 1. Single-resolution enhancement: Most current methods focus on the enhancement of single-resolution data and are difficult to utilize multi-resolution feature information simultaneously, resulting in limited enhancement effects at high resolutions (such as 10kb).
[0009] 2. Insufficient adaptability: Some methods such as HiCPlus and HiCNN are designed to process Bulk Hi-C data and have poor adaptability to sparse single-cell Hi-C data.
[0010] 3. Insufficient feature extraction: Existing feature extraction methods such as HiCRep / MDS are difficult to perform cell clustering and multi-dimensional analysis at the single-cell level.
[0011] 4. Challenges of data noise and sparsity: Single-cell Hi-C data is highly sparse and noisy, and most methods are not ideal in dealing with data imputation, especially limited in detecting key structures such as TADs and Loops.
[0012] In short, the contact matrix in the existing single-cell Hi-C data shows significant sparsity, which limits the in-depth understanding of the three-dimensional genomic structure. These deficiencies limit the application of existing methods in the enhancement of high-resolution single-cell Hi-C data and downstream analysis. Therefore, it is very necessary to develop a multi-scale graph neural network model to enhance high-resolution single-cell Hi-C data. Summary of the Invention
[0013] To solve the problem of sparsity in existing single-cell Hi-C data, the present invention proposes a single-cell Hi-C data enhancement method, system, and storage medium. By designing a multi-scale variational graph autoencoder (VGAE) model, the information of low-resolution and high-resolution data is effectively fused, thereby solving the problem of sparsity in single-cell Hi-C data and enhancing the integrity and quality of the single-cell Hi-C contact matrix.
[0014] On the one hand, the present invention provides a single-cell Hi-C data enhancement method, including the following steps:
[0015] S1. Preprocess the multi-resolution Hi-C data, extract the Hi-C contact matrices at different resolutions for each cell, and construct the Hi-C contact matrices at different resolutions into a graph structure to obtain the graph structure data at different resolutions;
[0016] S2. Construct a variational graph autoencoder (VGAE) model as a fusion unit model, input the graph structure data at different resolutions into the fusion unit model, perform feature extraction and matrix reconstruction, realize cross-scale fusion of the graph structure data at different resolutions, obtain the fusion latent space features at different resolutions, and reconstruct the contact matrix as the enhanced Hi-C contact matrix.
[0017] Preferably, the above enhancement method further includes the step of:
[0018] S3. Identify the domain boundaries in the Hi-C data through the enhanced Hi-C contact matrix.
[0019] On the other hand, the present invention provides a single-cell Hi-C data enhancement system for implementing the above enhancement method. The enhancement system includes the following modules:
[0020] A preprocessing module for preprocessing the multi-resolution Hi-C data, extracting the Hi-C contact matrices at different resolutions for each cell, and constructing the Hi-C contact matrices at different resolutions into a graph structure to obtain the graph structure data at different resolutions;
[0021] The fusion unit model construction module constructs a variational graph autoencoder (VGAE) model as the fusion unit model, inputs graph structure data at different resolutions into the fusion unit model, performs feature extraction and matrix reconstruction, realizes cross-scale fusion of graph structure data at different resolutions, obtains fusion latent space features at different resolutions, and reconstructs the contact matrix as the enhanced Hi-C contact matrix.
[0022] In another aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a computer, the computer is enabled to execute the above enhancement method.
[0023] Compared with the prior art, the technical effects achieved by the present invention include:
[0024] 1. Multi-resolution fusion strategy: The present invention combines single-cell Hi-C data at different resolutions (such as 100 kb, 50 kb, 10 kb) for the first time. By designing a multi-scale variational graph autoencoder (VGAE) model, the information of low-resolution and high-resolution data is effectively fused, thus solving the sparsity problem of single-cell Hi-C data. This multi-resolution fusion method not only enhances the integrity and quality of the single-cell Hi-C contact matrix, but also provides richer genomic structure information at higher resolutions, providing a more reliable data basis for the subsequent identification of TADs and Loops domains.
[0025] 2. Enhanced data quality: A model is designed based on a graph neural network, and a variational graph autoencoder (VGAE) is used as the core architecture. The graph structure characteristics of Hi-C data are fully utilized, the contact matrix is represented as the adjacency matrix of the graph structure, and a fusion unit model is created for feature extraction and matrix reconstruction. The fusion unit model can utilize data features at different scales, perform cross-scale feature fusion on low-resolution and high-resolution graph structure data, enhance the overall quality of the Hi-C contact matrix, realize information transfer and sharing between different-resolution features, improve the prediction performance of the model and the data enhancement effect; multi-resolution data fusion can effectively fill in the missing parts in sparse data, reduce noise, and improve the downstream analysis's ability to parse the genomic structure.
[0026] 3. Efficient identification of TADs and Loops structures: Based on the cross-domain analysis results of multi-scale features, through the enhanced contact matrix, the present invention can identify three-dimensional genomic features (such as TADs and Loops) at multiple resolutions at the single-cell level. Especially at the single-cell level, by accurately locating the boundaries of TADs and Loops structures, it can help researchers understand the spatial distribution of genomic functional regions and their changes in different cell states, which is of great significance for studying the spatial organization of chromosomes and gene regulation.
[0027] 4. Scalability and adaptability of the algorithm: The algorithm framework designed in the present invention is applicable to sparse and highly heterogeneous single-cell Hi-C data, and can be extended to other similar high-dimensional graph structure data at the same time, with good adaptability and generality. Brief Description of the Drawings
[0028] Figure 1 is a flowchart of the single-cell Hi-C data enhancement method based on multi-resolution and multi-scale in the embodiment of the present invention;
[0029] Figure 2 is a block diagram of the single-cell Hi-C data enhancement system based on multi-resolution and multi-scale in the embodiment of the present invention. Detailed Embodiment
[0030] The technical solutions of the present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto. Embodiment
[0031] The single-cell Hi-C data enhancement method based on multi-resolution and multi-scale in this embodiment has a core idea of enhancing single-cell Hi-C data at multiple resolutions, especially achieving effective fusion of feature information between low-resolution and high-resolution graphs.
[0032] To ensure the efficient operation of the model, an optimized data processing scheme is adopted in the data input part. The straw package is used to efficiently extract single-cell Hi-C data and convert it into a format suitable for graph neural network processing, so as to be able to quickly process data from multiple single cells and perform efficient training and analysis.
[0033] To effectively utilize single-cell Hi-C data at multiple resolutions (such as 100kb, 50kb, 10kb), it is first necessary to preprocess the data, and then enhance the data quality through a graph neural network (GNN). Through preprocessing, these data are converted into a format suitable for input into the fusion model, and then learned and reconstructed through the fusion model to further enhance the data quality.
[0034] Such asFigure 1 As shown in the figure, the single-cell Hi-C data enhancement method of this embodiment specifically includes the following steps:
[0035] S1. Preprocess the Hi-C data with multiple resolutions, extract the Hi-C contact matrices of each cell at different resolutions, construct the Hi-C contact matrices at different resolutions into a graph structure, and obtain the graph structure data at different resolutions.
[0036] This step specifically includes:
[0037] S11. Preprocess the single-cell Hi-C data with different resolutions, and extract the Hi-C contact matrices of each cell at different resolutions.
[0038] The preprocessing mainly includes normalization processing and filling of sparse matrices. The normalization processing ensures the numerical comparability of data with different resolutions, while the filling of sparse matrices aims to supplement the missing parts in the contact matrix to avoid the negative impact of sparsity on model training.
[0039] In this embodiment, the Hi-C data with multiple resolutions is first converted into single-cell Hi-C contact matrices with different resolutions, and then the single-cell Hi-C contact matrices with different resolutions are normalized. The Hi-C contact matrix at each resolution is mainly composed of chromosome segments (i.e., segmented genomic regions), and each matrix element in the contact matrix represents the contact frequency between the corresponding chromosome segments; after the single-cell Hi-C contact matrix is normalized, it is stored as a sparse matrix for subsequent graph neural network processing.
[0040] The purpose of the normalization processing is to make the Hi-C contact matrices with different resolutions comparable. Assume that a given contact matrix M, and its matrix elements represent the contact frequencies between chromosome segments. The normalization operation can be achieved through the following formula:
[0041] ;
[0042] where is the contact frequency of the i th row and the j th column in the contact matrix, is the mean value of all matrix elements in the contact matrix M, is the standard deviation of all elements in the contact matrix M, represents the normalization processing of the contact matrix.
[0043] Through the normalization processing, the values of the contact matrix will become a distribution with a mean of 0 and a standard deviation of 1, thereby improving the comparability between data with different resolutions.
[0044] In this embodiment, the data source of the Hi-C contact matrix is usually a public database, such as the mouse dataset in the GEO database. This dataset contains single-cell Hi-C data of various tissue types. The data enhancement method of this embodiment can be applied to other related single-cell Hi-C datasets to test its applicability and generalization ability in different cell types. During the research process of this embodiment, 1954 files with the suffix.hic were downloaded. These files contain Hi-C data from different single cells; the straw package was used to generate a list of these.hic files, and code was written to read the Hi-C contact matrix element values of 19 chromosomes for each cell, and finally 371,260 contact matrices were obtained. The acquisition process of the dataset can be summarized as follows: download Hi-C data from a public database; use the straw package to read the.hic files and extract the Hi-C contact matrix of each cell; use the extracted contact matrix data for subsequent model training and analysis.
[0045] S12. Construct the Hi-C contact matrix at each resolution into a graph structure.
[0046] Construct the preprocessed Hi-C contact matrix into graph structures of different scales, where the nodes of the graph structure represent chromosomal segments of the genome (such as the segmentation of genomic regions), and the edges of the graph structure represent the contact frequencies between different chromosomal segments. The node features adopt the features of topologically associating domains (TADs) and chromatin compartments (A / B compartments) at a resolution of 1 Mb.
[0047] The Hi-C contact matrix can be transformed into a graph structure, where the rows and columns of the matrix correspond to the nodes of the graph structure, and the matrix elements correspond to the edge weights (contact frequencies) between the nodes. For the normalized contact matrix M at each resolution, a graph structure G can be constructed. The node set of the graph structure is the set V of chromosomal segments of the genome, and the edge set of the graph structure is the contact frequencies between the chromosomal segments. The adjacency matrix A of the graph structure is determined by the elements in the contact matrix:
[0048] A(i,j)= M(i,j) ;
[0049] A(i,j) represents the edge weight between nodes i and nodes j in the graph structure. Therefore, for each contact matrix M, the indices with values are the edge set of the adjacency matrix A , and the values are the edge weights corresponding to the adjacency matrix A .
[0050] The Hi-C contact matrix at each resolution is regarded as a graph structure. The rows and columns in the contact matrix represent different regions of the genome, and the value of each matrix element represents the contact frequency between chromosomal segments.
[0051] Through preprocessing, graph structure data at different resolutions such as 10kb, 50kb, and 100kb can be obtained, thus transforming the contact matrix into graph structure data so that the subsequent model can better capture the complex spatial relationships between genomes.
[0052] S2. Construct a variational graph autoencoder (VGAE) model as the fusion unit model. Input the graph structure data at different resolutions into the fusion unit model for feature extraction and matrix reconstruction, realizing cross-scale fusion of the graph structure data at different resolutions, obtaining the fusion latent space features at different resolutions, and reconstructing the contact matrix as the enhanced Hi-C contact matrix.
[0053] This step constructs a variational graph autoencoder (VGAE) model based on graph neural networks. Input the graph structure data at different resolutions into the constructed VGAE model to obtain the low-dimensional space representations at different resolutions. The nodes of the graph neural network represent genomic regions, and the edges represent the interactions (i.e., contact frequencies) between these genomic regions. Through graph convolution operations, the graph neural network can effectively capture the complex relationships between nodes in the high-dimensional space and better capture the topological features of three-dimensional genomic data, which is particularly important for processing the multi-level spatial structures in Hi-C data.
[0054] The VGAE model adopts an encoder-decoder structure. The encoder is used to extract node features from the graph structures at different resolutions, while the decoder reconstructs the contact matrix through the inner product operation of the matrix to learn the genomic structure information at different scales.
[0055] In this step, the VGAE model can capture the genomic structure features at different scales by processing the graph structure data at different resolutions. Specifically, in the encoding stage, when inputting the graph structure data at multiple resolutions simultaneously, the VGAE model can learn the node features of the graph structure data at each scale and extract the latent structure information through graph convolution operations; then perform dimensionality transformation on the node features of the graph structures at different resolutions, and perform weighted fusion with appropriate weights to fuse the graph structure data at multiple resolutions. The fusion of multiple-resolution data can not only effectively improve the overall quality of the contact matrix but also enhance the biological significance of the data. In the decoding stage, the VGAE model reconstructs the contact matrix through the inner product operation of the matrix to generate a more accurate Hi-C contact matrix, thus better revealing the three-dimensional structure and spatial features of the genome.
[0056] This step specifically includes:
[0057] S21. Design the variational graph autoencoder (VGAE) model as the fusion unit model.
[0058] Taking advantage of the variational graph autoencoder (VGAE) in processing graph-structured data of sparse matrices, especially its excellent performance in link prediction and feature learning, the variational graph autoencoder (VGAE) based on graph neural network is selected to construct the fusion unit model.
[0059] S22. Set the encoder of the VGAE model, and input graph-structured data of multiple resolutions at the same time to enable the VGAE model to learn the graph-structured node features at each resolution, and extract the potential structural information through graph convolution operations; then perform dimensional transformation on the graph-structured node features at different resolutions, and strengthen the fusion of the graph-structured node features at multiple resolutions to output the fused latent space features at different resolutions.
[0060] Specifically, it includes:
[0061] S221. Set multiple convolutional layers and graph convolutional networks (GCNConv) in the encoder to perform convolutional processing on the graph-structured data at each resolution to obtain the node features after single-scale convolutional operations of the graph-structured data at each resolution.
[0062] First, for the graph-structured data at each resolution (10kb, 50kb, 100kb), let the adjacency matrix of the graph structure be A reso , and the input feature of each graph structure be X reso , and generate the hidden layer feature H reso through graph convolution operation:
[0063] H reso = GCNConv(X reso ,A reso );
[0064] The hidden layer feature H reso serves as the node feature after single-scale convolutional operation.
[0065] S222. Set multiple attention mechanism modules, specifically including the first attention mechanism module (S1AttInform), the second attention mechanism module (S2AttInform), and the third attention mechanism module (S3AttInform), which are respectively used to process the node features after single-scale convolutional operations at three different resolutions of 10kb, 50kb, and 100kb H reso, so as to obtain the node latent features with nodes and edges fused at different resolutions.
[0066] (1) First, the node features after single-scale convolution operation are processed by a multi-layer perceptron (MLP) to obtain the node features after multi-layer perceptron processing.
[0067] The MLP is a multi-layer fully connected network that processes the node features after single-scale convolution operation through a series of linear transformations and activation functions. The node features after single-scale convolution operation H reso are input into the multi-layer perceptron for multi-layer perceptron processing, which can be expressed as:
[0068] x node =MLP ([ H reso ) ;
[0069] where, x node is the node feature after multi-layer perceptron processing.
[0070] (2) Then, the node features after multi-layer perceptron processing are mapped through a node2edge module to generate the edge features of the graph structure.
[0071] During the node2edge transformation, the node features after multi-layer perceptron processing are first mapped to the edges to calculate the relationships between nodes in the graph structure, and the edge features are generated by calculating the differences between the receiving nodes and the sending nodes.
[0072] Specifically, given a node feature matrix , the receiving relationship matrix rel_rec between nodes, and the sending relationship matrix rel_send between nodes, there are:
[0073] ;
[0074] ;
[0075] ;
[0076] where, is the node feature matrix formed by the node features after multi-layer perceptron processing, and the dimension size is , where N is the number of nodes and D is the feature dimension; rel_rec is the receiving relationship matrix, which describes which nodes are "receiving" nodes; rel_send is the sending relationship matrix, which describes which nodes are "sending" nodes; receivers are the receiving node features calculated from the receiving relationship matrix and the node feature matrix; senders are the sending node features calculated from the sending relationship matrix and the node feature matrix; distance is the difference between the receiving node features and the sending node features. Finally, the edge features edges can be represented as a new matrix obtained by horizontally concatenating the receiving node features receivers and the sending node features senders:
[0077] ;
[0078] By combining the differences between the receiving node features and the sending node features, the edge features of the graph structure are generated.
[0079] (3) Next, the edge features edges are processed through a multi-layer perceptron (MLP) to obtain the edge features after multi-layer perceptron processing:
[0080] ;
[0081] Among them, are the edge features after multi-layer perceptron processing.
[0082] (4) Subsequently, the edge features after multi-layer perceptron processing are mapped through an edge2node module to obtain the updated node features.
[0083] The specific operation of edge2node is as follows: take the average of the edge features after multi-layer perceptron processing to obtain the new node features, which can help the node features aggregate the information of neighbors. Given the edge features after multi-layer perceptron processing and the receiving relationship matrix rel_rec between nodes, the edge-to-node mapping operation is as follows:
[0084] ;
[0085] Among them, are the edge features after multi-layer perceptron processing, and the dimension size is , N is the number of nodes, is the dimension of the edge features; rel_rec is the receiving relationship matrix, and rel_rec T is the transpose of the receiving relationship matrix; incoming is the incoming edge features of each node. Take the average of all incoming edge features to obtain the updated node feature as:
[0086] ;
[0087] Among them, incoming_size is the number of incoming edges of each node, which is used to average the incoming edge features. This mapping operation generates updated node features, representing the information obtained by aggregating the edge features of the node's neighbors.
[0088] (5) Next, the node features after the single-scale convolution operation H reso are concatenated with the node features updated by the edge-node conversion module along the last dimension to obtain the concatenated node features :
[0089] ;
[0090] (6) Finally, the concatenated node features are processed by a multi-layer perceptron (MLP) to obtain the fused node latent features of the nodes and edges after multi-layer perceptron processing :
[0091] ;
[0092] S223. Set multiple cross-scale fusion modules to perform weighted fusion processing on the fused node latent features of nodes and edges at different resolutions to obtain the fused latent space features.
[0093] The cross-scale fusion modules set in this embodiment include the fusion module S2_to_S1 from the second scale to the first scale, the fusion module S3_to_S1 from the third scale to the first scale, the fusion module S1_to_S2 from the first scale to the second scale, the fusion module S3_to_S2 from the third scale to the second scale, the fusion module S1_to_S3 from the first scale to the third scale, and the fusion module S2_to_S3 from the second scale to the third scale.
[0094] Each fusion module is a cross-scale information fusion module. The fusion module S2_to_S1 fuses the node latent features from the second resolution (e.g., 50kb) with the node latent features from the first resolution (e.g., 10kb). The fusion module S3_to_S1 fuses the node latent features from the third resolution (e.g., 100kb) with the node latent features from the first resolution. The fusion module S1_to_S2 fuses the node latent features from the first resolution with the node latent features from the second resolution. The fusion module S3_to_S2 fuses the node latent features from the third resolution with the node latent features from the second resolution. The fusion module S1_to_S3 fuses the node latent features from the first resolution with the node latent features from the third resolution. The fusion module S2_to_S3 fuses the node latent features from the third resolution with the node latent features from the second resolution.
[0095] The fusion process includes self-attention calculation of node features and weighted fusion of features. The specific fusion process includes the following steps:
[0096] (1) Process the node features after single-scale convolution operations at the first resolution (10kb), second resolution (50kb), and third resolution (100kb) through the first attention mechanism module, second attention mechanism module, and third attention mechanism module respectively, to obtain the node latent feature attention representations of the fused nodes and edges at the first resolution, the node latent feature attention representations of the fused nodes and edges at the second resolution, and the node latent feature attention representations of the fused nodes and edges at the third resolution.
[0097] For the node features after single-scale convolution operation at the first resolution , obtain the node latent feature attention representations of the fused nodes and edges at the first resolution through the first attention mechanism module :
[0098] ;
[0099] Among them, represents the receiving relationship matrix between nodes at the first resolution, represents the sending relationship matrix between nodes at the first resolution.
[0100] For the node features after single-scale convolution operation at the second resolution , obtain the node latent feature attention representations of the fused nodes and edges at the second resolution through the second attention mechanism module :
[0101] ;
[0102] Among them, represents the receiving relationship matrix between nodes at the second resolution, represents the sending relationship matrix between nodes at the second resolution.
[0103] For the node features after single-scale convolution operation at the third resolution , the node potential feature attention representation of the fusion of nodes and edges at the third resolution is obtained through the third attention mechanism module :
[0104] ;
[0105] Among them, represents the receiving relationship matrix between nodes at the third resolution, represents the sending relationship matrix between nodes at the third resolution.
[0106] (2) By calculating the inner product between the node potential feature attention representations of the fusion of nodes and edges at different resolutions, the self-attention weights are calculated to obtain the attention weight matrix for the fusion of each scale, including the attention weight matrix from the first resolution to the second resolution , the attention weight matrix from the first resolution to the third resolution , the attention weight matrix from the second resolution to the first resolution , the attention weight matrix from the second resolution to the third resolution , the attention weight matrix from the third resolution to the first resolution , the attention weight matrix from the third resolution to the second resolution .
[0107] Calculate the inner product between the node potential feature attention representation of the fusion of nodes and edges at the first resolution and that at the second resolution, and normalize it with softmax to obtain the attention weight matrix from the first resolution to the second resolution , the attention weight matrix from the second resolution to the first resolution :
[0108] ;
[0109] ;
[0110] Calculate the inner product between the node potential feature attention representation of the fusion of nodes and edges at the first resolution and that at the third resolution, and normalize it with softmax to obtain the attention weight matrix from the first resolution to the third resolution The attention weight matrix from the third resolution to the first resolution :
[0111] ;
[0112] ;
[0113] Calculate the inner product between the node latent feature attention representation of the fusion of nodes and edges at the third resolution and the node latent feature attention representation of the fusion of nodes and edges at the second resolution, and normalize it with softmax to obtain the attention weight matrix from the second resolution to the third resolution The attention weight matrix from the third resolution to the second resolution :
[0114] ;
[0115] ;
[0116] (3) Two-scale feature fusion. Each cross-scale fusion module fuses the node latent feature attention representations at the corresponding resolution according to the corresponding attention weight matrix to obtain the weighted fusion feature of the node latent feature at the corresponding resolution.
[0117] The fusion module from the first scale to the second scale uses the attention weight matrix from the first resolution to the second resolution to perform weighted fusion on the node latent feature attention representations of 10kb at the first resolution and 50kb at the second resolution:
[0118] ;
[0119] where is the weighted fusion feature of the node latent feature at the first resolution. Finally, the weighted fusion feature of the node latent feature at the first resolution is output through the fusion module from the first scale to the second scale, which contains the shared information from 10kb at the first resolution and 50kb at the second resolution.
[0120] The fusion module from the first scale to the third scale uses the attention weight matrix from the first resolution to the third resolution to perform weighted fusion on the node latent feature attention representations of 10kb at the first resolution and 100kb at the third resolution:
[0121] ;
[0122] where It is the weighted fusion feature of the node latent features at the first resolution. Finally, the weighted fusion feature of the node latent features at the first resolution is output through the fusion module from the first scale to the third scale, which contains the shared information from 10kb at the first resolution and 100kb at the third resolution.
[0123] The fusion module from the second scale to the first scale uses the attention weight matrix from the second resolution to the first resolution to perform weighted fusion on the attention representation of the node latent features at 10kb of the first resolution and the attention representation of the node latent features at 50kb of the second resolution:
[0124] ;
[0125] where It is the weighted fusion feature of the node latent features at the second resolution. Finally, the weighted fusion feature of the node latent features at the second resolution is output through the fusion module from the second scale to the first scale, which contains the shared information from 10kb at the first resolution and 50kb at the second resolution.
[0126] The fusion module from the second scale to the third scale uses the attention weight matrix from the second resolution to the third resolution to perform weighted fusion on the attention representation of the node latent features at 100kb of the third resolution and the attention representation of the node latent features at 50kb of the second resolution:
[0127] ;
[0128] where It is the weighted fusion feature of the node latent features at the second resolution. Finally, the weighted fusion feature of the node latent features at the third resolution is output through the fusion module from the second scale to the third scale, which contains the shared information from 100kb at the third resolution and 50kb at the second resolution.
[0129] The fusion module from the third scale to the first scale uses the attention weight matrix from the third resolution to the first resolution to perform weighted fusion on the attention representation of the node latent features at 10kb of the first resolution and the attention representation of the node latent features at 100kb of the third resolution:
[0130] ;
[0131] where It is the weighted fusion feature of the node latent features at the third resolution. Finally, the weighted fusion feature of the node latent features at the third resolution is output through the fusion module from the third scale to the first scale, which contains the shared information from 10kb at the first resolution and 100kb at the third resolution.
[0132] The fusion module from the third scale to the second scale uses the attention weight matrix from the third resolution to the second resolution to perform weighted fusion on the node latent feature attention representation at the third resolution of 100 kb and the node latent feature attention representation at the second resolution of 50 kb:
[0133] ;
[0134] where is the weighted fusion feature of the node latent features at the third resolution. Finally, the weighted fusion feature of the node latent features at the third resolution is output through the fusion module from the third scale to the second scale, which contains the shared information from 100 kb at the third resolution and 50 kb at the second resolution.
[0135] (4) Three-scale feature fusion.
[0136] In the three-scale feature fusion stage, based on the results of the two-scale feature fusion, the VGAE model further performs weighted fusion on the weighted fusion features of the node latent features at different resolutions to obtain the fusion latent space features at different resolutions. Let respectively represent the node features after single-scale convolution operations at three different resolutions (10 kb, 50 kb, 100 kb), and the weights are used to control the fusion ratio, usually set to 0.3, to balance the contributions of features at different resolutions. The three-scale feature fusion operation is achieved through the following formula:
[0137] ;
[0138] ;
[0139] ;
[0140] In the formula represents the fusion latent space feature at the first resolution of 10 kb, represents the fusion latent space feature at the second resolution of 50 kb, represents the fusion latent space feature at the third resolution of 100 kb. Through three-scale feature fusion, the spatial features at different resolutions are weighted to obtain the final fusion latent space features at different resolutions .
[0141] S23. Set the decoder of the VGAE model. According to the fused latent space features at different resolutions output by the encoder, calculate the similarity between the nodes of the graph structure through the inner product operation of matrices, so as to reconstruct and generate the predicted values of the Hi-C contact matrix at different resolutions, and obtain the predicted contact matrix.
[0142] In the decoding stage, the fused latent space features at different resolutions output by the encoder 、 、 are used to reconstruct the contact matrix. Through the inner product operation of matrices, calculate the similarity between the nodes of the graph structure, and generate the predicted values of the Hi-C contact matrix at different resolutions:
[0143] ;
[0144] ;
[0145] ;
[0146] where represents the predicted value of the Hi-C contact matrix at the first resolution, represents the predicted value of the Hi-C contact matrix at the second resolution, represents the predicted value of the Hi-C contact matrix at the third resolution; the fused latent space feature , represents the fused node features obtained through the encoder; the Sigmoid function is used to ensure that the values of the contact matrix are in the range of [0,1], indicating the connection probability between nodes.
[0147] The inner product operation of the matrix is used to map the input fused latent space features through a low-dimensional space and reconstruct them back to the original contact matrix to calculate the similarity between node features (i.e., contact frequency). After being processed by the Sigmoid function, a connection probability representing between nodes is obtained as the final predicted value of the Hi-C contact matrix, so as to ensure that the matrix reconstruction result can maintain the spatial structure characteristics of the graph structure data.
[0148] S24. Based on the reconstruction loss and KL divergence loss, construct the loss function of the VGAE model.
[0149] The loss function mainly includes two parts: one part is the reconstruction loss, which is used to measure the difference between the original contact matrix and the predicted contact matrix;
[0150] ;
[0151] where is the reconstruction loss, N represents the number of samples, that is, the total number of elements in the contact matrix; is an element of the original Hi-C contact matrix, representing the contact frequency between nodes i and j; is the contact probability value of the Hi-C contact matrix predicted by the model, representing the predicted contact probability between nodes i and j, and the value of the contact probability ranges from 0 to 1.
[0152] Another part is the KL divergence loss, which is used to measure the difference between the distribution of the latent space features output by the encoder and the standard normal distribution:
[0153] ;
[0154] where represents the KL divergence loss, 、 are the mean and variance of the fused latent space features output by the encoder respectively.
[0155] These two parts of the loss are added together with weights to obtain the final loss function, which is used to optimize the training process of the VGAE model (i.e., the fusion unit model).
[0156] S25. Use the trained fusion unit model for feature extraction to obtain an enhanced Hi-C contact matrix.
[0157] After the model training is completed, the original Hi-C data is input into the trained model to obtain a new contact matrix with enhanced graph structures for each. Then, the edge information and edge weights in the new contact matrix are extracted, and through the pre tool in juicertools, the new contact matrix is saved in the Hi-C contact matrix format. This process provides a basis for the subsequent identification of TADs and Loops domain boundaries, enabling further analysis of the spatial features of the genomic structure.
[0158] S3. Identification of domain boundaries.
[0159] Through the enhanced Hi-C contact matrix, in this embodiment, by using the arrowhead and hiccups tools in juicertools, the TADs and Loops domain boundaries in the Hi-C data can be accurately identified. The identification of these domain boundaries provides important spatial information for studying gene regulation and disease mechanisms.
[0160] Specifically, in this embodiment, for the identified TADs and Loops, the number of TADs and Loops and the Jaccard index of the boundaries are calculated to obtain an enhanced contrast effect at the TADs level. Then, the corresponding image is obtained by Aggregate Peak Analysis, and the enhanced effect of each TAD and Loop is represented in an average sense.
[0161] Based on the same inventive concept, this embodiment also provides a single-cell Hi-C data enhancement system based on multi-resolution and multi-scale for implementing the enhancement method of this embodiment, such as Figure 2 , and the enhancement system includes the following modules:
[0162] A preprocessing module 100 for preprocessing the multi-resolution Hi-C data, extracting the Hi-C contact matrix at different resolutions of each cell, constructing the Hi-C contact matrices at different resolutions into a graph structure, and obtaining the graph structure data at different resolutions;
[0163] A fusion unit model construction module 200 constructs a variational graph autoencoder (VGAE) model as the fusion unit model, inputs the graph structure data at different resolutions into the fusion unit model for feature extraction and matrix reconstruction, realizes cross-scale fusion of the graph structure data at different resolutions, obtains the fusion latent space features at different resolutions, and reconstructs the contact matrix as the enhanced Hi-C contact matrix.
[0164] Furthermore, the enhancement system of this embodiment may further include an identification module 300 for identifying the domain boundaries in the Hi-C data through the enhanced Hi-C contact matrix.
[0165] In addition, this embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a computer, the computer is enabled to execute the enhancement method of this embodiment to realize the enhancement of single-cell Hi-C data.
[0166] Among them, the computer program includes one or more computer instructions. When the computer instructions are loaded and executed on the computer, the processes or functions described in this embodiment are fully or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. The computer-readable storage medium may be any available medium that the computer can access, or a data storage device such as a server or a data center that includes one or more integrated available media.
[0167] This embodiment can identify and analyze domains such as TADs and Loops at the single-cell level, which is an innovation compared to traditional bulk Hi-C data methods. Because traditional bulk data usually cannot capture the sparsity and genomic structure variability at the single-cell level. This embodiment can reveal the spatial dynamic characteristics of the genome in a single cell through a multi-resolution graph neural network, especially the changes in TADs and Loops under different cell types or physiological states. Thus, it effectively solves the sparsity problem of single-cell Hi-C data, especially in the identification of TADs and Loops. Through a multi-resolution fusion method, it provides more detailed and reliable genomic structure information, which is helpful for genomics research, especially providing new technical means when exploring gene regulation and disease mechanisms.
[0168] In summary, the present invention uses a variational graph autoencoder (VGAE) model based on the fusion of multiple resolutions to train the graph structure of single-cell Hi-C data. By introducing the information of A / B compartments, chromatin loops Loops, and topologically associated domains TADs in single-cell Hi-C data, the overall performance of the Hi-C contact matrix is enhanced, so that the boundaries of Loops and TADs domains in the contact matrix can be identified and analyzed more systematically. Through the innovative multi-resolution deep learning method, the present invention overcomes the sparsity problem of single-cell Hi-C data, provides new ideas for three-dimensional genomics research, and effectively enhances the ability to analyze genomic structure, especially in the identification of TADs and Loops domains, with significant technical advantages.
[0169] The above embodiments are one of the implementation manners of the present invention. The implementation manners of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement manners and are all included in the protection scope of the present invention.
Claims
1. A single cell Hi-C data enhancement method, characterized in that: The following steps are involved: S1. Preprocess the multi-resolution Hi-C data, extract the Hi-C contact matrix of each cell at different resolutions, construct the Hi-C contact matrix at different resolutions into a graph structure, and obtain graph structure data at different resolutions; S2. Construct a variational graph autoencoder VGAE model as a fusion unit model, input graph structure data at different resolutions into the fusion unit model, perform feature extraction and matrix reconstruction, realize cross-scale fusion of graph structure data at different resolutions, obtain fusion latent space features at different resolutions, and reconstruct the contact matrix as the enhanced Hi-C contact matrix; Step S2 includes: S21. Design the variational graph autoencoder VGAE model as the fusion unit model. S22, setting the encoder of the VGAE model, inputting multi-resolution graph structure data at the same time so that the VGAE model learns the graph structure node features at each resolution, and extracts the potential structural information through graph convolution operation; then transforming the graph structure node features at different resolutions, enhancing the fusion of the graph structure node features at multiple resolutions, and outputting the fused potential space features at different resolutions; S23, setting a decoder of the VGAE model, calculating the similarity between nodes of the graph structure through an inner product operation of a matrix according to the fused latent space features at different resolutions output by the encoder, so as to reconstruct and generate predicted values of the Hi-C contact matrix at different resolutions, and obtain a predicted contact matrix; S24. Based on the reconstruction loss and the KL divergence loss, a loss function of the VGAE model is constructed to optimize the training process of the fusion unit model; wherein the reconstruction loss is used to measure the difference between the original contact matrix and the predicted contact matrix, and the KL divergence loss is used to measure the difference between the latent space feature distribution output by the encoder and the standard normal distribution; The reconstruction loss is: ; in is the reconstruction loss, N represents the total number of elements in the contact matrix; is the element of the original Hi-C contact matrix, indicating the contact frequency between node i and node j; is the contact probability value of the Hi-C contact matrix predicted by the model, indicating the predicted contact probability between node i and node j; The KL divergence loss is: ; in represents the KL divergence loss, , are the mean and variance of the fused latent space features output by the encoder, respectively; S25. Use the trained fusion unit model to perform feature extraction to obtain the enhanced Hi-C contact matrix.
2. The single cell Hi-C data enhancement method according to claim 1, characterized in that: Step S1 includes: S11, preprocessing single-cell Hi-C data at different resolutions to extract Hi-C contact matrices of each cell at different resolutions; the preprocessing includes normalization processing and filling of sparse matrices; S12. Construct the Hi-C contact matrix at each resolution into a graph structure, where the nodes of the graph structure represent the chromosome fragments of the genome, and the edges of the graph structure represent the contact frequencies between different chromosome fragments.
3. The single cell Hi-C data enhancement method according to claim 1, characterized in that: In step S2, the variational graph autoencoder VGAE model adopts an encoder-decoder structure. The encoder is used to extract node features from graph structures of different resolutions, while the decoder reconstructs the contact matrix through the inner product operation of the matrix to learn the genome structure information at different scales.
4. The single cell Hi-C data enhancement method according to claim 1, characterized in that: Step S22 includes: S221, setting a plurality of convolution layers and a graph convolution network in the encoder to perform convolution processing on the graph structure data at each resolution, and obtaining node features after the single-scale convolution operation of the graph structure data at each resolution; S222, setting multiple attention mechanism modules, respectively used to process node features after single-scale convolution operations at different resolutions, so as to obtain node potential features fused with nodes and edges at different resolutions; S223, setting multiple cross-scale fusion modules, performing weighted fusion processing on the node potential features fused by the nodes and edges at different resolutions, and obtaining fused potential space features.
5. The single cell Hi-C data enhancement method according to claim 4, characterized in that: In step S221, for the graph structure data at each resolution, the adjacency matrix of the graph structure is A reso , the input features of each graph structure are X reso , generate hidden layer features through graph convolution operation: < i>H reso = GCNConv(X reso , A reso ); The hidden layer features H reso As node features after single-scale convolution operation; Step S222 includes: (1) The node features after the single-scale convolution operation are processed by a multi-layer perceptron to obtain the node features after multi-layer perceptron processing; (2) The node features after multi-layer perception processing are mapped through the node edge conversion module to generate edge features of the graph structure; In the process of node-edge conversion, the node features after multi-layer perception processing are first mapped to the edges, the relationship between the nodes in the graph structure is calculated, and the edge features are generated by calculating the difference between the receiving node and the sending node; (3) Process the edge features edges through a multi-layer perceptron to obtain edge features after multi-layer perceptron processing; (4) The edge features after multi-layer perception processing are mapped through the edge node conversion module to obtain updated node features; In the process of edge node conversion, given the edge features processed by multi-layer perception And the receiving relationship matrix rel_rec between nodes, perform edge-to-node mapping operation: ; Where rel_rec is the receiving relationship matrix, rel_rec T is the transpose of the receiving relationship matrix; incoming is the incoming edge feature of each node; all incoming edge features are averaged to obtain the updated node features for: ; Among them, incoming_size is the number of incoming edges of each node; (5) Concatenate the node features after the single-scale convolution operation and the node features updated by the edge node conversion module along the last dimension to obtain the concatenated node features; (6) The concatenated node features are processed by a multi-layer perceptron to obtain the node potential features that are a fusion of the nodes and edges after multi-layer perceptron processing.
6. The single cell Hi-C data enhancement method according to claim 4, characterized in that: The attention mechanism module of step S222 includes a first attention mechanism module, a second attention mechanism module, and a third attention mechanism module; The cross-scale fusion module set in step S223 includes a fusion module from the second scale to the first scale, a fusion module from the third scale to the first scale, a fusion module from the first scale to the second scale, a fusion module from the third scale to the second scale, a fusion module from the first scale to the third scale, and a fusion module from the second scale to the third scale; The fusion process of step S223 includes: (1) The first attention mechanism module, the second attention mechanism module, and the third attention mechanism module process the node features after the single-scale convolution operation at the first resolution, the second resolution, and the third resolution respectively, and obtain the node potential feature attention representation of the node and edge fusion at the first resolution , node latent feature attention representation of nodes and edges fused at the second resolution , Node latent feature attention representation of nodes and edges fused at the third resolution ; (2) By calculating the inner product between the node potential feature attention representations of nodes and edges at different resolutions, the self-attention weight is calculated to obtain the attention weight matrix of each scale fusion, including the attention weight matrix from the first resolution to the second resolution , the attention weight matrix from the first resolution to the third resolution , the attention weight matrix from the second resolution to the first resolution , the attention weight matrix from the second resolution to the third resolution , the attention weight matrix from the third resolution to the first resolution , the attention weight matrix from the third resolution to the second resolution ; (3) Two-scale feature fusion: Each cross-scale fusion module fuses the attention representation of the node potential features at the corresponding resolution according to the corresponding attention weight matrix to obtain the weighted fusion features of the node potential features at the corresponding resolution; The fusion module from the first scale to the second scale uses the attention weight matrix from the first resolution to the second resolution The node latent feature attention representation of the first resolution and the node latent feature attention representation of the second resolution are weighted fused: ; in, is the weighted fusion feature of the node potential features at the first resolution; The fusion module from the first scale to the third scale uses the attention weight matrix from the first resolution to the third resolution The node latent feature attention representation of the first resolution and the node latent feature attention representation of the third resolution are weighted fused: ; in, is the weighted fusion feature of the node potential features at the first resolution; The fusion module from the second scale to the first scale uses the attention weight matrix from the second resolution to the first resolution The node latent feature attention representation of the first resolution and the node latent feature attention representation of the second resolution are weighted fused: ; in, It is the weighted fusion feature of the node potential features of the second resolution; The fusion module from the second scale to the third scale uses the attention weight matrix from the second resolution to the third resolution The node latent feature attention representation of the third resolution and the node latent feature attention representation of the second resolution are weighted fused: ; in, It is the weighted fusion feature of the node potential features of the second resolution; The third-scale to first-scale fusion module uses the attention weight matrix from the third resolution to the first resolution The node latent feature attention representation of the first resolution and the node latent feature attention representation of the third resolution are weighted fused: ; in, It is the weighted fusion feature of the node potential features at the third resolution; The third-scale to second-scale fusion module uses the attention weight matrix from the third resolution to the second resolution The node latent feature attention representation of the third resolution and the node latent feature attention representation of the second resolution are weighted fused: ; in, It is the weighted fusion feature of the node potential features at the third resolution; (4) Three-scale feature fusion: Based on the result of two-scale feature fusion, the weighted fusion features of node potential features at different resolutions are further weighted fused to obtain fused latent space features at different resolutions: ; ; ; In the formula Represents the node features after single-scale convolution operation at three different resolutions, and the weights Used to control the ratio of fusion; represents the fused latent space features of the first resolution, represents the fused latent space features of the second resolution, Represents the fused latent space features at the third resolution.
7. The single cell Hi-C data enhancement method according to claim 1, characterized in that: The enhancement method further comprises the steps of: S3. Identify domain boundaries in Hi-C data using the enhanced Hi-C contact matrix.
8. A single cell Hi-C data enhancement system, characterized in that: For implementing the enhancement method according to any one of claims 1 to 6, the enhancement system comprises the following modules: A preprocessing module is used to preprocess multi-resolution Hi-C data, extract the Hi-C contact matrix at different resolutions for each cell, construct the Hi-C contact matrix at different resolutions into a graph structure, and obtain graph structure data at different resolutions; The fusion unit model construction module constructs a variational graph autoencoder VGAE model as the fusion unit model, inputs graph structure data at different resolutions into the fusion unit model, performs feature extraction and matrix reconstruction, realizes cross-scale fusion of graph structure data at different resolutions, obtains fusion latent space features at different resolutions, and reconstructs the contact matrix as the enhanced Hi-C contact matrix.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed on a computer, the computer is enabled to execute the enhancement method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Hi-C data resolution enhancement method and device
CN113628112A