Multi-layer heterogeneous network unicellular organism network inference method based on meta-path enhancement

By integrating scRNA-seq, scATAC-seq, and ST data through a multi-layer heterogeneous biological network inference method based on metapath enhancement, the heterogeneity and sparsity problems in single-cell multi-omics data fusion are solved, and more accurate cell clustering and biological network inference are achieved.

CN121811981APending Publication Date: 2026-04-07HEBEI UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing single-cell multi-omics data fusion methods are insufficient in handling cellular heterogeneity and data sparsity, making it difficult to effectively integrate multi-omics information, resulting in insufficient accuracy and reliability of single-cell biological network inference.

Method used

We employ a multi-layer heterogeneous biological network inference method based on metapath enhancement. By constructing a single-cell multi-omics heterogeneous network and integrating scRNA-seq, scATAC-seq, and ST data, we utilize metapath enhancement technology to preserve biological prior knowledge and common information while capturing complex dependencies between features, thereby improving cell clustering performance.

Benefits of technology

It achieves effective integration of multi-omics data, improves the accuracy and reliability of single-cell biological network inference, can be extended to the fusion of other types of omics data, and enhances cell clustering effect and biological network inference performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811981A_ABST
    Figure CN121811981A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-layer heterogeneous network unicellular organism network inference method based on meta-path enhancement, which mainly comprises a gene regulation knowledge base enhanced multi-layer heterogeneous network construction module for integrating an external gene interaction network and multiple omics data such as scRNA-seq, scATAC-seq, ST and the like; constructing a single-cell multi-omics multilayer heterogeneous network containing cell-cell, cell-gene and gene-gene relationships, and fusing spatial constraints to consider cell positions and tissue structures; and the feature enhancement module based on the meta-path explores complex semantics of the network by designing a multi-hop meta-path mode, designs an adaptive multi-view learning framework and a multi-round enhancement mechanism, and optimizes feature representation by using cell-gene interaction and cross-modal attention fusion. The unicellular biological network can be effectively deduced, the deduction accuracy and biological interpretation are remarkably improved, the method plays an important role in understanding the cell biological process, developing and treating diseases and the like, has good expandability, and can further integrate multi-modal omics data such as proteomics and metabonomics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of bioinformatics and machine learning. Specifically, it relates to a method for inferring multilayer heterogeneous biological networks based on metapath enhancement, which is mainly applied to single-cell sequencing data analysis, cell type identification and functional annotation, cancer stage analysis and other fields. Background Technology

[0002] Single-cell biological networks refer to molecular interaction networks constructed at the single-cell level, including gene regulatory networks, chromatin accessibility networks, and cell-cell interaction networks. These networks are used to describe the complex regulatory relationships and functional connections between molecules within the cell. They can reveal the dynamic processes of gene expression regulation, signal transduction pathways, and the molecular mechanisms of cell state transitions, which are of great significance for understanding biological processes such as cell differentiation, disease development, and drug action mechanisms. Accurately deducing single-cell biological networks helps identify key regulatory factors, predict cell fate, discover disease-related molecular biomarkers, and provide a theoretical basis for precision medicine and personalized treatment.

[0003] The advent of single-cell sequencing technology has greatly enhanced researchers' understanding of cells, enabling the construction of expression profiles for each cell at the individual cell level and revealing more molecular-level cellular information, such as transcriptomics. However, in single-cell biological network inference, single-omics data typically only provide a single dimension of information about cellular activity, while multi-omics data can provide more comprehensive and detailed biological information. Therefore, effectively integrating different types of omics data has become an important research task for single-cell biological network inference.

[0004] Multi-omics data, including gene transcriptomics (scRNA-seq), chromatin accessibility omics (e.g., scATAC-seq), and spatial transcriptomics (ST), can simultaneously analyze information from multiple biological layers. Single-cell multi-omics data fusion can simultaneously capture the molecular characteristics of cells from multiple dimensions, effectively compensating for the limitations of single-omics data and providing key technical support for constructing more accurate and complete single-cell biological networks. Through multi-omics fusion, researchers can reveal the heterogeneity of transcriptomics, chromatin state, and protein expression within the same cell type, gain a deeper understanding of the differences between individual cells within a cell population, thereby revealing the diversity of cell function and phenotype, and obtaining more accurate network inference and prediction results.

[0005] Existing single-cell multi-omics data fusion methods primarily focus on integrating two types of omics data. However, they still have limitations in handling the heterogeneity of single-cell multi-omics data, preserving spatial information, and elucidating downstream regulatory mechanisms. For example, methods integrating scRNA-seq and scATAC-seq are insufficient in handling cellular heterogeneity and data sparsity; the lack of spatial information prevents the revelation of gene regulation and expression distribution within tissue spatial structures. Researchers can more easily obtain rich data containing two or more omics layers through simple cell sequencing technology. Therefore, single-cell multi-omics fusion methods capable of integrating information from multiple omics layers present a new challenge.

[0006] While deep learning algorithms, such as variational autoencoders and generative adversarial networks, have been applied to integrate and analyze single-cell multi-omics data, these methods still face numerous challenges in handling the heterogeneity and scale differences of multi-omics data. In particular, during data integration, problems such as incomplete information integration and poor scalability exist, often making it difficult to simultaneously maintain the unique characteristics of each omics dataset. Furthermore, the integration process is prone to missing biological information. These challenges severely restrict the accuracy and reliability of multi-omics data analysis, necessitating the development of new methods to effectively integrate and analyze multi-omics data and improve the inference performance of single-cell biological networks. Summary of the Invention

[0007] To address the aforementioned problems in existing technologies, this invention proposes a multi-layer heterogeneous biological network inference method based on meta-path enhancement. By constructing a single-cell multi-omics heterogeneous network and employing meta-path enhancement technology, it can be effectively applied to the field of single-cell biological network inference, including applications such as cell function annotation, cancer analysis, and revealing complex biological processes. It is suitable for processing high-dimensional heterogeneous multi-omics data and achieves effective integration of scRNA-seq, scATAC-seq, and ST omics data. Furthermore, this invention has strong scalability, extending to the fusion of other types of omics data, preserving prior biological knowledge, fusing common information between different omics, and retaining the specific information of each omics data. It effectively captures complex dependencies between features, improves cell clustering effects, and enhances the performance of single-cell biological network inference.

[0008] The present invention provides a method for inferring single-cell biological networks in multilayer heterogeneous networks based on meta-path enhancement, comprising the following steps: Step 1: Acquire single-cell multi-omics data, including scRNA-seq data. E scATAC-seq data P Using spatial transcriptome data (ST), external gene interaction networks were obtained from the STRING database. G inter And construct a gene relationship strength matrix Rinter The specific steps are as follows: Step S11: For scRNA-seq data, the gene expression matrix is ​​represented as follows: ,in and These represent genes and cells, respectively. To obtain a qualitative representation of each gene in all cells, a Hidden Markov Model (HMM) is used. In this model, each gene... The state can be discretized as Each state is defined and generated by the HMM. A discrete state, representing Genes in a cell Different regulatory signals or expression levels. HMM is applied to generate different regulatory signals or expression levels for each gene. The discrete expression patterns of genes in different cells were captured. To handle regions with excessively low expression values, rows or columns with non-zero values ​​less than 0.1% were removed from the matrix. The gene expression matrices were then log-normalized to identify the top 2000 highly variable genes for each matrix. If a matrix contained fewer than 2000 genes, all usable genes were retained, resulting in a preprocessed matrix. .

[0009] Step S12: For scATAC-seq data, the chromatin accessibility matrix is ​​represented as follows: ,in and These represent peak values ​​and cells, respectively. To integrate scRNA-seq and scATAC-seq data, Transformed into a gene-cell matrix First, calculate the peak value. To genes Regulation potential weight : in, yes Center to gene The distance from the transcription start site (TSS), This is a user-defined half-life distance, typically set to 10kb. For a given gene... If with distance If it exceeds 150kb, then when When the weight is 10kb, Less than 0.0005. In this method, to optimize computational efficiency, if If it exceeds 150kb, then Set to 0.

[0010] For peaks located in exon regions, the weights are adjusted to account for gene length bias. If the peak... Located in genes The weights of exon regions will be normalized by the total exon length to correct for the problem of higher peak detection probability in longer genes. Conversely, peaks located in nearby gene promoters or exon regions will be excluded from the calculation of the gene's regulatory potential.

[0011] Step S13: Traverse the non-repeating genes of the three omics systems and encode the gene relationships into a binary adjacency matrix. : in, It is the number of genes in all non-repetitive sequences across the three omics studies.

[0012] Step a4: This study uses a contrastive learning approach to learn the relationships between genes, for each gene in the gene network. Define the set of positive samples (Including genes) Directly linked genes) and negative sample set (Including genes) Genes that are not directly connected. Each gene Initial representation in the network The low-dimensional representation of each gene is obtained by a multilayer perceptron (MLP) and through contrastive learning. Based on these representations, arbitrary gene pairs can be computed. Relationship strength The corresponding comparative learning objective function is: in, Cosine similarity is used to measure the degree of similarity between the representations of two genes. The temperature parameter is set to 0.07 to control the difficulty of the comparative learning. Step 2: Obtain the cell representation matrix from the spatial transcriptome data (ST) using dual adjacency relationships and multi-level information aggregation mechanisms. Simultaneously, the gene relationship strength matrix obtained in step one is applied. R inter Update the gene expression matrix of the ST data to obtain the updated gene expression matrix. Constructing a gene regulation knowledge base to enhance the ST heterogeneous network G 2. The specific steps are as follows: Step S21: Based on the heterogeneity plot of ST data GA 1:1 cell representation matrix is ​​used to construct two different types of neighbor relationship graphs for cells: a location adjacency graph based on spatial distance and a tissue similarity graph based on expression similarity. For the location adjacency graph, the Euclidean distance between cells is calculated, and the six nearest cells to each cell are selected as its spatial neighbors to construct a spatial adjacency matrix. For the tissue similarity map, the expression similarity between cells is calculated, and the six cells with the highest similarity are selected as expression neighbors to construct a tissue similarity matrix. .

[0013] Step S22: Based on the cell neighbor relationship map obtained in step S21 Similar to tissue diagram An adaptive multi-level information aggregation mechanism is used to update the cell expression matrix, and information aggregation is performed separately for each relationship type: in, and Representing cells respectively exist and The six nearest neighbor set in the middle, attention weight and Calculated using the self-attention mechanism: in, and For learnable transformation matrices, , This is a learnable attention vector. This self-attention mechanism can adaptively determine the importance of different neighboring nodes.

[0014] Step S23: Adaptively fuse spatial and organizational information through a gating mechanism to obtain the cell representation matrix of the ST heterogeneous network. : in, It is spatial information Organizational information and original representation The candidate cells generated by the fusion represent, and The weight matrix is ​​a learnable matrix. This determines the relative importance of various pieces of information when merging them to generate a new cell representation. Used to learn how to control the extent to which information is updated. This is a gating vector used to control the fusion ratio of new and original information. The sigmoid activation function maps values ​​to the interval [0,1]. This indicates element-wise multiplication, and this gating mechanism can adaptively determine how much original information and newly aggregated information to retain based on the data.

[0015] Step S24: Use the relation strength matrix obtained from contrastive learning heterogeneous plot of ST data G The gene expression matrix of the ST network is obtained by updating the gene expression matrix of the 1-bit network. For each gene It adaptively aggregates information from the initial ST data and external networks through a self-attention mechanism: in, For genes In gene networks The initial representation in For attention weights, This is a learnable weight matrix.

[0016] Step S25: Use the cell representation matrix of the ST heterogeneous network and gene representation matrix Constructing ST heterogeneous networks including cell-cell, cell-gene, and gene-gene relationships. G 2. The cell represents a matrix. This represents the relationship between cells in the diagram. When the value is greater than 0, it indicates... G 1 There are corresponding cell-cell edges, and similarly, the gene representation matrix is ​​used. and gene networks To determine cell-gene edges and gene-gene edges respectively.

[0017] Step 3: Apply the two matrices obtained in Step 1 E , P Integrative analysis to construct heterogeneous network structures G 3. Encoding learning and gene relationship strength through deep graph neural networks R inter The update guide yields the cell representation matrix of the scRNA-seq and scATAC-seq heterogeneous networks. and gene representation matrix This leads to the construction of a heterogeneous scRNA-seq / scATAC-seq network with an enhanced gene regulation knowledge base. G 4. The specific steps are as follows: Step S31: Calculate the intercellular Euclidean distance from scRNA-seq data. For each cell , and each other cell Euclidean distance between The calculation is as follows: Step S32: Calculate the Euclidean distance between cells in the scATAC-seq data. For each cell With each and every other cell Euclidean distance between : Step S33: Calculate the intercellular combinatorial distance between scRNA-seq and scATAC-seq. .cell Combined distance It is obtained by weighting and calculating the various distances: in, The parameter is set to 0.5 in this invention.

[0018] Step S34: By combining distances Construct a new nearest neighbor matrix . The entries in the matrix represent the nearest neighbors based on a weighted combination of gene expression and regulatory potential distances. The integrated WNN matrix... Defined as: Step S35: Integrate using the Weighted Nearest Neighbor (WNN) method to obtain the comprehensive matrix. For each cell Its comprehensive expression It is through its nearest neighbors from two datasets (by...) The expression (defined) is obtained by averaging: in, Represents cells in a WNN diagram The nearest neighbor set, Cells derived from RNA data Gene expression, Cells from ATAC data Gene expression. The resulting comprehensive matrix. It is a gene expression matrix that integrates scRNA-seq and scATAC-seq data; Step S36: Based on the comprehensive matrix Calculate the pairwise cosine similarity between cells to obtain the cell similarity matrix. Comprehensive matrix ,in Indicates the number of genes. This represents the number of cells. To understand the relationships between cells, this invention calculates the cosine similarity of their expressed features. Cells and cells Cosine similarity between The definition is as follows: Step S37: Construct an inter-cell similarity matrix using these distances. ,in Represents cells and cells The similarity between them. To convert these similarities into binary values, this invention employs a threshold method, the calculation process of which is shown below: Step S38: Use the synthesis matrix and inter-cell similarity matrix Building heterogeneous networks And a graph autoencoder is used to obtain low-dimensional representations of cells and genes. and .

[0019] In heterogeneous networks The algorithm contains two types of nodes: cells and genes. Edges represent the similarity between cells and the expression relationships between cells and genes. A cell similarity matrix is ​​used. Intercellular relationships within cells serve as the initial representation of cell nodes. and using a comprehensive matrix The expression levels of genes in all cells serve as the initial representation of gene nodes. These initial representations are used as inputs to the encoder in a graph autoencoder (GAE). The role of the initial representations is to provide the encoder with meaningful features derived directly from biological data, ensuring that the initial representation of each node can capture important information about its gene expression.

[0020] In order to see the diagram To derive a meaningful low-dimensional representation, this invention employs a Gaussian Image Array (GAE). The GAE consists of two main components: an encoder and a decoder. The encoder converts the image into a meaningful low-dimensional representation. The nodes in the array are mapped to a low-dimensional space. For each node... (This can be a cell or a gene), the encoder generates a low-dimensional representation matrix. Its formal definition is .

[0021] GAE is trained by minimizing a reconstruction loss, which measures the decoder's ability to reconstruct the original graph structure from the representation generated by the encoder. Loss function The definition is as follows: in, It is the edge set in the original graph. This represents a pair of nodes. After training, the encoder outputs the cell representation matrices of the scRNA-seq and scATAC-seq heterogeneous networks. and the initial gene representation matrix .

[0022] Step S39: Introduce gene networks Gene representation matrix The gene representation matrix is ​​obtained by updating. : in, For genes In gene networks The initial representation in For attention weights, The weight matrix is ​​learnable. This update method enables heterogeneous networks to integrate gene-gene functional associations from the STRING database, forming a more complete biological regulatory network structure; Step S310: Use cells to represent the matrix and gene representation matrix Constructing a heterogeneous network of scRNA-seq / scATAC-seq that includes cell-cell, cell-gene, and gene-gene relationships. . It includes two types of nodes: cells and genes. Cells represent matrices. This represents the intercellular relationships; a value greater than 0 indicates the existence of a corresponding cell-cell edge. Similarly, based on the gene representation matrix... and gene networks The same strategy was used to determine the presence or absence of cell-gene edges and gene-gene edges, thereby constructing a scRNA-seq / scATAC-seq heterogeneous network. .

[0023] This invention is scalable and can be extended to other multi-omics data. Following the heterogeneous network construction method of scRNA-seq and scATAC-seq data, other omics data, such as proteomics and metabolomics, can also construct heterogeneous maps using the same method, and then integrate the heterogeneous maps to achieve the fusion of multi-omics data.

[0024] Step 4: Design a multi-layer heterogeneous network integration scheme to better integrate the ST heterogeneous network and the scRNA-seq / scATAC-seq heterogeneous network, resulting in a multi-layer heterogeneous network. The specific steps are as follows: Step S41: and The connection between two cell layers is established by calculating the cosine similarity between cells: in, and These are the learning rates that control cross-layer similarity and gene embedding contributions, respectively. This study contributes to the gene embedding rate. A data-driven, dynamic adjustment strategy was implemented instead of using fixed values.

[0025] Step S42: For the shared gene layer, an adaptive integration strategy based on a two-layer heterogeneous network was designed. This integration method preserves gene representation information from different sources, laying the foundation for subsequent meta-path feature learning. Step 5: Design a meta-path-based feature enhancement strategy. Through an adaptive multi-view learning framework and a multi-round feature enhancement mechanism, systematically explore the complex semantic relationships in the network to obtain three meta-path views that fully exploit heterogeneous network information. The specific steps are as follows: Step S51: Generate three parallel metapath subgraph views, learn network features from different perspectives, and adaptively determine the metapath length under each view based on data features. For each subgraph Define the probability distribution of its maximum path length: in, It is a view The global representation is obtained by pooling the representations of all nodes: The maximum path length is obtained through sampling: Step S52: Calculate the meta-path score for different lengths in each view, and use an attention mechanism to calculate the attention score for paths of different lengths. For each view, the path attention score is as follows: Step S53: Calculate the attention of the nodes within each meta-path: in, ; Step S54: Using the path attention and node attention obtained in steps S42 and S43, the information for different meta-path lengths is weighted and combined for updating. For each view, the expression matrices of the three types of nodes are updated through adaptive aggregation as follows: in, Indicates the first On the path of the element, the node The set of neighboring nodes. This update mechanism also considers path attention. and node attention weights This ensures the accuracy of information aggregation and allows important biological connections to receive more attention.

[0026] Step Six: Design a two-level fusion and adaptive multi-round feature enhancement mechanism to fuse the meta-path view obtained in Step Five to generate a unified cell expression matrix. This method is used for single-cell biological network inference tasks to improve the model's ability to characterize cellular heterogeneity. The specific steps are as follows: Step S61: Perform hierarchical feature fusion on multiple meta-path views, including meta-path level fusion and view level fusion. Fusion is performed at the meta-path level, where the representations of meta-paths of different lengths in each sub-graph view are fused into a single view feature. in, Indicates the first In each view, The intermediate feature representation obtained after the skip path passes through the attention mechanism; Step S62: Subsequently, at the view level, features from the three views are fused together using MLP to calculate view weights based on their global information, resulting in a unified representation. in, View importance weights: Step S63: Repeat steps S61 to S62 to perform adaptive multi-round enhancement. During the process, an adaptive convergence judgment mechanism based on relative change rate is adopted. The relative change rate of the feature representation after each round of enhancement is calculated, and a dynamic benchmark is used to judge convergence. When the current change rate is significantly lower than the historical change, it is considered to have reached convergence. The minimum number of rounds is set accordingly. and maximum number of rounds As a boundary constraint.

[0027] Step S64: In each round, first enhance cellular expression characteristics through cell-gene interactions: in, Indicates with cells A set of genes that interact. It is an attention-based aggregation function: Step S65: Integrating primitive cell expression and gene-enhanced cell expression: Step S66: Generate a unified cell expression matrix through a cross-modal attention fusion mechanism: The final unified cell expression matrix This is the final representation of cellular biological networks, revealing the essential connections between cells and between cells and genes. It can then be used for clustering tasks to perform cell type annotation, cancer prediction and diagnosis, and many other practical applications.

[0028] The present invention has the following features: (1) The present invention innovatively constructs a multi-layer heterogeneous network with enhanced gene regulation knowledge base. The gene-gene interaction relationship verified in the gene interaction network is introduced into the scRNA-seq and scATAC-seq heterogeneous network layer and the ST heterogeneous network layer respectively. The edges between genes are added to the heterogeneous network to enrich the semantic information between omics genes and significantly enhance the biological significance of the network. This integration strategy not only makes up for the sparsity of experimental data, but also provides more reliable support for gene regulation relationship. (2) The present invention proposes a novel multi-layer network architecture based on spatial constraints. When constructing the ST heterogeneous network, the spatial location and tissue structure similarity of cells are considered at the same time, and it is connected to the scRNA-seq and scATAC-seq layers through shared gene nodes. This hierarchical network structure not only retains the characteristics of each omics data, but also realizes the organic unity of cell spatial organization information and molecular features. (3) The present invention proposes a feature enhancement strategy based on metapath, which integrates multi-hop metapath mode and random subgraph sampling method to systematically explore the complex indirect dependency relationship in the network. This method not only captures indirect connections between different types of nodes, but also learns richer cellular representations, thereby improving the accuracy and biological interpretability of network inference.

[0029] This invention is applied to the field of single-cell biological network inference. By constructing a multi-omics, multi-layered heterogeneous network enhanced by a gene regulation knowledge base, it achieves the organic integration of three omics data: scRNA-seq, scATAC-seq, and ST. It is also scalable and can be extended to the fusion of other multi-omics data, improving cell clustering effects. This lays the foundation for subsequent analysis of regulatory mechanisms, cancer stage analysis, and the revelation of complex biological processes in single-cell biological network inference. Attached Figure Description

[0030] The present invention will be further described below with reference to the accompanying drawings and embodiments; Figure 1 This is a schematic diagram of the structure of the multilayer heterogeneous network single-cell biological network inference method based on meta-path enhancement of the present invention.

[0031] Figure 2 This is a data processing flowchart of the single-cell biological network inference method for multilayer heterogeneous networks based on meta-path enhancement according to the present invention. Detailed Implementation

[0032] To provide a clearer understanding of the technical features, objectives, and effects of this invention, specific embodiments are now described in detail with reference to the accompanying drawings. The effectiveness of this invention is demonstrated on matched multi-omics datasets in the field of single-cell sequencing data analysis.

[0033] Figure 1This is a schematic diagram of the single-cell biological network inference method based on meta-path enhancement for multi-layer heterogeneous networks in this invention. It mainly includes a single-cell heterogeneous network construction module enhanced with a gene regulation knowledge base and a meta-path-based feature enhancement module. The single-cell heterogeneous network construction module with enhanced gene regulation knowledge base fully extracts ST data and constructs an ST heterogeneous network enhanced with a gene regulation knowledge base. Simultaneously, it constructs heterogeneous graphs of scRNA-seq and scATAC-seq data. Through deep learning and external gene networks, it obtains enhanced gene expression matrices and cell representation matrices, constructs a scRNA-seq / scATAC-seq heterogeneous network enhanced with a regulation knowledge base, and finally achieves the construction of the heterogeneous network enhanced with a gene regulation knowledge base through a shared gene layer. The meta-path-based feature enhancement module designs an adaptive multi-view learning framework, generating three parallel meta-path subgraph views. Each subgraph automatically determines the maximum path length based on data features and uses a two-layer attention mechanism to aggregate node information. During multiple rounds of feature enhancement, this invention continuously optimizes cell expression features through cell-gene interaction relationships and cross-modal attention fusion mechanisms, ultimately obtaining a unified cell expression matrix. This design preserves the characteristics of different omics data while achieving an organic unity between cellular spatial organization information and molecular characteristics.

[0034] To verify the effectiveness of this method, this invention used a publicly available mouse breast cancer dataset (GSE212482) from the GEO website. This dataset contains matched scRNA-seq, scATAC-seq, and ST data from the same breast tissue slices at different stages of cancer progression. This invention selected subsets C3, L2, and S2, covering three stages: non-breast cancer, pre-breast cancer, and late-stage breast cancer, respectively. Detailed information about the dataset is shown in Table 1. Cell count refers to the number of cells in the scRNA-seq and scATAC-seq data, and spatial locus count refers to the number of spatially distributed loci in the ST data.

[0035] Table 1. Statistics of Dataset Information To obtain relatively stable experimental results, all methods were repeated 10 times, and the average value was used to represent the final performance of the model on different datasets. The model learning rate was set to 0.001. In the heterogeneous networks of scRNA-seq and scATAC-seq, the half-life distance for calculating the peak-gene weights was set. The learning rate is 10kb, and the number of node samples is 2000. In a multi-layer heterogeneous network, the cross-layer similarity learning rate... Set to 0.4, gene representation contribution rate A data-driven, dynamic adjustment strategy is adopted.

[0036] The experiment used three common unsupervised clustering evaluation metrics to assess the clustering effect: silhouette coefficient (SC), Davis-Bordin index (DBI), and Kalinsky-Hallabas index (CH). These metrics comprehensively evaluated the clustering effectiveness from different perspectives, ensuring a robust and reliable evaluation of the clustering results.

[0037] SC measures how similar an object is to its own cluster compared to other clusters. For each data point... SC value Defined as: in, yes The average distance between the point and all other points in the same cluster. yes The minimum average distance between points in other clusters (i.e., the nearest cluster). The overall SC is the average profile value of all points. SC ranges from -1 to 1, with higher values ​​indicating clearer clusters.

[0038] DBI is a method used to quantify the average similarity ratio between each cluster and its most similar cluster. Its definition is as follows: in, It is the number of clusters. It is the first Each point in the cluster is related to the first... The average distance between the centroids of each cluster It is the first The and the first The distance between the centroids of each cluster. A lower DBI indicates better clustering performance, as it signifies compact clusters and good separation.

[0039] CH, also known as the variance ratio criterion, is used to evaluate the ratio of the dispersion between clusters to the sum of the dispersion within clusters. It is defined as follows: in, It is the distance between each cluster center and the overall mean of all data points. It is the distance between data points within each cluster and the cluster center. It represents the total number of data points. A higher CH index indicates a more defined clustering pattern.

[0040] This invention was tested on three commonly used single-cell multi-omics datasets and compared with a contrastive multi-omics deep fusion clustering algorithm based on dual data augmentation and six commonly used single-cell clustering algorithms. The specific comparison algorithms are described below: DeepMAPS: A deep learning-based multi-omics data integration and analysis method, proposed by Ma A et al. in the paper "Single-cell biological network inference using a heterogeneous graph transformer" in Nature Communications 2023, 14(1):1-14.

[0041] Seurat: A multi-omics data integration method improved on the single-cell data analysis toolkit Seurat, proposed by Hao Y et al. in the paper "Dictionary learning for integrative, multimodal and scalable single-cell analysis" in Nature Biotechnology 2024, 42(2):293-304.

[0042] MOFA+: A multi-omics data integration analysis method based on factor analysis, proposed by Argelaguet R et al. in the paper "MOFA+: a statistical framework for comprehensive integration of multi-modal single-cell data" in Genome Biology 2020, 21(1):1-17.

[0043] STLearn: An analytical tool specifically developed for spatial transcriptome data, proposed by Pham D et al. in their paper "Robust mapping of spatiotemporaltrajectories and cell–cell interactions in healthy and diseased tissues" in Nature Communications 2023, 14(1): 7739.

[0044] CCST: A spatial transcriptomics analysis method based on graph convolutional networks, proposed by Li J et al. in the paper "Cell clustering for spatial transcriptomics data with graph neural networks" in Nature Computational Science 2022, 2(6): 399-408.

[0045] SEDR: A method for spatial transcriptomics data analysis using deep autoencoders, proposed by Xu H et al. in the paper "Unsupervised spatially embedded deep representation of spatial transcriptomics" in GenomeMedicine 2024, 16(1):1-12.

[0046] The comparative experimental results of the three datasets are shown in Tables 2-4. The underlined result indicates the suboptimal result, and the bold result indicates the optimal result.

[0047] Table 2. SC values ​​of our method and six comparison methods on three datasets. Table 3. DBI values ​​of our method and six comparison methods on three datasets. Table 4. CH values ​​of our method and six comparison methods on three datasets. This method consistently achieves higher clustering performance. For the SC metric, this method achieves optimal results on datasets C3, S2, and L2, outperforming the second-best results by 180.3%, 103.3%, and 48.6%, respectively. This indicates that the clusters generated by this method are more cohesive and better reflect the underlying biological structure of cells. For the DBI metric, a lower DBI value indicates tighter relationships within clusters, suggesting more reasonable cluster division. This method achieves optimal performance on all three datasets, reducing the DBI by 81.8%, 48.1%, and 23.6% compared to the second-best results, reflecting the rationality of the clustering results obtained by this invention. For the CH metric, which measures the ratio of inter-cluster dispersion to the sum of intra-cluster dispersion, a higher value indicates a clearer cluster definition. This method improves performance by 39.1%, 94.1%, and 45.0% compared to the optimal results on the three datasets, respectively. These data demonstrate that this method can form tighter intra-cluster structures and clearer inter-cluster boundaries.

[0048] This method represents a significant advancement in the effective integration of multi-omics data. Models such as Deepmaps, Seurat, and MOFA+ focus on scRNA-seq and scATAC-seq data, while models like Stlearn, CCST, and SEDR focus on ST data. For example, on the C3 dataset, this method achieved the highest SC and CH scores and the lowest DBI score, indicating strong intra-cluster relationships and clear inter-cluster relationships, demonstrating its superior clustering performance. These experimental results demonstrate that this method possesses a superior ability to effectively integrate multi-omics data and perform single-cell biological network inference, surpassing the performance of many multi-omics data fusion analysis models.

[0049] As can be seen from the above embodiments, the present invention performs excellently in single-cell multi-omics biological network inference. It also further illustrates that the present invention makes full use of the characteristics of multi-omics, takes into account the rich data information in multi-omics data, effectively extracts the specific and shared information of single-cell multi-omics, and effectively solves the technical difficulties in the practical application process in the field of single-cell biological network inference.

[0050] The embodiments of the present invention have been described above with reference to the accompanying drawings. For those skilled in the art, the present invention is not limited to the details of the exemplary embodiments described above, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Therefore, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0051] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0052] The technologies, shapes, and structures not described in detail in this invention are all known technologies.

Claims

1. A method for inferring single-cell biological networks in multilayer heterogeneous networks based on meta-path enhancement, characterized in that, The method includes the following steps: Step 1: Acquire single-cell multi-omics data, including scRNA-seq data (E), scATAC-seq data (P), and spatial transcriptome data (ST). Obtain the external gene interaction network (Ginter) from the STRING database and construct the gene relationship strength matrix (Rinter). The specific construction method is as follows: Traverse the non-repeating genes of the three omics systems and encode gene relationships into a binary adjacency matrix. : in, This refers to the number of genes in all non-repetitive sequences across the three omics frameworks. This study employs a contrastive learning approach to learn the relationships between genes, specifically for each gene in the gene network. Define the set of positive samples (Including genes) Directly linked genes) and negative sample set (Including genes) (Genes that are not directly connected) Each gene Initial representation in the network The low-dimensional representation of each gene is obtained by a multilayer perceptron (MLP) and through contrastive learning. Based on these representations, arbitrary gene pairs can be computed. Relationship strength The corresponding comparative learning objective function is: in, Cosine similarity is used to measure the degree of similarity between the representations of two genes. The temperature parameter is set to 0.07 to control the difficulty of the comparative learning. Step 2: Obtain the cell representation matrix from the spatial transcriptome data (ST) using dual adjacency relationships and multi-level information aggregation mechanisms. Simultaneously, the gene relationship strength matrix Rinter obtained in step one is used to update the gene expression matrix of the ST data, resulting in the updated gene expression matrix. Construct an ST heterogeneous network G2 with an enhanced gene regulation knowledge base; Step 3: Using the two matrices E and P obtained in Step 1, construct a heterogeneous network structure G3 through a deep graph neural network. Guided by graph autoencoder learning and gene relationship strength Rinter updates, obtain the cell representation matrices of the scRNA-seq and scATAC-seq heterogeneous networks. and gene representation matrix Furthermore, a heterogeneous scRNA-seq / scATAC-seq network G4, enhanced with a gene regulation knowledge base, is constructed. Step 4: Design a multi-layer heterogeneous network integration scheme to integrate the ST heterogeneous network and the scRNA-seq / scATAC-seq heterogeneous network to construct a single-cell multi-omics multi-layer heterogeneous network G. The specific integration process is as follows: The connection between the two cell layers G2 and G4 is established by calculating the cosine similarity between cells: in, and These are the learning rates that control cross-layer similarity and gene embedding contributions, respectively. This study contributes to the gene embedding rate. A data-driven, dynamic adjustment strategy was implemented instead of using fixed values. For the shared gene layer, an adaptive integration strategy based on a two-layer heterogeneous network is designed to retain gene representation information from different sources as the basis for subsequent meta-path feature learning: Step 5: Design a meta-path-based feature enhancement strategy. Through an adaptive multi-view learning framework and a multi-round feature enhancement mechanism, systematically explore the complex semantic relationships in the network and obtain three meta-path views that fully exploit heterogeneous network information. Step Six: Design a two-level fusion and adaptive multi-round feature enhancement mechanism to fuse the meta-path view obtained in Step Five to generate a unified cell expression matrix. It is used for single-cell biological network inference tasks to improve the model's ability to characterize cellular heterogeneity.

2. The method for inferring single-cell biological networks in multilayer heterogeneous networks based on meta-path enhancement according to claim 1, characterized in that, Step two involves constructing the ST heterogeneous network G2 with an enhanced gene regulation knowledge base. The specific steps are as follows: Step S21: Based on the cell representation matrix of the heterogeneous graph G1 represented by ST data, construct two different types of neighbor relationship graphs for cells: a location adjacency graph based on spatial distance and a tissue similarity graph based on expression similarity. For the location adjacency graph, calculate the Euclidean distance between cells, and select the 6 nearest cells of each cell as its spatial neighbors to construct a spatial adjacency matrix. For the tissue similarity map, the expression similarity between cells is calculated, and the six cells with the highest similarity are selected as expression neighbors to construct a tissue similarity matrix. ; Step S22: Based on the cell neighbor relationship graph obtained in step S21, update the cell expression matrix using an adaptive multi-level information aggregation mechanism. First, perform information aggregation for each relationship type separately: in, and Representing cells respectively exist and The six nearest neighbor set in the middle, attention weight and Calculated using the self-attention mechanism: in, and For learnable transformation matrices, , As a learnable attention vector, the self-attention mechanism adaptively determines the importance of different neighboring nodes; Step S23: Adaptively fuse spatial and organizational information through a gating mechanism to obtain the cell representation matrix of the ST heterogeneous network. : in, It is spatial information Organizational information and original representation The candidate cells generated by the fusion represent, and The weight matrix is ​​a learnable matrix. This determines the relative importance of various pieces of information when merging them to generate a new cell representation. Used to learn how to control the extent of information updates. This is a gating vector used to control the fusion ratio of new and original information. The sigmoid activation function maps values ​​to the interval [0,1]. This indicates element-wise multiplication, and this gating mechanism can adaptively determine how much original information and newly aggregated information to retain based on the data; Step S24: Use the relation strength matrix obtained from contrastive learning The gene expression matrix of the ST data is updated to obtain the gene representation matrix of the ST network. For each gene It adaptively aggregates information from the initial ST data and external networks through a self-attention mechanism: in, For genes In gene networks The initial representation in For attention weights, The weight matrix is ​​a learnable matrix; Step S25: Use the cell representation matrix of the ST heterogeneous network and gene representation matrix Construct the ST heterogeneous network G2, which includes cell-cell, cell-gene, and gene-gene relationships.

3. The method for inferring single-cell biological networks in multilayer heterogeneous networks based on meta-path enhancement according to claim 1, characterized in that, Step three involves constructing the G4 heterogeneous scRNA-seq / scATAC-seq network with an enhanced gene regulation knowledge base. The specific steps are as follows: Step S31: Using the two matrices E and P obtained in Step 1, calculate the intercellular Euclidean distance of scRNA-seq. Intercellular Euclidean distance with scATAC-seq The combined distance between scRNA-seq and scATAC-seq cells The result is obtained by weighted summation of distances: ,in The parameter is set to 0.5; the combined distance displays the nearest neighbor relationships between cells, thereby constructing a weighted nearest neighbor matrix. ; Step S32: Use the weighted nearest neighbor method to integrate and obtain the comprehensive matrix. For each cell Its comprehensive expression It is obtained by averaging the expressions of its nearest neighbors from two omics datasets, by definition: in, Represents cells in the WNN matrix The nearest neighbor set, Cells with scRNA-seq data Gene expression, Cells from scATAC-seq data The gene expression of this comprehensive matrix is ​​obtained. Gene expression matrix integrating scRNA-seq and scATAC-seq data; Step S33: Based on the comprehensive matrix Calculate the pairwise cosine similarity between cells to obtain the inter-cell similarity matrix, which reflects the similarity relationship between cells from the two omics systems. ; Step S34: Using the synthesis matrix and inter-cell similarity matrix A heterogeneous network graph G4, fusion of scRNA-seq and scATAC-seq, was constructed. The nodes consist of two types: cells and genes. The edges connecting the nodes differ from those used in other heterogeneous network construction methods. Besides the common cell-gene edges found in other heterogeneous networks, this method also utilizes an inter-cell similarity matrix. Construct cell-to-cell edges; Step S35: Input X and S into the graph autoencoder, minimize the reconstruction loss for training, and output the cell representation matrix of the scRNA-seq and scATAC-seq heterogeneous networks after training. Gene representation matrix, introducing gene interconnection network The gene representation matrix is ​​obtained by updating the gene representation matrix. This allows the heterogeneous network to integrate gene-gene functional associations from the STRING database, forming a more complete biological regulatory network structure. Step S36: Represent the matrix using cells and gene representation matrix The method constructs a scRNA-seq / scATAC-seq heterogeneous network G4 that includes cell-cell, cell-gene, and gene-gene relationships. This method can also be extended to the construction of other multi-omics data biological networks by constructing fusion heterogeneous maps of other omics in the same way, and has strong scalability.

4. The method for inferring single-cell biological networks in multilayer heterogeneous networks based on meta-path enhancement according to claim 1, characterized in that, In step four, a single-cell multi-omics multilayer heterogeneous network is constructed, which integrates the ST heterogeneous network and the scRNA-seq / scATAC-seq heterogeneous network.

5. The method for inferring single-cell biological networks in multilayer heterogeneous networks based on meta-path enhancement according to claim 1, characterized in that, The feature enhancement strategy based on meta-paths in step five is as follows: Step S51: Generate three parallel metapath subgraph views, learn network features from different perspectives, and adaptively determine the metapath length under each view based on data features. For each subgraph Define the probability distribution of its maximum path length: in, It is a view The global representation is obtained by pooling the representations of all nodes: The maximum path length is obtained through sampling: Step S52: Calculate the meta-path score for different lengths in each view, and use an attention mechanism to calculate the attention score for paths of different lengths. For each view, the path attention score is as follows: Step S53: Calculate the attention of the nodes within each meta-path: in, ; Step S54: Using the path attention and node attention obtained in steps S42 and S43, the information for different meta-path lengths is weighted and combined for updating. For each view, the expression matrices of the three types of nodes are updated through adaptive aggregation as follows: in, Indicates the first On the path of the element, the node The set of neighboring nodes, and the update mechanism also considers path attention. and node attention weights To ensure the accuracy of information aggregation, attention should be paid to important biological connections.

6. The method for inferring single-cell biological networks in multilayer heterogeneous networks based on meta-path enhancement according to claim 1, characterized in that, Step six involves a multi-layer heterogeneous network fusion unified representation and an adaptive multi-round feature enhancement mechanism. The specific steps are as follows: Step S61: Perform hierarchical feature fusion on multiple meta-path views, including meta-path level fusion and view level fusion. Fusion is performed at the meta-path level, fusing the representations of meta-paths of different lengths into a single view feature in each sub-graph view. in, Indicates the first In each view The intermediate feature representation obtained after the skip path passes through the attention mechanism; Step S62: Subsequently, at the view level, features from the three views are fused together using MLP to calculate view weights based on their global information, resulting in a unified representation. in, View importance weights: Step S63: Repeat steps S51 to S52 to perform adaptive multi-round enhancement. During the process, an adaptive convergence judgment mechanism based on relative change rate is adopted. The relative change rate of the feature representation after each round of enhancement is calculated, and a dynamic benchmark is used to judge convergence. When the current change rate is significantly lower than the historical change, it is considered to have reached convergence. The minimum number of rounds is set accordingly. and maximum number of rounds As a boundary constraint; Step S64: Generate a unified cell expression matrix through a cross-modal attention fusion mechanism. : The method for inferring single-cell biological networks in multilayer heterogeneous networks based on meta-path enhancement according to claim 1 is characterized in that, In step S53, adaptive multi-round enhancement is performed, and the specific steps are as follows: Step S71: In each round, first enhance cell expression characteristics through cell-gene interactions: in, Indicates with cells A set of genes that interact. It is an attention-based aggregation function: Step S62: Integrating primitive cell expression and gene-enhanced cell expression: The final unified cell expression matrix It is the final representation of cellular biological networks, revealing the connections between cells and between cells and genes, and is used for cell clustering tasks and cell type annotation.

Citation Information

Cited By

  • Inplanatable decision-making system for single-cell multi-omics data integration analysis

    CN122117020A