ScRNA-seq missing data prediction method based on depth map contrast learning

Through the method based on deep map comparison learning, the problem of insufficient efficiency and accuracy of missing value interpolation in single-cell RNA sequencing is solved, and more efficient and accurate recovery of missing value is achieved, promoting cell subpopulations clustering and cell development trajectory inference.

CN120144927APending Publication Date: 2025-06-13GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510328873.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existence of missing values ​​in single-cell RNA sequencing technology poses an obstacle to downstream analysis of the data set, and the existing interpolation methods are insufficient in terms of efficiency and accuracy.

Method used

Using a method based on deep map comparison learning, a cell map is constructed, and a cell embedding representation is learned using the contrast learning module. The neighbor cells of each cell are determined, and the relevant weights are calculated by the least squares method, and the missing values ​​of each cell are finally filled.

Benefits of technology

It improves the accuracy and efficiency of missing value interpolation, can better restore missing values, promote the clustering of cell subpopulations, and help in the inference of cell development trajectory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144927A_ABST
    Figure CN120144927A_ABST
Patent Text Reader

Abstract

The invention discloses an scRNA-seq missing data prediction method based on depth map contrast learning. The scRNA-seq missing data prediction method comprises the following steps: 1, acquiring an scRNA-seq data matrix; 2, the scRNA-seq matrix is preprocessed; and 3, constructing a cell map by using the preprocessed scRNA-seq data. And 4, inputting the cell map data into a contrast learning module, and learning and obtaining the embedded representation of the cells. And 5, determining a neighbor cell of each cell based on the cell embedding representation, calculating a related weight of the neighbor cell and the neighbor cell through a least square method, and finally filling a missing value of each cell. And 6, evaluating the accuracy of the predicted value through downstream analysis. According to the method, the missing value can be well recovered, the clustering of the cell subpopulation is promoted, and the deduction of the cell development trajectory is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of predicting missing data in single-cell RNA sequencing in bioinformatics. Specifically, it relates to a method for predicting scRNA-seq missing data based on deep graph contrastive learning. Background Art

[0002] The rapid development of single-cell RNA sequencing (scRNA-seq) technology has enhanced the study of cell-to-cell heterogeneity and dynamics in complex tissues. Compared with the traditional bulk RNA-seq technology, scRNA-seq can resolve gene expression profiles at single-cell resolution, thus being widely used in fields such as tumor research, immunotherapy, and drug development. Although the single-cell RNA sequencing (scRNA-seq) technology has great potential, a major challenge it faces is the occurrence of missing events, that is, due to technical factors such as low RNA capture rate, mRNA capture conditions, amplification bias, and sequencing depth, some expressed transcripts are wrongly recorded as zero. The presence of missing values in gene expression data poses a considerable obstacle to the downstream analysis of scRNA-seq datasets. Therefore, it is particularly important to develop effective missing value imputation methods.

[0003] In recent years, a variety of imputation methods have been developed, and these methods can be roughly divided into four categories. The first category is the imputation method based on smoothing similarity, and its specific idea is to use the similarity information between cells or genes in the data to recover the missing values. For example, DrImpute uses a clustering method for estimation. It first obtains multiple groups of cell clusters by setting different numbers of clusters to identify similar cells, and then estimates the missing values by averaging the gene expression values of similar cells. The final imputation value is obtained by averaging multiple clustering estimations. However, DrImpute requires multiple iterative clusterings and averaging, which makes its estimation process time-consuming and consumes a large amount of memory resources. MAGIC is based on an affinity graph of Markov chains, and it recovers the expression values of missing genes by sharing information between similar cells through data diffusion technology. However, accurately determining the diffusion time is a major challenge faced by MAGIC, which may lead to overestimation of missing values.

[0004] The second type of imputation method is to use probability models to model sparsity and directly calculate the missing values in the data. For example, SAVER models the unique molecular identifier (UMI) as a negative binomial random variable and uses the gamma-Poisson mixture distribution to model the true expression level of each gene in each cell. bayNorm adopts a binomial model based on the mRNA capture mechanism, infers the prior through empirical Bayesian methods, and estimates the gene expression level using cross-cell expression data. VIPER selects a set of the most similar cells from sparse local neighborhood cells through a non-negative sparse regression model, and then uses the gene expression data of these neighboring cells to impute the missing values of the target cells. scImpute uses a gamma-Poisson mixture model to estimate the missing probability of each gene in each cell, and selects similar cells according to the genes with low missing probabilities to estimate the missing data. However, most of these methods rely on assumptions about the relationships between cells and may limit their performance in cases involving fewer cell types.

[0005] The third type uses matrix factorization techniques to estimate missing data. For example, ALRA is an imputation method based on adaptive threshold low-rank approximation. ALRA uses singular value decomposition (SVD) to transform the high-dimensional gene expression matrix into a low-dimensional approximate representation, thereby preserving the intrinsic correlations between cells or genes. This method automatically determines the approximate rank to identify the main biological signal components. After low-rank approximation, the zero values in the original matrix are filled with non-zero values. By setting a threshold to distinguish and recover biological zero values and rescaling these values so that their mean and standard deviation are consistent with the original matrix, the imputation of technical zero values and the denoising of original non-zero values are achieved. However, ALRA can usually only capture the linear relationships in the original expression data and may ignore some more complex non-linear biological signals.

[0006] The fourth type of imputation method uses deep learning models for estimation. For example, deepImpute constructs multiple sub-neural networks using the "divide and conquer" idea, trains the sub-neural networks using the rejection layer and weighted mean squared error loss function, and estimates missing values. However, deepImpute mainly focuses on learning the similarity between genes and may not be able to fully capture the biological differences at the cellular level. GE-Impute is a graph embedding imputation method that constructs an original cell similarity network using Euclidean distance, then uses breadth-first search (BFS) and depth-first search (DFS) strategies to simulate the embedding representation of each cell with a random walk of a fixed length, and subsequently reconstructs the cell similarity network, estimating the missing data of each cell by averaging the expression values of neighboring cells. However, GE-Impute depends on the average expression values of neighboring cells and may not be able to fully capture the global cell relationships. CL-Impute is a contrastive learning-based imputation method that enhances the model's ability to capture data features by constructing positive and negative sample pairs, thereby generating the embedding representation of cells, and selecting the gene expression values of similar cells based on these embedding representations to impute missing data. However, since CL-Impute uses a multi-head attention network to capture the relationships between cells, which requires the entire matrix to be input into the network for learning, it may consume a large amount of memory.

[0007] In recent years, graph representation learning based on graph neural networks (GNNs) has received extensive attention. GNNs effectively enhance the representation ability of node features by exploring the correlation between the target node and its neighboring nodes in graph-structured data. Most existing GNN models adopt supervised learning methods and require a large amount of labeled data for training. However, the limited data types of scRNA-seq and the difficulty in obtaining sufficient labels have, to a certain extent, restricted the application and development of GNNs in the field of scRNA-seq. As a self-supervised learning method, contrastive learning does not require labeled data, and its core idea is to learn feature representations by maximizing the similarity between positive samples and minimizing the similarity between negative samples. Therefore, it is reasonable and promising to apply contrastive learning to GNNs to learn the embedding features of cell nodes in scRNA-seq data in a self-supervised manner.

[0008] In view of this, it is of great significance to study the imputation method for scRNA-seq missing data. The present invention proposes a method for imputing scRNA-seq missing data based on deep graph contrastive learning. Summary of the Invention

[0009] The present invention proposes a method for predicting scRNA-seq missing data based on deep graph contrastive learning, and the main steps are as follows:

[0010] Step 1: Obtain the scRNA-seq matrix.

[0011] Obtain scRNA-seq from the NCBI GEO or 10x Genomics database and name it , where represents the scRNA-seq data, and represent the number of genes and cells respectively.

[0012] Step 2: Preprocess the scRNA-seq matrix.

[0013] To reduce the technical biases introduced during the sequencing process, the present invention performs data preprocessing on the dataset. Specifically, it is divided into three steps. First, filter the genes expressed in less than three cells and the cells with an expression count less than fifty. Then, since the data in the count matrix is discrete, size factors are used to normalize the filtered gene expression matrix to eliminate the difference limitations between batches. Finally, the log function is used to transform the normalized gene expression matrix, thereby obtaining the preprocessed gene expression matrix.

[0014] Step 3: Construct a cell graph using the preprocessed scRNA-seq data.

[0015] First, use Principal Components Analysis (PCA) to reduce the dimension of the preprocessed single-cell RNA sequencing data. Then, use the cosine distance to calculate the similarity between cells for the dimension-reduced gene expression matrix. Finally, use the K-nearest neighbor (KNN) algorithm to screen the neighbor cells of each cell and construct a cell adjacency graph.

[0016] Step 4: Input the cell graph data into the contrastive learning module to learn and obtain the embedded representation of cells.

[0017] Contrastive learning aims to learn the invariant features between similar and dissimilar data pairs. Since contrastive learning requires two data pairs as input, a graph data augmentation method needs to be adopted to construct data pairs. Specifically, it includes two methods, namely edge dropping and feature masking. Then, use deep graph contrastive learning based on cell node level to learn cell embeddings.

[0018] Step 5: Determine the neighbor cells of each cell based on the cell embedded representation, calculate its correlation weights with the neighbor cells by the least squares method, and finally fill in the missing values of each cell.

[0019] After the network training is completed, the original graph data is fed into the network to obtain the cell embedding matrix. Then, based on the cell embedding matrix, the cosine similarity method is used to find the k most similar cells for each cell. Finally, based on the similar cells, the least squares method is used to estimate the missing values of each cell.

[0020] Step 6: Evaluate the accuracy of the predicted values through downstream analysis.

[0021] To measure the accuracy of the imputed values, the present invention evaluates from three aspects respectively, including the accuracy analysis of the imputed values, clustering, and the inference of cell developmental trajectories. Brief Description of the Drawings

[0022] Figure 1 It is a schematic diagram of the method flow of the present invention.

[0023] Figure 2 It is a comparison graph of the imputation performance between the present invention and other imputation methods.

[0024] Figure 3 It is a comparison graph of the clustering performance between the present invention and other imputation methods.

[0025] Figure 4 It is a clustering graph of cell subsets before and after imputation of the Usoskin dataset by the present invention and other imputation methods.

[0026] Figure 5 It is a clustering graph of cell subsets before and after imputation of the SimData dataset by the present invention and other imputation methods.

[0027] Figure 6 It is a clustering graph of cell subsets before and after imputation of the Baron dataset by the present invention and other imputation methods.

[0028] Figure 7 It is an inference graph of cell developmental trajectories before and after imputation of the E3E7 dataset by the present invention and other imputation methods.

[0029] Figure 8 It is the running time of the present invention and other imputation methods on all datasets. Detailed Embodiments

[0030] The present invention is a method for predicting scRNA-seq missing data based on deep graph contrast learning. The following further elaborates on the present invention in combination with specific embodiments and simulation experiments. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of evidence collection of the present invention.

[0031] As Figure 1As shown in the figure, a method for predicting scRNA-seq missing data based on deep graph contrast learning specifically includes the following steps:

[0032] Preferably, the obtaining of the scRNA-seq association matrix in step 1 is specifically as follows:

[0033] For the real dataset, five publicly available real datasets, Pbmc, Juraket-293t, Baron, Usoskin, and E3E7, were obtained from the NCBI GEO or 10x Genomics database. All five datasets can be represented by where represents the scRNA-seq data, and represent the number of genes and cells, respectively. For the simulated dataset, a simulated dataset simData containing 3000 cells and 1500 genes was generated using the Splatter R package. The cells were divided into seven groups, accounting for 10%, 10%, 10%, 10%, 20%, 20%, and 20% of the total number of cells, respectively.

[0034] Preferably, the preprocessing of the scRNA-seq matrix in step 2 is specifically as follows:

[0035] To reduce the technical bias introduced during sequencing, genes expressed in fewer than three cells and cells with fewer than fifty expression counts were filtered. The filtered matrix is where , . Since the data in the count matrix is discrete, the filtered gene expression matrix is normalized using the size factor to eliminate the difference limitations between batches.

[0036]

[0037] where represents the expression value of gene in cell , represents the sum of all gene expression values in the th cell, represents the number of filtered cells, represents the number of filtered genes, represents the normalized gene expression matrix, is the size factor (referenced from the Seurat R package).

[0038] The log function is used for the normalized gene expression matrix Perform a transformation to obtain the processed gene expression matrix .

[0039]

[0040] Among them, represents the normalized gene expression matrix.

[0041] By normalizing and performing log transformation on each gene expression value, the influence of extreme values caused by sample data with overly large differences is reduced. Finally, since genes with weak expression capabilities in different types do not have a great impact on cell heterogeneity, and genes with high variability need to be paid more attention to, therefore, the top 2000 highly variable genes are selected for subsequent research.

[0042] Preferably, constructing a cell graph using the preprocessed scRNA-seq data in step 3 is specifically as follows:

[0043] Use a graph neural network to perform feature learning on the cell-gene expression matrix to identify similar cells and impute missing values. For this purpose, it is necessary to convert the gene expression data into graph-structured data to capture the association relationships between cells. A cell adjacency graph is constructed through the following three steps.

[0044] (1) Dimensionality reduction of the expression matrix

[0045] Single-cell RNA sequencing data has a high degree of sparsity and often contains a large amount of redundant information, which further masks the true biological differences. Although highly variable genes have been screened, high-dimensional data will still affect the computational efficiency when calculating cell similarity. Therefore, after feature selection, the principal component analysis (PCA) dimensionality reduction algorithm is used to further compress the expression matrix of highly variable genes. PCA captures the main variance of the data through linear combinations of genes, achieving both maximization of variance and the effect of dimensionality reduction.

[0046] (2) Calculate cell similarity

[0047] In the dimensionality-reduced expression matrix, the cosine distance is used to calculate the similarity between cells.

[0048] (3) Construct a k-nearest neighbor graph

[0049] After completing the calculation of cell similarity, sort according to cell similarity. Subsequently, use the K-nearest neighbor (KNN) algorithm to screen the adjacent cells of each cell and construct a cell adjacency graph for subsequent research. Among them represents the cell node, Represents an edge, The element value in is determined by a formula. Represents the preprocessed single-cell RNA sequencing data.

[0050]

[0051] Where Represents a cell And cell The edge relationship between them, Represents a cell And The correlation of is within Sorting ranges, an edge relationship is established between the two nodes, otherwise outside Sorting ranges, there is no edge relationship between the two nodes.

[0052] Preferably, in step 4, the cell graph data is input into the contrast learning module to learn and obtain the embedded representation of the cells. Specifically:

[0053] Contrast learning aims to learn invariant features between similar and dissimilar data pairs. Since contrast learning requires two similar data pairs as input, a graph data augmentation method is adopted to construct data pairs. Research shows that data augmentation can effectively improve the performance of the contrast model. It mainly includes two graph augmentation techniques: edge dropping and feature masking.

[0054] Edge dropping: Randomly The established edges in the graph Are randomly deleted with probability Then, an indicator matrix Is created to determine which edges in the graph Need to be deleted, and finally, Represents the graph After deleting the edges. Among them, the indicator matrix And the edge Can be expressed as:

[0055]

[0056]

[0057] Feature masking: For feature masking, an indicator vector Is generated, and each feature item Where Is the given feature masking probability. The masked feature matrix Is expressed as:

[0058]

[0059] Among them, represents the Hadamard product operator, represents the vector concatenation operator.

[0060] When the edge deletion and feature masking of the graph are both completed, two different but related graphs and can be obtained for subsequent contrast learning.

[0061] When two augmented data and are ready, node-level deep graph contrast learning is used to learn cell embeddings. Since GATv2Conv provides a dynamic attention mechanism, compared with graph neural networks (GNNs) and graph convolutional neural networks (GCNs), it can automatically calculate the importance of neighbor nodes and update the feature information of its own node according to the degree of importance; compared with graph attention networks (GATs), it can calculate the importance degree of neighbor nodes to its own node with a dynamic attention mechanism, making the attention score conditional on the query node. The GATv2Conv network not only depends on node features but also on adjacency relationships when learning cell embedding features. Therefore, it is used as a feature extractor to learn the latent representation of cell nodes. Specifically, the GATv2Conv network performs a self-attention mechanism operation on the graph and iteratively updates the node representation by passing information between neighborhoods. Taking the topological structure and feature matrix of the graph as inputs, low-dimensional node representations can be learned through a multi-head GATv2Conv network, and the specific steps are as follows:

[0062] (1) Input the augmented data and into the encoder network for dimensionality reduction. Taking as an example, its feature matrix is , and input it into for dimensionality reduction. The dimensionality-reduced matrix can be represented by , where is the dimensionality-reduced feature dimension, and the dimension is set to 128.

[0063]

[0064] (2) Input the dimensionality-reduced matrix into the GATv2Conv network to construct the attention coefficient matrix , where each element represents the attention score between node and node :

[0065]

[0066] wherein and respectively represent the nodes (cells) in Figure and the eigenvector of the node (cell) and the node (cell) . The role of and is a trainable parameter. The role of is to reduce the dimension of the feature matrix after splicing and . The role of is to convert the eigenvector into a constant

[0067] (3) Perform softmax normalization on the attention coefficient matrix to obtain the attention matrix , where each element represents the attention score of the normalized node and the node :

[0068]

[0069] wherein represents the set of neighbor nodes (cells) of the node (cell)

[0070] (4) Update the node features using the attention matrix:

[0071]

[0072] where is the Relu activation function, is the updated node feature matrix for the augmented data

[0073] For two augmented graph data, a weight - shared GATv2Conv network is used to generate the embedding matrices and , where is the dimension of the shallow - layer features. Then, the obtained shallow - layer features and are fed into a weight - shared projection head for projecting the embedding features of the two views into a common feature space and calculating the contrast loss in this space. Among them, ​​It is implemented by a two-layer Multilayer Perceptron (MLP). By calculating the contrastive loss in the projection space, better representation ability can be obtained. The projected embedding matrices are respectively denoted as and . Next, maximize the feature representations of the same node in the two generated views of the same node, while minimizing its feature representations with the remaining nodes. Taking a node (cell) as an example, the embedding feature from the node (cell) in Figure node (cell) is . Regarding it as an anchor point, then the embedding feature from the node (cell) in Figure node (cell) is . and are positive samples. For the anchor point , the remaining node features in Figure and Figure are all negative samples. Through training, it is expected to maximize the consistency between the features of positive sample pairs and minimize the consistency between the features of negative sample pairs. Here, the cosine similarity is used to measure the similarity between samples, and InfoNCE (Information Noise-Contrastive Estimation) is used as the loss function, and its formula is:

[0074]

[0075] where is the temperature parameter, is the indicator function, , is the negative sample pair obtained from the same amplified data, is the negative sample pair obtained from different amplified data. Since the two amplified data can be interchanged, finally an overall contrastive loss can be obtained, and its formula is:

[0076]

[0077] where, is the number of samples (number of cells). By minimizing the contrastive loss , the positive sample nodes are pulled closer and the negative sample nodes are pushed away, so that the features of each cell node are effectively retained.

[0078] Preferably, in step 5, the neighbor cells of each cell are determined based on the cell embedding representation, and the correlation weights with the neighbor cells are calculated by the least squares method, and finally the missing values of each cell are filled. Specifically:

[0079] The purpose of contrastive learning is to learn the effective features of each cell. After the network training is completed, the original graph data is fed into the network to obtain the cell embedding matrix , where is the feature dimension, is the number of cells. Based on the cell embedding matrix , the cosine similarity method is used to find the most similar cells for each cell, and then the least squares method is used to estimate the missing values of each cell based on the similar cells. The calculation formula is as follows:

[0080]

[0081] where represents the cell to be imputed. By using the method of scImpute and scLink to identify missing zeros, the missing probability of each gene in each cell is calculated, and the positions with a missing probability greater than 0.5 are marked as cells to be imputed . represents the set of similar cells of cell , represents the correlation weight between the similar cells and the cell to be imputed, which is calculated from their non-zero values. Finally, the gene expression values of the similar cells are used to replace the missing values in the cell to be imputed . The imputed matrix is denoted by , where is the number of genes, is the number of cells.

[0082] Preferably, the accuracy of the predicted values is evaluated through downstream analysis in step 6. Specifically:

[0083] To measure the accuracy of the imputed values, the present invention is evaluated from three aspects, including the accuracy analysis of the imputed values, clustering, and the inference of cell developmental trajectories.

[0084] The technical effects of the present invention are further described below through experimental verification:

[0085] 1. Experimental conditions and content:

[0086] The experiments of the present invention are all completed on a 2.50GHz CPU, NVIDIA GeForce RTX 4060, and a windows11 operating system.

[0087] 2. Analysis of experimental results:

[0088] Evaluations conducted on five real scRNA-seq datasets and one simulated dataset consistently show that, compared with existing imputation methods, the imputation method of the present invention exhibits superior performance. It can well recover missing values, facilitate the clustering of cell subsets, and contribute to cell developmental trajectory inference.

[0089] The experimental comparison results conducted by the present invention are as Figures 2 - 5 shown.

[0090] The above description is only a specific example of the present invention and does not constitute any limitation to the present invention. Obviously, for professionals in the field, understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention, and all of them should be included in the protection scope of the claims of the present invention.

Claims

1. A method for predicting missing data in scRNA-seq based on deep graph contrast learning, characterized in that: The following steps are involved: Step 1: Obtain scRNA-seq data; Step 2: Preprocessing scRNA-seq data; Step 3: Construct a cell map using the preprocessed scRNA-seq data; Step 4: Input the cell map data into the contrastive learning module to learn and obtain the embedded representation of the cell; Step 5: Determine the neighbor cells of each cell based on the cell embedding representation, and calculate the relevant weights between it and the neighbor cells by the least squares method, and finally fill the missing values ​​of each cell; Step 6: Evaluate the accuracy of the predicted values ​​through downstream analysis.

2. The scRNA-seq missing data prediction method based on deep graph contrast learning according to claim 1, characterized in that: In step 1, specifically: Get scRNA-seq from NCBI GEO or 10x Genomics database and name it ,in represents scRNA-seq data, and represent the number of genes and cells, respectively.

3. The scRNA-seq missing data prediction method based on deep graph contrast learning according to claim 1, characterized in that: In step 2, specifically: In order to reduce the technical deviation introduced in the sequencing process, the present invention performs data preprocessing on the data set; specifically, it is divided into three steps. First, genes expressed in less than three cells and cells with less than fifty expressions are filtered out; then, since the data in the count matrix is ​​discrete, the filtered gene expression matrix is ​​standardized using a size factor to eliminate the difference restrictions between batches; finally, the normalized gene expression matrix is ​​transformed using a log function to obtain a preprocessed gene expression matrix.

4. The scRNA-seq missing data prediction method based on deep graph contrast learning according to claim 1, characterized in that: In step 3, specifically: First, principal component analysis (PCA) was used to reduce the dimension of the preprocessed single-cell RNA sequencing data. Then, the cosine distance was used to calculate the similarity between cells in the gene expression matrix after dimensionality reduction. Finally, the K-nearest neighbor (KNN) algorithm was used to screen the neighbor cells of each cell and construct a cell adjacency graph.

5. The scRNA-seq missing data prediction method based on deep graph contrast learning according to claim 1, characterized in that: In step 4, specifically: Contrastive learning aims to learn the invariant features between similar and dissimilar data pairs. Since contrastive learning requires two data pairs as input, a graph data augmentation method is needed to construct the data pairs. Specifically, there are two ways, namely discarding edges and feature masks. Then, deep graph contrastive learning based on cell node level is used to learn cell embedding.

6. The scRNA-seq missing data prediction method based on deep graph contrast learning according to claim 1, characterized in that: In step 5, specifically: After the network training is completed, the original graph data is sent into the network to obtain the cell embedding matrix; then the cosine similarity method is used to find the k most similar cells for each cell based on the cell embedding matrix; finally, the least squares method is used to estimate the missing value of each cell based on similar cells.

7. The scRNA-seq missing data prediction method based on deep graph contrast learning according to claim 1, characterized in that: In step 6, specifically: In order to measure the accuracy of the interpolation values, the present invention evaluates from three aspects, including accuracy analysis of the interpolation values, clustering, and inference of cell development trajectories.

Citation Information

Cited By

  • Single-cell RNA sequencing data interpolation method based on hypergraph contrast learning and application thereof

    CN121075421A