A single-cell RNA sequencing data clustering method

By constructing a ZINB-based auto-coding module and graph auto-coding module, and performing attention fusion and self-supervised training, the transition smoothing problem of single-cell RNA sequencing data clustering method when there are too many layers is solved, and the clustering performance is significantly improved.

CN118675623BActive Publication Date: 2025-05-27YUNNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410787672.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-18
Publication Date
2025-05-27
Estimated Expiration
2044-06-18

AI Technical Summary

Technical Problem

Existing single-cell RNA sequencing data clustering methods are prone to transition smoothing when there are too many layers, and the key patterns of gene expression are lost, resulting in poor clustering performance.

Method used

A single-cell RNA sequencing data clustering method is adopted to construct a ZINB-based auto-coding module and a graph auto-coding module, and connect them layer by layer through attention fusion mechanism, fusing gene expression information and cell structure information, and end-to-end clustering training and synchronous optimization through self-supervised strategies.

Benefits of technology

Effectively fusing gene expression information and cell structure information improves clustering performance, avoids the problem of transition smoothing, and significantly improves the accuracy of cell clustering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118675623B_ABST
    Figure CN118675623B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of cell clustering, and discloses a method for clustering single-cell RNA sequencing data. The present invention extracts the content information of single-cell RNA sequencing data based on the self-encoding module of ZINB, and reconstructs the gene expression matrix and data distribution through ZINB decoding; secondly, uses the graph self-encoding module to capture the high-order structural information between cells, and extracts the structural information between cells by performing graph convolutional network processing on the cell-cell graph; through the attention fusion mechanism, the self-encoding module based on ZINB and the graph self-encoding module are fused and embedded layer by layer, effectively fusing the feature representation of gene expression information and structural information into the same representation, enhancing the clustering performance; finally, the present invention adopts a self-supervised strategy to integrate the self-encoding module based on ZINB and the graph self-encoding module into a unified framework, and effectively performs end-to-end clustering training and synchronous optimization on these two modules, thereby improving the clustering performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of RNA sequencing data clustering, and in particular, to a method for clustering single-cell RNA sequencing data. Background Art

[0002] Cell clustering is one of the most important tasks in single-cell RNA sequencing data analysis. However, due to the limitations of sequencing technology, single-cell RNA sequencing data has high sparsity and complex noise patterns. Traditional clustering methods such as kmeans and hierarchical clustering rely on similarity metrics and cannot well meet the requirements of single-cell data clustering.

[0003] To better reveal the specific expression patterns of scRNA-seq data, deep embedding clustering algorithms have been proposed and used for cell type identification and clustering tasks, such as DESC, scDCC, scDeepCluster, etc.

[0004] These methods usually use an autoencoder to learn the latent representation of the data and the clustering assignment. However, these methods only focus on the learning of the data itself and ignore the relationships between cells, that is, the structural information of the data. Due to the sparsity of single-cell data, it is not ideal to learn the latent representation only from gene expression information to guide cell clustering. The structural information of cells (i.e., the relationships between cells) must be embedded into the deep clustering method. Based on this idea, deep graph embedding clustering algorithms have been proposed to solve the cell clustering problem, such as cGNN, scGAE, scDSC, etc. They usually use a graph autoencoder to capture cell structural information to guide clustering. However, these methods tend to be overly smooth when there are many layers, thus losing the key patterns of gene expression and ultimately resulting in poor cell clustering performance.

[0005] Therefore, the prior art still needs to be improved and developed. Summary of the Invention

[0006] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a method for clustering single-cell RNA sequencing data, aiming to solve the problem that the existing clustering methods based on cell RNA sequencing data are prone to overly smooth when there are too many layers, thus losing the key patterns of gene expression and resulting in poor clustering performance.

[0007] In the first aspect of the present invention, a method for clustering single-cell RNA sequencing data is provided, including: obtaining single-cell RNA sequencing data, preprocessing the single-cell RNA sequencing data to obtain a single-cell RNA gene expression matrix X; constructing a cell-cell graph adjacency matrix A using the KNN method based on the single-cell RNA gene expression matrix X; constructing a ZINB-based autoencoder module and a graph autoencoder module; connecting the ZINB-based autoencoder module and the graph autoencoder module layer by layer through an attention fusion mechanism, and effectively fusing the feature representation of gene expression information and cell structure information into the same representation through this layer-by-layer embedding operation; integrating the ZINB-based autoencoder module and the graph autoencoder module into a unified framework through a self-supervised strategy, and performing end-to-end clustering training and synchronous optimization on these two modules based on the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A to obtain a clustering prediction model; processing the single-cell RNA sequencing data to be clustered to obtain a single-cell RNA gene expression matrix X and a cell-cell graph adjacency matrix A to be clustered, inputting them into the clustering prediction model, outputting an embedded representation, and obtaining the predicted cell categories based on the embedded representation to complete clustering.

[0008] Optionally, in the first implementation manner of the first aspect of the present invention, constructing a ZINB-based autoencoder module includes the steps: the autoencoder module consists of a symmetric encoder and decoder, the encoder of the autoencoder module has a total of L layers, the single-cell RNA gene expression matrix X is used as the initial input of the autoencoder module, and the input of the l-th layer is H l-1 , then its output representation is: H l =φ(W l H l-1 +b i ), where φ is an activation function, W l represents the weight of the l-th layer, and b i represents the bias vector of the l-th layer; three independent fully connected layers are respectively connected to the last layer of the decoder to estimate the three parameters of ZINB: the probability parameter π, the dispersion parameter θ, and the mean parameter μ, so as to integrate the ZINB model into the autoencoder module; the ZINB distribution uses the three parameters to simulate the single-cell RNA sequencing data and reconstruct its data distribution: ZINB(X|π, μ, θ)=πδ 0 (X)+(1 - π)NB(X|μ, θ); the final loss function of the ZINB-based autoencoder module is defined as the negative log-likelihood estimate of the ZINB distribution:

[0009] Optionally, in the second implementation manner of the first aspect of the present invention, the steps of constructing the graph autoencoder module include: the graph autoencoder module uses several layers of graph convolutional networks as the backbone network; the graph autoencoder module has a total of L layers, and the cell-cell graph adjacency matrix A is used as the initial input of the first layer of the graph autoencoder module, and the input of the l-th layer is Z l-1 , and its output is expressed as: where I is the identity diagonal matrix, is the degree matrix, U l-1 represents the weight of the (l - 1)-th layer, and the normalized adjacency matrix propagates Z l-1 to obtain a new representation Z l .

[0010] Optionally, in the third implementation manner of the first aspect of the present invention, the steps of effectively fusing the feature representation of gene expression information and the cell structure information into the same representation by layer-by-layer connection of the ZINB-based autoencoder module and the graph autoencoder module through an attention fusion mechanism include: passing the output of the ZINB-based autoencoder module and the output of the graph autoencoder module through an attention fusion module for layer-by-layer heterogeneous structure fusion embedding to obtain an integrated representation integrating structural information and content information: R l-1 = F att (αH l-1 +(1 - α)Z l-1 ), where α is an adjustable weight parameter, and F att represents the attention fusion operation; using the integrated representation R l-1 as the input of the graph convolutional network to learn high-order discriminant information and generate a new representation

[0011] Optionally, in the fourth implementation manner of the first aspect of the present invention, integrating the ZINB-based autoencoder module and the graph autoencoder module into a unified framework through a self-supervised strategy, and performing end-to-end clustering training and synchronous optimization on these two modules based on the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A, the steps of obtaining the clustering prediction model include: performing k-means clustering on the intermediate layer output of the ZINB-based autoencoder module to obtain a set of initial clustering centers where k is the number of clusters; using the Student's t-distribution to calculate the soft assignment between the embedding representation and the clustering center , and the formula is as follows: where h i is the embedding representation The i-th sample, λ represents the degrees of freedom of the Student's t-distribution, and the soft assignment q ij ∈Q, where Q is the soft clustering distribution; based on the soft clustering distribution Q, an auxiliary target distribution is set to supervise the learning of the soft clustering distribution. By learning the high-confidence assignments of the target distribution, the clustering is improved, and the soft clustering frequencies are used to calculate the target distribution. The calculation method is as follows: where is the soft clustering frequency, and the target distribution p ij ∈P; the KL divergence loss between the soft clustering distribution Q and the target distribution P is used as the optimization objective to obtain higher-quality clustering: The target distribution P is used to supervise the learning of the graph autoencoder module: where z ij ∈Z pre , Through the above two loss functions, the optimization objective is integrated into a target distribution P, making the learned representation suitable for the clustering task.

[0012] The second aspect of the present invention provides a single-cell RNA sequencing data clustering device, including: a first data processing module for obtaining single-cell RNA sequencing data, preprocessing the single-cell RNA sequencing data to obtain a single-cell RNA gene expression matrix X; a second data processing module for constructing a cell-cell graph adjacency matrix A using the KNN method according to the single-cell RNA gene expression matrix X; a construction module for constructing a ZINB-based autoencoder module and a graph autoencoder module; a fusion module for layer-by-layer connecting the ZINB-based autoencoder module and the graph autoencoder module through an attention fusion mechanism, effectively fusing the feature representations of gene expression information and cell structure information into the same representation through this layer-by-layer embedding operation; a training module for integrating the ZINB-based autoencoder module and the graph autoencoder module into a unified framework through a self-supervised strategy, performing end-to-end clustering training and synchronous optimization on these two modules based on the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A to obtain a clustering prediction model; a clustering module for processing the single-cell RNA sequencing data to be clustered to obtain a single-cell RNA gene expression matrix X and a cell-cell graph adjacency matrix A to be clustered, inputting them into the clustering prediction model, outputting an embedded representation, and obtaining the predicted cell categories according to the embedded representation to complete the clustering.

[0013] In a third aspect of the present invention, a single-cell RNA sequencing data clustering device is provided, including: a memory and at least one processor. Computer-readable instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor invokes the computer-readable instructions in the memory to cause the single-cell RNA sequencing data clustering device to execute each step of the single-cell RNA sequencing data clustering method as described above.

[0014] In a fourth aspect of the present invention, a computer-readable storage medium is provided. Computer-readable instructions are stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute each step of the single-cell RNA sequencing data clustering method as described above.

[0015] Beneficial effects: The present invention provides a single-cell RNA sequencing data clustering method. A multi-layer graph convolutional network (GCN) in the graph auto-encoding module is used to capture high-order structural relationships in single-cell RNA sequencing data; in order to alleviate the problem of over-smoothing of the graph GCN, a ZINB-based auto-encoding module is introduced to extract content information of single-cell RNA sequencing data and learn the latent representation of gene expression data; then, an attention fusion mechanism is used to fuse and embed the above two modules layer by layer and use it as the input of each layer of the GCN to guide the final clustering direction. Therefore, the method of the present invention cross-fuses gene expression information and cell structure information, thereby improving the clustering performance. Description of the Drawings

[0016] Figure 1 It is a flowchart of the single-cell RNA sequencing data clustering method provided by the embodiment of the present invention.

[0017] Figure 2 It is a schematic diagram of the principle of the single-cell RNA sequencing data clustering method of the present invention.

[0018] Figure 3 It is a schematic diagram for comparing the clustering analysis evaluation indexes of the method of the present invention (scASDC) and six other baseline methods on four datasets.

[0019] Figure 4 It is a schematic diagram for comparing the results of dimensionality reduction visualization of the final embedded representation using UMAP for the method of the present invention and other baseline methods on the QS Limb Muscle dataset.

[0020] Figure 5 It is a schematic diagram of the structure of the single-cell RNA sequencing data clustering device provided by the embodiment of the present invention.

[0021] Figure 6 It is a schematic diagram of the structure of the single-cell RNA sequencing data clustering device provided by the embodiment of the present invention. Detailed implementation manners

[0022] An embodiment of the present invention provides a method for clustering single-cell RNA sequencing data. Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0023] Please refer to Figure 1 , Figure 1 which is a flowchart of a preferred embodiment of a method for clustering single-cell RNA sequencing data provided by the present invention. As shown in the figure, it includes the steps:

[0024] S10. Obtain single-cell RNA sequencing data, preprocess the single-cell RNA sequencing data to obtain a single-cell RNA gene expression matrix X;

[0025] S20. According to the single-cell RNA gene expression matrix X, use the KNN method to construct a cell-cell graph adjacency matrix A;

[0026] S30. Construct a ZINB-based autoencoder module and a graph autoencoder module;

[0027] S40. Connect the ZINB-based autoencoder module and the graph autoencoder module layer by layer through an attention fusion mechanism, and effectively fuse the feature representation of gene expression information and cell structure information into the same representation through this layer-by-layer embedding operation;

[0028] S50. Integrate the ZINB-based autoencoder module and the graph autoencoder module into a unified framework through a self-supervised strategy, and perform end-to-end clustering training and synchronous optimization on these two modules based on the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A to obtain a clustering prediction model;

[0029] S60. After processing the single-cell RNA sequencing data to be clustered, obtain a single-cell RNA gene expression matrix X and a cell-cell graph adjacency matrix A to be clustered, input them into the clustering prediction model, output an embedded representation, and obtain the predicted cell categories according to the embedded representation to complete clustering.

[0030] Specifically, as Figure 2 shown, the present invention takes the preprocessed single-cell RNA sequencing data matrix X as the input of the ZINB-based autoencoder module, which can learn a mapping from the data space to the low-dimensional feature space, capture the content information in the data, obtain a low-dimensional latent representation containing feature information, and reconstruct the gene expression matrix and data distribution through ZINB decoding; meanwhile, the present invention constructs a cell-cell graph adjacency matrix A based on the single-cell RNA sequencing data matrix X and takes it as the initial input of the graph autoencoder module, which is composed of a multi-layer graph convolutional network and can effectively extract the high-order structural information between cells; it should be noted that the present invention creatively connects the above two modules layer by layer through an attention fusion mechanism and takes them as the input of each layer of the graph convolutional network. Through this layer-by-layer embedding operation, the feature representation and structural information of the gene expression information can be effectively fused into the same representation; finally, a self-supervised strategy is used to integrate the ZINB-based autoencoder module and the graph autoencoder module into a unified framework, and the two modules are effectively subjected to end-to-end clustering training and synchronous optimization to obtain a clustering prediction model; based on the clustering prediction model, the gene expression information and cell structure information are cross-fused to improve the clustering performance. Therefore, the method of the present invention can learn effective latent representations of cells and genes, and this latent representation contains both the content information and structural information of the original single-cell RNA sequencing data, and has better robustness in downstream analysis tasks such as single-cell clustering.

[0031] In some embodiments, after obtaining the single-cell RNA sequencing data, taking the scRNA-seq gene expression count matrix as the input, where n is the number of cells and g is the number of genes, using the scanpy toolkit to perform data screening and preprocessing. First, filter out the genes that are not expressed in the cells, that is, remove the all-zero columns in to reduce the high dropout rate in the scRNA-seq data (the phenomenon that the expression of some genes is not detected in some cells); secondly, convert the discrete gene expression data into continuous data to ensure the stability of the size factor, that is, perform data normalization, defined as follows:

[0032] where represents the median of the total expression of all cells; finally, calculate the normalized dispersion value of each gene, sort it, and use the top p highly variable genes with high information content for research; finally, obtain the preprocessed single-cell RNA gene expression matrix where n represents the number of cells, is the representation of the i-th cell, and p represents the number of genes.

[0033] In some embodiments, after obtaining the single-cell RNA gene expression matrix X, the KNN method is used to construct a cell-cell graph adjacency matrix For capturing the structural information in the data, in the KNN graph, each node represents a cell, and the edges between the nodes represent the relationships between the cells. First, the heat kernel method is used to calculate the similarity between cells: where t represents the variance scale parameter. Then, using the similarities between all the calculated cells, the top k cells with the highest similarity to each cell are extracted and they are connected, that is: The cell-cell graph adjacency matrix is obtained

[0034] In some embodiments, for the single-cell clustering task, it is important to learn an effective data representation to simulate the data distribution of single-cell RNA sequencing data. The present invention adopts a ZINB autoencoder module for feature extraction and data reconstruction. The zero-inflated negative binomial (ZINB) distribution can effectively simulate highly sparse and over-dispersed count data for capturing the overall probability structure of scRNA-seq data. In this embodiment, the autoencoder module is composed of a symmetric encoder and decoder. Specifically, assume that the autoencoder encoding module has a total of L layers, and the input of the l-th layer is H l-1 , then its output can be expressed as: H l = φ(W l Hx -1 + b i ), where the activation function φ can be flexibly selected according to specific applications. The present invention selects the ReLU function as the activation function. W l represents the weight of the l-th layer, and b i represents the bias vector of the l-th layer. It should be noted that the input H 0 of the first layer is the preprocessed single-cell RNA sequencing data matrix X, and the output H L of the L-th layer is the reconstructed gene expression matrix X, which has the same dimension as the original input matrix X

[0035] Different from the traditional autoencoder, in order to capture the global probability structure of the data, this embodiment integrates the ZINB model into the decoder of the autoencoder module. Specifically, in this embodiment, three independent fully connected layers are respectively connected to the last layer of the autoencoder module to estimate the three parameters of the ZINB distribution: the probability parameter π, the dispersion parameter θ, and the mean parameter μ: where the size factor S iIt is obtained by dividing the sum of the total expression levels of each cell by the total expression level of the reference cell (usually the cell with the median or average expression level), and is calculated during the data preprocessing stage, W π , W μ , W θ represents the weight parameter. The ZINB distribution uses the above three parameters to simulate the original scRNA-seq data and reconstruct its data distribution: ZINB(x|π, μ, θ) = πδ 0 (X) + (1 - π)NB(x|μ, θ),

[0036] Furthermore, in order to guide the model to learn more effective representations, this embodiment hopes that the reconstructed ZINB distribution is as close as possible to the original data distribution. Therefore, this embodiment defines the final loss function of the ZINB-based autoencoder module as the negative log-likelihood estimation of the ZINB distribution:

[0037]

[0038] In some embodiments, although the ZINB-based autoencoder module can learn useful representations from gene expression data and reconstruct the data distribution of the gene expression matrix, it ignores the structural information between cells. This embodiment introduces a graph autoencoder module to solve this problem. This module consists of several layers of graph convolutional networks (GCNs) as the backbone network. The graph convolutional network takes the graph structure as the input, can effectively extract the structural information in the data, and perform heterogeneous structure fusion on the content information and the structural information. Specifically, this embodiment first constructs a cell-cell graph adjacency matrix A through KNN as the initial input of the GCN. Assume that the graph autoencoder encoding module has L layers, and the input of the l-th layer is Z l-1 , then its output can be expressed as: where I is the identity diagonal matrix, is the degree matrix, U l-1 represents the weight of the (l - 1)-th layer, and the normalized adjacency matrix propagates Z l-1 to obtain a new representation Z l .

[0039] In some embodiments, in order to obtain a more abundant representation, this embodiment performs layer-by-layer heterogeneous structure fusion embedding on the output of the ZINB-based autoencoder and the output of the graph autoencoder through an attention fusion module to obtain a more robust integrated representation R l-1 , R l-1 = F att (αH l-1 + (1 - α)Zl-1 ), where α is an adjustable weight parameter, and F att represents the attention fusion operation. To enable the network to learn more abundant representations and reduce the interference of noise, the present invention uses the integrated representation R l-1 as the input of the GCN to learn high-order discriminant information and generate new representations: It should be noted that the input R 0 to the first layer of the graph autoencoder module is the preprocessed single-cell RNA sequencing data matrix X.

[0040] Among the representations learned by the GCN, the middle layer is selected for subsequent clustering, and a softmax operation is performed on the middle layer to predict the probability distribution of the single-cell data: where the obtained representation z ij ∈Z pre can be regarded as the probability that the i-th sample belongs to the j-th cluster center, and is used for the end-to-end model learning and training of the subsequent self-supervised module.

[0041] Furthermore, in order to guide the model training direction and make the learned representations more reliable, the present embodiment uses the following two reconstruction errors as the loss function of the graph autoencoder module: where the output Z of the last layer of the graph autoencoder module is transformed into a reconstructed adjacency matrix L through the inner product operation By minimizing the formula the graph autoencoder module can capture more effective structural information, which is beneficial to improving the subsequent clustering performance. And the minimization of the formula expects the final output Z L of the graph autoencoder module to retain some expression information of the original single-cell RNA sequencing data, so that the final data representation still contains some specific expression patterns of the original single-cell data.

[0042] In this embodiment, the two modules, namely the ZINB-based autoencoder module and the graph autoencoder module, can respectively learn the content information and structural information of the data. How to fuse and embed the outputs of these two modules is a very important issue. A simple way is to simply add the two outputs to obtain a fused representation. However, this approach will lose some important expression patterns or structural information. Therefore, in this embodiment, an attention mechanism is used to fuse the content information and structural information and reduce noise to obtain a richer representation. Briefly speaking, the attention mechanism maps a query and a set of key-value pairs to obtain an output, where the query, key-value, and output are all vectors. Specifically, we calculate the dot product of the query and the key and apply the softmax function to obtain the relevant weights of the values, and then multiply these weights by the values to get the final result. To simplify the calculation, in the actual operation process, a group of key-value pairs are calculated simultaneously each time, and the vector calculation is transformed into a matrix operation. The specific calculation method of the attention mechanism is as follows: Attention(Q, K, V) = softmax(QK T )V, where Q, K, and V are the query matrix, key matrix, and value matrix respectively. On this basis, the present invention further introduces a multi-head attention mechanism, which can simultaneously focus on the information in each subspace at different positions of the data, so that the obtained attention fusion representation contains richer and more robust information.

[0043] Specifically, the present invention performs the attention operation on Q, K, and V M times respectively with different attention operations, and then executes the attention function in parallel, connects them and projects them again to obtain the final result. The i-th attention operation is expressed as follows: Among them, is a transformation matrix responsible for projecting the corresponding three matrices respectively.

[0044] Assume that the attention module has a total of L layers, and the output of its l-th layer is:

[0045] R l = F att (Q, K, V) = F att (W q Y l-1 , W k Y l-1 , W v Y l-1 ) = W · Concat(head 1 ,..., head M ), where Y l-1 = αH l-1 + (1 - α)Z l-1 , W q , W k , W vis the transformation matrix, W is a weight matrix, and Concat(·) represents the matrix concatenation operation.

[0046] In some embodiments, the ultimate goal of model training is to obtain a better cell clustering effect. Therefore, it is necessary to introduce a self-supervised module to supervise the model and make the intermediate layer output learned finally be an optimal clustering representation as much as possible. This module mainly includes three important distributions: the soft clustering distribution Q, the target distribution P, and the probability distribution Z for clustering. These three distributions are all calculated during the training process and change continuously as the training progresses to obtain a more suitable representation for clustering. Therefore, this module is called the self-supervised module.

[0047] Specifically, the present invention first performs K-means clustering on the intermediate layer output of the ZINB-based autoencoder module to obtain a set of initial clustering centers where k is the number of clusters for clustering, which is explicitly given before training; then it is necessary to calculate the soft assignment between the embedding representation and the clustering center The present invention uses the Student's t-distribution to calculate the similarity between the two to obtain its soft assignment. For the i-th sample and the j-th clustering center, its soft assignment q ij ∈Q (which can be understood as the probability that the i-th sample is assigned to the j-th cluster) is calculated as follows:

[0048] where h i is the i-th sample of the embedding representation and λ represents the degree of freedom of the Student's t-distribution, which is a modifiable parameter and is set to 1 in the present invention.

[0049] Based on the soft clustering distribution Q, the present invention also requires an auxiliary target distribution to supervise the learning of the soft clustering distribution. By learning the high-confidence classification of the target distribution, the clustering is improved. A suitable target distribution is crucial for the final clustering quality.

[0050] The present invention uses the soft clustering frequency to calculate a target distribution p ij ∈P, and its calculation method is as follows:

[0051] where, is the soft clustering frequency.

[0052] To make the distributions Q and P closer, the present invention uses the KL divergence loss between these two distributions as the optimization target for this part to obtain higher-quality clustering:

[0053]

[0054] This self-supervised strategy can guide the ZINB-based autoencoder module to learn a data representation more suitable for clustering, while the representation learned by the graph autoencoder module contains richer content information and structural information. We can also adopt the same self-supervised strategy and use the target distribution P to supervise the learning of the graph autoencoder module:

[0055] where z ij ∈Z pre . Through and these two loss functions, we integrate the optimization objective into a distribution P, making the learned representation more suitable for the clustering task.

[0056] Based on this, the total loss function of the entire model of the present invention is as follows:

[0057] where λ 1 ,λ 2 ,λ 3 ,λ 4 ,λ 5 are hyperparameters measuring the importance of each loss function.

[0058] In the present invention, through the above self-supervised strategy, the ZINB-based autoencoder module and the graph autoencoder module are integrated into a unified framework, and end-to-end clustering training and synchronous optimization are performed on these two modules based on the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A to obtain a clustering prediction model; after processing the single-cell RNA sequencing data to be clustered, the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A to be clustered are obtained, and they are input into the clustering prediction model, and the embedded representation Z pre is output, and the index of the maximum value in each row (or each sample) of Z pre is found, that is, the category of each cell predicted by the model is obtained, and the prediction label is obtained to complete the clustering operation.

[0059] In summary, the present invention proposes a clustering method for single-cell RNA sequencing data with enhanced attention. This method introduces a ZINB-based autoencoder module to extract the content information of the data. This module captures the content information in the data by learning the mapping from the data space to the low-dimensional feature space, and reconstructs the gene expression matrix and data distribution through ZINB decoding. This method can effectively simulate highly sparse and over-dispersed count data, thereby better capturing the overall probability structure of single-cell RNA sequencing data. Secondly, the present invention designs a graph autoencoder module, which uses a graph convolutional network to capture the high-order structural information between cells. By performing graph convolutional network processing on the cell-cell graph, the structural information between cells can be effectively extracted, thereby performing heterogeneous structure fusion of the content information and the structural information. Thirdly, the present invention designs an attention fusion mechanism to fuse and embed the ZINB-based autoencoder module and the graph autoencoder module layer by layer, guiding the final clustering direction. This layer-by-layer embedding operation can effectively fuse the feature representation of gene expression information and the structural information into the same representation, enhancing the clustering performance. Finally, the present invention adopts a self-supervised strategy to integrate the ZINB-based autoencoder module and the graph autoencoder module into a unified framework, and effectively performs end-to-end clustering training and synchronous optimization on these two modules, thereby improving the clustering performance.

[0060] In the deep learning environment based on PyTorch, we experimentally compared the capabilities and robustness of the method of the present invention (scASDC) with other mainstream algorithms (DESC, SDCN, scDeepCluster, DCA, scDSC, DEC). The experimental datasets are shown in Table 1. In this embodiment, 4 real and publicly available single-cell RNA sequencing datasets are collected on multiple platforms. These datasets mainly come from different tissues of humans and mice. These datasets include Pollen (tissue), Adam (kidney), QSLimb Muscle (limb muscle), QSDiaphragm (diaphragm). Among them, the two 'QS' datasets come from the mouse single-cell RNA sequencing data generated by Smart-seq2 sequencing in the Stanford University research. The Pollen (SRP041736) and Adam (GSE94333) datasets come from the National Center for Biotechnology Information of the United States (https: / / www.ncbi.nlm.nih.gov / ). Through clustering performance evaluation metrics and chart analysis, we evaluated each group of data in detail.

[0061] Table 1 Single-cell RNA sequencing data information

[0062] Data set Number of cells Number of genes Number of categories Platform Pollen 301 21721 11 SMARTer QS Limb Muscle 1090 23341 6 Smart-seq2 Adam 3660 23797 8 Drop-seq QSDiaphragm 870 23341 5 Smart-seq2

[0063] When evaluating unsupervised clustering algorithms, NMI (Normalized Mutual Information) and ARI (Adjusted Rand Index) are two commonly used metrics to measure the similarity between the clustering results and the true labels. The following will introduce their definitions and calculation formulas in detail:

[0064] 1. NMI (Normalized Mutual Information)

[0065] NMI is a metric used to evaluate the quality of clustering. It is based on the concept of information theory and measures the normalized value of the mutual information between the clustering results and the true labels.

[0066] The calculation formula of NMI is as follows: where \(C\) represents the set of clusters in the clustering results, \(|C|\) represents the number of clusters, \(K\) represents the number of classes in the true labels, \(n_{ij}\) ij represents the number of samples that belong to both cluster \(i\) and true label \(j\), \(n_{i\cdot}\) i· represents the number of samples in cluster \(i\), \(n_{\cdot j}\) ·j represents the number of samples in true label \(j\), and \(N\) represents the total number of samples.

[0067] 2. ARI (Adjusted Rand Index)

[0068] ARI is a metric used to compare the similarity between two clustering results, taking into account the consistency between pairs of samples in the clustering results.

[0069] The calculation formula of ARI is as follows:

[0070] where \(a_{l}\) i represents the number of pairs of samples that belong to the same cluster in the \(l\)-th clustering result, \(b_{j}\) j represents the number of pairs of samples that belong to the same class in the \(j\)-th true label, \(n_{ij}\) ij represents the number of pairs of samples that are assigned to both the \(i\)-th clustering result and the \(j\)-th true label, \(\binom{n}{k}\) represents the combination number of choosing \(k\) elements from \(n\) elements, and \(N\) is the total number of samples. The above two metrics are both values between 0 and 1, and the higher the value, the higher the similarity between the clustering results and the true labels.

[0071] Figure 3 Figure shows the comparison of clustering analysis evaluation metrics between the method of the present invention (scASDC) and six other baseline methods on four datasets. Each method is run 10 times and the average value is taken. It can be seen from Figure 3 this that the method of the present invention achieves the optimal or sub-optimal performance in the clustering of the four datasets.

[0072] Figure 4Schematic diagram for comparing the results of dimensionality reduction visualization of the final embedded representation using UMAP for the method of the present invention and other baseline methods on the QSLimb Muscle dataset. The results show that compared with traditional single-cell RNA sequencing data analysis methods, the method of the present invention exhibits stronger ability and excellent performance in cell clustering and has significant robustness.

[0073] The single-cell RNA sequencing data clustering method in the embodiments of the present invention is described above. Next, the single-cell RNA sequencing data clustering device in the embodiments of the present invention will be described. Please refer to Figure 5 One embodiment of the single-cell RNA sequencing data clustering device in the embodiments of the present invention includes:

[0074] The first data processing module 10 is used to obtain single-cell RNA sequencing data, preprocess the single-cell RNA sequencing data, and obtain a single-cell RNA gene expression matrix X;

[0075] The second data processing module 20 is used to construct a cell-cell graph adjacency matrix A using the KNN method according to the single-cell RNA gene expression matrix X;

[0076] The construction module 30 is used to construct a ZINB-based autoencoder module and a graph autoencoder module;

[0077] The fusion module 40 is used to layer-connect the ZINB-based autoencoder module and the graph autoencoder module through an attention fusion mechanism, and effectively fuse the feature representation of gene expression information and cell structure information into the same representation through this layer-by-layer embedding operation;

[0078] The training module 50 is used to integrate the ZINB-based autoencoder module and the graph autoencoder module into a unified framework through a self-supervised strategy, and perform end-to-end clustering training and synchronous optimization on these two modules based on the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A to obtain a clustering prediction model;

[0079] The clustering module 60 is used to process the single-cell RNA sequencing data to be clustered to obtain a single-cell RNA gene expression matrix X and a cell-cell graph adjacency matrix A to be clustered, input them into the clustering prediction model, output an embedded representation, and obtain the predicted cell categories according to the embedded representation to complete clustering.

[0080] Above Figure 5 The single-cell RNA sequencing data clustering device in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the single-cell RNA sequencing data clustering device in the embodiments of the present invention will be described in detail from the perspective of hardware processing.

[0081] Figure 6 This is a schematic structural diagram of a single-cell RNA sequencing data clustering device provided by an embodiment of the present invention. The single-cell RNA sequencing data clustering device 100 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 11 (for example, one or more processors) and a memory 12, and one or more storage media 13 for storing application programs 133 or data 132 (for example, one or more mass storage devices). Among them, the memory 12 and the storage media 13 may be transient storage or persistent storage. The program stored in the storage media 13 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the single-cell RNA sequencing data clustering device 100. Further, the processor 11 may be configured to communicate with the storage media 13 and execute a series of instruction operations in the storage media 13 on the super-resolution gene expression map prediction device 100.

[0082] The single-cell RNA sequencing data clustering device 100 may further include one or more power supplies 14, one or more wired or wireless network interfaces 15, one or more input / output interfaces 16, and / or one or more operating systems 131, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 6 The device structure shown does not limit the single-cell RNA sequencing data clustering device 100, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0083] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium, and when the instructions are run on a computer, the computer is caused to execute the steps of the single-cell RNA sequencing data clustering method.

[0084] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system or device and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0085] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0086] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A single-cell RNA sequencing data clustering method, characterized in that: Includes steps: Acquire single-cell RNA sequencing data, and preprocess the single-cell RNA sequencing data to obtain a single-cell RNA gene expression matrix X; According to the single cell RNA gene expression matrix X, a cell-cell graph adjacency matrix A is constructed using a KNN method; Construct a ZINB-based self-encoding module and a graph self-encoding module; The ZINB-based autoencoder module and the graph autoencoder module are connected layer by layer through an attention fusion mechanism, and the feature representation of gene expression information and cell structure information are effectively fused into the same representation through this layer-by-layer embedding operation; The ZINB-based autoencoder module and the graph autoencoder module are integrated into a unified framework through a self-supervised strategy, and the two modules are end-to-end clustered trained and synchronously optimized based on the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A to obtain a clustering prediction model; After processing the single-cell RNA sequencing data to be clustered, the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A to be clustered are obtained, which are input into the clustering prediction model, and an embedded representation is output. According to the embedded representation, the predicted cell category is obtained and clustering is completed; Among them, building a ZINB-based self-encoding module includes the following steps: The autoencoder module is composed of a mutually symmetrical encoder and decoder. The encoder of the autoencoder module has L layers. The single-cell RNA gene expression matrix X is used as the initial input of the autoencoder module. The input of the lth layer is H l-1 , then its output is expressed as: H l =φ(W l H l-1 +b l ), where φ is the activation function, W l represents the weight of the lth layer, b l Represents the bias vector of the lth layer; The last layer of the decoder is connected to three independent fully connected layers, which are used to estimate the three parameters of ZINB: probability parameter π, dispersion parameter θ and mean parameter μ, so as to integrate the ZINB model into the autoencoder module. The ZINB distribution uses the three parameters to simulate the single-cell RNA sequencing data and reconstruct its data distribution: ZINB(X|π,μ,θ)=πδ0(X)+(1-π)NB(X|μ,θ); The final loss function of the ZINB-based autoencoder module is defined as the negative log-likelihood estimate of the ZINB distribution: The steps to build a graph autoencoder module include: The graph autoencoder module is composed of several layers of graph convolutional networks as the backbone network; The graph autoencoder module has a total of L layers. The cell-cell graph adjacency matrix A is used as the initial input of the first layer of the graph autoencoder module. The input of the lth layer is Z l-1 , then its output is expressed as: Where I is the unit diagonal matrix, is the degree matrix, U l-1 Represents the weight of the l-1th layer, normalized adjacency matrix Z l-1 Propagate and get the new representation Z l ; The step of constructing the graph self-encoding module further includes: using the following two reconstruction errors as the loss function of the graph self-encoding module: in The output Z of the last layer of the graph autoencoder module is converted into L Convert to reconstructed adjacency matrix The steps of connecting the ZINB-based autoencoder module and the graph autoencoder module layer by layer through an attention fusion mechanism, and effectively fusing the feature representation of gene expression information and cell structure information into the same representation through this layer-by-layer embedding operation include: The output of the ZINB-based autoencoder module and the output of the graph autoencoder module are embedded in a layer-by-layer heterogeneous structure fusion through an attention fusion module to obtain an integrated representation that integrates structural information and content information: l-1 =F att (αH l-1 +(1-α)Z l-1 ) where α is an adjustable weight parameter, F att represents the attention fusion operation; Using integrated representation R l-1 As input to graph convolutional networks to learn high-level discriminative information and generate new representations The ZINB-based autoencoder module and the graph autoencoder module are integrated into a unified framework through a self-supervised strategy, and the two modules are end-to-end clustered trained and synchronously optimized based on the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A. The steps of obtaining the cluster prediction model include: Output of the intermediate layer of the ZINB-based autoencoder module Perform k-means clustering to obtain a set of initial cluster centers Where k is the number of clusters; Compute embedding representation using Student's t distribution and cluster centers The soft allocation between is as follows: where h i is the embedding representation The i-th sample of , λ represents the degrees of freedom of Student's t distribution, and the soft distribution q ij ∈Q, Q is the soft cluster distribution; On the basis of the soft cluster distribution Q, an auxiliary target distribution is set to supervise the learning of the soft cluster distribution. The clustering is improved by learning the high confidence distribution of the target distribution. The target distribution is calculated using the soft cluster frequency. The calculation method is as follows: in is the soft clustering frequency, the target distribution p ij ∈P; Use the KL divergence loss between the soft clustering distribution Q and the target distribution P as the optimization target to obtain higher quality clustering: The target distribution P is used to supervise the learning of the graph self-encoding module: Among them, z ij ∈Z pre , Through the above two loss functions, the optimization objective is integrated into a target distribution P, making the learned representation suitable for clustering tasks.

2. A single cell RNA sequencing data clustering device for implementing the method of claim 1, characterized in that: include: A first data processing module is used to obtain single-cell RNA sequencing data, pre-process the single-cell RNA sequencing data, and obtain a single-cell RNA gene expression matrix X; A second data processing module is used to construct a cell-cell graph adjacency matrix A using a KNN method according to the single-cell RNA gene expression matrix X; Construction module, used to construct ZINB-based autoencoding module and graph autoencoding module; A fusion module, used to connect the ZINB-based autoencoder module and the graph autoencoder module layer by layer through an attention fusion mechanism, and effectively fuse the feature representation of gene expression information and cell structure information into the same representation through this layer-by-layer embedding operation; A training module, used to integrate the ZINB-based autoencoder module and the graph autoencoder module into a unified framework through a self-supervised strategy, and perform end-to-end clustering training and synchronous optimization on the two modules based on the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A to obtain a clustering prediction model; The clustering module is used to process the single-cell RNA sequencing data to be clustered to obtain the single-cell RNA gene expression matrix X and the cell-cell graph adjacency matrix A to be clustered, input them into the clustering prediction model, output the embedded representation, obtain the predicted cell category according to the embedded representation and complete the clustering.

3. A single-cell RNA sequencing data clustering device, characterized in that: comprising a memory and at least one processor, wherein the memory has computer-readable instructions stored therein; The at least one processor calls the computer-readable instructions in the memory to execute the various steps of the single-cell RNA sequencing data clustering method as described in claim 1.

4. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the various steps of the single-cell RNA sequencing data clustering method as claimed in claim 1 are implemented.

Citation Information

Patent Citations

  • Single-cell RNA-seq data clustering method based on dual self-supervision

    CN114022693A

  • Protein SNO site prediction method of deep learning network fusing features

    CN117976035A