An Asymmetric Autoencoder-based Single-cell Distillation Discriminant Clustering Method, System, Electronic Device and Medium

Through asymmetric autoencoder training and domain alignment technology, the problem of insufficient clustering accuracy and robustness in single-cell RNA sequencing data analysis is solved, and efficient clustering and domain adaptability on complex data sets are achieved.

CN119724374BActive Publication Date: 2025-07-29QUFU NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411781657.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-07-29
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

The existing single-cell RNA sequencing data analysis methods are insufficient in clustering accuracy and robustness when processing sparseness, noise, high-dimensional and cross-domain data sets, especially when dealing with complex samples with strong heterogeneity, which are prone to misclassification.

Method used

A single-cell distillation discriminant clustering method based on asymmetric autoencoder is adopted. By training asymmetric autoencoder, the source data label information and characteristic distance information are used to gather similar cells and different cell clusters away, and the clustering center is calculated through the low-dimensional representation of source data and target data, so that cells near the boundary of the cell cluster are approached to the clustering center, implicitly completing the domain alignment task.

Benefits of technology

It improves the accuracy and robustness of single-cell data clustering, and can more accurately infer the complex cell composition and functional characteristics of cell tissues, adapting to the stable clustering performance under different data sets and labeling situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724374B_ABST
    Figure CN119724374B_ABST
Patent Text Reader

Abstract

The present invention provides a single-cell distillation discriminant clustering method, system, electronic device and medium based on an asymmetric autoencoder, belonging to the technical field of single-cell clustering, and comprising the following steps: Step S1, training the original data based on the asymmetric autoencoder; wherein the original data includes source data and target data; Step S2, using the source data label information and the feature distance information to make the cells of the same class gather and the different cell clusters move away; Step S3, calculating the clustering center by using the low-dimensional representations of the source data and the target data, so that the cells near the boundary of the cell cluster move closer to the clustering center, implicitly completing the domain alignment task. By adopting the above single-cell distillation discriminant clustering method, system, electronic device and medium based on an asymmetric autoencoder, the present invention performs excellently in cell clustering performance and helps to infer the complex cell composition and functional characteristics inside the cell tissue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of single-cell clustering, and in particular, to a single-cell distillation discriminant clustering method, system, electronic device and medium based on an asymmetric autoencoder. Background Art

[0002] Single-cell RNA sequencing (scRNA-seq) technology can measure the gene expression levels of individual cells, providing extremely high cell differential resolution for researchers to better study the cell heterogeneity within tissues. In the past decade, scRNA-seq has been widely applied in multiple research fields such as oncology, microbiology, and developmental biology. Cell clustering is a key step in scRNA-seq data analysis, which groups cells according to their gene expression patterns and further completes tasks such as cell function recognition and cell type annotation. However, the sparsity, batch effect, and high dimensionality of scRNA-seq data pose great challenges to cell clustering. Technical biases lead to low RNA capture rates, extremely sparse scRNA-seq transcriptional expression profiles, and a large number of dropout events, increasing the noise and missing values in the data, thus affecting the accuracy of cell clustering analysis. At the same time, due to differences in tissues in different batches, different times, different operators, etc., the measured gene expression levels are different, resulting in errors in data analysis. In addition, scRNA-seq data usually has extremely high dimensionality, which makes it very complex to extract cell type-specific information and describe the characteristic relationships between cells, further increasing the difficulty of cell clustering analysis.

[0003] To overcome the above challenges, researchers have proposed many traditional unsupervised clustering methods for single-cell data in the early stage. These methods mainly learn important features from the original data through dimensionality reduction techniques, then calculate the similarity or distance between cells in different ways, and use clustering methods such as K-means and Louvain to complete the clustering task. SC3 generates multiple cell similarity matrices through different feature extraction methods, then uses different clustering methods to cluster each similarity matrix, and obtains the final stable clustering through consensus clustering. SIMLR initializes the similarity matrix based on methods such as Euclidean distance and Gaussian kernel after reducing the dimensionality of single-cell data using common dimensionality reduction methods, and learns the local and global information of the relationships between cells through an alternating optimization strategy to achieve the final clustering. CIDR imputes the missing gene expression values to reduce the impact of data sparsity, then performs principal component analysis on the imputed data, and uses hierarchical clustering to identify cell subpopulations. Although these methods can effectively reveal the diversity between cells, due to the lack of prior knowledge, the clustering results are less robust to data noise and high-dimensional sparsity, and are prone to misclassification when dealing with complex samples with strong heterogeneity.

[0004] To make up for the deficiencies of unsupervised methods in clustering accuracy, many research teams have introduced semi-supervised learning strategies to guide clustering by leveraging a small amount of well-annotated cell type information. Single-cell semi-supervised clustering methods can be divided into two categories: machine learning-based and deep learning-based. Machine learning-based single-cell clustering methods usually use some known cell labels to train machine learning models and then predict the target data based on the trained models. Scmap compares the gene expression patterns of target data and reference data using multiple similarity functions to find the most similar cell types, thus achieving the final clustering. Seurat 3.0 constructs a shared network using reciprocal nearest neighbors to capture the local neighborhood relationships between cells and then uses the Louvain algorithm to identify cell subsets. Such methods can effectively improve the ability to identify rare cell types and enhance the robustness to noise. However, the dependence of such methods on labeled data makes them vulnerable to domain shift when the number of labels is insufficient or the label distribution differs significantly from the target data, resulting in a decline in the model's generalization ability.

[0005] With the rise of deep learning technology, researchers have combined deep representation learning with semi-supervised strategies and proposed semi-supervised deep clustering methods. These methods use the non-linear mapping ability of deep neural networks for feature extraction, which can capture the complex expression patterns of scRNA-seq data, thereby improving clustering performance. scDeepCluster effectively captures the non-linear structure in the data and denoises the data using an autoencoder based on the zero-inflated negative binomial (ZINB) model, and then uses deep embedded clustering to move data points closer to the cluster centers, thus completing the clustering task. ItClust learns the features of the data through an autoencoder and constructs a similarity matrix between cells using a similarity metric. Finally, by combining iterative optimization and graph-based clustering strategies, it can effectively identify cell subpopulations (achieve high-quality cell clustering) in complex high-dimensional data. scSemiAAE combines the ideas of semi-supervised learning and adversarial autoencoders to achieve better feature learning and clustering effects. scSemiCluster adds structural similarity regularization to the reference data to constrain the clustering results of the target data and introduces pairwise constraints to enhance the compactness of clustering. scMCKC, based on the zero-inflated negative binomial model autoencoder, considers the association between similar cells, constructs cell-level constraints, and uses prior information to construct pairwise constraints, and uses a weighted soft K-means algorithm for final clustering. scDECL introduces a hybrid data augmentation strategy and interpolation loss to improve the diversity of the dataset and the robustness of the model, and transforms prior information into enhanced pairwise constraints to guide clustering. scTPC is based on a zero-inflated negative binomial distribution pre-trained denoising autoencoder. Then it uses triple constraints and pairwise constraints generated by partially labeled cells for deep clustering. Although these semi-supervised deep clustering methods can improve the clustering effect to a certain extent, they still face the problem of inconsistent distributions between the source domain and the target domain when dealing with cross-domain datasets. In addition, most existing methods rely on complex feature alignment strategies and ignore the in-depth exploration of the internal structure of target domain data, resulting in insufficient discriminative ability of target data. Summary of the Invention

[0006] The object of the present invention is to provide a single-cell distillation discriminative clustering method, system, electronic device and medium based on an asymmetric autoencoder, which performs excellently in cell clustering performance and helps to infer the complex cell composition and functional characteristics inside the cell tissue.

[0007] To achieve the above object, the present invention provides a single-cell distillation discriminative clustering method based on an asymmetric autoencoder, including the following steps:

[0008] Step S1, training the original data based on an asymmetric autoencoder; wherein, the original data includes source data and target data;

[0009] Step S2: Utilize the source data label information and feature distance information to cluster similar cells and separate different cell clusters;

[0010] Step S3: Calculate the clustering centers using the low-dimensional representations of the source data and the target data, and make the cells near the boundaries of the cell clusters approach the clustering centers, implicitly completing the domain alignment task.

[0011] Preferably, in step S1, training the original data based on an asymmetric autoencoder includes the following steps:

[0012] Step S11: Input the gene expression matrix X = {x1, x2,..., x n} ∈ R n×m of the original data into the three parameter matrices for estimating the ZINB distribution by the asymmetric autoencoder, including the zero-inflation parameter matrix Π = {π1, π2,..., π n} ∈ R n×m , the mean parameter matrix Μ = {μ1, μ2,..., μ n} ∈ R n×m and the dispersion parameter matrix Θ = {θ1, θ2,..., θ n} ∈ R n×m ;

[0013] where n represents the number of samples of all data; m represents the dimension of the input data; x n represents the gene expression vector of the nth cell; π n represents the zero-inflation parameter vector corresponding to the nth cell; μ n represents the mean parameter vector corresponding to the nth cell; θ n represents the dispersion parameter vector corresponding to the nth cell;

[0014] Step S12: Let x ij represent the expression value of the jth gene in the ith cell, and π ij , μ ij , θ ij represent the zero-inflation parameter, mean parameter, and dispersion parameter corresponding to the jth gene in the ith cell respectively, and make it follow the ZINB distribution:

[0015] p ZINB (x ij ; π ij , μ ij , θ ij ) = π ij δ0(x ij ) + (1 - π ij )p NB (x ij ; μ ij , θ ij );

[0016]

[0017] Among them, Γ(·) is the Gamma function; δ0(·) is the Dirac function, and when x ij = 0, δ0(x ij ) takes the value of 1, and when x ij ≠ 0, δ0(x ij ) takes the value of 0;

[0018] Step S13: Reconstruct the loss of the data according to the negative log-likelihood function of the ZINB distribution:

[0019]

[0020] Among them, L zinb represents the reconstruction loss based on the ZINB distribution; p ZINB (x ij ) represents the ZINB distribution function;

[0021] Step S14: Calculate the reconstructed data according to the mean parameter matrix M and the size factor vector s The calculation formula is as follows:

[0022]

[0023] Among them, s is the size factor vector corresponding to each cell, and the calculation method is the total gene expression of cell i divided by the median of the total gene expression of all cells; diag(·) represents the function of converting a vector into the diagonal elements in a square matrix;

[0024] For each sample x i , construct a positive pair and 2n - 2 negative pairs

[0025]

[0026] Step S15: Calculate the contrast loss according to the constructed positive pairs and negative pairs:

[0027]

[0028] Among them, represents the contrast loss constructed from the original data of the i-th cell; represents the contrast loss constructed from the reconstructed data of the i-th cell; L con represents the overall contrast loss; x a represents the gene expression vector of the a-th cell; sim(·) is the cosine similarity function,

[0029] Preferably, in step S2, using the source data label information and the feature distance information to cluster the same type of cells and separate different cell clusters includes the following steps:

[0030] Step S21: Add a classifier after the encoder in the asymmetric autoencoder, use the gene expression matrix of the source data as the input, and obtain the predicted probability matrix of the reference data from the classifier where n s represents the number of samples in the source data, K represents the number of cell types, represents the predicted probability that the i-th cell in the source data belongs to the k-th cell type; use the standard cross-entropy function as the classification loss function:

[0031]

[0032] where L ce represents the classification loss; represents the true label of cell i in the source data; I(·) represents the indicator function;

[0033] Step S22: Use the gene expression matrix of the source data as the input, and obtain the feature representation of the source data through the asymmetric autoencoder Calculate the cluster center according to the label information of the source data:

[0034]

[0035] where φ k represents the cluster center of cell type k; represents the low-dimensional feature vector of the i-th cell in the source data; Φ = {φ1, φ2,..., φ K} ∈ R K×d represents the cluster center matrix; d represents the dimension of the low-dimensional representation; the cluster center is updated in each iteration of the network training, and the specific formula is as follows:

[0036]

[0037] where represents the cluster center of cell type k in the epoch-th iteration, represents the cluster center of cell type k in the (epoch - 1)-th iteration; represents the sample set of cell type k in the source data, α ∈ [0, 1] represents the moving average coefficient; represents the low-dimensional feature vector of the a-th cell in the source data;

[0038] Step S23: Use the standard cross-entropy as the loss function for distilled clustering, and the specific formula is as follows,

[0039]

[0040] Among them, τ represents the temperature coefficient; represents the distance vector between the low-dimensional representation of cell i in the source data and the cluster center; represents the cluster center φ in the source data k and the distance vector between other cluster centers.

[0041] Preferably, in step S3, the cluster centers are calculated using the low-dimensional representations of the source data and the target data, so that the cells near the cell cluster boundary move closer to the cluster centers, implicitly completing the domain alignment task, including the following steps:

[0042] Step S31: Use the source data and the target data as the input of the asymmetric autoencoder to obtain the low-dimensional features of the source data and the low-dimensional features of the target data and their corresponding pseudo-labels Combined with the true labels of the source data, calculate the final cluster centers:

[0043]

[0044] Among them, c k represents the cluster center of cell type k in the fine-tuning stage, C = {c1, c2,..., c K} are the cluster centers of the source data and the target data in the feature space; n s represents the number of cells in the source data; n t represents the number of cells in the target data, n = n s + n t ; represents the low-dimensional feature vector of the i-th cell in the source data; represents the low-dimensional feature vector of the i-th cell in the target data;

[0045] Step S32: Use the t-distribution to calculate the similarity between the low-dimensional feature Z and the cluster center C, that is, the soft clustering assignment probability:

[0046]

[0047] Among them, β represents the degree of freedom of the t-distribution; represents the soft clustering assignment probability distribution; represents the soft clustering assignment probability that the i-th cell belongs to the k-th cell type; z i represents the low-dimensional feature vector of the i-th cell; c l represents the l-th cluster center; Use cells with high confidence to improve the clustering quality, and define an auxiliary probability distribution based on the distribution of the soft clustering assignment probability:

[0048]

[0049] Among them, represents the soft clustering assignment probability that cell i belongs to cell type l; represents the auxiliary probability distribution, represents the auxiliary probability that cell i belongs to cell type k;

[0050] Step S33: Use the standard cross-entropy function as the clustering loss:

[0051]

[0052] where L dec represents the clustering loss.

[0053] A single-cell distillation discriminant clustering system based on an asymmetric autoencoder, comprising:

[0054] A contrast pre-training module for training the original data based on an asymmetric autoencoder;

[0055] A distillation discriminant clustering module for using the source data label information and feature distance information to cluster the same type of cells together and separate different cell clusters;

[0056] A deep embedding clustering fine-tuning module for calculating the clustering centers using the low-dimensional representations of the source data and the target data, and making the cells near the boundaries of the cell clusters approach the clustering centers to implicitly complete the domain alignment task.

[0057] A computer device, comprising: a memory and a processor; the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned single-cell distillation discriminant clustering method based on an asymmetric autoencoder are implemented.

[0058] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned single-cell distillation discriminant clustering method based on an asymmetric autoencoder are implemented.

[0059] Therefore, the present invention adopts the above-mentioned single-cell distillation discriminant clustering method, system, electronic device and medium based on an asymmetric autoencoder, and the beneficial technical effects are as follows: Through the training of the asymmetric autoencoder, this method effectively captures the complex expression patterns of single-cell data, and combines the label information and feature distance information of the source data to cluster similar cells and separate different cell clusters. At the same time, the clustering centers are calculated using the low-dimensional representations of the source data and the target data, further optimizing the boundaries of the cell clusters and implicitly completing the domain alignment task. Under the combined action of these technical characteristics, this method performs excellently in cell clustering, helping to more accurately infer the complex cell composition and functional characteristics within the cell tissue, and providing strong support for the analysis of single-cell RNA sequencing (scRNA-seq) data. Description of the Drawings

[0060] Figure 1 It is a flowchart of a single-cell distillation discriminant clustering method based on an asymmetric autoencoder of the present invention;

[0061] Figure 2 It is a comparison of the ACC and ARI values of the scAADC method and other methods on a simulated dataset; among them, Figure 2 in (a) is a comparison of the ACC values of the scAADC method and other methods on the larger_balance simulated dataset; Figure 2 in (b) is a comparison of the ACC values of the scAADC method and other methods on the larger_imbalance simulated dataset; Figure 2 in (c) is a comparison of the ACC values of the scAADC method and other methods on the smaller_balance simulated dataset; Figure 2 in (d) is a comparison of the ACC values of the scAADC method and other methods on the smaller_imbalance simulated dataset; Figure 2 in (e) is a comparison of the ARI values of the scAADC method and other methods on the larger_balance simulated dataset; Figure 2 in (f) is a comparison of the ARI values of the scAADC method and other methods on the larger_imbalance simulated dataset; Figure 2 in (g) is a comparison of the ARI values of the scAADC method and other methods on the smaller_balance simulated dataset; Figure 2 in (h) is a comparison of the ARI values of the scAADC method and other methods on the smaller_imbalance simulated dataset;

[0062] Figure 3 It is a comparison of the ACC and ARI values of the scAADC method and other methods on a single dataset; among them,Figure 3 In (a), it is the comparison of the ACC values of the scAADC method and other methods on the Adam dataset; Figure 3 In (b), it is the comparison of the ACC values of the scAADC method and other methods on the Young dataset; Figure 3 In (c), it is the comparison of the ARI values of the scAADC method and other methods on the Adam dataset; Figure 3 In (d), it is the comparison of the ARI values of the scAADC method and other methods on the Young dataset;

[0063] Figure 4 It is the comparison of the ACC and ARI values of the scAADC method and other methods on cross-datasets; among them, Figure 4 In (a), it is the comparison of the ACC values of the scAADC method and other methods on cross-datasets; Figure 4 In (b), it is the comparison of the ARI values of the scAADC method and other methods on cross-datasets. Detailed implementation manner

[0064] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0065] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meaning understood by those of ordinary skill in the field to which the present invention belongs.

[0066] Embodiment 1

[0067] As Figure 1 shown, it is a flowchart of a single-cell distillation discriminant clustering method (scAADC) based on an asymmetric autoencoder of the present invention, including the following steps:

[0068] Step S1: Train the original data based on the asymmetric autoencoder; wherein, the original data includes source data and target data.

[0069] Step S11: Input the gene expression matrix X = {x1, x2,..., x n} ∈ R n×m of the original data into the three parameter matrices for estimating the ZINB distribution by the asymmetric autoencoder, including the zero-inflation parameter matrix Π = {π1, π2,..., π n} ∈ R n×m , the mean parameter matrix Μ = {μ1, μ2,..., μ n} ∈ R n×m and the dispersion parameter matrix Θ = {θ1, θ2,..., θ n} ∈ R n×m ;

[0070] where n represents the sample size of all data; m represents the dimension of the input data; x n represents the gene expression vector of the nth cell; π n represents the zero-inflation parameter vector corresponding to the nth cell; μ n represents the mean parameter vector corresponding to the nth cell; θ n represents the dispersion parameter vector corresponding to the nth cell;

[0071] Step S12: Let x ij represent the expression value of the jth gene in the ith cell, and π ij , μ ij , θ ij represent the zero-inflation parameter, mean parameter, and dispersion parameter corresponding to the jth gene in the ith cell respectively, and make it follow the ZINB distribution:

[0072] p ZINB (x ij ; π ij , μ ij , θ ij ) = π ij δ0(x ij ) + (1 - π ij )p NB (x ij ; μ ij , θ ij );

[0073]

[0074] where Γ(·) is the Gamma function; δ0(·) is the Dirac function, and when x ij = 0, δ0(x ij ) takes the value of 1, and when x ij ≠ 0, δ0(x ij ) takes the value of 0;

[0075] Step S13: Reconstruct the loss of the data according to the negative log-likelihood function of the ZINB distribution:

[0076]

[0077] where L zinb represents the reconstruction loss based on the ZINB distribution; p ZINB (x ij ) represents the ZINB distribution function;

[0078] Step S14: Calculate the reconstructed data according to the mean parameter matrix M and the size factor vector s The calculation formula is as follows:

[0079]

[0080] Among them, s is the size factor vector corresponding to each cell, and the calculation method is the total gene expression of cell i divided by the median of the total gene expression of all cells; diag(·) represents a function that converts a vector into the diagonal elements in a square matrix;

[0081] For each sample x i , construct a positive pair and 2n - 2 negative pairs

[0082]

[0083] Step S15: Calculate the contrastive loss according to the constructed positive pairs and negative pairs:

[0084]

[0085] Among them, represents the contrastive loss constructed from the original data of the i-th cell; represents the contrastive loss constructed from the reconstructed data of the i-th cell; L con represents the overall contrastive loss; x a represents the gene expression vector of the a-th cell; sim(·) is the cosine similarity function,

[0086] Step S2: Use the source data label information and feature distance information to cluster similar cells and separate different cell clusters.

[0087] Step S21: Add a classifier behind the encoder in the asymmetric autoencoder, use the gene expression matrix of the source data as the input, and obtain the predicted probability matrix of the reference data from the classifier Among them, n s represents the number of samples in the source data, K represents the number of cell types, represents the predicted probability that the i-th cell in the source data belongs to the k-th cell type; use the standard cross-entropy function as the classification loss function:

[0088]

[0089] Among them, L ce represents the classification loss; represents the true label of cell i in the source data; I(·) represents the indicator function;

[0090] Step S22: Use the gene expression matrix of the source data as the input, and obtain the feature representation of the source data through the asymmetric autoencoder Calculate the clustering center according to the label information of the source data:

[0091]

[0092] Among them, φ k represents the clustering center of cell type k; represents the low-dimensional feature vector of the i-th cell in the source data; Φ = {φ1, φ2,..., φ K} ∈ R K×d represents the clustering center matrix; d represents the dimension of the low-dimensional representation; the clustering center is updated in each iteration of network training, and the specific formula is as follows:

[0093]

[0094] Among them, represents the clustering center of cell type k in the epoch-th iteration, represents the clustering center of cell type k in the (epoch - 1)-th iteration; represents the sample set of cell type k in the source data, and α ∈ [0, 1] represents the moving average coefficient; represents the low-dimensional feature vector of the a-th cell in the source data;

[0095] Step S23: Use the standard cross-entropy as the loss function of the distilled clustering, and the specific formula is as follows,

[0096]

[0097] Among them, τ represents the temperature coefficient; represents the distance vector between the low-dimensional representation of cell i in the source data and the clustering center; represents the clustering center φ k in the source data and the distance vector between other clustering centers.

[0098] Step S3: Calculate the clustering center using the low-dimensional representations of the source data and the target data, so that the cells near the cell cluster boundary move closer to the clustering center, implicitly completing the domain alignment task.

[0099] Step S31: Use the source data and the target data as the input of the asymmetric autoencoder to obtain the low-dimensional features of the source data and the low-dimensional features of the target data

[0100]

[0101] Among them, c k represents the clustering center of cell type k in the fine-tuning stage, C = {c1, c2,..., c K} is the clustering center of the source data and the target data in the feature space; n s represents the number of cells in the source data; n t represents the number of cells in the target data, n = n s +n t ; represents the low-dimensional feature vector of the i-th cell in the source data; represents the low-dimensional feature vector of the i-th cell in the target data;

[0102] Step S32: Calculate the similarity between the low-dimensional feature Z and the clustering center C using the t-distribution, that is, the soft clustering assignment probability:

[0103]

[0104] where β represents the degree of freedom of the t-distribution; represents the soft clustering assignment probability distribution; represents the soft clustering assignment probability that the i-th cell belongs to the k-th cell type; z i represents the low-dimensional feature vector of the i-th cell; c l represents the l-th clustering center; Use cells with high confidence to improve the clustering quality, and define an auxiliary probability distribution based on the distribution of the soft clustering assignment probability:

[0105]

[0106] where, represents the soft clustering assignment probability that cell i belongs to cell type l; represents the auxiliary probability distribution, represents the auxiliary probability that cell i belongs to cell type k;

[0107] Step S33: Use the standard cross-entropy function as the clustering loss:

[0108]

[0109] where, represents the auxiliary probability that cell i belongs to cell type l; L dec represents the clustering loss.

[0110] To verify the effectiveness of the present invention, the performance of the present invention is tested from the aspect of cell type clustering below.

[0111] 1. Dataset and evaluation metrics

[0112] (1) Simulated dataset

[0113] Simulated datasets with different numbers of clusters, cluster sizes, and data sizes were constructed using the R package Splatter. Specifically, for different numbers of clusters, the simulated datasets were set to include different numbers of cell types {4, 5, 6, 7}. For the cluster size, two cases were considered: one is "balanced", that is, the number of samples for each cell type is the same; the other is "unbalanced", that is, the number of samples for each cell type decreases by a ratio of 0.8, that is, the number of cells in the latter cell type is 80% of the previous type. In terms of data size setting, considering the relative sample sizes of the source data and the target data, two control cases were constructed: one is "larger", that is, the number of cells in the source data is twice that of the target data; the other is "smaller", that is, the number of cells in the source data is only one-half of the target data. Based on the above configurations, a total of 16 types of simulated datasets were generated, and 10 independent datasets were constructed for each type of simulated dataset by setting different random seeds. Each dataset contains 6300 cells and 2500 genes, thus ensuring the stability of the experimental results and the reliability of the comparative analysis.

[0114] (2) Single dataset, cross-dataset

[0115] The specific information of the real data used for single datasets and cross-datasets is shown in Table 1.

[0116] Table 1 Information table of single datasets and cross-datasets

[0117] Dataset organ Celltypenum Cellnum Zeropercent Baronhuman Pancreas 13 8569 90.62% Enge Pancreas 6 2282 86.05% Muraro Pancreas 9 2122 73.02% 10X_Tongue Tongue 3 7538 86.71% Smart-seq2_Tongue Tongue 2 1350 79.61% 10X_Spleen Spleen 5 9552 94.34% Smart-seq2_Spleen Spleen 3 1697 92.19% 10X_Trachea Trachea 5 11269 93.66% Smart-seq2_Trachea Trachea 4 1350 85.48% Adam Kidney 8 3660 92.33% Young Kidney 11 5685 94.70%

[0118] (3) Evaluation metrics

[0119] In the experiment, the adjusted Rand index (ARI) and annotation accuracy (ACC) were selected as evaluation metrics to evaluate the clustering performance of the model. The clustering performance of the algorithm is related to ARI and ACC. The higher the values of ARI and ACC, the better the clustering performance of the model.

[0120] 2. Comparative analysis of experimental results

[0121] (a) Simulated datasets

[0122] Under different settings of simulated datasets, there are significant differences in the ACC and ARI values of different methods, such as Figure 2As shown in the box plot. As the number of clusters increases, the ACC and ARI values of all methods show a downward trend. However, scAADC always maintains the optimal performance, and the performance degradation is not obvious, indicating that this method can accurately identify the cell types of each cluster and shows extremely high robustness to different numbers of cell types. scAADC can outperform other methods under imbalanced data, indicating that it can effectively learn the representative features of each cluster in complex data distributions. When the sample size of the source data is smaller than that of the target data, the clustering performance of scAADC can still be better than other methods, effectively extracting the representative features of each cluster and maintaining a stable clustering effect. The experimental results show that scAADC can maintain excellent performance regardless of the data scale, the balance of cluster structure, or the change in the number of clusters, demonstrating its stability and robustness in processing complex single-cell data.

[0123] (b) Single dataset

[0124] Comparison of the ACC and ARI results of scAADC and other comparison methods (Seurat, scSemiCluster, ItClust, scmap, scSemiAAE, scMCKC, scDECL) with the change of the number of known labels on the Adam and Young datasets, as Figure 3 shown in the line chart in. It can be seen from the figure that in the Adam dataset, as the number of known labeled samples increases, the clustering accuracy of scAADC always remains at the highest level (except when the labeled samples are 3), and it slightly increases with the increase of the number of labels, indicating its robustness in processing scarce labeled data and its effective ability to capture potential features. In the Young dataset, we obtained similar results. The experimental results show that scAADC can better handle the situation of scarce labeled samples and maintain stable and efficient clustering performance in few-shot learning.

[0125] (c) Cross-dataset

[0126] Comparison results of the ACC and ARI of eight clustering methods on six groups of datasets, as Figure 4 shown in the histogram in. It can be seen from the figure that scAADC achieves the best clustering performance on different cross-datasets, indicating that this method has good robustness on different datasets. Generally speaking, scAADC performs excellently in clustering performance, especially showing higher robustness and stability when processing different tissue datasets.

[0127] Example 2

[0128] The present invention also provides a single-cell distillation discriminant clustering system based on an asymmetric autoencoder, including:

[0129] A contrast pre-training module for training the original data based on an asymmetric auto-encoder;

[0130] A distillation discriminative clustering module for using the source data label information and feature distance information to cluster the same type of cells and separate different cell clusters;

[0131] A deep embedding clustering fine-tuning module for calculating the clustering center using the low-dimensional representations of the source data and the target data, making the cells near the cell cluster boundary approach the clustering center, and implicitly completing the domain alignment task.

[0132] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0133] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0134] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer diskettes (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.

[0135] It should be noted that the content not elaborated in detail in the present invention is prior art and well-known to those skilled in the art.

[0136] Therefore, by adopting the above-mentioned single-cell distillation discriminant clustering method, system, electronic device and medium based on an asymmetric autoencoder, the present invention can effectively perform clustering discrimination on target domain data, significantly improving the accuracy and robustness of clustering.

[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. An asymmetric autoencoder-based single-cell distillation discriminant clustering method, characterized in that, It includes the following steps: Step S1: Training the original data based on an asymmetric autoencoder; wherein, the original data includes source data and target data; including: Step S11: Input the gene expression matrix of the original data into the three parameter matrices for estimating the ZINB distribution by the asymmetric autoencoder, including the zero-inflation parameter matrix , the mean parameter matrix and the dispersion parameter matrix ; Among them, represents the sample size of all data; represents the dimension of the input data; represents the gene expression vector of the th cell; represents the zero-inflation parameter vector corresponding to the th cell; represents the mean parameter vector corresponding to the th cell; represents the dispersion parameter vector corresponding to the Step S12: Set to represent the expression value of the -th gene in the -th cell. Let respectively represent the zero-inflation parameter, mean parameter, and dispersion parameter corresponding to the -th gene in the -th cell, and make it follow the ZINB distribution; Step S13: Reconstructing the loss of data according to the negative log-likelihood function of the ZINB distribution; Step S14. Calculate the reconstructed data according to the mean value parameter matrix and the size factor vector . The calculation formula is as follows: ; Among them, is the size factor vector corresponding to each cell, and the calculation method is to divide the total gene expression of the cell by the median of the total gene expression of all cells; represents a function that converts a vector into the diagonal elements in a square matrix; For each sample , construct a positive pair and negative pairs ; Step S15: Calculating the contrastive loss according to the constructed positive pairs and negative pairs: Step S2: Using the source data label information and feature distance information to cluster similar cells and separate different cell clusters, including: Step S21: Add a classifier after the encoder in the asymmetric autoencoder, take the gene expression matrix of the source data as input, and obtain the predicted probability matrix of the reference data from the classifier , , where represents the number of samples of the source data, represents the number of cell types, represents the predicted probability that the -th cell in the source data belongs to the -th cell type; Use the standard cross-entropy function as the classification loss function: ; Among them, represents the classification loss; represents the true label of cells in the source data ; represents the indicator function; Step S22: Taking the gene expression matrix of the source data as input, obtaining the feature representation of the source data through an asymmetric autoencoder, and calculating the clustering center according to the label information of the source data. The clustering center is updated in each iteration of network training; Step S23: Using standard cross-entropy as the loss function for distilled clustering; Step S3: Calculating the clustering center using the low-dimensional representations of the source data and target data, making the cells near the cell cluster boundary approach the clustering center, and implicitly completing the domain alignment task, including: Step S31: Taking the source data and target data as the input of the asymmetric autoencoder, obtaining the low-dimensional features of the source data and the low-dimensional features of the target data and their corresponding pseudo-labels, and combining the true labels of the source data to calculate the final clustering center; Step S32: Use distributed computing to calculate the low-dimensional features and the similarity between the clustering centers i.e., the soft clustering assignment probability; use cells with high confidence to improve the clustering quality, and define an auxiliary probability distribution based on the distribution of the soft clustering assignment probability: Step S33: Using the standard cross-entropy function as the clustering loss.

2. The single-cell distillation discriminant clustering method based on an asymmetric autoencoder according to claim 1, wherein In step S1, when training the original data based on an asymmetric autoencoder, it further includes the following steps: Step S12, set denote the expression value of the -th gene in the -th cell. respectively denote the zero-inflation parameter, mean parameter and dispersion parameter corresponding to the -th gene in the -th cell, and make it follow the ZINB distribution: ; ; Among them, is the Gamma function; is the Dirac function, when it takes the value of 1, and when it takes the value of 0; Step S13: Reconstructing the loss of data according to the negative log-likelihood function of the ZINB distribution: ; Among them, represents the reconstruction loss based on the ZINB distribution; represents the ZINB distribution function; Step S15: Calculating the contrastive loss according to the constructed positive pairs and negative pairs: ; ; ; Among them, represents the contrast loss of the original data structure of the th cell; represents the contrast loss of the reconstructed data structure of the th cell; represents the overall contrast loss; represents the gene expression vector of the th cell; is the cosine similarity function, , , , , , .

3. The single-cell distillation discriminant clustering method based on an asymmetric autoencoder according to claim 2, wherein In step S2, when using the source data label information and feature distance information to cluster similar cells and separate different cell clusters, it further includes the following steps: Step S22: Using the gene expression matrix of the source data as input, obtain the feature representation of the source data through an asymmetric autoencoder, and calculate the cluster centers according to the label information of the source data: , ; Among them, represents the cluster center with the cell type of ; represents the low-dimensional feature vector of the -th cell in the source data; represents the cluster center matrix; represents the dimension of the low-dimensional representation; the cluster center is updated in each iteration of network training, and the specific formula is as follows: ; Among them, represents that the cell type of the th iteration is the cluster center of ; represents that the cell type of the th iteration is the cluster center of ; represents the sample set in the source data where the cell type is ; represents the moving average coefficient; represents the low-dimensional feature vector of the nd cell in the source data; Step S23: Using standard cross-entropy as the loss function for distilled clustering, and the specific formula is as follows, ; Among them, represents the temperature coefficient; represents the distance vector between the low-dimensional representation of cells in the source data and the cluster center; represents the distance vector between the cluster center in the source data and other cluster centers.

4. The single-cell distillation discrimination clustering method based on an asymmetric autoencoder according to claim 3, wherein, In step S3, when calculating the clustering center using the low-dimensional representations of the source data and target data, making the cells near the cell cluster boundary approach the clustering center, and implicitly completing the domain alignment task, it includes the following steps: Step S31: Use the source data and the target data as the input of the asymmetric autoencoder to obtain the low-dimensional features of the source data and the low-dimensional features of the target data and their corresponding pseudo-labels , and combine with the true labels of the source data to calculate the final clustering centers: ; Among them, represents the clustering center of cell type in the fine-tuning stage, which is the clustering center of source data and target data in the feature space; represents the number of cells in the source data; represents the number of cells in the target data, ; represents the low-dimensional feature vector of the th cell in the source data; represents the low-dimensional feature vector of the th cell in the target data; Step S32. Use distributed computing to calculate the similarity between low-dimensional features and the clustering centers i.e., the soft clustering assignment probability: ; Among them, represents the degree of freedom of the distribution; represents the soft clustering assignment probability distribution; represents the th cell belonging to the th cell type's soft clustering assignment probability; represents the low-dimensional feature vector of the th cell; represents the th clustering center; High-confidence cells are used to improve the clustering quality, and an auxiliary probability distribution is defined based on the distribution of the soft clustering assignment probability: ; Among them, , represents the soft clustering assignment probability that a cell belongs to the cell type ; ; represents the auxiliary probability distribution, and represents the auxiliary probability that a cell belongs to the cell type Step S33: Using the standard cross-entropy function as the clustering loss: ; Among them, represents the clustering loss.

5. A single-cell distillation discriminant clustering system based on an asymmetric autoencoder, characterized in that For implementing the single-cell distilled discriminative clustering method based on an asymmetric autoencoder according to any one of claims 1-4, including: A contrastive pre-training module for training the original data based on an asymmetric autoencoder; A distilled discriminative clustering module for using the source data label information and feature distance information to cluster similar cells and separate different cell clusters; A deep embedding clustering fine-tuning module for calculating the clustering center using the low-dimensional representations of the source data and target data, making the cells near the cell cluster boundary approach the clustering center, and implicitly completing the domain alignment task.

6. A computer device, comprising: A memory and a processor; The memory stores a computer program, characterized in that when the processor executes the computer program, it implements the steps of the single-cell distilled discriminative clustering method based on an asymmetric autoencoder according to any one of claims 1-4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the single-cell distilled discriminative clustering method based on an asymmetric autoencoder according to any one of claims 1-4.

Citation Information

Patent Citations

  • Single-cell RNA sequencing clustering method based on adversarial autoencoder

    CN111785329A

  • Method for unsupervised identification of single-cell morphological profiling based on deep learning

    US20240112026A1