Lightweight optimal clustering selection method for scRNA-seq data

By using the MLP encoder and simplified graph convolution module in single-cell RNA sequencing data clustering, dimensionality reduction and feature matrix updates, combined with self-optimized clustering and clustering performance judgment, the problems of low clustering efficiency and poor effect in the existing technology are solved, and more efficient and accurate cell subpopulations are achieved.

CN120145083APending Publication Date: 2025-06-13GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510328867.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing single-cell RNA sequencing data clustering methods based on deep learning have problems such as inefficiency and poor clustering effects in dimensionality reduction and training strategies, resulting in erroneous division of cell subpopulations and inaccurate biological conclusions.

Method used

Multi-layer perceptron (MLP) encoder is used to reduce dimensionality layer by layer, and the feature matrix is ​​iteratively updated using the simplified graph convolution module. Combined with self-optimized clustering and cluster performance judgment, the round with the best clustering performance in multiple iterative updates is automatically selected to obtain more accurate clustering results.

Benefits of technology

It achieves more efficient model operation and lower hardware resource requirements, while ensuring the accuracy of clustering effects and better analyzing the complex biological characteristics of single-cell data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145083A_ABST
    Figure CN120145083A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight optimal clustering selection method for scRNA-seq data. The lightweight optimal clustering selection method comprises the following steps: 1, acquiring and preprocessing an scRNA-seq data matrix; 2, performing layer-by-layer dimensionality reduction on data by using a multi-layer perceptron encoder to obtain a low-dimensional feature matrix; and 3, constructing an undirected cell map, and enhancing and removing cell map noise through an NE network. And 4, reconstructing features by using graph convolution, inputting the low-dimensional feature matrix and the cell graph into a simplified graph convolution module, and updating the reconstructed feature matrix round by round. And 5, carrying out cell clustering on the characteristic matrix after each round of updating by using self-optimization clustering to obtain cell representation. And 6, judging and evaluating the clustering performance of the current round, storing an optimal clustering result, and carrying out feature reconstruction and self-optimization clustering of the next round. And 7, outputting and analyzing final optimal feature information. According to the method, scRNA-seq data can be well processed, low-dimensional embedded features can be extracted, tag prediction is more accurate, the requirement for hardware resources is low, and an efficient method is provided for analyzing biological characteristics of single cell data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of single-cell RNA sequencing data clustering in bioinformatics. Specifically, it relates to a lightweight optimal clustering selection method for scRNA-seq data. Background Art

[0002] With the rapid development of single-cell RNA sequencing (scRNA-seq) technology, it plays an increasingly important role in the field of biological research, providing a powerful tool for analyzing cell heterogeneity and functions at the single-cell level. ScRNA-seq can comprehensively analyze the transcriptome of a single cell, generating a large amount of complex data. Effective clustering analysis of these data to identify different cell subpopulations is a key step in exploring cell biological characteristics, understanding developmental processes, and disease mechanisms.

[0003] In previous studies, clustering methods based on deep learning have achieved certain results in scRNA-seq data processing, but many problems have also emerged. On the one hand, in the dimensionality reduction stage of traditional deep learning-based clustering methods, multiple rounds of autoencoder modules are often used. Since hundreds of rounds of training are required, it not only occupies a large amount of hardware resources, has extremely high requirements for the memory and processor performance of computing devices, but also has a long training time, greatly limiting the research efficiency. On the other hand, most existing clustering methods have defects in the training strategy. Usually, a fixed number of rounds of training is adopted, and the clustering results of the last round are output. However, in the actual training process of scRNA-seq data, as each round of training progresses, the feature matrix is in a dynamic update and change, and the final training result may not be the one with the best clustering effect among all rounds. This limitation makes the clustering results unable to fully reflect the true structure of the data, which may lead to misclassification of cell subpopulations and affect the accuracy of subsequent biological conclusions.

[0004] To overcome these challenges, there is an urgent need for a more efficient and accurate scRNA-seq data clustering method to meet the urgent needs of current biological research for single-cell data analysis. Summary of the Invention

[0005] The purpose of the present invention is to provide a lightweight optimal clustering selection method for scRNA-seq data. The present invention reduces the data dimension layer by layer through a multi-layer perceptron (MLP) encoder, uses a simplified graph convolutional module to iteratively update the feature matrix, and performs cell clustering on the updated feature matrix using self-optimizing clustering to obtain cell representations. The clustering performance is judged for each round of iterative update. By calculating the Davies-Bouldin index and the silhouette coefficient, the round with the best clustering performance among multiple rounds of iterative updates is determined, and the clustering information of the optimal round is saved to achieve a more accurate clustering effect.

[0006] The advantage of this method is that both the MLP autoencoder and the simplified graph convolutional network used are lightweight modules, with a short model running time and low requirements for hardware resources. The self-optimizing clustering and clustering performance discriminator ensure the clustering effect, select the optimal clustering information in multiple rounds of clustering updates, provide a new method for single-cell RNA sequencing clustering, and help to analyze the complex biological characteristics of single-cell data.

[0007] The present invention proposes a lightweight optimal clustering selection method for scRNA-seq data, and the main steps are as follows:

[0008] Step 1: Obtain and preprocess scRNA-seq data: Obtain the data matrix from the NCBI GEO database, screen the gene expression matrix, extract highly variable genes, perform logarithmic transformation and normalization to reduce noise and improve data quality.

[0009] Step 2: Feature dimensionality reduction and extraction: Use the MLP encoder to reduce the dimensionality of the data layer by layer, reduce the dimensionality of the data while maintaining the data information content, and obtain a low-dimensional feature matrix.

[0010] Step 3: Cell graph construction: Calculate the similarity matrix between cells, construct an undirected cell graph, and enhance and remove noise through the NE network.

[0011] Step 4: Graph convolutional reconstruction of features: Input the low-dimensional feature matrix and the cell graph into the simplified graph convolutional module to reconstruct the feature matrix and the cell graph.

[0012] Step 5: Self-optimizing clustering: Use self-optimizing clustering to perform cell clustering on the updated feature matrix to obtain cell representations.

[0013] Step 6: Clustering performance judgment: Evaluate the current clustering performance by calculating the Davies-Bouldin Index (DBI) and the Silhouette coefficient. When the clustering performance is better, save the current clustering result as the optimal clustering, and perform the next round of feature reconstruction and self-optimizing clustering.

[0014] Step 7: Output the optimal features: Perform clustering in multiple loops, evaluate the performance in each round of clustering. When the clustering performance index is better than the optimal clustering, update the optimal clustering. After the loop ends, output the optimal clustering result, and use t-SNE for clustering result analysis. Description of the Drawings

[0015] Figure 1 It is a schematic diagram of the method flow of the present invention.

[0016] Figure 2ARI index comparison chart between the present invention and other clustering methods.

[0017] Figure 3 NMI index comparison chart between the present invention and other clustering methods.

[0018] Figure 4 Comparison between predicted labels and true labels of different datasets in the t-SNE clustering visualization of the present invention. Detailed implementation manners

[0019] The present invention is a lightweight optimal clustering selection method for scRNA-seq data. The present invention will be further described in detail below in combination with specific implementation manners and simulation experiments. Those skilled in the art should understand that these implementation methods are only used to explain the technical principle of the present invention and are not intended to limit the scope of evidence collection of the present invention.

[0020] As Figure 1 shown, a lightweight optimal clustering selection method for scRNA-seq data specifically includes the following steps:

[0021] Preferably, the obtaining and preprocessing of scRNA-seq data described in step 1 are specifically as follows:

[0022] Six publicly available real datasets Goolam, Kolod, Diaphragm, Muraro, Kelin, and Limb_Muscle are obtained from the NCBI GEO database. All six datasets can be represented by indicating, and respectively representing the number of cells and genes. Quality control is performed on the original count matrix to filter out all cells or genes with zero counts; the top 2500 genes with the most specific gene expression are selected as the gene expression matrix ; then the obtained matrix is logarithmically transformed and fractionally normalized to obtain the normalized data matrix for subsequent feature extraction and graph construction.

[0023] Preferably, the feature extraction and dimensionality reduction described in step 2 are specifically as follows:

[0024] The dimension of the feature matrix is gradually compressed from high dimension to the target low-dimensional space through each hidden layer in the MLP. In each layer, the input features are linearly transformed and non-linearly activated and then lower-dimensional feature representations are output. In the single-cell RNA sequencing data clustering task, the MLP can effectively compress the high-dimensional gene expression matrix by layer-by-layer dimensionality reduction, retain key features and reduce noise interference.

[0025] For the input feature matrix , the process of reducing the dimension layer by layer through the MLP is as follows:

[0026]

[0027] Among them, is the feature matrix of the layer, is the feature dimension of the layer. and are the weight matrix and the bias term respectively, is the non-linear activation function.

[0028] Each layer of the MLP gradually compresses the feature dimension from to through linear transformation and non-linear activation, and finally outputs the feature matrix of the target dimension.

[0029] Specifically, the present invention adopts a three-layer MLP to gradually reduce the dimension of the feature matrix from to layer by layer, and the formula is as follows:

[0030]

[0031]

[0032]

[0033]

[0034] The obtained low-dimensional embedding matrix will be used in the simplified graph convolution module of step 4.

[0035] Preferably, the cell graph construction described in step 3 is specifically as follows:

[0036] Since there are a large number of dropouts and noises in scRNA-seq data, directly calculating the similarity based on the original data and constructing a cell graph may lead to an unreliable graph structure, which in turn affects the clustering effect. Therefore, the present invention adopts Network Enhancement (NE) to denoise the cell similarity matrix and construct a more reliable cell graph.

[0037] The specific steps are as follows:

[0038] The preprocessed gene expression matrix is used as the input for graph construction. In order to quantify the similarity between cells, the present invention adopts the Pearson correlation coefficient to calculate the initial similarity matrix . The calculation formula of the Pearson correlation coefficient is as follows:

[0039]

[0040] Among them, represents the expression level of gene in cell , is the average expression value of cell .

[0041] Introduce the Network Enhancement (NE) technology to denoise the similarity matrix , eliminate weak edges and enhance the actual connections between cells, and convert the initial similarity matrix into an enhanced similarity matrix . The specific steps are as follows:

[0042] First, perform NE processing on the similarity matrix to obtain a reweighted similarity matrix . Then, by setting a threshold t, perform denoising to obtain a denoised similarity matrix , and the formula is as follows:

[0043]

[0044] After NE processing, an enhanced similarity matrix is obtained, and the adjacency relationship matrix of the undirected cell graph is constructed using the K-Nearest Neighbor (KNN) algorithm.

[0045] Preferably, the graph convolution reconstruction features described in step 4 are specifically:

[0046] A simplified graph convolution model is used to update the feature matrix. The Simplified Graph Convolution Network (SGC) removes the non-linear activation and weight matrix, and only retains the core process of neighborhood information aggregation, thereby greatly improving the calculation efficiency while ensuring the clustering accuracy. The propagation formula for obtaining the feature matrix is as follows:

[0047]

[0048] Among them, is the adjacency matrix, is 's degree matrix, , is the The feature matrix of the layer directly aggregates the information of neighboring nodes onto the target node through the normalization operations of the adjacency matrix and the degree matrix, effectively reducing the computational complexity.

[0049] To better capture the potential associations between cells, the present invention integrates a multi-head attention mechanism to enhance the utilization of neighborhood information. The improved propagation formula is:

[0050]

[0051]

[0052] Among them, is the similarity matrix enhanced by NE denoising, represents the Hadamard product, is the identity matrix.

[0053] Preferably, the self-optimizing clustering described in step 5 is as follows:

[0054] Perform self-optimizing clustering on the feature matrix in each round. While clustering the cell labels, the self-optimizing soft clustering continuously iterates the soft clustering distributions Q and P, self-optimizes the cell labels, and improves the final clustering result. The soft assignment probability of the membership matrix Q is as follows:

[0055]

[0056] Among them, are different cells participating in the clustering assignment, is the clustering center. Based on the Q matrix, a more optimized membership matrix P is constructed. The formula for the optimized high-confidence distribution P is as follows:

[0057]

[0058] The membership matrix P is used as the target of Q to optimize and reassign the relationship structure diagram between cells, obtaining a new and better cell clustering.

[0059] Preferably, the clustering performance judgment described in step 6 is as follows:

[0060] The present invention designs an optimal clustering discrimination method. Instead of directly outputting the clustering result of the final round, a clustering performance evaluation is performed every time the feature matrix is reconstructed by graph convolution. By calculating the Davies-Bouldin index (DBI) and the Silhouette coefficient index of the current clustering result, it is judged whether the current round is the optimal clustering.

[0061] The specific method is as follows:

[0062] (1)Convolution and clustering round by round: Starting from the first round of convolution, perform graph convolution round by round, and cluster the updated feature matrix after each round.

[0063] (2)Performance evaluation: After each round of convolution, calculate the Davies-Bouldin index (DBI) and Silhouette coefficient index (silhouette coefficient) of the current clustering result.

[0064] DBI focuses on the average distance between samples and the centroid, while the silhouette coefficient measures the average distance between a sample and other samples. Its calculation formula is:

[0065]

[0066] where is the average distance from all samples in cluster to the centroid of this cluster (intra-cluster scatter), is the distance between the centroids of cluster and cluster (inter-cluster separation), is the number of clustering categories.

[0067] The silhouette coefficient evaluates the compactness and separation of clustering by calculating the relationship between the distance of a sample to samples in the same cluster and the distance of the sample to the nearest sample in a different cluster. Its calculation formula is:

[0068]

[0069] where is the average distance from sample to other samples in the same cluster, is the average distance from sample to the nearest sample in a different cluster.

[0070] Finally, the silhouette coefficients of all samples are averaged to obtain the average silhouette coefficient of the entire dataset :

[0071]

[0072] Since the larger it is, the better the clustering effect, while it is the opposite for , so we set the clustering scoring metric of the model as

[0073]

[0074] represents the clustering score of the th round, , Similarly, when the value is larger, it indicates a better clustering effect.

[0075] (3) Optimal clustering judgment: Specifically, in the first round, , is the optimal score, and the clustering information of the first round (including the feature matrix, clustering labels, etc.) is saved; entering the second round, if , then save and replace the clustering information of the second round as the optimal clustering information, and regard the second round as the optimal clustering, ; if , then still regard the information of the first round as the optimal clustering information and enter the next round. Perform 30 rounds of graph convolution updates in this way, output the saved optimal clustering information, and obtain the optimal clustering result. In this way, the model can automatically judge and output the optimal clustering. Even if the clustering effect deteriorates due to too many rounds, the output is still the clustering result of the optimal clustering round before the decline.

[0076] Preferably, the output of the optimal features and analysis are described in step 7. Specifically:

[0077] Output the optimal features and optimal clustering information obtained in step 6, and use t-distributed stochastic neighbor embedding (t-SNE) to project the optimal feature matrix and optimal prediction labels into a two-dimensional space to visually display the clustering structure between cells and perform clustering result analysis.

[0078] To measure the accuracy of the clustering result, the present invention is verified from three aspects: the ARI index, the NMI index, and clustering visualization.

[0079] The technical effects of the present invention are further described below through experimental verification:

[0080] 1. Experimental conditions and content:

[0081] The experiments of the present invention are all completed on a 2.50GHz CPU, NVIDIA GeForce RTX 4060, and the windows11 operating system.

[0082] 2. Experimental result analysis:

[0083] Evaluations conducted on six real scRNA-seq datasets consistently show that, compared with existing clustering methods, the clustering method of the present invention exhibits superior performance, can better process scRNA-seq data, has more accurate predicted labels, and helps to analyze the complex biological characteristics of single-cell data.

[0084] The experimental comparison results of the present invention are as Figures 2 to 4 shown.

[0085] The above description is only a specific example of the present invention and does not constitute any limitation to the present invention. Obviously, for professionals in the field, understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention, and all of them should be included in the protection scope of the claims of the present invention.

Claims

1. A lightweight optimal clustering selection method for scRNA-seq data, characterized in that: The following steps are involved: Step 1: Obtain scRNA-seq data and preprocess: Obtain data matrix from NCBI GEO database, screen, extract high-variable, logarithmically transform and normalize the gene expression matrix to reduce noise and improve data quality; Step 2: Feature dimensionality reduction and extraction: Use the multi-layer perceptron encoder to reduce the dimensionality of the data layer by layer, while maintaining the information content of the data, and reduce the dimensionality of the data to obtain a low-dimensional feature matrix; Step 3: Cell graph construction: Calculate the similarity matrix between cells, construct an undirected cell graph, and remove noise through NE network enhancement; Step 4: Graph convolution to reconstruct features: Input the low-dimensional feature matrix and the cell map into the simplified graph convolution module to reconstruct the feature matrix; Step 5: Self-optimizing clustering: Use self-optimizing clustering to perform cell clustering on the feature matrix after each round of update to obtain cell representation; Step 6: Clustering performance judgment: Evaluate the clustering performance of the current round by calculating the Davies-Bouldin Index (DBI) and the Silhouette coefficient. When the clustering performance is better, save this clustering result as the optimal clustering and proceed to the next round of feature reconstruction and self-optimization clustering.

2. A lightweight optimal clustering selection method for scRNA-seq data according to claim 1, characterized in that: Step 1 is as follows: Step 1-1: Obtain scRNA-seq data from the NCBI GEO database and name it ,in represents the scRNA-seq data matrix, and denote the number of cells and genes, respectively; Step 1-2: Raw count matrix Perform quality control and filter out all cells or genes with zero counts; then screen highly variable genes and extract the top 2000 most specific genes as the gene expression matrix ; Step 1-3: Logarithm transformation and fraction normalization are performed on the obtained matrix to obtain the normalized data matrix , used for subsequent feature extraction and graph construction.

3. A lightweight optimal clustering selection method for scRNA-seq data according to claim 1, characterized in that: Step 2 is as follows: A multi-layer perceptron (MLP) is used to process highly sparse and over-dispersed gene expression data, extract denoised embeddings, and capture low-dimensional gene expression features; the input feature matrix , through three layers of MLP layer by layer to reduce the dimension .

4. A lightweight optimal clustering selection method for scRNA-seq data according to claim 1, characterized in that: Step 3 is as follows: Calculate the initial similarity matrix between cells based on the Pearson Correlation Coefficient , and introduces the network enhancement (NE) technology to denoise the similarity matrix, eliminate weak edges and enhance the actual connection between cells, and transform the initial similarity matrix into an enhanced similarity matrix ; Use KNN algorithm to construct the adjacency matrix of undirected cell graph .

5. A lightweight optimal clustering selection method for scRNA-seq data according to claim 1, characterized in that: Step 4 is as follows: A simplified graph convolution model is used to update the feature matrix. The simplified graph convolution network (SGC) removes nonlinear activation and weight matrices and only retains the core process of neighborhood information aggregation, thereby greatly improving computational efficiency while ensuring clustering accuracy. The propagation formula is as follows: ; in, is the adjacency matrix, yes The degree matrix of , It is The feature matrix of the layer directly aggregates the information of the neighboring nodes to the target node through the normalization operation of the adjacency matrix and the degree matrix, which effectively reduces the computational complexity; In order to better capture the potential correlation between cells, the present invention integrates a multi-head attention mechanism to enhance the utilization of neighborhood information; the improved propagation formula is: ; ; in, is the similarity matrix enhanced by NE denoising, represents the Hadamard product, is the identity matrix.

6. A lightweight optimal clustering selection method for scRNA-seq data according to claim 1, characterized in that: Step 5 is as follows: For each round of feature matrix Perform self-optimizing clustering; while clustering cell labels, self-optimizing soft clustering continuously iterates the soft clustering distribution Q and P, optimizes cell labels by itself, and improves the final clustering results; Soft assignment probability of membership matrix Q The formula is as follows: ; in, For different cells participating in clustering assignment, is the cluster center; a more optimized membership matrix P is constructed based on the Q matrix, and the optimized high confidence distribution P formula is as follows: ; The membership matrix P is used as the target of Q to optimize and redistribute the relationship structure diagram between cells to obtain new and better cell clustering.

7. A lightweight optimal clustering selection method for scRNA-seq data according to claim 1, characterized in that: Step 6 is as follows: The present invention designs an optimal clustering discrimination method, which does not directly output the clustering result of the final round, but performs a clustering performance evaluation each time the feature matrix is ​​reconstructed by graph convolution, and judges whether the current round is the optimal clustering by calculating the Davies-Bouldin index (DBI) and Silhouette coefficient of the current clustering result; and performs a clustering score each time the feature matrix is ​​reconstructed by graph convolution to judge whether the current round is the optimal clustering. If so, the clustering information of the current round is saved and replaced as the optimal clustering information, and the next round is performed. If not, the clustering information of the current round is discarded and the next round is entered until the preset number of rounds is completed.

8. A lightweight optimal clustering selection method for scRNA-seq data according to claim 1, characterized in that: Step 7 is as follows: After all the training rounds are completed, the optimal clustering information is finally saved and output. The optimal clustering features and the optimal prediction labels are projected into two-dimensional space using t-distributed stochastic neighbor embedding (t-SNE) to intuitively display the clustering structure between cells and analyze the clustering results.

Citation Information

Cited By

  • Power equipment operation and maintenance method and system fusing inspection data

    CN121301894A