Gene state cancer clustering analysis method based on deep learning
Through the deep learning-based genetic cancer clustering analysis method, combined with grid-based preprocessing and deep-optimized pan-cancer clustering algorithm, the problems of outliers and batch effects in single-cell sequencing data analysis are solved, achieving higher clustering accuracy and batch effect reduction.
Patent Information
- Application Number
- CN202510148954.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-02-05
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-27
AI Technical Summary
There are outlier effects and batch effects problems in the analysis of existing single-cell sequencing data, resulting in reduced clustering accuracy.
A deep learning-based genetic cancer clustering analysis method is used, combined with grid-based preprocessing and deep optimization of pan-cancer clustering algorithm (G-DESC-E), to reduce the impact of outliers through grid segmentation and isolated point removal, and the objective function is constructed using KL divergence and label entropy to optimize the clustering results.
It improves the accuracy and stability of clustering, reduces the impact of batch effects, and significantly improves the recognition performance of key genes in pan-cancer.
Smart Images

Figure CN120220798A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics technology, and particularly relates to a method for predicting multi-organ drug-induced pathology based on deep learning. Background Art
[0002] The development of single-cell sequencing technology has opened up new perspectives for cancer research. It can explore the changes in gene expression levels at the single-cell level, thereby revealing the interactions between genes and the key genes closely related to cancer occurrence. However, traditional gene sequencing methods face significant technical challenges in in-depth research on cancer cell marker genes, including problems such as the high dimensionality, sparsity of data, and batch effects, which make the systematic analysis of cancer-related genes complex and resource-intensive.
[0003] In recent years, with the gradual maturity of single-cell sequencing technology, this technology has been widely applied to various cancer research fields. For example, in a 2020 study, Greenwald et al. used single-cell sequencing technology to analyze a pan-cancer dataset, identified patterns of cellular heterogeneity, and first introduced single-cell sequencing into the framework of pan-cancer research. Subsequently, Zhou et al. revealed the lineage differentiation of human lung adenocarcinoma cells through single-cell sequencing technology, providing new possibilities for the early diagnosis of lung adenocarcinoma and the study of metastasis at specific stages. In addition, relevant research teams have also achieved preliminary results in the research on cancer recurrence prediction, prognosis modeling, and treatment response using single-cell sequencing technology.
[0004] Although single-cell sequencing technology has great potential in cancer research, existing research still has some limitations, especially in clustering analysis. Clustering is an important tool for single-cell sequencing data analysis, which can help researchers discover cell populations with similar gene expression patterns. However, existing clustering methods face two major challenges: one is the influence of outliers. Traditional methods such as the K-means and Louvain algorithms are easily interfered by outliers, reducing the accuracy of clustering; the other is the batch effect problem, that is, the technical differences introduced by different sequencing batches will mask biological changes, thereby affecting the reliability of clustering results.
[0005] To address these issues, researchers have proposed some improved algorithms. For example, Goldwasser et al. developed a forest fire clustering method in 2022. By combining iterative label propagation and Monte Carlo simulation, it effectively reduces the impact of outliers on clustering accuracy while enhancing the stability and scalability of the algorithm. However, this method is costly in terms of resource consumption. Other studies have focused on batch effect correction. For example, traditional methods such as Harmony and Seurat reduce the interference of batch effects through data preprocessing, while the DESC algorithm developed by Lyu et al. optimizes the objective function using Kullback-Leibler (KL) divergence, further improving clustering accuracy. However, these methods still perform sub-optimally when dealing with data with weak or no batch effects.
[0006] Based on this, researchers have gradually realized that focusing solely on batch effect correction cannot fully solve the clustering accuracy problem. To improve clustering accuracy and mitigate the impact of batch effects simultaneously, it is particularly important to develop an innovative clustering algorithm that combines data preprocessing and deep optimization. The G-DESC-E algorithm emerged. It eliminates outliers through a grid-based preprocessing strategy and designs the objective function by combining KL divergence and label entropy, achieving a balance between batch effect correction and clustering optimization. Experiments on multiple cancer datasets show that the G-DESC-E algorithm not only outperforms traditional algorithms in clustering accuracy but also demonstrates outstanding performance in the identification of pan-cancer key genes, providing important support for cancer gene state research and clinical applications. Summary of the Invention
[0007] The object of the present invention is to address the deficiencies in the analysis of single-cell sequencing data in the prior art. A gene state cancer clustering analysis method based on deep learning is proposed, which for the first time realizes a pan-cancer clustering algorithm (G-DESC-E) that combines grid-based preprocessing and deep optimization.
[0008] The present invention is realized by the following technical solutions:
[0009] A gene state cancer clustering analysis method based on deep learning according to the present invention includes:
[0010] S1: Construct a dataset to be analyzed, including a systemic lupus erythematosus dataset, a liver immune microenvironment dataset, and a single-cell sequencing dataset of multiple cancer types with introduced batch effects, where the single-cell sequencing dataset contains a gene expression matrix;
[0011] S2: Construct a deep neural network consisting of an input layer, a hidden layer, and an output layer. Receive the gene expression matrix as the input of the deep neural network in the input layer, perform normalization processing on the gene data in the gene expression matrix, preliminarily screen to obtain high-variation gene data, and perform feature extraction on the high-variation gene data to obtain high-variation gene features. The hidden layer includes multiple stacked encoder and decoder layers. The high-variation gene features are processed by the encoder to obtain low-dimensional high-variation gene expression features, and then the low-dimensional high-variation gene expression features are restored to the original input gene expression matrix by the decoder. The output layer will output the gene expression matrix.
[0012] S3: Construct a grid preprocessing model, input the low-dimensional high-variation gene expression features into the grid preprocessing model, and perform network preprocessing on the low-dimensional high-variation gene expression features, including grid segmentation and outlier removal. Construct a clustering algorithm model, input the preprocessed low-dimensional high-variation gene expression features into the clustering algorithm model, and perform clustering using KL divergence and label entropy to obtain the clustering distribution result of the low-dimensional high-variation gene expression features. Use the objective function and optimization strategy to iteratively optimize the clustering result. By iteratively optimizing the objective function, gradually adjust the allocation of cells to the clustering centers, and update the clustering result until the convergence condition is met to achieve the global optimality of the final clustering result, which is used as the output of the clustering algorithm model.
[0013] S4: Perform testing and evaluation on the clustering algorithm model in step 3.
[0014] In some embodiments, the iterative optimization of the clustering result using the objective function and optimization strategy in step S3 further includes: using the objective function and optimization strategy, the clustering algorithm model strengthens the label entropy in terms of the KL divergence metric, and amplifies the ratio weight of the auxiliary probability distribution of gene features to the clustering distribution. Complete the optimization of the objective function through the stochastic gradient descent method, that is, in each iteration, calculate the gradients of the sample embedding features and the clustering centers, and gradually optimize the parameters of the model until the change in the loss function is less than the set threshold. Obtain the clustering distribution of each gene feature sample.
[0015] Further, after multiple tests, determine the objective function obtained by squaring the ratio of the current probability distributions of the auxiliary distribution and the clustering gene features.
[0016] In some embodiments, the calculation formula of the KL divergence metric is as follows:
[0017]
[0018] where p c represents the proportion of the number of cells in the c-th batch to the total number, and q cRepresents the proportion of the c-th batch of cells in a given region based on the results of the clustering algorithm, which is determined by the results of the clustering algorithm and satisfies the following constraints:
[0019]
[0021] In some embodiments, the label entropy is calculated by the following formula:
[0022]
[0023] where P i,j represents the probability distribution that cell i is assigned to cluster j. The larger the label entropy value, the more chaotic the clustering assignment information is, and the greater the difference between the assignment result and the correct label.
[0024] In some embodiments, the objective function is defined as follows:
[0025]
[0026] where t i,j represents the auxiliary probability distribution of gene features, s i,j represents the current probability distribution of clustering gene features, K represents the number of clustering gene features, and n represents the number of cells.
[0027] In some embodiments, the definition of the auxiliary probability distribution of gene features is:
[0028]
[0029] Using the auxiliary probability distribution t of gene features i,j to enhance the purity of clustering, that is, to optimize the clustering result by preferentially considering cells with high confidence. The auxiliary probability distribution of gene features is used to measure the confidence of cell clustering assignment.
[0030] In some embodiments, an optimization strategy is designed, including:
[0031] First, use the Student's distribution as the kernel function to evaluate the similarity between the embedding point f of each cell i i and the center k of cluster j j :
[0032]
[0033] where f i = f W (m i ) is located in the feature space F, and m i represents the element of the embedded matrix M, and α represents the degree of freedom of the Student's t-distribution;
[0034] The coding layer optimizes the clustering result by minimizing the Loss in the objective function, and improves the clustering purity by focusing on high-confidence cells;
[0035] The optimization of the objective function is achieved by stochastic gradient descent, mainly focusing on the data point f i and the clustering center k i Feature space embedding gradient calculation:
[0036]
[0037] During the iteration process, if the loss does not reach the set epoch number threshold, the auxiliary probability distribution of the gene features is updated, otherwise the iteration stops.
[0038] In some embodiments, the default value of tol is 0.005, and the calculation formula is:
[0039]
[0040] where X curr is the cluster number obtained from the current maximum cluster assignment probability, X prev is the cluster number obtained from the previous step, |X curr ≠X prev | is the number of different cells in X curr and X prev .
[0041] In some embodiments, step S2 further includes: the hidden layer adopts a fully connected network, and each hidden layer uses an activation function to extract non-linear features in the single-cell gene expression data, and introduces a Dropout mechanism to randomly discard some neurons of the deep neural network according to a set probability to prevent overfitting of the data.
[0042] In some embodiments, the method further includes the step of providing downstream analysis services, that is, providing a visualization image classification service and a visualization analysis service.
[0043] Compared with the prior art, the present invention can achieve the following beneficial technical effects:
[0044] 1) By introducing the methods of grid division and density threshold, outliers are effectively removed, thereby reducing the analysis error;
[0045] 2) Combine KL divergence and label entropy to construct an objective function, design an optimization strategy with the objective function as the core, and improve the clustering accuracy in an iterative optimization manner while reducing the influence of batch effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1Overall flowchart of the gene-state cancer clustering analysis method based on deep learning of the present invention;
[0047] Figure 2 Block diagram of an embodiment of the present invention;
[0048] Figure 3 Graph of the comparative test results in the embodiment;
[0049] Figure 4 Visualization result graph of gene KDM5C in the embodiment;
[0050] Figure 5 Visualization result graph of pan-cancer in the embodiment;
[0051] Figure 6 Bar graph of pan-cancer GO analysis in the embodiment;
[0052] Figure 7 Visualization result graph of lung cancer in the embodiment;
[0053] Figure 8 Visualization result graph of colorectal cancer in the embodiment;
[0054] Figure 9 Visualization result of brain cancer in the embodiment. Detailed implementation manners
[0055] Next, the technical solution of the present invention will be described in detail in conjunction with the accompanying drawings and specific embodiments.
[0056] As Figure 1 shown, the gene-state cancer clustering analysis method based on deep learning of the present invention (abbreviated as G-DESC-E) realizes a pan-cancer clustering algorithm that combines grid-based preprocessing and a single-cell sequencing data clustering model for deep optimization based on a deep neural network model. The specific process includes the following steps:
[0057] Step 1, construct a dataset to be analyzed, including a lupus erythematosus dataset, a liver immune microenvironment dataset, and a single-cell sequencing dataset of multiple cancer types including introduced batch effects. Each of the single-cell sequencing datasets contains three files: a gene list, a cell identifier, and a gene expression matrix;
[0058] Step 1-1, prepare the dataset. The dataset includes: a lupus erythematosus dataset, a liver immune microenvironment dataset, and colorectal cancer, ovarian cancer, breast cancer, lung cancer, and brain cancer datasets; for example, use the lupus erythematosus dataset and liver immune microenvironment dataset from the public dataset NCBI database and the lung cancer, brain cancer, ovarian cancer, colorectal cancer, and breast cancer datasets from the 10x platform;
[0059] In step 1-1, the cancer dataset was downloaded from https: / / www.10xgenomics.com / company, and the lupus dataset and the liver immune microenvironment dataset were downloaded from the NCBI database, National Center for Biotechnology Information;
[0060] In step 1-1, the datasets are all single-cell data obtained using single-cell sequencing technology. Each dataset contains three files: a gene list, cell identifiers, and a gene expression matrix. The gene list provides annotation information for genes, listing the names or identifiers of all genes detected in the experiment, which is used to match the rows in the gene expression matrix to ensure that each row index corresponds to a specific gene; the cell identifier file marks the identity of each cell, providing unique identifiers for all single cells in the dataset, which is used to match the columns in the gene expression matrix to ensure that each column index corresponds to a specific cell; the gene expression matrix file is the core data file, containing the expression levels of genes in each cell, providing basic data for downstream analysis (such as clustering, dimensionality reduction, and differential gene analysis);
[0061] Step 1-2, mix and organize the datasets in step 1-1;
[0062] In step 1-2, first, the single-cell sequencing datasets of brain cancer, lung cancer, breast cancer, colorectal cancer, and ovarian cancer are mixed. The mixed single-cell sequencing dataset contains cells from multiple cancer types, thus introducing batch effects. The single-cancer-type datasets before mixing do not have batch effects themselves. The dataset with introduced batch effects can be better used to evaluate the performance of the algorithm in reducing batch effects and improving clustering accuracy.
[0063] At the same time, the lupus dataset and the liver immune microenvironment dataset are split and organized. These two original datasets contain a large number of records, and direct use may lead to low analysis efficiency. In addition, it is found that the number of columns in the gene expression matrix of these two datasets is inconsistent with the number of cells in the barcode file, which may be caused by not capturing all cells during the sequencing process or omissions during data preprocessing. To solve this problem, by matching the gene expression matrix and the barcode file, the records that do not appear in both are deleted, thus ensuring the consistency and integrity of the data;
[0064] Step 2: Construct a deep neural network composed of an input layer, a hidden layer, and an output layer. At the input layer, receive the gene expression matrix as the input of the deep neural network, perform standardization processing on the gene data in the gene expression matrix, screen out high-variation gene data, and extract gene features from the preliminarily screened high-variation gene data. The hidden layer includes multiple stacked encoder and decoder layers. After being processed by the encoder, low-dimensional high-variation gene expression features are obtained, and then the low-dimensional high-variation gene expression features are restored to the original input gene expression matrix through the decoder. The gene expression matrix will be output through the output layer. The specific description is as follows:
[0065] Step 2-1: Construct a deep neural network model composed of an input layer, a hidden layer, and an output layer. At the input layer, use the low-dimensional single-cell gene expression data as the input. At the hidden layer, implement grid preprocessing (G) and a single-cell sequencing data clustering algorithm (DESC-E). At the output layer, output a visualization image and clustering results. That is, the input layer receives the data of the gene expression matrix, and the feature dimension of the data can be adjusted according to the result of the grid preprocessing in Data Step 2. The hidden layer adopts a fully connected method. Each hidden layer uses an activation function (such as ReLU) to extract non-linear features in the single-cell gene expression data. A Dropout mechanism is introduced in the hidden layer to prevent overfitting. The hidden layer includes multiple stacked encoder and decoder layers, which map the high-dimensional single-cell gene expression matrix to a low-dimensional embedding space, extract non-linear features in the single-cell gene expression data layer by layer through the encoder, and reconstruct the data in the low-dimensional embedding space to the original data space of the single-cell gene expression data to minimize the reconstruction error;
[0066] Step 2-2: Standardize the gene expression matrix of each cell in the dataset to be analyzed. The specific description is as follows: First, perform normalization processing on the gene expression matrix of each cell, standardize the unique molecular identifier (UMI) count of each gene. The normalization steps include dividing the UMI count by the total UMI count of each cell, multiplying by a fixed factor (such as 10,000), and converting to a logarithmic scale. The expression of the normalized result (Normalized Value) of the gene expression matrix of each cell is as follows:
[0067]
[0068] where UMI Count represents the UMI count, and Total UMI Count represents the total UMI count of each cell.
[0069] Subsequently, perform standardization at the gene level, including subtracting the mean from the expression value of each gene in all cells and dividing by the standard deviation.
[0070] Step 2-3: Perform feature selection on the dataset to be analyzed, and use the highly variable genes of the cell population as the selected features. Specific description: Use the "filter_genes_dispersion" function in the Scanpy tool to screen for highly variable genes; this function identifies genes with significant expression differences in the cell population, which helps to reveal cell heterogeneity.
[0071] Step 2-4: Perform dimensionality reduction on the data of the preliminarily screened highly variable gene features to obtain low-dimensional highly variable gene expression features.
[0072] Step 2-5: Perform grid processing on the data points corresponding to the low-dimensional gene expression matrix to remove outlier data points. The specific description is as follows:
[0073] Step 2-3-1: Perform feature extraction on the selected features and perform dimensionality reduction. Specific description: Use a Stacked Autoencoder to perform dimensionality reduction on high-dimensional gene expression data. The autoencoder consists of an encoder and a decoder. The encoder is used to transform the input from high-dimensional to low-dimensional feature space, and the original input is restored through the decoder. The expression is as follows:
[0074] Encoding: h = f(Wx + b)
[0075] Decoding:
[0076] where f and g represent activation functions (such as ReLU or tanh), W and b represent the weight and bias parameters of the network, h represents the low-dimensional representation after encoding, c represents the bias vector of the encoder, represents the output of the decoder. The low-dimensional features after being processed by the encoder are the input for subsequent clustering, which helps to improve the efficiency and accuracy of the clustering algorithm.
[0077] S3: Construct a grid preprocessing model, input the low-dimensional highly variable gene expression features into the grid preprocessing model, and perform network preprocessing including grid segmentation and outlier removal on the low-dimensional highly variable gene expression features through the grid preprocessing model; construct a clustering algorithm model, input the preprocessed low-dimensional highly variable gene expression features into the clustering algorithm model, use KL divergence and label entropy to obtain an optimized clustering result of the low-dimensional highly variable gene expression features, and use the objective function and optimization strategy to iteratively optimize the clustering result as the output of the clustering algorithm model. The specific description is as follows:
[0078] Step 3-1: Construct a grid preprocessing module. Use the grid preprocessing module to perform grid preprocessing on the input low-dimensional gene expression matrix, including defining grid division parameters, density classification, and removing isolated points. According to the distribution of gene data points in the low-dimensional space, the dimensionality-reduced data points are divided into a set of grids, and appropriate grid numbers and density thresholds d are set. t If the density of a certain grid is higher than the density threshold d t it is marked as a "dense area". If the grid density is lower than the density threshold and the Euclidean distance from the grid to the nearest dense area grid is greater than half of the grid diagonal, the data points in the grid are marked as "isolated points". Use an iterative algorithm to continuously detect and remove isolated points until no new isolated points are detected in the dataset, thereby ensuring the purity and resource efficiency of the input data.
[0079] Step 3-2: Construct a clustering algorithm model. Calculate the gene feature discrete random variable and gene feature continuous random variable according to the KL divergence formula, and then measure the information loss when fitting the true probability distribution of gene features based on the gene feature discrete random variable and gene feature continuous random variable to determine the objective function to minimize the KL divergence, thereby optimizing the clustering model to make its output closer to the true label.
[0080] First, determine the objective function. The specific description is as follows: Optimize the entire clustering algorithm by minimizing the loss function. The objective function used is defined as:
[0081]
[0082] where t i,j is the auxiliary probability distribution of gene features, s i,j is the current probability distribution of gene features, K is the number of clustering gene features, n is the number of cells, and the definition of the auxiliary probability distribution of gene features is:
[0083]
[0084] Use the auxiliary probability distribution t of gene features i,j to enhance the purity of clustering, that is, optimize the clustering results by giving priority to cells with high confidence. Use the auxiliary probability distribution of gene features to measure the confidence of cell clustering assignments.
[0085] The objective function is the core of the optimization strategy. The core of the objective function design is to combine the Kullback-Leibler (KL) divergence and label entropy to optimize the clustering results.
[0086] Secondly, the optimization strategy is designed as follows:
[0087] First, use the Student's distribution as the kernel function to evaluate the embedding point f of each cell i i with the cluster j center k j for similarity:
[0088]
[0089] where f i = f W (m i ) lies in the feature space F, m i is derived from the embedded matrix M, and α represents the degrees of freedom of the Student's distribution.
[0090] The encoding layer optimizes the clustering results by minimizing the loss function Loss in the objective function. The role of the auxiliary distribution t is to improve the clustering purity by focusing on high-confidence cells, and this probability can be used to measure the confidence of the cell's clustering assignment.
[0091] The optimization of the objective function is achieved through Stochastic Gradient Descent (SGD), mainly focusing on the data point f i and the cluster j center k j Feature space embedding gradient calculation:
[0092]
[0093] The neural network uses these formulas to calculate the parameter gradients during the standard backpropagation process, and the model is trained using Keras. During the iteration process, if the loss does not reach the set epoch number threshold, the auxiliary distribution t is updated, otherwise the iteration stops. The default value of tol is 0.005, and the calculation formula is
[0094]
[0095] where X curr is the cluster number obtained from the current maximum cluster assignment probability, X prev is the cluster number obtained from the previous step, |X curr ≠ X prev | is the number of different cells in X curr and X prev .
[0096] Using the objective function and the optimization strategy, the clustering algorithm model strengthens the label entropy based on the KL divergence formula by amplifying the ratio weight of the auxiliary probability distribution of gene features to the clustering distribution. The optimization of the objective function is completed by the Stochastic Gradient Descent (SGD) method. That is, in each iteration, by calculating the gradients of the sample embedding features and the clustering centers, the parameters of the model are gradually optimized until the change in the loss function is less than the set threshold; the clustering distribution of each gene feature sample is obtained, and the output dimension is consistent with the predefined number of clustering categories.
[0097] For a dataset composed of multiple batches, denote these batches as C, and the calculation formula of the KL divergence index is as follows:
[0098]
[0099] where p c represents the proportion of the number of cells in the c-th batch to the total number, and q c represents the proportion of the cells in the c-th batch in the given region based on the results of the clustering algorithm. This proportion is determined by the results of the clustering algorithm and satisfies the following constraint conditions:
[0100]
[0101] The lower the KL divergence value, the better the elimination effect of the batch effect;
[0102] The label entropy is an important indicator for the clustering results of the optimization algorithm. The label entropy of each cell i is calculated by the following formula:
[0103]
[0104] where, P i,j represents the probability that cell i is assigned to cluster j, and the calculation method is the same as t i,j . The larger the label entropy value, the more chaotic the clustering assignment information is, and the greater the difference between the assignment result and the correct label.
[0105] The calculation of the label entropy E only depends on the probability P i,j . By combining the label entropy with the KL divergence index and comparing it with the initial objective function of DESC, the importance of the weight factor t i,j can be seen, and it is enhanced. Through extensive testing, this study optimized the weight enhancement ratio and finally designed the objective function.
[0106] Furthermore, after multiple tests, it is determined that the objective function obtained by squaring the ratio of the current probability distribution of the auxiliary distribution and the clustering gene features is used as the final clustering result;
[0107] Step 4, test and evaluate the clustering algorithm model in Step 3, where;
[0108] Step 4-1, Cluster result analysis, specific description: Use the clustering result analysis of the single-cell sequencing dataset obtained in Step 3. At the same time, for the mixed pan-cancer dataset and single-cancer dataset, the processing ability and accuracy of the model can be tested respectively;
[0109] Step 4-2, Conduct comparative experiment design, specific description: Select the traditional clustering algorithm (Leiden algorithm) and the unoptimized DESC algorithm as the comparison groups for the dataset in Step 1-1. By comparing the clustering results of different methods on the same dataset, verify the advantages of the G-DESC-E model in removing batch effects and improving clustering accuracy;
[0110] Step 4-3, Determine evaluation metrics, specific description: Use the Adjusted Rand Index (ARI) to evaluate the clustering effect of the model. The higher the ARI, the higher the matching degree between the clustering result and the true label.
[0111] The present invention further includes:
[0112] Step 5, Provide downstream analysis services, where;
[0113] Step 5-1, Provide visualization image classification services, specific description: In the visualization images output by the deep neural network model in Step 3-3, the present invention uses t-SNE technology to visualize the clustering results to intuitively display the classification effects of different cancer datasets. Specifically, after the single-cell gene expression data is processed by the G-DESC-E algorithm, the obtained low-dimensional embedded features are used as input to project the distribution of data points in a two-dimensional space. The clustered categories are distinguished by different colors, and each color corresponds to an independent cancer subtype or cell state. This method can not only show the clustering performance of the model, but also facilitate observing potential patterns in the data and the relationships between cells.
[0114] Step 5-2, Provide visualization analysis services, specific description: Judge the gene expression level according to the change in the shading of the classified images in Step 5-1, and deeply explore the biological significance of the clustering results according to the cell distribution and gene expression. The present invention combines the GO analysis results with the clustering visualization graph to generate a bar chart and a classification relationship graph of gene function classification. We select the top 3000 genes with the highest expression levels for each clustering category, and conduct GO enrichment analysis based on three levels: molecular function (MF), cellular component (CC), and biological process (BP) to generate a bar chart, showing the interrelationships and hierarchical relationships between GO functions, and revealing the role of gene functions in cancer occurrence and development.
[0115] The above overall process is further described in combination with embodiments: The gene expression matrix obtained in step 2 is input into the input layer of the deep neural network in step 3-2, and clustering and evaluation are performed using the model according to the method in step 4 above. The evaluation index is selected as the ARI value of the clustering evaluation index. After clustering, according to the method in step 5, a visualization image is obtained according to the requirements to complete the downstream analysis. The calculation method of the clustering evaluation index ARI is as follows:
[0116]
[0117] where n uv represents the number of cells that are assigned to cluster u in the true label and to cluster v in the clustering label. a u is the number of cells assigned to cluster u in the true label, and b v is the number of cells assigned to cluster v in the clustering label.
[0118] According to the comparative test method in step 4-2 above, the Leiden algorithm and the original DESC algorithm are compared. The experimental data of the three methods show that the G-DESC-E index of the present invention is the best. As shown in Table 1, the deep learning-based single-cell data clustering method has a high prediction accuracy.
[0119] Table 1 Comparative test results
[0120]
[0121]
[0122] According to the downstream analysis process in step 5 above, the visualization image shows that the expression levels of KDM5C are all up-regulated in lung cancer, brain cancer, rectal cancer, and ovarian cancer. Combining that this gene has been confirmed to be one of the genes that synergistically promote the growth of breast cancer cells, it can be inferred that this gene may be able to be used as a key gene for cancer prediction; all four members of the PBX family are related to cancer, and the image shows that PBX2 is highly expressed in all five cancers, which may also be a direction for cancer prediction.
[0123] Read out in combination with the bubble chart, the biological processes that are more affected are detection of chemical stimulus involved in sensory perception, detection of chemical stimulus involved in sensory perception of smell, sensory perception of smell. Among the molecular functions, the more affected ones are olfactory receptor activity and odorant binding, both of which are related to sensory perception, especially olfaction. This result shows on the one hand the significant damage caused by cancer to the human body, and on the other hand, whether the change in olfaction can predict the occurrence of cancer is worthy of further research.
[0124] The highly expressed genes are clearly shown in the visualization images of single cancers, and these genes all have the potential to be candidate key genes for the occurrence, mutation, and prognosis of this cancer.
[0125] The above gives an exemplary description of the present invention. The neural network is only a typical representative and does not limit the present invention. It should be noted that without departing from the core of the present invention, any simple deformation, modification, or equivalent replacement that can be made by those skilled in the art without creative labor falls within the protection scope of the present invention.
Claims
1. A method for clustering analysis of genotype cancer based on deep learning, characterized in that: include: S1: Constructing the datasets to be analyzed, including a lupus erythematosus dataset, a liver immune microenvironment dataset, and a mixed single-cell sequencing dataset of multiple cancer types with batch effects introduced, wherein the single-cell sequencing dataset includes a gene expression matrix; S2: constructing a deep neural network consisting of an input layer, a hidden layer and an output layer, receiving the gene expression matrix at the input layer as the input of the deep neural network, performing standardization processing on the gene data in the gene expression matrix, preliminarily screening to obtain high-variation gene data, and performing feature extraction on the high-variation gene data to obtain high-variation gene features; The hidden layer includes a plurality of stacked encoder and decoder layers, the high-variability gene features are processed by the encoder to obtain low-dimensional high-variability gene expression features, and then the low-dimensional high-variability gene expression features are restored to the original input gene expression matrix through the decoder, and the gene expression matrix is output through the output layer; S3: constructing a grid preprocessing model, inputting the low-dimensional high-variability gene expression features into the grid preprocessing model, and performing network preprocessing including grid segmentation and isolated point removal on the low-dimensional high-variability gene expression features; Constructing a clustering algorithm model, inputting the low-dimensional high-variability gene expression features after preprocessing into the clustering algorithm model, performing clustering using KL divergence and label entropy to obtain low-dimensional high-variability gene expression feature clustering distribution results, iteratively optimizing the clustering results using the objective function and optimization strategy, gradually adjusting the allocation of cells to cluster centers through iterative optimization of the objective function, updating the clustering results until the convergence conditions are met, and achieving the final global optimal clustering result as the output of the clustering algorithm model; S4: Perform clustering algorithm model testing and evaluation in step 3.
2. A method for clustering analysis of genotype cancer based on deep learning according to claim 1, characterized in that: The step S3 further comprises: using the objective function and the optimization strategy to iteratively optimize the clustering results, the clustering algorithm model strengthens the label entropy on the KL divergence index, and amplifies the ratio weight of the auxiliary probability distribution of the gene feature to the cluster distribution; optimizing the objective function by the stochastic gradient descent method, that is, in each iteration, by calculating the gradient of the sample embedding feature and the cluster center, the parameters of the model are gradually optimized until the loss function changes less than the set threshold; obtaining the cluster distribution of each gene feature sample; Furthermore, after multiple tests, the objective function obtained by squaring the ratio of the auxiliary distribution to the current probability distribution of the clustering gene features is determined.
3. A method for clustering analysis of genotype cancer based on deep learning according to claim 1, characterized in that: The calculation formula of the KL divergence indicator is as follows: Among them, pc represents the proportion of the number of cells in the cth batch to the total number, and qc represents the proportion of the cth batch of cells in a given area based on the results of the clustering algorithm. This proportion is determined by the results of the clustering algorithm and meets the following constraints:
4. A method for clustering analysis of genotype cancer based on deep learning according to claim 1, characterized in that: The label entropy calculation formula is as follows: Among them, P i,j It represents the probability distribution of cell i being assigned to cluster j. The larger the label entropy value, the more chaotic the clustering assignment information is, and the greater the difference between the assignment result and the correct label.
5. The method for clustering analysis of genotype cancer based on deep learning according to claim 1, characterized in that: The objective function is defined as follows: Among them, t i,j represents the auxiliary probability distribution of gene features, s i,j represents the current probability distribution of clustered gene features, K represents the number of clustered gene features, and n represents the number of cells.
6. A method for clustering analysis of genotype cancer based on deep learning according to claim 1, characterized in that: The auxiliary probability distribution of the gene feature is defined as: Using auxiliary probability distribution of genetic features i,j Enhance the purity of clustering, that is, optimize the clustering results by giving priority to cells with high confidence, and use the auxiliary probability distribution of gene features to measure the confidence of cell clustering allocation.
7. The method for clustering analysis of genotype cancer based on deep learning according to claim 1, characterized in that: Design optimization strategies, including: First, the embedding point f of each cell i is evaluated using the Student distribution as the kernel function i With cluster j center k j Similarity: Among them, f i =f W (m i ) is located in the feature space F, and m i represents the elements of the embedded matrix M, α represents the degrees of freedom of the Student's t distribution; The encoding layer optimizes the clustering results by minimizing the Loss in the objective function and improves the clustering purity by focusing on high-confidence cells; The optimization of the objective function is achieved by stochastic gradient descent, focusing on the data point f i and cluster center k i Feature space embedding gradient calculation: During the iteration, if the loss does not reach the set epoc If h is the domain value, the auxiliary probability distribution of the gene feature is updated, otherwise the iteration stops.
8. The method for clustering analysis of genotype cancer based on deep learning according to claim 7, characterized in that: The default value of tol is 0.005, and the calculation formula is: Among them, X curr The cluster number obtained by the current maximum cluster allocation probability, X prev is the cluster number obtained in the previous step, |X curr ≠X prev | for X curr and X prev Different numbers of cells in .
9. The method for clustering analysis of genotype cancer based on deep learning according to claim 1, characterized in that: Step S2 further includes: the hidden layer adopts a fully connected network, each hidden layer adopts an activation function to extract nonlinear features in the single-cell gene expression data, and introduces a Dropout mechanism to randomly discard some neurons of the deep neural network according to a set probability to prevent data overfitting.
10. The method for clustering analysis of genotype cancer based on deep learning according to claim 1, characterized in that: The method also includes the step of providing downstream analysis services, namely, providing visual image classification services and providing visual analysis services.
Citation Information
Cited By
Method and system for constructing cancer early screening model based on difference statistics
CN120853693A
A method and system for constructing a cancer early screening model based on difference statistics
CN120853693B