Single cell clustering statistical significance test method and system based on manifold learning

Through the statistical significance test method based on manifold learning and Monte Carlo simulation, the limitations and noise problems in single-cell cluster significance test are solved, and the accuracy and stability of the test are improved.

CN120072052AInactive Publication Date: 2025-05-30XIAMEN UNIV

Patent Information

Application Number
CN202510528024.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has limitations and noise problems in single-cell cluster significance tests, resulting in low accuracy of significance tests.

Method used

Single-cell data were dimensionalized by manifold learning-based method, and statistical significance tests were performed through Monte Carlo simulation and Gaussian distribution fitting, and cluster clusters were merged from bottom to up.

Benefits of technology

The accuracy and stability of cluster significance test are improved, excessive clustering and false positive problems are avoided, and the differences between cell populations are strictly statistically verified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072052A_ABST
    Figure CN120072052A_ABST
Patent Text Reader

Abstract

The invention discloses a single cell clustering statistics significance test method and system based on manifold learning, and the method comprises the following steps: obtaining original single cell data, carrying out the standardization processing of the original single cell data, and removing the deviation and noise in the data; performing dimension reduction on the standardized data by using a manifold learning method; dividing the dimension reduction data based on a pre-calculated clustering label, and performing hierarchical clustering on the divided data to obtain a hierarchical tree; hypothesis testing is carried out step by step from the bottom to the top of the hierarchical tree by adopting a Monte Carlo simulation method, and whether the clusters are combined or not is determined according to a hypothesis testing result; and finally, taking the combined hierarchical tree as a single cell clustering correction result. According to the method, the reliability and the accuracy of the clustering statistical test can be improved, and unnecessary group splitting and misjudgment can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and particularly relates to a method and system for statistical significance testing of single-cell clustering based on manifold learning. Background Art

[0002] In recent years, single-cell transcriptomics technology has made remarkable progress and become an important tool for revealing cell diversity. Among them, clustering analysis is one of the most traditional and crucial methods. Through clustering analysis algorithms, purposes such as cell type classification can be achieved. Currently, mainstream clustering algorithms tend to over-segment known cell types into more subpopulations. These segmentations are often achieved through differential expression gene analysis, but lack sufficient statistical verification, making it impossible to clearly distinguish whether these subpopulations are real biological variations, technical noises, or different states of the same cell type. This over-segmentation may not only lead to the discovery of false subpopulations but also affect subsequent biological analysis and experimental design.

[0003] Currently, some methods for reliability testing of clustering algorithm results have been proposed. For example, the SigClust method evaluates the clustering significance based on a Gaussian distribution model.

[0004] The Chinese patent application with the publication number CN112233796A discloses a research method for immune-enhanced molecular subtypes in early liver cancer, which includes the following steps: S1, data download; S2, data preprocessing; S3, screening of the expression profiles of immune genes; S4, screening of molecular subtypes; S5, expression profile clustering analysis; S6, relationship with clinical features; S7, relationship with immunity; S8, prognostic difference analysis; S9, relationship with mutations; S10, relationship with the expression of immune checkpoint genes; S11, WGCNA analysis and mining; S12, verification using external datasets. When screening molecular subtypes in step S4, the R software package SigClus is used to analyze the clustering significance between every two of these five subtypes, and it is found that there are significant differences in expression distribution p < 0.05 between C1-vs-C5, C4-vs-C5, and C2-vs-C3, and marginal significance is shown between C1-vs-C2, C1-vs-C3, C2-vs-C4, and C3-vs-C5.

[0005] The above solution uses a Gaussian distribution-based method to evaluate the significance of clustering, but it has limitations in multi-clustering analysis and can only compare two clusters pairwise. Some improved solutions have been extended on this basis by introducing hierarchical comparison, which solves the problem of not being able to compare multi-clustering results. However, the comparison order is unreasonable. It is top-down, which is prone to false judgments. At the same time, it directly uses the original sequencing data for statistical tests, resulting in relatively large noise. Cytocipher focuses on enrichment scoring and two-sample t-tests, but there are also problems in accurately identifying gene sets and establishing a reliable scoring method. There are also some improved methods that focus on enrichment scoring and two-sample t-tests, but there are also deficiencies in accurately identifying gene sets and establishing a reliable scoring method. Summary of the Invention

[0006] The present invention provides a method and system for statistical significance testing of single-cell clustering based on manifold learning, aiming to solve the problems such as low accuracy of significance testing caused by the limitations and relatively large noise of the existing technology.

[0007] To solve the above technical problems, the single-cell clustering statistical significance testing method proposed by the present invention includes the following steps: Obtain the original single-cell data, perform standardization processing on the original single-cell data to remove the bias and noise in the data; Use the manifold learning method to reduce the dimension of the standardized data; Based on the pre-computed clustering labels, divide the dimension-reduced data, and perform hierarchical clustering on the divided data to obtain a hierarchical tree; Adopt the Monte Carlo simulation method to gradually perform hypothesis testing from the bottom to the top of the hierarchical tree, and decide whether to merge clusters according to the hypothesis testing results; The finally merged hierarchical tree is used as the single-cell clustering correction result.

[0008] Preferably, the specific method of the hypothesis testing is as follows: Merge the data of the two clusters to be compared, and perform Gaussian distribution fitting on the merged data; Randomly extract samples with the same size as the data of the two clusters to be compared from the fitted Gaussian distribution, and perform hierarchical clustering to divide them into two sub-clusters; According to the preset number of simulations, simulate the two sub-clusters, calculate the variance increment of the two sub-clusters in each simulation, construct the null hypothesis distribution and calculate the corresponding p-value; If the p-value is less than the set significance level, reject the merger of the two clusters to be compared; otherwise, merge the two clusters to be compared.

[0009] Preferably, the calculation method of the variance increment is as follows: For the total sum of squares of any cluster C, it is defined as follows:

[0010] Wherein, represents the total sum of squares of the clustering cluster C, represents a sample point in the cluster C, is the centroid of the clustering cluster C; The variance increment caused by merging two sub - clusters and is expressed as follows:

[0011] Wherein, represents the variance increment of two sub - clusters and and represents the total sum of squares of two sub - clusters and and represents the total sum of squares of a sub - cluster and represents the total sum of squares of another sub - cluster and

[0012] Preferably, constructing the null hypothesis distribution and calculating the corresponding p - value are specifically as follows: Calculate the true variance increment of two clustering clusters to be compared; Count the number of times that the variance increment of two sub - clusters in the statistical simulation distribution is greater than or equal to the true variance increment, and the ratio of this number to the total number of simulations is the p - value.

[0013] Preferably, the set significance level is adaptively adjusted each time of comparison, and the adjustment method is:

[0014] Wherein, is the significance level after adaptive adjustment, is the initial significance level, are the sample numbers of two clustering clusters to be compared respectively, is the total number of samples in the hierarchical tree structure data.

[0015] Preferably, the initial significance level is set to 0.05.

[0016] Preferably, the number of simulations is 1000 times.

[0017] Preferably, the Gaussian distribution fitting adopts the maximum likelihood estimation method.

[0018] Preferably, the manifold learning adopts the PHATE algorithm.

[0019] Correspondingly, the present invention further provides a single-cell clustering statistical significance test system based on manifold learning. The system is used to implement the above-mentioned significance test method, and includes: A data preprocessing module, configured to obtain original single-cell data and perform normalization processing on the original single-cell data to remove biases and noises; A dimensionality reduction module, configured to perform dimensionality reduction on the data after normalization processing by using a manifold learning algorithm, and the manifold learning algorithm includes the PHATE algorithm; An initial clustering module, configured to divide the dimensionality-reduced data according to pre-computed clustering labels and perform hierarchical clustering to construct a hierarchical clustering tree; A hypothesis testing module, configured to traverse the hierarchical tree from bottom to top based on the Monte Carlo simulation method and perform a statistical significance hypothesis test on each pair of clusters to be compared; A clustering optimization module, configured to merge the clustering clusters that pass the significance hypothesis test based on the above tests and output the finally optimized clustering hierarchy as a correction result.

[0020] Compared with the prior art, the present invention has the following technical effects: 1. The clustering statistical significance test method proposed by the present invention uses manifold learning technology to perform dimensionality reduction on high-dimensional single-cell data, avoiding the fine-grained information that may be lost in high-dimensional data by traditional linear dimensionality reduction methods. Through non-linear dimensionality reduction, the present invention can better reveal the true relationship between cell clusters, provide a more accurate data representation for subsequent clustering and statistical tests, and ensure the accuracy of subsequent clustering and statistical tests.

[0021] 2. The clustering statistical significance test method proposed by the present invention performs a statistical significance test in a bottom-up manner based on the hierarchical structure, improving the traditional top-down clustering test process. Through this optimization, the occurrence of misjudgments is avoided, and the stability and scientificity of the statistical test results are significantly improved. This method can more accurately identify the significant differences between cell populations and improve the reliability of the test process.

[0022] 3. For the clustering statistical significance test method proposed by the present invention, the present invention performs Gaussian distribution fitting on the dimensionality-reduced data, avoiding the fitting difficulty problem caused by the low-sample and high-dimensional characteristics in the original single-cell expression data, ensuring the accuracy of the fitting result, and effectively overcoming the fitting problem in the case of low samples.

[0023] 4. The clustering statistical significance test method proposed by the present invention uses the Monte Carlo simulation method to construct the null hypothesis distribution and calculate the p-value, which can effectively detect the significance of cell populations. This method ensures that the differences between cell populations are strictly statistically verified, thus avoiding the problems of over-clustering and false positives. Through this method, the reliability and accuracy of the test results are significantly improved, which helps to avoid unnecessary population splitting and misjudgment.

[0024] 5. The clustering statistical significance test method proposed by the present invention ensures the reliability of the statistical results by adaptively adjusting the significance level, and maintains high precision and consistency in all tests. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic flow chart of the significance test method described in the present invention; Figure 2 is a schematic diagram of the manifold learning method described in the embodiment of the present invention; Figure 3 is a schematic diagram of the hierarchical tree structure described in the embodiment of the present invention; Figure 4 is a schematic diagram of the hypothesis verification sequence described in the embodiment of the present invention; Figure 5 is a schematic diagram of simulating two sub-clusters described in the embodiment of the present invention; Figure 6 is a schematic diagram of the variance increment and p-value described in the embodiment of the present invention; Figure 7 is a schematic diagram of the result after the test combination described in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present application and with reference to the accompanying drawings.

[0027] Embodiment 1 This embodiment is a single-cell clustering statistical significance test method based on manifold learning. As Figure 1 shown, it includes the following steps 1 to 5: Step 1, obtain the original single-cell data, perform standardization processing on the original single-cell data, and remove the bias and noise in the data.

[0028] Step 2, use the manifold learning method to reduce the dimension of the standardized data to a suitable latent space to better retain the non-linear relationship and local features between cells. The process of the manifold learning method is as Figure 2As shown, when using the manifold learning method in this embodiment, the PHATE (Persistence Homology Adaptive t-SNE) algorithm is adopted. When the PHATE algorithm is used for single-cell transcriptome data analysis, by combining local similarity and global relationship encoding, it can preserve the local and global structures of the data in the low-dimensional space, thus providing more accurate visualization results of biological data.

[0029] In some other embodiments of the present invention, the manifold learning method can be one or a combination of more than one of the t-distributed Stochastic Neighbor Embedding (t-SNE), Uniform Manifold Approximation and Projection (UMAP), or Locally Linear Embedding (LLE) algorithms, all of which can be used to achieve data dimensionality reduction while maintaining the local structure.

[0030] A specific dimensionality reduction step of this embodiment is as follows: Using the processed standardized expression matrix, the PHATE algorithm first calculates the similarity between cells according to the Euclidean distance or cosine distance between cells to generate a distance matrix; in this embodiment, the Euclidean distance is used. After converting the distance matrix into a local similarity matrix, a Markov transition matrix is constructed to represent the diffusion process of cell states; the diffusion time is set to the default value of 1 or automatically optimized through information entropy. The local structure information is propagated in the similarity graph through the diffusion process to generate a potential low-dimensional embedding space. The diffusion distance matrix is reduced in dimension by the multi-dimensional scaling embedding method in the potential space, and finally the original high-dimensional data is projected into a two-dimensional or three-dimensional space. The output low-dimensional embedding matrix is used for subsequent clustering structure division and statistical significance testing.

[0031] Specifically, the implementation of the above PHATE algorithm comes from the phate library in Python.

[0032] Through the dimensionality reduction process optimized by manifold learning, the present invention can more accurately preserve the non-linear relationships between cells, thus avoiding the loss of subtle differences between cell clusters caused by traditional linear dimensionality reduction methods and improving the accuracy of constructing the hierarchical tree. The PHATE algorithm can effectively reveal the potential structure of high-dimensional single-cell data, reveal non-linear distribution characteristics such as developmental trajectories and cell subpopulations, and provide high-quality input for subsequent hierarchical clustering and significance testing.

[0033] Step 3: Based on the pre-computed clustering labels, partition the dimensionality-reduced data, and perform hierarchical clustering on the partitioned data to obtain a hierarchical tree. The hierarchical tree structure is as shown in Figure 3 shown.

[0034] On the premise that the original single-cell data obtained in Step 1 has been clustered, the corresponding clustering labels can be directly used to perform hierarchical clustering on the partitioned data based on the corresponding clustering labels.

[0035] Alternatively, use a clustering algorithm such as the K-means algorithm to perform an initial partition on the low-dimensional embedding space, implemented using KMeans in scikit-learn, and output a label vector indicating the initial cluster to which each cell belongs.

[0036] According to the clustering labels, partition the dimensionality-reduced embedding data into K sub-cluster sets, and perform a hierarchical clustering operation on the above K clusters: use the centroid of each initial clustering cluster to represent the sub-cluster; the clustering distance metric can use the Euclidean distance, average linkage, etc. Use the linkage function in SciPy to output a linkage matrix for hierarchical clustering, and the linkage matrix is used to construct a tree structure. The constructed hierarchical tree can be regarded as a binary tree, where each leaf node represents an initial clustering cluster, and each non-leaf node in the tree represents the merging relationship of two sub-clusters. This tree provides a basis for subsequent bottom-up significance merging tests.

[0037] Step 4: As shown in Figure 4 shown, use the Monte Carlo simulation method to gradually perform hypothesis testing from the bottom to the top of the hierarchical tree, and decide whether to merge the clustering clusters according to the hypothesis testing results. The hypothesis testing is based on the premise that if two clusters come from the same parent cluster, then their distributions should conform to the Gaussian distribution.

[0038] In this step, the specific method of the hypothesis testing includes steps S41 - S44: S41: Merge the data of the two clustering clusters to be compared, and perform Gaussian distribution fitting on the merged data. The Gaussian distribution fitting uses the maximum likelihood estimation method.

[0039] From the hierarchical tree constructed in Step 3, select two leaf nodes (clusters) that need to be compared currently, extract the corresponding sub-matrices from the dimensionality-reduced embedding matrix, and after merging the two clusters, obtain a merged dataset. Fit the merged data to a d-dimensional multivariate Gaussian distribution, and calculate the maximum likelihood estimation value for the sample points in the merged dataset. The fitting result defines a multivariate Gaussian distribution as a parameter pair, and this parameter pair is used for simulated sampling in the subsequent step S42 to generate simulated clusters for constructing the null hypothesis distribution.

[0040] S42: Randomly sample from the fitted Gaussian distribution a sample with the same size as the data of the two clusters to be compared, and perform hierarchical clustering to dichotomize them into two sub-clusters.

[0041] Suppose the original number of samples in the two clusters are , and a total of simulated cells need to be sampled; after simulation, they need to be divided into two sub-clusters, and the sizes of the two sub-clusters are equal to the sizes of the original clusters, which are also respectively. Then use numpy or scipy.stats.multivariate_normal to generate samples from the estimated multivariate Gaussian distribution, and output the simulated data matrix. Then perform dichotomous hierarchical clustering on the generated simulated data and divide it into two sub-clusters: use the Euclidean distance as the similarity metric, and the clustering method adopts the single-linkage or average-linkage method; use the scipy.cluster.hierarchy module to implement clustering and truncation. According to the clustering results, divide the simulated samples into two sub-clusters, and these two sub-clusters will be used for the next variance increment calculation.

[0042] S43: As Figure 5 shown, according to the preset number of simulations, simulate the two sub-clusters, calculate the variance increment of the two sub-clusters in each simulation, construct the null hypothesis distribution and calculate the corresponding p-value; the specific steps of constructing the null hypothesis distribution and calculating the corresponding p-value are as follows: calculate the true variance increment of the two clusters to be compared. Count the number of times that the variance increment of the two sub-clusters in the simulated distribution is greater than or equal to the true variance increment, and the ratio of this number to the total number of simulations is the p-value. In one embodiment of the present invention, the number of simulations is 1000 times. The null hypothesis is that "the distributions of the two clusters conform to the Gaussian distribution", and the alternative hypothesis is that "the distributions of the two clusters do not conform to the Gaussian distribution".

[0043] In this embodiment, based on the simulated sub-cluster pairs obtained in step S42, repeat the sampling and clustering operations 1000 times (i.e., set the number of simulations T = 1000), calculate the variance increment of the simulated sub-clusters each time, construct the null hypothesis distribution in this way, and calculate the p-value with the true variance increment as a reference. Perform the simulated variance increment calculation for each simulation to obtain a variance increment list with a length of T = 1000; regard it as the statistical quantity distribution under the null hypothesis; visualize the statistical quantity distribution as the histogram shown in the flow chart, as Figure 6 shown, where each bar corresponds to the variance increment calculated in one simulation. Calculate the variance increment of the two real clusters, and its variance increment is shown as the vertical red dashed line in Figure 6 . Reflected in the flow chart shown in Figure 6 , this process corresponds to the right area part above the histogram where the red dashed line falls, that is, the probability interval represented by the p-value.

[0044] S44: If the p-value is less than the set significance level, reject the merger of the two clusters to be compared; otherwise, merge the two clusters to be compared.

[0045] This step is based on the p-value result obtained in step S43 and compares it with the preset significance level (or the adaptively adjusted significance level) to decide whether to retain the original cluster structure or merge it into a new cluster, and applies the judgment result to the hierarchical tree structure, thus realizing the bottom-up cluster merger correction. Among them, the significance level can be a fixed value or an adaptive adjustment. As Figure 6 shown, the p-value of the statistical test in the upper left part is 0.0062. Assuming that the preset significance level is 0.05, since p = 0.0062 < 0.05, the null hypothesis is rejected, indicating that there is a significant difference between these two clusters, and the two clusters to be compared cannot be merged. As Figure 6 shown in the lower left part, the p-value is 0.9327. Since p = 0.9327 > 0.05, the null hypothesis is accepted, indicating that there is no significant difference between these two clusters, and the two clusters can be merged. In the case where the p-value is equal to the set significance level, it can be set to accept or reject the null hypothesis by oneself.

[0046] Furthermore, in the above hypothesis testing method, the calculation method of the variance increment is as follows: For the total sum of squares of any cluster C, it is defined as follows:

[0047] In the formula, represents the total sum of squares of the cluster C, represents a sample point in the cluster C, is the centroid of the cluster C; The variance increment caused by merging two sub-clusters and is expressed as follows:

[0048] In the formula, represents the variance increment of the two sub-clusters and , represents the total sum of squares of the two sub-clusters and , represents the total sum of squares of one sub-cluster , represents the total sum of squares of the other sub-cluster .

[0049] During multiple hypothesis testing, there may be a risk of obtaining significant results by chance. To control this risk, in some embodiments of the present invention, the family error rate is controlled by adjusting the significance level, that is, the set significance level is adaptively adjusted. The adaptive adjustment timing is at each calculation, and the adjustment method is as follows:

[0050] In the formula, is the significance level after adaptive adjustment, is the initial significance level, are the sample numbers of the two clusters to be compared respectively, is the total sample tree in the hierarchical tree structure data. In this embodiment, the initial significance level is set to 0.05; by adaptively adjusting the significance level, the reliability of the statistical results is ensured, and high precision and consistency are maintained in all tests.

[0051] In other embodiments of the present invention, the initial significance level can be set differently according to the original data.

[0052] Starting from the bottom of the hierarchical tree, each pair of clusters to be merged is tested in turn, and S41 - S44 are repeatedly executed until the entire tree is traversed or the termination condition is reached.

[0053] Through the Monte Carlo simulation and the hypothesis testing method based on the Gaussian distribution in this step, better stability and reliability are achieved when dealing with low - sample and high - dimensional data. By dynamically adjusting the significance level, the phenomenon of over - clustering is avoided, making the clustering results more in line with biological reality.

[0054] Step Five, the finally merged hierarchical tree is used as the single - cell clustering correction result. The correction result corresponding to the hierarchical tree structure described in Step Three is as Figure 7 shown. After completing the tests and structure updates on the entire hierarchical tree, the final clustering results represented by the leaf nodes of the current clustering tree are output. These clustering results are the cell subset division results corrected by statistical tests and supported by significance.

[0055] Through the above steps, the present invention can quickly identify and merge the over - clustered cell clusters from a statistical perspective in complex clustering results, providing a more scientific and accurate tool for the interpretation of single - cell transcriptome data, effectively avoiding over - clustering and false positives.

[0056] Embodiment Two This embodiment is a single - cell clustering statistical significance test system based on manifold learning. The system is used to implement the significance test method as described in Embodiment One, and includes: A data preprocessing module, which is used to obtain original single-cell data and perform normalization processing on the original single-cell data to remove biases and noises; A dimensionality reduction module, which is used to perform dimensionality reduction on the data after normalization processing by using a manifold learning algorithm, and the manifold learning algorithm includes the PHATE algorithm; An initial clustering module, which is used to divide the dimensionality-reduced data according to pre-computed clustering labels and perform hierarchical clustering to construct a hierarchical clustering tree; A hypothesis testing module, which is used to traverse the hierarchical tree from bottom to top based on the Monte Carlo simulation method and perform a statistical significance hypothesis test on each pair of clusters to be compared; A clustering optimization module, which is used to merge the clustering clusters that pass the significance hypothesis test on the basis of the above tests and output the finally optimized clustering hierarchy as the correction result.

[0057] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A single-cell clustering statistical significance test method based on manifold learning, characterized in that: The following steps are involved: Acquire raw single-cell data, perform standardization on the raw single-cell data, and remove deviation and noise in the data; Use manifold learning methods to reduce the dimension of the standardized data; The dimension-reduced data is divided based on the pre-calculated clustering labels, and the divided data is hierarchically clustered to obtain a hierarchical tree; Monte Carlo simulation method is used to conduct hypothesis testing step by step from the bottom to the top of the hierarchical tree, and whether to merge clusters is determined based on the hypothesis test results; The completed hierarchical trees were finally merged as the single-cell clustering correction results.

2. The single-cell clustering statistical significance test method based on manifold learning according to claim 1, characterized in that: The specific method of hypothesis testing is: Merge the two cluster data to be compared, and perform Gaussian distribution fitting on the merged data; Randomly extract samples of the same size as the two clusters to be compared from the fitted Gaussian distribution, and perform hierarchical clustering to divide them into two sub-clusters; According to the preset number of simulations, the two subclusters are simulated, the variance increment of the two subclusters in each simulation is calculated, the null hypothesis distribution is constructed and the corresponding p-value is calculated; If the p-value is less than the set significance level, the two clusters to be compared are rejected. Otherwise, merge the two clusters to be compared.

3. The single-cell clustering statistical significance test method based on manifold learning according to claim 2, characterized in that: The calculation method of the variance increment is: For any cluster C, the total sum of squares is defined as follows: In the formula, represents the total sum of squares of cluster C, represents a sample point in cluster C, is the centroid of cluster C; Merge two subclusters and The resulting variance increase is expressed as follows: In the formula, Represents two subclusters and The variance increment of Represents two subclusters and The total sum of squares, Represents a subcluster The total sum of squares, Represents another subcluster The total sum of squares.

4. The method for statistical significance test of single cell clustering based on manifold learning according to claim 2, characterized in that: The construction of the null hypothesis distribution and calculation of the corresponding p-value are specifically as follows: Calculate the true variance increment of the two clusters to be compared; The number of times the variance increment of two subclusters in the statistical simulation distribution is greater than or equal to the true variance increment, and the ratio of the number to the total number of simulations is the p-value.

5. The method for statistical significance test of single-cell clustering based on manifold learning according to claim 2, characterized in that: The significance level is adaptively adjusted in each comparison by: In the formula, is the adaptively adjusted significance level, is the initial significance level, are the number of samples of the two clusters to be compared, is the total sample tree in the hierarchical tree structure data.

6. The method for statistical significance test of single-cell clustering based on manifold learning according to claim 5, characterized in that: The initial significance level was set at 0.

05.

7. The method for statistical significance test of single-cell clustering based on manifold learning according to claim 2, characterized in that: The number of simulations is 1000.

8. The method for statistical significance test of single cell clustering based on manifold learning according to claim 2, characterized in that: The Gaussian distribution fitting adopts the maximum likelihood estimation method.

9. The method for statistical significance test of single cell clustering based on manifold learning according to claim 1, characterized in that: The manifold learning adopts the PHATE algorithm.

10. Single-cell clustering statistical significance test system based on manifold learning, characterized by: The system is used to implement the significance testing method according to any one of claims 1 to 9, comprising: A data preprocessing module, used to obtain raw single-cell data and perform standardization on the raw single-cell data to remove deviation and noise; A dimension reduction module, used for reducing the dimension of the standardized data by using a manifold learning algorithm, wherein the manifold learning algorithm includes a PHATE algorithm; An initial clustering module that partitions the dimension-reduced data according to pre-computed cluster labels and performs hierarchical clustering to build a hierarchical clustering tree; Hypothesis testing module, which is used to traverse the hierarchical tree from bottom to top based on the Monte Carlo simulation method and perform statistical significance hypothesis tests on each cluster to be compared; The clustering optimization module is used to merge the clusters that pass the significance hypothesis test based on the above test, and output the final optimized clustering hierarchy as the correction result.

Citation Information

Patent Citations

  • Research method of immunopotentiated molecular subtype in early liver cancer

    CN112233796A

  • Clustering analysis method for single-cell omics data

    CN115527610A

  • Semi-supervised single cell RNA sequencing data clustering method based on iterative screening

    CN119763673A

Cited By

  • Cell subset division optimization method and device based on single cell clustering result

    CN121601027A