Metacell-based spatial transcriptome cell type annotation methods, apparatuses, and media

By generating meta-cells and using non-negative matrix factorization to calculate cell type weights, the computational complexity problem of large-scale single-cell data analysis is solved, and efficient and accurate cell type annotation is achieved.

CN120319326BActive Publication Date: 2025-10-10HANGZHOU LC BIOTECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510814919.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-10
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Existing single-cell data analysis tools have high computational complexity when processing large-scale data, and existing meta-cell inference methods lack accuracy and efficiency on large-scale data sets, making it difficult to efficiently annotate cell types.

Method used

By clustering single-cell data to generate meta-cells, and using the meta-cell feature count matrix and non-negative matrix decomposition, the cell type weights of each spatial position in the spatial omics data are calculated to achieve cell type annotation.

Benefits of technology

The computational complexity of the analysis process is greatly reduced, and the analysis efficiency and the accuracy and stability of the annotations are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120319326B_ABST
    Figure CN120319326B_ABST
Patent Text Reader

Abstract

The application discloses a metacell-based spatial transcriptome cell type annotation method, device and medium, and belongs to the technical field of bioinformatics. The method comprises the following steps: splitting single cell omics data with cell type annotation according to cell types, clustering cells of the same cell type and generating metacells, and obtaining a metacell feature count matrix; calculating marker features specific to each cell type relative to other cell types; calculating a basis matrix of the marker features of each cell type, and based on a feature count matrix of spatial omics data to be annotated, calculating a cell type weight value corresponding to each spatial position, and taking the cell type with the highest weight value as the cell type annotation of the spatial position. The spatial cell type annotation method, device and medium are used to greatly reduce the calculation amount of the analysis process, improve the analysis efficiency, and increase the annotation accuracy and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of bioinformatics, and in particular to a meta-cell-based spatial transcriptome cell type annotation method, device, and medium. Background Art

[0002] In recent years, spatial omics and single-cell omics sequencing technologies have developed rapidly and become popular. Single-cell omics technology can reveal the heterogeneity of gene expression or chromatin openness, protein kurtosis, etc. at the resolution of a single cell, while retaining relatively rich gene expression or other information. However, it requires the isolation of single cells through means such as tissue dissociation, resulting in the complete loss of the spatial location information of the cells. Spatial omics technology, on the other hand, can retain the in situ information of cells, combining gene expression or proteins with spatial coordinates, and includes various methods based on microdissection, spatial barcoding, or imaging (such as 10x Visium, MERFISH, Stereo-seq, 10x Visium HD and other products). However, spatial omics has defects such as low resolution or insufficient detection depth, resulting in the need to use single-cell omics data to assist in the identification of cells during its analysis process. Taking transcriptomes as an example, there are currently a variety of methods that use single-cell transcriptome data to assist in annotating spatial transcriptome cell types, but they are usually calculated using complete single-cell transcriptome data, which usually contains tens of thousands, hundreds of thousands, or even tens of millions of cells, resulting in huge analysis and calculation costs and a long time consumption; or single-cell data are randomly sampled and analyzed using a small number of cells, but conventional single-cell transcriptome expression profiles are sparse matrices and may contain large noise, which will make the analysis results easily affected by random factors (such as noise or failure to capture low-expression landmark genes), especially in low-quality data, which is prone to large errors.

[0003] To cope with the computational overhead brought by large-scale single-cell sequencing data, researchers have proposed a variety of efficient single-cell analysis tools, mainly used for tasks such as data interpolation, integration, clustering, and cell type annotation. However, these tools are usually designed specifically for specific tasks and are difficult to directly integrate into existing single-cell data analysis frameworks. In order to achieve more general and efficient single-cell data processing, one solution is to compress the raw data, thereby reducing data redundancy and enabling traditional analysis tools to process large-scale sequencing data more efficiently. For single-cell data compression, a representative method is metacell inference, which compresses several single cells into a single representative metacell by aggregating biologically similar cell populations, thereby effectively reducing the number of cells while retaining biological information to the greatest extent.

[0004] Metacell inference methods offer significant advantages in large-scale data processing. First, the data compression provided by metacells reduces the computational overhead of sequencing data analysis. Second, by aggregating cells with similar characteristics, metacells mitigate data sparsity, improving the robustness of downstream analyses (such as cell type annotation and developmental trajectory inference). However, while metacell inference methods have achieved promising results in some application scenarios, their accuracy and efficiency on large datasets remain limited. For example, the currently state-of-the-art SEACell algorithm clusters single cells by constructing a global adjacency matrix and inferring metacells based on the clustering results. While this algorithm achieves good results with smaller datasets, it takes more than a day to process 100,000 single cells and struggles with larger single-cell datasets due to its exponential memory overhead. In other words, existing metacell inference methods essentially shift the computational bottleneck from downstream analysis to the metacell inference stage, without truly addressing the computational complexity issue. Summary of the Invention

[0005] In order to solve at least one of the above technical problems, the technical solution adopted in this application is as follows.

[0006] The first aspect of the present application provides a method for annotating cell types based on spatial omics data of metacells, comprising the following steps:

[0007] S101, splitting the single-cell omics data with cell type annotations by cell type, clustering cells of the same cell type and generating meta-cells to obtain a meta-cell feature count matrix;

[0008] S102, calculating the marker features specific to each cell type relative to other cell types based on the meta-cell feature count matrix, and combining the marker features of all cell types to generate a marker feature set;

[0009] S103, calculating a basis matrix of marker features of each cell type based on the meta-cell feature count matrix;

[0010] S104, based on the feature count matrix of the spatial omics data to be annotated, using the base matrix, calculate the cell type weight value corresponding to each spatial position, and take the cell type with the highest weight value as the cell type annotation of the spatial position.

[0011] Metacells are a computational strategy that generates "supercellular" units by clustering single-cell data. Metacells aim to reduce the high noise and sparsity of single-cell data by aggregating similar cells while preserving biological heterogeneity.

[0012] In some embodiments of the present application, step S101 generates a meta-cell feature count matrix by the following steps:

[0013] S1011, for the single-cell omics data of the same cell type, perform dimensionality reduction processing on the cell-feature count matrix to obtain a cell dimensionality reduction matrix;

[0014] S1012, based on the cell dimensionality reduction matrix, calculate the distance between any two cells in the matrix, and cluster the cells based on the distance between the cells, and further calculate the cell with the smallest distance, that is, the most similar cell. k cells, of which k =10~30;

[0015] S1013, randomly select N cells, where the N cells do not overlap and are distributed in all clusters, where N = 100-1000, and for any feature of any cell in the N cells, calculate the feature of the cell and the cell most similar to the cell. k The average counts in cells are obtained, thereby obtaining the N feature average count matrix.

[0016] For each cell type, steps S1011 to S1013 are repeated to obtain the meta-cell feature count matrix.

[0017] In some embodiments of the present application, principal component analysis (PCA) is used to perform dimensionality reduction processing on the cell-feature count matrix.

[0018] In some specific embodiments of the present application, cells are clustered into 10 categories.

[0019] In this application, the value of N is adjusted according to the number of cells, usually taking 1% of the number of cells of the cell type; however, when the number of cells of the cell type is less than 10,000, N is set to 100; when the number of cells of the cell type is greater than 100,000, N is set to 1000.

[0020] In some embodiments of the present application, for any one cell among the N cells, the average count of each feature in the cell is calculated using the following formula:

[0021]

[0022] in, Indicates that the cell i The average count of features; Indicates that the cell i The count of features, Indicates that the cell j The most similar cell i The count of features, j =1~ k , thus obtaining the N feature average expression count matrix.

[0023] In some possible implementation schemes of the present application, in step S102, the average difference and p-value of each feature of each cell type relative to other cell types in the meta-cell feature count matrix are calculated, and in each cell type, the features are sorted from small to large according to the p-value. For features with the same p-value, they are sorted from large to small according to the average difference, and only the top 10 features of each cell type are retained as marker features.

[0024] In some embodiments of the present application, the average difference represents the average value of the difference in the count of the feature in the cell type relative to the count in other cell types, and is characterized by the difference size log2FC, that is, the value of the difference fold (FC) after Log2 processing, so the average difference is represented by avg_log2FC.

[0025] In some possible implementation schemes of the present application, step S103 is specifically as follows:

[0026] Eliminate all non-mark feature sets in the meta-cell feature count matrix to obtain the meta-cell mark feature matrix , and decompose to obtain the base matrix and the first coefficient matrix.

[0027] in, , m Indicates the number of signature features, represents the number of meta-cells,

[0028] Will Decompose into one dimensional basis matrix W and The first coefficient matrix of ,Right now ,in, i Indicates the type of cell.

[0029] In some specific embodiments of the present application, a non-negative matrix decomposition algorithm is used to Decompose into basis matrices W and the first coefficient matrix Non-negative Matrix Factorization (NMF) is an unsupervised learning algorithm that aims to transform a A non-negative matrix of dimension t is decomposed into two non-negative matrices (one -dimensional non-negative matrix and a The product of a 2-dimensional non-negative matrix (a matrix with a 2-dimensional matrix). By constraining all matrix elements to be non-negative, the decomposition results are interpretable and more relevant to biological scenarios. Furthermore, non-negative matrix factorization can capture key patterns in data, remove noise, and identify significant features.

[0030] In some embodiments of the present application, step S104 is specifically as follows:

[0031] all features in the set of non-marker features in the feature count matrix of the spatial omics data to be annotated are removed to obtain a spatial marker feature matrix a second coefficient matrix is obtained based on the spatial marker feature count matrix using the basis matrix, and the weight values of various cell types at each spatial position are represented using the second coefficient matrix.

[0032] wherein, , represents the number of spatial positions,

[0033] In some embodiments of the present application, the second coefficient matrix is calculated using the following formula ,

[0034] ,

[0035] wherein, is a matrix of dimension, representing the weight values of various cell types at each spatial position.

[0036] In some embodiments of the present application, in step S104, the second coefficient matrix is calculated using a least squares method .

[0037] In the present application, the omics data is selected from one of epigenomics data, transcriptomics data, and proteomics data.

[0038] In some embodiments of the present application, the omics data is transcriptomics data, the features are genes, and correspondingly, the feature count is gene expression count.

[0039] The second aspect of the present application provides a computer device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the cell type annotation method for spatial omics data based on metacells according to any one of the first aspect of the present application.

[0040] The third aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the cell type annotation method for spatial omics data based on metacells according to any one of the first aspect of the present application.

[0041] Compared with the prior art, the present application has the following beneficial effects:

[0042] The method, device and medium of the present application, by clustering single cells to generate meta-cells, and performing cell type feature calculation and non-negative matrix decomposition based on the meta-cells, calculate the cell type rights of each spatial position in spatial omics, and further obtain the cell type annotation results of each spatial position, greatly reduce the computational complexity of the analysis process, improve the analysis efficiency, and increase the accuracy and stability of the annotation.

[0043] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The above and other objects, features and advantages of the exemplary embodiments of the present application will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present application are shown in an illustrative and non-limiting manner, in which:

[0045] Figure 1 Schematic diagram of the implementation process of the method for annotating cell types based on spatial transcriptomics data of metacells in Example 1 of the present application is shown;

[0046] Figure 2 The figure shows the distribution of astrocyte weight values ​​in spatial positions in Example 2 of the present application;

[0047] Figure 3 The figure shows the distribution of endothelial cell weight values ​​in spatial positions in Example 2 of the present application;

[0048] Figure 4 The figure shows the distribution of the weight values ​​of neurons in spatial positions in Example 2 of the present application;

[0049] Figure 5 The figure shows the distribution of the weight values ​​of microglial cells in spatial positions in Example 2 of the present application;

[0050] Figure 6 The figure shows the spatial distribution of oligodendrocyte weight values ​​in Example 2 of the present application;

[0051] Figure 7 The cell type identification results for each spatial location in Example 2 of the present application are shown. DETAILED DESCRIPTION

[0052] Unless otherwise indicated, implied from the context, or customary in the art, all parts and percentages in this application are based on weight, and the test and characterization methods used are current as of the filing date of this application. Where applicable, the contents of any patents, patent applications, or publications referred to in this application are incorporated herein by reference in their entirety, and their equivalent patent families are also incorporated by reference, particularly for definitions of relevant terms in the art disclosed therein. If the definition of a specific term disclosed in the prior art is inconsistent with any definition provided in this application, the definition of the term provided in this application shall prevail.

[0053] In order to make the technical problems, technical solutions and beneficial effects solved by the present application clearer and more understandable, the present application is further described in detail below in conjunction with the embodiments.

[0054] The following examples are provided herein to illustrate preferred embodiments of the present application. Those skilled in the art will appreciate that the techniques disclosed in the following examples represent techniques discovered by the inventors that can be used to implement the present application and, therefore, can be considered preferred embodiments of the present application. However, those skilled in the art will appreciate, based on this specification, that many modifications may be made to the specific embodiments disclosed herein while still achieving the same or similar results without departing from the spirit or scope of the present application.

[0055] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs, and the disclosures and materials cited therein are hereby incorporated by reference.

[0056] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many technical equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the claims.

[0057] The experimental methods in the following examples, unless otherwise specified, are all conventional methods. The instruments and equipment used in the following examples, unless otherwise specified, are all conventional laboratory instruments and equipment; the experimental materials used in the following examples, unless otherwise specified, are all purchased from conventional biochemical reagent stores.

[0058] Example 1 Meta-cell-based spatial transcriptome cell type annotation method

[0059] This embodiment provides a method for annotating cell types using a spatial transcriptome based on metacells. The flowchart is as follows: Figure 1 shown.

[0060] It mainly includes steps S101 to S104, which are described in detail as follows:

[0061] S101: Split the single-cell transcriptome data containing cell type annotations by cell type, cluster cells of the same cell type, and generate meta-cells. This is achieved through the following steps S1011 to S1014:

[0062] S1011. Single-cell transcriptome data of the same cell type (e.g., T cells) were selected. The original cell-gene expression count matrix (shown in Table 1) was normalized, standardized, highly variable genes were calculated, and PCA dimensionality reduction was performed using the NormalizeData, FindVariableFeatures, ScaleData, and RunPCA functions of the Seurat software to obtain the cell-PCA dimensionality reduction matrix (Table 2).

[0063] Table 1 Original cell-gene expression count matrix (partial)

[0064]

[0065] Table 2 Cell-PCA dimensionality reduction matrix (partial)

[0066]

[0067] S1012, based on the cell-PCA dimensionality reduction matrix, calculate the Euclidean distance between each cell in the matrix. Based on the Euclidean distance between cells, use the k-means algorithm to cluster the cells into 10 categories, and use the KNN (k-nearest neighbor) algorithm to calculate the most similar (i.e., the smallest Euclidean distance) of each cell. k cells ( k It is a hyperparameter that can be adjusted according to the number of cells and set to any integer greater than 1, usually set to 25).

[0068] S1013, generate meta-cells. Set the expected number of meta-cells N (N ranges from 100 to 1000, taking 1% of the number of cells of the cell type. However, when the number of cells of the cell type is less than 10,000, N is set to 100; when the number of cells of the cell type is greater than 100,000, N is set to 1000). Randomly select N cells from the matrix (N cells do not overlap with each other and must be distributed in the 10 clusters obtained in step S1012). Among the N cells, each cell is compared with its most similar cell. k The original gene expression count matrices of the cells were merged, as shown in Table 3.

[0069] Table 3 Randomly selected cells and their nearest neighbors k Original expression profile of cells (partial)

[0070]

[0071] The average expression count of each gene in randomly selected cells was calculated using the following formula:

[0072]

[0073] in, Indicates the first i The average expression count of each gene; Indicates the first i The expression counts of genes, Represents the randomly selected cell j The most similar cell i The expression counts of genes, j =1~ k .

[0074] Thus, N cells generate N gene average expression count matrices, and each gene average expression count matrix is ​​the meta-cell gene expression count matrix, that is, N meta-cell gene expression count matrices are generated, as shown in Table 4.

[0075] Table 4. Gene expression count matrix of N meta-cells (partial)

[0076]

[0077] S1014, repeat steps S1011 to S1013 for the single-cell transcriptome data of each cell type, that is, obtain N meta-cell gene expression count matrices for each cell type, and after merging, obtain the complete meta-cell gene expression count matrix corresponding to the single-cell transcriptome data of this embodiment, that is, the meta-cell count matrix.

[0078] S102: After normalizing the meta-cell count matrix and correcting the sequencing depth, the marker genes of each cell type are calculated.

[0079] The matrix was normalized and corrected for sequencing depth using the NormalizeData function in Seurat software. The resulting matrix is ​​referred to as the meta-cell data matrix. The FindAllMarkers function in Seurat software and the Wilcox difference test were used to calculate the avg_log2FC (log2 fold difference) and p-value (i.e., significance level) for each gene in each cell type within the meta-cell data matrix. Essentially, a Wilcox rank sum test was performed on each gene in a single cell type compared to all other cell types to identify genes that were highly expressed specifically in each cell type. Within each cell type, genes were sorted by p-value in ascending order. For genes with the same p-value, genes were sorted by avg_log2FC in descending order. Only the top 10 marker genes for each cell type were retained (as shown in Table 5), and marker genes from all cell types were combined to create a marker gene set.

[0080] Table 5 TOP10 marker genes for each cell type (partial)

[0081]

[0082] S103, remove all genes in the meta-cell data matrix that are not marker genes, which is called the meta-cell marker gene data matrix. dimensional matrix ,in m Indicates the number of marker genes, Represents the number of meta-cells. Use the non-negative matrix factorization algorithm to transform the matrix Decompose into one dimensional basis matrix W and The coefficient matrix of ,Right now .in, i Indicates the type of cell.

[0083] S104, the spatial position-gene expression count matrix of the spatial transcriptome data is normalized and corrected for sequencing depth using the NormalizeData function of the seurat software, and all genes that are not in the marker gene set are also removed to generate a spatial marker gene data matrix, i.e., a dimensional matrix ,in Represents the number of spatial locations. Using non-negative least squares, calculate The coefficient matrix in .in for dimensional matrix, representing the weight values ​​of various cell types at each spatial position. For each spatial position, the cell type with the highest weight value is used as the cell type annotation of that spatial position.

[0084] Example 2 Application of the Metacell-Based Spatial Transcriptome Cell Type Annotation Method

[0085] This example provides an application of a meta-cell-based spatial transcriptome cell type annotation method. Single-cell transcriptome data are used to annotate high-definition spatial transcriptome data of the whole mouse brain with cell types, as described in Example 1.

[0086] The original data of the spatial transcriptome are shown in Table 6. Then, the NormalizeData function of the seurat software was used for normalization and sequencing depth correction. All genes in the non-marker gene set were also removed to generate the spatial marker gene data matrix, that is, the matrix As shown in step S104, the weight values ​​of various cell types at each spatial position are calculated using non-negative least squares method (as shown in Table 7 and Figures 2 to 6 ), and the cell type with the highest weight value is used as the cell type annotation of the spatial position (as shown in Table 8 and Figure 7 ).

[0087] Table 6 Original spatial position-gene expression count matrix of spatial transcriptome

[0088]

[0089] Table 7 Weight values ​​of various cell types at each spatial position

[0090]

[0091] Table 8 Identification results of cell types at each spatial location

[0092]

[0093] All documents mentioned in this application are incorporated herein by reference, just as if each document were incorporated herein by reference individually. It should also be understood that after reading the above teachings of this application, those skilled in the art may make various changes or modifications to this application, and that such equivalents also fall within the scope of the claims appended hereto.

Claims

1. A method for annotating cell types based on spatial omics data of metacells, characterized by: The following steps are involved: S101, splitting the single-cell omics data with cell type annotations by cell type, clustering cells of the same cell type and generating meta-cells to obtain a meta-cell feature count matrix; S102, calculating the marker features specific to each cell type relative to other cell types based on the meta-cell feature count matrix, and combining the marker features of all cell types to generate a marker feature set; S103, calculating a basis matrix of marker features of each cell type based on the meta-cell feature count matrix; S104, based on the feature count matrix of the spatial omics data to be annotated, using the base matrix, calculate the cell type weight value corresponding to each spatial position, and take the cell type with the highest weight value as the cell type annotation of the spatial position, Step S103 is specifically as follows: The non-negative matrix decomposition algorithm is used to remove all non-mark feature sets in the meta-cell feature count matrix to obtain the meta-cell mark feature matrix , , m Indicates the number of signature features, represents the number of meta-cells, Will Decompose into one dimensional basis matrix W and The first coefficient matrix of ,Right now ,in, i Indicates the type of cell, Step S104 is specifically as follows: Eliminate all non-mark feature sets in the feature count matrix of the spatial omics data to be annotated to obtain the spatial marker feature count matrix ,in, , represents the number of spatial locations, Based on the spatial landmark feature count matrix , calculate the second coefficient matrix using the following formula , , in, for dimensional matrix, representing the weight values ​​of various cell types at each spatial position.

2. A method for annotating cell types based on meta-cell spatial omics data according to claim 1, characterized in that: Step S101 is implemented by the following steps: S1011, for the single-cell omics data of the same cell type, perform dimensionality reduction processing on the cell-feature count matrix to obtain a cell dimensionality reduction matrix; S1012, based on the cell dimensionality reduction matrix, calculate the distance between any two cells in the matrix, and cluster the cells based on the distance between the cells, and further calculate the cell with the smallest distance, that is, the most similar cell. k cells, of which k =10~30; S1013, randomly select N cells, where the N cells do not overlap and are distributed in all clusters, where N = 100-1000, and for any feature of any cell in the N cells, calculate the feature of the cell and the cell most similar to the cell. k The average counts in cells are obtained, thus obtaining the N feature average count matrix, For each cell type, steps S1011 to S1013 are repeated to obtain the meta-cell feature count matrix.

3. The method for annotating cell types based on meta-cell spatial omics data according to claim 1, characterized in that: In step S102, the average multiple and p-value of each feature of each cell type relative to other cell types in the meta-cell feature count matrix are calculated. In each cell type, the features are sorted from small to large according to the p-value. For features with the same p-value, they are sorted from large to small according to the average difference. Only the top 10 features of each cell type are retained as marker features.

4. A method for annotating cell types based on meta-cell spatial omics data according to any one of claims 1 to 3, characterized in that: The omics data is selected from epigenomic data, transcriptomics data, and proteomics data.

5. A computer device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method according to any one of claims 1 to 4 when executing the computer program.

6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Spatial transcriptome cell clustering and analyzing method

    CN114091603A

  • Single-cell RNA sequencing data clustering method based on Transform model prediction

    CN118380048A