A spatial transcriptome-based method and system for unbiased cell type localization

By employing Moran's index screening and spatial clustering methods, combined with spatial transcriptomics and single-cell data, the bias problem of deconvolution algorithms in spatial transcriptomics is addressed, achieving unbiased localization of cell types and improving the accuracy of spatial transcriptomics analysis.

CN116364180BActive Publication Date: 2026-02-10NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310339311.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2026-02-10
Estimated Expiration
2043-03-30

AI Technical Summary

Technical Problem

Existing deconvolution algorithms are biased in spatial transcriptomics, failing to accurately identify spatial clustering patterns of most cell types, resulting in inaccurate spatial distribution of cell types.

Method used

A cell type unbiased localization method based on Moran's index was adopted. Through spatial clustering, deconvolution and enrichment score calculation, cell types with spatial heterogeneity were screened and matched with tissue regions. Multiple algorithms were used to integrate spatial transcriptome data and single-cell data to achieve unbiased localization of cell types.

Benefits of technology

It enables spatial aggregation pattern recognition of most cell types, improves the accuracy of cell type localization in space, and provides a foundation for further research on cell heterogeneity and interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116364180B_ABST
    Figure CN116364180B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of spatial transcriptome sequencing data analysis, and provides a cell type unbiased positioning method and system based on spatial transcriptome, which comprises the following steps: dividing sampling points in spatial transcriptome data into several categories through a spatial clustering method; obtaining the proportion of cell types of each sampling point through a deconvolution method, and taking the proportion of a certain cell type in a certain sampling point as the enrichment score of the cell type at the sampling point; screening out cell types with spatial heterogeneity based on the enrichment score; for each cell type with spatial heterogeneity, screening out sampling points with spatial heterogeneity based on the enrichment score, obtaining a spatial heterogeneity point set, and matching the spatial heterogeneity point set with all tissue regions to obtain the best matching tissue region. Most of the cell types in the spatial transcriptome can find their spatial aggregation patterns, and the co-localization of cell types in space can be discovered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of spatial transcriptome sequencing data analysis technology, and particularly relates to a cell type unbiased localization method and system based on spatial transcriptome. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Spatial transcriptomics is an interdisciplinary field combining life sciences and computer science. Breakthroughs in this area have led to new discoveries in the study of diseases and biological processes. However, current sequencing technologies have limitations: spatial transcriptomics can measure the location of transcript production, but with low resolution, meaning it cannot distinguish which cell produced the transcript. Single-cell sequencing, on the other hand, can obtain transcripts (mRNA) at single-cell resolution, but the necessary dissociation step results in the loss of location information.

[0004] Some analytical tools propose deconvolutional methods to integrate single-cell and spatial transcriptome data, treating cells at each sampling point as a mixture of multiple cell types. It builds a model based on cell type expression data from single-cell sequencing data, takes experimental data from each sampling point in the spatial transcriptome as input, and outputs a maximum a posteriori estimate of the distribution of cell subpopulations within a tissue region given the gene expression distribution at a given sampling point.

[0005] However, spatial heterogeneity of cell types is widespread, and the spatial distribution of cell types obtained by deconvolution algorithms mostly does not have a certain clustering pattern, which is caused by the bias of current deconvolution algorithms. Summary of the Invention

[0006] To address the technical problems mentioned above, this invention provides a cell type unbiased localization method and system based on spatial transcriptomics, enabling most cell types in the spatial transcriptome to find their spatial aggregation patterns, thereby further discovering cell type co-localization in space.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] The first aspect of the present invention provides a cell type unbiased localization method based on spatial transcriptomics, comprising:

[0009] Acquire spatial transcriptomic data from tissue sections;

[0010] Spatial clustering methods are used to divide the sampling points in the spatial transcriptome data into several categories, with each category representing a tissue region in a tissue slice;

[0011] Based on spatial transcriptome data, the proportion of cell types at each sampling point is obtained by deconvolution, and the proportion of a certain cell type at a certain sampling point is used as the enrichment score of that cell type at that sampling point.

[0012] Based on the enrichment score, a global Moran index is calculated for each cell type to screen out cell types with spatial heterogeneity;

[0013] For each cell type with spatial heterogeneity, sampling points with spatial heterogeneity are selected based on enrichment scores to obtain a spatial heterogeneous point set. The spatial heterogeneous point set is then matched with all tissue regions to obtain the best matching tissue region.

[0014] Furthermore, the step of matching the spatially heterogeneous point set with all tissue regions is as follows:

[0015] (1) The set of points consisting of all the sampling points contained in an organization region is called an organization region point set;

[0016] (2) For a certain cell type, calculate the F1 score of the spatial heterogeneity point set of the cell type and the point set of each tissue region, and select the tissue region corresponding to the highest F1 score as the best matching tissue region for the cell type.

[0017] (3) For a certain cell type, if there are multiple best matching tissue regions, each best matching tissue region is expanded as a new tissue region and the process returns to step (2).

[0018] Furthermore, the method for expanding the best-matching organizational region is as follows: for a certain best-matching organizational region, randomly select one of the organizational regions other than the best-matching organizational region and attach it to the best-matching organizational region.

[0019] Furthermore, for a certain spatially heterogeneous cell type, the spatially heterogeneous sampling point i satisfies:

[0020]

[0021] Where, x i This represents the enrichment fraction of the spatially heterogeneous cell type at sampling point i. w is the average of all enrichment scores. ij Z represents the distance between sampling point i and its neighboring sampling point j, where n represents the number of neighboring sampling points of sampling point i. i pval represents the significance of the local Moran index at sampling point i for this spatially heterogeneous cell type, pval represents calculating the p value, and ζ is a set value.

[0022] Furthermore, the spatial clustering method is one or more of the following: Leiden algorithm, BayesSpace method, SpaGCN method, and stlearn method.

[0023] Furthermore, the deconvolution method is one or more of the following: SPOTlight method, spacexr method, and stereoscope method.

[0024] Furthermore, it also includes: calculating the Pearson correlation coefficient between each pair of cell types based on the best-matching tissue regions.

[0025] A second aspect of the present invention provides a cell type unbiased localization system based on spatial transcriptomics, comprising:

[0026] The data acquisition module is configured to acquire spatial transcriptome data from tissue slices.

[0027] The clustering module is configured to divide the sampling points in the spatial transcriptome data into several categories using a spatial clustering method, where each category represents a tissue region in a tissue slice.

[0028] The deconvolution module is configured to: obtain the proportion of cell types at each sampling point based on spatial transcriptome data through deconvolution, and use the proportion of a certain cell type at a certain sampling point as the enrichment score of that cell type at that sampling point;

[0029] The screening module is configured to: calculate the global Moran index for each cell type based on the enrichment score, and screen out cell types with spatial heterogeneity;

[0030] The localization module is configured to: for each cell type with spatial heterogeneity, filter out sampling points with spatial heterogeneity based on enrichment scores to obtain a spatial heterogeneous point set, and match the spatial heterogeneous point set with all tissue regions to obtain the best matching tissue region.

[0031] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the cell type unbiased localization method based on spatial transcriptomics as described above.

[0032] A fourth aspect of the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the cell type unbiased localization method based on spatial transcriptomics as described above.

[0033] Compared with the prior art, the beneficial effects of the present invention are:

[0034] This invention provides a cell type unbiased localization method based on spatial transcriptomics, which develops a cell type spatial localization algorithm based on Moran's index, overcoming the bias of previous deconvolution tools by selecting the best match between the spatial domain and the cell type.

[0035] This invention provides a cell type unbiased localization method based on spatial transcriptomics, which integrates the main analysis workflow of spatial transcriptomics, including data preprocessing, spatial clustering and deconvolution, and can automatically realize joint analysis workflow, meeting the analysis needs of the field of spatial transcriptomics; the tools can be freely configured in each processing step. Attached Figure Description

[0036] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0037] Figure 1 This is a flowchart of a cell type unbiased localization method based on spatial transcriptomics according to Embodiment 1 of the present invention. Detailed Implementation

[0038] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0039] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0040] Terminology Explanation:

[0041] Spatial heterogeneity refers to the non-uniformity of cell types in spatial distribution and their aggregation patterns.

[0042] Bias refers to the fact that a deconvolution algorithm can effectively identify the spatial heterogeneity of some cell types, but cannot for most other cell types.

[0043] Example 1

[0044] This embodiment provides a cell type unbiased localization method based on spatial transcriptomics, such as... Figure 1 As shown, it includes the following steps:

[0045] S1, Data Preprocessing: Obtain spatial transcriptome data from tissue sections and preprocess the spatial transcriptome data.

[0046] As one implementation method, if the spatial transcriptome data is 10X Visium sequencing data, then the 10X Visium sequencing data is read and preprocessed.

[0047] Spatial transcriptome data include: spatial expression matrix (with sampling points as rows and gene expression as columns, where each element in the matrix represents the expression level of a gene at the sampling point), location information of the sampling points in tissue sections, and histological images (tissue sections).

[0048] 10X Visium sequencing data includes: expression matrices consisting of gene expression data from sampling points on tissue sections, tissue section images at different resolutions, and the specific location of each sampling point in the tissue section image.

[0049] In this embodiment, reading 10X Visium sequencing data specifically includes: using computer instructions to read the spatial transcriptome data expression matrix, sampling point information, gene information, tissue slice images at different resolutions, the specific location of each sampling point in the tissue slice image, and the scaling factor between the original high-resolution image and the low-resolution image from the hard disk for subsequent analysis.

[0050] In this embodiment, the spatial transcriptome data is preprocessed, specifically including: filtering out sampling points in the spatial transcriptome data that do not appear on the tissue; filtering out genes that are lowly expressed in all sampling points and mitochondrial genes; performing normalization, standardization and principal component analysis on the expression matrix in sequence; and screening out genes with high expression.

[0051] Specifically, preprocessing of spatial transcriptome data includes:

[0052] (1) Filter out sampling points that do not appear on the tissue, genes that are lowly expressed in all sampling points, and mitochondrial genes; among them, mitochondrial genes refer to the genetic information carried in mitochondria, which are marked with the prefix "MT-" in the gene information. Spatial transcriptome data refers to an expression matrix with sampling points as rows and gene expression as columns, where each element in the matrix represents the expression level of a gene at the sampling point.

[0053] (2) Screen out the genes with high expression and obtain the expression matrix after screening;

[0054] (3) The normalized expression matrix is ​​obtained by standardizing and normalizing the filtered expression matrix, as follows:

[0055] X = log(X + 1)

[0056]

[0057] Where X represents the expression matrix after screening, μ represents the average expression value of each gene, and σ represents the variance of the expression value of each gene. Standardization was performed using log normalization, and normalization was performed using Z-score normalization.

[0058] (4) Perform principal component analysis on the normalized expression matrix, specifically including:

[0059] Calculate the covariance matrix S of the sample (a dataset consisting of several sampling points):

[0060]

[0061] in n represents the number of sampling points. For a matrix S, there is a set of eigenvectors v. By orthogonalizing and normalizing this set of vectors, we can obtain a set of orthogonal unit vectors.

[0062] Eigenvalue decomposition is performed, and S is decomposed into the following equation:

[0063] S=PΣP -1

[0064] Where P is a matrix composed of the eigenvectors of matrix S, and Σ is a diagonal matrix whose diagonal elements are the eigenvalues. The eigenvalues ​​are sorted in descending order, and the k largest eigenvalues ​​c1, c2, ..., c3 are selected. k Then, the corresponding k eigenvectors are used as row vectors to form an eigenvector matrix P0. Finally, the data is transformed into a new space constructed from the k eigenvectors:

[0065] X′=P0X

[0066] In this embodiment, as Figure 1 As shown, FFPE mouse kidney tissue slices using 10X Visium sequencing technology were used as a spatial transcriptome dataset. The expression matrix of the samples was preprocessed to obtain 2264 sampling points with 19465 gene types in the tissue.

[0067] S2, Spatial Clustering: Using the preprocessed spatial representation matrix, the location information of sampling points in tissue slices, and histological images, the sampling points in the spatial transcriptome data are divided into several categories by clustering methods. Each category represents a tissue region in the tissue slice image.

[0068] Using the preprocessed spatial expression matrix, the location information of sampling points in tissue sections, and histological images, the sampling points in the spatial transcriptome data were divided into several tissue regions using clustering methods, specifically including:

[0069] For each row of the spatial representation matrix, a community clustering algorithm is used. Specifically, the Leiden algorithm is used to cluster communities, treat tissue slices as communities and sampling points as nodes. By calculating the relationship between nodes and edges, the spatial domain is quickly found and refined. Then, network aggregation is performed based on the refined partitions. The non-refined partitions are used to create initial partitions for the aggregated network. Finally, the iteration steps are repeated until convergence.

[0070] Alternatively, the genes in the spatial expression matrix can be dimensionality reduced, and then each dimension after dimensionality reduction can be modeled and clustered using a multivariate t-distribution model. Finally, the parameters can be updated. Specifically, the BayesSpace method is used to first reduce the dimensionality of the genes in the decontaminated spatial expression matrix, then model and cluster each dimension after dimensionality reduction using a multivariate t-distribution model, and finally update the parameters using the Metropolis-Hastings algorithm, which uses the two-dimensional spatial information integrated by the Potts model as its prior distribution.

[0071] Alternatively, the SpaGCN method can be used to integrate spatial location information and histological imaging information (H&E staining map), calculate the physical distance between each sampling point in the preprocessed spatial expression matrix, and use a graph convolutional neural network to integrate distance information with gene expression levels. Finally, based on the output of the graph convolutional network, an unsupervised deep embedding clustering analysis method is used to divide the sampling points in the spatial transcriptome data into several tissue regions.

[0072] Alternatively, the stlearn method can be used to normalize gene expression by analyzing the H&E staining map region and neighborhood information corresponding to each sampling point, followed by unsupervised clustering.

[0073] The spatial clustering steps in this embodiment provide the above four methods. Users can choose to use one or more methods for spatial clustering to make full use of two-dimensional spatial location information and histological images to obtain multiple tissue region segmentation results as candidates for spatial domains.

[0074] In this embodiment, spatial clustering is performed on the preprocessed 10X Visium sequencing data. Specifically, this includes: using the Leiden algorithm to partition the spatial domains, resulting in 19 spatial domains (organ regions), denoted as Leiden spatial domain candidates; using the BayesSpace algorithm to partition the spatial domains, resulting in 10 spatial domains, denoted as BayesSpace spatial domain candidates; using the SpaGCN algorithm to partition the spatial domains, resulting in 11 spatial domains, denoted as SpaGCN candidates; and using the stlearn algorithm to partition the spatial domains, resulting in 16 spatial domains, denoted as stlearn candidates. A total of four spatial domain candidates are obtained.

[0075] S3, Deconvolution: Based on the expression of cell types in the single-cell dataset, the gene expression of all sampling points is used as input, and one or more methods can be selected as needed to obtain the distribution of cell types in various tissue regions.

[0076] The deconvolution method treats cells at each sampling point as a mixture of multiple cell types. Based on the expression patterns of cell subpopulations in the single-cell dataset, it uses the gene expression of all cells at the sampling point as input to obtain the distribution (enrichment score) of cell types in various tissue regions. Specifically, this includes:

[0077] The SPOTlight method was adopted, which integrates spatial data with single-cell reference data, initializes the data using non-negative matrix factor regression, and then uses non-negative least squares to obtain the proportion of cell types at each sampling point based on the marker genes and spatial gene expression data of each cell type.

[0078] Alternatively, the SpaceXR method can be used. This method treats spatial transcriptome data as conforming to a Poisson distribution. Given the provided single-cell data, it reads the average expression of all genes in each cell type, and then uses a stepwise fitting method to find one or two cell types that best match the gene expression of the sampling point. Then, it uses the maximum likelihood estimation method to fit the parameters of the Poisson distribution, thereby inferring the proportion of cell types in the sampling point.

[0079] Alternatively, the stereoscope method can be used, which treats both single-cell reference expression data and spatial transcriptome data as conforming to a negative binomial distribution. Given the provided single-cell data, the values ​​of specific parameters of the cell type distribution are obtained by searching for the maximum likelihood estimate (MLE). Finally, based on the observed spatial data, the proportion of cell types at each sampling point is estimated using the prior distribution of cell types.

[0080] The proportion of a certain cell type in a certain sampling point is used as the enrichment score of that cell type at that sampling point.

[0081] The deconvolution step in this embodiment provides the above three methods. Users can choose to use one or more methods to perform deconvolution and obtain the spatial localization of each cell type. Each method provides a set of cell type spatial localization schemes, called candidates for enrichment scores, that is, to obtain candidates for spatial enrichment scores of multiple cell types.

[0082] In this embodiment, the Tabula single-cell dataset, containing 7 cell types, is introduced. The preprocessed 10X Visium sequencing data is deconvolved, specifically including: deconvolution using the SPOTlight algorithm, with the resulting enrichment scores filtered using the global Moran index to identify spatial enrichment scores for 7 cell types, designated as SPOTlight enrichment score candidates; deconvolution using the spacexr algorithm to obtain spatial enrichment scores for the 7 cell types, designated as spacexr spatial domain candidates; and deconvolution using the StereoScope algorithm to obtain spatial localization for the 7 cell types, designated as StereoScope spatial domain candidates. A total of three enrichment score candidates are obtained.

[0083] S4, spatial localization of cell types, includes: screening for spatially heterogeneous cell types using the global Moran index; for each cell type, finding cell type enrichment regions using the local Moran index; finding the optimal spatial domain match using a greedy strategy; testing for differentially expressed genes at the optimal match using hypergeometric distribution; and calculating the Pearson correlation coefficient between each pair of cell types using the optimal spatial localization of cell types.

[0084] (1) For each candidate cell type enrichment score obtained by the deconvolution method, calculate the global Moran index.

[0085] For each cell type enrichment score obtained by the deconvolution method, calculate the global Moran index. Assume a given cell type has an enrichment score of x at each sampling unit i. i The mean enrichment score is The global Moran index I is then calculated using the following formula:

[0086]

[0087] Where N represents the number of sampling points; w ij w represents the distance between sampling points i and j. When two sampling units are adjacent, w ij =1, otherwise w ij =0; W represents all w ij The sum of the values ​​of I and I is given. The range of I is [-1, 1]. The larger the value, the greater the spatial clustering of the cell type. Preferably, when the global Moran index I value is greater than 0.5, the spatial heterogeneity of the cell type can be considered.

[0088] (2) Using the local Moran index, the best match between cell type and spatial domain is found, including:

[0089] (201) For a cell type with spatial heterogeneity, calculate the local Moran index of its enrichment fraction at each sampling point to find the sampling points with spatial heterogeneity.

[0090] For a cell type exhibiting spatial heterogeneity, calculate its enrichment fraction x. i The local Moran's index at each sampling point. The formula for calculating the local Moran's index is as follows:

[0091]

[0092] in, Use the Z-test to determine the significance of the local Moran index:

[0093]

[0094] Since spatial transcriptomics primarily explores cell aggregation in tissues, sampling points that meet the following three conditions are included in the spatial heterogeneity point set P:

[0095]

[0096] The first condition indicates that the enrichment score of i is greater than the average; the second condition indicates that the average enrichment score of the surrounding sampling points is also greater than the average; and the third condition indicates that sampling points with an adjusted p-value (pval) less than the set value ζ (taken as 0.05) are selected as points with spatial heterogeneity. Where x i This represents the enrichment fraction of the spatially heterogeneous cell type at sampling point i. w is the average of all enrichment scores. ij Z represents the distance between sampling point i and its neighboring sampling point j, where n represents the number of neighboring sampling points of sampling point i. i pval represents the significance of the local Moran index at sampling point i for this spatially heterogeneous cell type, pval represents calculating the p value, and ζ is a set value.

[0097] (202) For each spatial clustering method, a candidate spatial domain segmentation scheme will be given, assuming the tissue is segmented into k regions. For the spatially heterogeneous point set P of cell type i with spatial heterogeneity. i and each spatial domain point set Q1, Q2, ..., Q k The F1 score is calculated as follows:

[0098]

[0099]

[0100]

[0101] Among them, Q j Let Q represent the set of points containing all sampled points within the j-th tissue region (called the j-th spatial domain point set or the j-th tissue region point set). β represents the weight between precision p and recall r, defaulting to 1. Based on the F1-score, the most suitable spatial domain point set Q for cell type i can be found. d(i) :

[0102] d(i) = argmax(F1score(Q) j )), 1≤j≤k

[0103] However, cell type i may be enriched in multiple tissue regions, meaning multiple regions should match a single cell type. Therefore, a greedy strategy is implemented to expand the optimal match. Other regions Q are repeatedly considered. j Attach to Q d(i) The F1 score is then recalculated to determine whether Q is included. j :

[0104]

[0105]

[0106] in, This represents the initial set of points for the iteration. Describes the set of points in the nth iteration. The condition for the iteration to end is It no longer changes with the increase of n; this operation means that if the region Q is increased... j To increase the F1 score, these areas are added step by step. Up, until If it stops changing, it indicates that the best match expansion is complete.

[0107] In summary, the steps for matching spatially heterogeneous point sets with all organizational regions are as follows:

[0108] (1) The set of points consisting of all the sampling points contained in an organization region is called an organization region point set;

[0109] (2) For a certain cell type, calculate the F1 score of the spatial heterogeneity point set of the cell type and the point set of each tissue region, and select the tissue region corresponding to the highest F1 score as the best matching tissue region for the cell type.

[0110] (3) For a certain cell type, if there are multiple best matching tissue regions, each best matching tissue region is expanded as a new tissue region and the process returns to step (2).

[0111] The method for expanding the best-matching organization region is as follows: for a certain best-matching organization region, randomly select an organization region other than the best-matching organization region and attach it to the best-matching organization region.

[0112] (203) These sampling points are optimally matched with the spatial domain candidates, and the differentially expressed genes at the optimal match are tested using hypergeometric distribution. That is, the differentially expressed genes of cell type and spatial domain at the optimal match are found using hypergeometric distribution test.

[0113] Among these methods, the hypergeometric distribution test was used to identify differentially expressed genes for cell types and spatial domains at the optimal match, specifically including:

[0114] After finding the best matching tissue region d′(i) for cell type i, Q is detected. d′(i) Abundance of cell types on Q. Investigating the abundance of cell types on Q. d′(i) Differentially expressed genes in cell type i and cell type i (adjusted p-value < 10 in Wilcoxon rank-sum test) -5 (Cohen's d effect size > 0.2). Then, the intersection between the differentially expressed genes was found, and the significance of the intersection was tested using the hypergeometric cumulative distribution:

[0115]

[0116] Where N represents the total number of background genes; m and n represent the number of genes in the point set Q. d′(i) The number of differentially expressed genes in cell type i and cell type i, where k represents the size of the intersection.

[0117] (204) Explore co-localization of cell types in space using Pearson correlation. That is, calculate the Pearson correlation coefficient between each pair of cell types using the optimal spatial localization of cell types. Specifically, add a spatial lag term to the cell type enrichment score. For each sampling point k and each cell type i, let y′ ik Given the original cell type composition fractions, the Pearson correlation r between the two cell types u and v is calculated as follows:

[0118]

[0119] Hierarchical clustering was applied to the correlation matrix to reveal colocalization of cell types and to visualize significant correlations among cell types (adjusted p-value < 0.05).

[0120] In this embodiment, the spatial localization of cell types specifically includes:

[0121] Taking StereoScope candidates as an example, among the seven cell types, spatial heterogeneity was found in six cell types using the global Moran index. For these six cell types, the corresponding spatially heterogeneous sampling points were found using the local Moran index. Using the F1 score, a greedy strategy was employed to find the best-matching spatial domain among the four candidate spatial domains. For each cell type among the four enrichment score candidates, the best spatial domain match was found. Finally, for the six cell types, the best match from the four candidates was again selected using the F1 score as the spatial localization of the cell type. Furthermore, differentially expressed genes in the spatial localization region were identified, and their intersection with the differentially expressed genes of the corresponding cell type in the single-cell dataset was taken as the differentially expressed genes for the spatial localization of the cell type.

[0122] The method in this embodiment overcomes the bias of the previous deconvolution method, enabling most cell types in the spatial transcriptome to find their spatial aggregation patterns, thereby further discovering the colocalization of cell types in space.

[0123] The method described in this embodiment can be applied to fields such as spatial transcriptomics and single-cell transcriptomics. It can combine spatial transcriptomics data and single-cell transcriptomics data to obtain unbiased spatial localization of cell types in tissue regions. Specifically, the unbiasedness is reflected in the fact that by organically integrating existing methods, the spatial localization of the vast majority of cell types can be found. This approach overcomes the disadvantage of using a single deconvolution method, which can only locate a subset of cell types. The conclusions drawn from the method in this embodiment are of great significance for exploring cellular spatial heterogeneity and cell interactions within the microenvironment, and form the basis for further downstream spatial transcriptomics analyses.

[0124] Example 2

[0125] This embodiment provides a cell type unbiased localization system based on spatial transcriptomics, which specifically includes:

[0126] The data acquisition module is configured to acquire spatial transcriptome data from tissue slices.

[0127] The clustering module is configured to divide the sampling points in the spatial transcriptome data into several categories using a spatial clustering method, where each category represents a tissue region in a tissue slice.

[0128] The deconvolution module is configured to: obtain the proportion of cell types at each sampling point based on spatial transcriptome data through deconvolution, and use the proportion of a certain cell type at a certain sampling point as the enrichment score of that cell type at that sampling point;

[0129] The screening module is configured to: calculate the global Moran index for each cell type based on the enrichment score, and screen out cell types with spatial heterogeneity;

[0130] The localization module is configured to: for each cell type with spatial heterogeneity, filter out sampling points with spatial heterogeneity based on enrichment scores to obtain a spatial heterogeneous point set, and match the spatial heterogeneous point set with all tissue regions to obtain the best matching tissue region.

[0131] It should be noted that each module in this embodiment corresponds one-to-one with each step in Embodiment 1, and their specific implementation processes are the same, so they will not be repeated here.

[0132] Example 3

[0133] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the cell type unbiased localization method based on spatial transcriptomics as described in Embodiment 1 above.

[0134] Example 4

[0135] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the cell type unbiased localization method based on spatial transcriptomics as described in Embodiment 1 above.

[0136] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A cell type unbiased localization method based on spatial transcriptomics, characterized in that, include: Acquire spatial transcriptomic data from tissue sections; Spatial clustering methods are used to divide the sampling points in the spatial transcriptome data into several categories, with each category representing a tissue region in a tissue slice; Based on spatial transcriptome data, the proportion of cell types at each sampling point is obtained by deconvolution, and the proportion of a certain cell type at a certain sampling point is used as the enrichment score of that cell type at that sampling point. The deconvolution method is one or more of the following: SPOTlight method, spacexr method, and stereoscope method; The SPOTlight method integrates spatial data with single-cell reference data, initializes it using non-negative matrix factor regression, and then uses non-negative least squares to obtain the proportion of cell types at each sampling point based on the marker genes and spatial gene expression data of each cell type. The SpaceXR method treats spatial transcriptome data as conforming to a Poisson distribution. Given the provided single-cell data, it reads the average expression of all genes in each cell type, and then uses a stepwise fitting method to find one or two cell types that best match the gene expression at the sampling point. Then, it uses the maximum likelihood estimation method to fit the parameters of the Poisson distribution, thereby inferring the proportion of cell types in the sampling point. The stereoscope method treats both single-cell reference expression data and spatial transcriptome data as conforming to a negative binomial distribution. Given the provided single-cell data, it obtains the values ​​of specific parameters of cell type distribution by searching for maximum likelihood estimation. Finally, based on the observed spatial data, it estimates the proportion of cell types at each sampling point using the prior distribution of cell types. Based on the enrichment score, a global Moran index is calculated for each cell type to screen out cell types with spatial heterogeneity; For each cell type with spatial heterogeneity, sampling points with spatial heterogeneity are selected based on enrichment scores to obtain a spatial heterogeneous point set. The spatial heterogeneous point set is then matched with all tissue regions to obtain the best matching tissue region. The steps for matching the spatially heterogeneous point set with all tissue regions are as follows: (1) The set of points consisting of all the sampling points contained in an organization region is called an organization region point set; (2) For a certain cell type, calculate the F1 score of the spatial heterogeneity point set of the cell type and the point set of each tissue region, and select the tissue region corresponding to the highest F1 score as the best matching tissue region for the cell type. (3) For a certain cell type, if there are multiple best matching tissue regions, each best matching tissue region is expanded as a new tissue region and returned to step (2).

2. The cell type unbiased localization method based on spatial transcriptomics as described in claim 1, characterized in that, The method to expand the best-matching organization region is as follows: for a certain best-matching organization region, randomly select an organization region other than the best-matching organization region and attach it to the best-matching organization region.

3. The cell type unbiased localization method based on spatial transcriptomics as described in claim 1, characterized in that, For a given cell type exhibiting spatial heterogeneity, the sampling points exhibiting spatial heterogeneity i satisfy: in, This indicates that the spatially heterogeneous cell type is located at the sampling point. i enrichment fraction at the location The average of all enrichment scores. Indicates sampling point i and its neighboring sampling points j The distance between them n Indicates sampling point i The number of neighborhood sampling points, This indicates that the spatially heterogeneous cell type is located at the sampling point. i The significance of the local Moran index at the location, This indicates that the value of p is being sought. This is the set value.

4. The cell type unbiased localization method based on spatial transcriptomics as described in claim 1, characterized in that, The spatial clustering method is one or more of the following: Leiden algorithm, BayesSpace method, SpaGCN method, and stlearn method.

5. The cell type unbiased localization method based on spatial transcriptomics as described in claim 1, characterized in that, Also includes: Based on the best-matched tissue regions, the Pearson correlation coefficient between each pair of cell types is calculated.

6. A cell type unbiased localization system based on spatial transcriptomics, characterized in that, include: The data acquisition module is configured to acquire spatial transcriptome data from tissue slices. The clustering module is configured to divide the sampling points in the spatial transcriptome data into several categories using a spatial clustering method, where each category represents a tissue region in a tissue slice. The deconvolution module is configured to: obtain the proportion of cell types at each sampling point based on spatial transcriptome data through deconvolution, and use the proportion of a certain cell type at a certain sampling point as the enrichment score of that cell type at that sampling point; The deconvolution method is one or more of the following: SPOTlight method, spacexr method, and stereoscope method; The SPOTlight method integrates spatial data with single-cell reference data, initializes it using non-negative matrix factor regression, and then uses non-negative least squares to obtain the proportion of cell types at each sampling point based on the marker genes and spatial gene expression data of each cell type. The SpaceXR method treats spatial transcriptome data as conforming to a Poisson distribution. Given the provided single-cell data, it reads the average expression of all genes in each cell type, and then uses a stepwise fitting method to find one or two cell types that best match the gene expression at the sampling point. Then, it uses the maximum likelihood estimation method to fit the parameters of the Poisson distribution, thereby inferring the proportion of cell types in the sampling point. The stereoscope method treats both single-cell reference expression data and spatial transcriptome data as conforming to a negative binomial distribution. Given the provided single-cell data, it obtains the values ​​of specific parameters of cell type distribution by searching for maximum likelihood estimation. Finally, based on the observed spatial data, it estimates the proportion of cell types at each sampling point using the prior distribution of cell types. The screening module is configured to: calculate the global Moran index for each cell type based on the enrichment score, and screen out cell types with spatial heterogeneity; The localization module is configured to: for each cell type with spatial heterogeneity, filter out sampling points with spatial heterogeneity based on enrichment scores to obtain a spatial heterogeneous point set, and match the spatial heterogeneous point set with all tissue regions to obtain the best matching tissue region; The steps for matching the spatially heterogeneous point set with all tissue regions are as follows: (1) The set of points consisting of all the sampling points contained in an organization region is called an organization region point set; (2) For a certain cell type, calculate the F1 score of the spatial heterogeneity point set of the cell type and the point set of each tissue region, and select the tissue region corresponding to the highest F1 score as the best matching tissue region for the cell type. (3) For a certain cell type, if there are multiple best matching tissue regions, each best matching tissue region is expanded as a new tissue region and returned to step (2).

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the cell type unbiased localization method based on spatial transcriptomics as described in any one of claims 1-5.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the cell type unbiased localization method based on spatial transcriptomics as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Spatial transcriptome cell clustering and analyzing method

    CN114091603A

  • Method and system for deducing expression pattern of cell subset in spatial transcriptome

    CN114944194A