A system for uncoupling exhausted t cells in a tumor microenvironment
By integrating single-cell and spatial transcriptomics analysis workflows and using deconvolution methods to locate the spatial position of exhausted T cells in the tumor microenvironment, this method solves the problem of difficulty in analyzing the differences in cell behavior in different regions of tumor tissue in existing technologies, and achieves high-resolution localization and proportional analysis of exhausted T cells.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies make it difficult to analyze differences in cell behavior in different regions of the same tumor tissue, especially the spatial localization of exhausted T cells. Single-cell sequencing technology is also unable to achieve high-resolution spatial transcriptome sequencing.
This study integrates single-cell transcriptomics and spatial transcriptomics analysis workflows. Through data acquisition, preprocessing, clustering, marker gene identification, and deconvolution methods, it locates the spatial position of exhausted T cells in the tumor microenvironment and uses deconvolution methods to identify the spatial domains and proportions of exhausted T cells.
It enables accurate localization and proportional analysis of exhausted T cells in tumor tissues, overcomes the limitations of existing technologies in spatial localization, and provides a more refined immunomics analysis method.
Smart Images

Figure CN116364188B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, and in particular to a system for decoupling exhausted T cells in the tumor microenvironment. Background Technology
[0002] The statements in this section merely refer to the background art related to this invention and do not necessarily constitute prior art.
[0003] Single-cell transcriptomics and spatial transcriptomics both fall under the category of transcriptomics, representing an interdisciplinary field between life sciences and computer science. Breakthroughs in transcriptomics have led to new discoveries in the study of diseases and biological processes. Immunomics primarily studies a complete library of immune-related molecules, their target molecules, and their functions. Specific T cells in the tumor microenvironment are crucial for cancer development and prognosis, and they can effectively promote the function of tumor immunotherapy.
[0004] The inventors discovered that current research on immunomics primarily revolves around single-cell technology. The typical experimental method involves sequencing single cells from normal and tumor tissues, comparing differences in single-cell expression profiles to uncover molecular biological mechanisms. A limitation of this method is its inability to analyze cell behavior in different regions within the same tumor tissue; that is, it can only compare the average differences in cell behavior between normal and tumor tissues, not the differences in cell behavior between different regions within the tumor tissue itself.
[0005] Meanwhile, for certain types of cells, especially exhausted T cells, it is difficult to identify their spatial location from the complex microenvironment using single-cell sequencing technology. However, spatial transcriptomics can directly provide the spatial location of exhausted T cells in the tumor microenvironment, but current spatial transcriptomics sequencing technologies cannot achieve single-cell resolution. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a system for decoupling exhausted T cells in the tumor microenvironment. This system fully utilizes the high resolution of single-cell transcriptomics and the spatial location information preserved by the spatial transcriptome, enabling the spatial localization of exhausted T cells in the tumor microenvironment, thereby further analyzing their structure and function in different tumor microenvironments.
[0007] In a first aspect, the present invention provides a system for decoupling exhausted T cells in a tumor microenvironment;
[0008] A system for decoupling exhausted T cells in a tumor microenvironment, comprising:
[0009] The data acquisition module is configured to acquire single-cell data and spatial transcriptome data of tumor tissue.
[0010] The preprocessing module is configured to preprocess single-cell data and spatial transcriptome data separately.
[0011] The extraction module is configured to: extract tumor-infiltrating T cell data from the preprocessed single-cell data by performing two clustering and marker gene calculations on the preprocessed single-cell data, and extract exhausted T cell data from the tumor-infiltrating T cell data.
[0012] The spatial domain partitioning module is configured to cluster the preprocessed tumor tissue spatial transcriptome data to obtain several spatial domains.
[0013] The T-cell localization module is configured to: use deconvolution methods to identify the spatial domain of tumor-infiltrating T cells in several spatial domains, taking single-cell data and tumor-infiltrating T-cell data as references.
[0014] The T cell enrichment analysis module is configured to: use deconvolution methods to identify the spatial domain of exhausted T cells in the spatial domain of tumor-infiltrating T cells, using tumor-infiltrating T cell data and exhausted T cell data as references, and obtain the proportion of exhausted T cells in tumor-infiltrating T cells.
[0015] Furthermore, the preprocessing includes gene filtering, standardization, normalization, and principal component analysis.
[0016] Furthermore, the extraction module includes a first marker gene recognition module;
[0017] The first marker gene identification module is configured to: cluster the preprocessed single-cell data to obtain several first clusters, and calculate the marker gene for each first cluster.
[0018] Furthermore, the first marker gene recognition module includes a filtering module;
[0019] The filtering module is configured to: for a certain first group, calculate the percentage and logarithmic expression of each gene in the first group and other first groups, perform gene filtering, and obtain the filtered genes.
[0020] Furthermore, the first marker gene recognition module also includes a verification module;
[0021] The testing module is configured to: perform testing on the filtered genes for a certain first group, obtain p-values, correct the obtained p-values, sort the p-values from smallest to largest, and select the top-ranked genes as the marker genes for the first group.
[0022] Furthermore, the extraction module also includes a tumor-infiltrating T-cell extraction module;
[0023] The tumor-infiltrating T-cell extraction module is configured to extract tumor-infiltrating T-cell data from preprocessed single-cell data based on the marker genes of all first-group cells and the marker genes of different cells stored in the marker gene database.
[0024] Furthermore, the extraction module includes a second marker gene recognition module;
[0025] The second marker gene recognition module is configured to: cluster all tumor-infiltrating T cell data to obtain several second clusters, and calculate the marker gene for each second cluster.
[0026] Furthermore, the extraction module also includes an exhaustive T-cell extraction module;
[0027] The exhausted T cell extraction module is configured to extract exhausted T cell data from tumor-infiltrating T cell data based on the marker genes of all second groups and the marker genes of different cells stored in the marker gene database.
[0028] Secondly, the present invention also provides an electronic device, comprising:
[0029] Memory, used for non-transitory storage of computer-readable instructions; and
[0030] Processor, for executing the computer-readable instructions,
[0031] When the computer-readable instructions are executed by the processor, the following steps are performed:
[0032] Acquire single-cell data and spatial transcriptome data from tumor tissues;
[0033] Single-cell data and spatial transcriptome data were preprocessed separately;
[0034] By performing two clustering and marker gene calculations on the preprocessed single-cell data, tumor-infiltrating T-cell data were extracted from the preprocessed single-cell data, and exhausted T-cell data were extracted from the tumor-infiltrating T-cell data.
[0035] The preprocessed spatial transcriptome data of tumor tissue were clustered to obtain several spatial domains;
[0036] Using deconvolution methods, single-cell data and tumor-infiltrating T-cell data are used as references to identify the spatial domain of tumor-infiltrating T-cells in several spatial domains.
[0037] Using the deconvolution method, tumor-infiltrating T cell data and exhausted T cell data are used as references to identify the spatial domain of exhausted T cells in the spatial domain of tumor-infiltrating T cells, and to obtain the proportion of exhausted T cells in tumor-infiltrating T cells.
[0038] Thirdly, the present invention also provides a storage medium for non-transitory storage of computer-readable instructions, wherein when the non-transitory computer-readable instructions are executed by a computer, the following steps are performed:
[0039] Acquire single-cell data and spatial transcriptome data from tumor tissues;
[0040] Single-cell data and spatial transcriptome data were preprocessed separately;
[0041] By performing two clustering and marker gene calculations on the preprocessed single-cell data, tumor-infiltrating T-cell data were extracted from the preprocessed single-cell data, and exhausted T-cell data were extracted from the tumor-infiltrating T-cell data.
[0042] The preprocessed spatial transcriptome data of tumor tissue were clustered to obtain several spatial domains;
[0043] Using deconvolution methods, single-cell data and tumor-infiltrating T-cell data are used as references to identify the spatial domain of tumor-infiltrating T-cells in several spatial domains.
[0044] Using the deconvolution method, tumor-infiltrating T cell data and exhausted T cell data are used as references to identify the spatial domain of exhausted T cells in the spatial domain of tumor-infiltrating T cells, and to obtain the proportion of exhausted T cells in tumor-infiltrating T cells.
[0045] Compared with the prior art, the beneficial effects of the present invention are:
[0046] This invention integrates the main analytical workflows of single-cell transcriptomics and spatial transcriptomics, enabling automated joint immunomics analysis. It fully utilizes the high resolution of single-cell transcriptomics and the spatial location information preserved by spatial transcriptomics, accurately locating the region of T cells in tumor tissue, and also analyzing the proportion of exhausted T cells in all T cell regions.
[0047] The results obtained by this invention can be used to analyze the differences in T cell activity between different cancer tissues, especially carcinoma in situ and invasive carcinoma tissues, using immunomics analysis, overcoming the spatial limitations of previous research paradigms and technologies.
[0048] The advantages of additional aspects of the invention will be set forth in part in the description which follows, or may be learned by practice of the invention. Attached Figure Description
[0049] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0050] Figure 1 This is a schematic diagram of a system for decoupling exhausted T cells in a tumor microenvironment, as described in Example 1. Detailed Implementation
[0051] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0052] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0053] All data acquisition in this embodiment is carried out in accordance with laws and regulations and with user consent, and the data is used legally.
[0054] Terminology Explanation:
[0055] Exhausted T cells: Under continuous exposure to antigens (such as tumors and chronic infections), T cells eventually fail to differentiate into an immune memory phenotype and are in a state of functional exhaustion, which is called T cell exhaustion. T cells in a state of functional exhaustion are called exhausted T cells.
[0056] Example 1
[0057] This embodiment provides a system for decoupling exhausted T cells in the tumor microenvironment;
[0058] like Figure 1 The system shown is for decoupling exhausted T cells in the tumor microenvironment, comprising: a data acquisition module, a preprocessing module, an extraction module (including a first marker gene recognition module, a tumor-infiltrating T cell extraction module, a second marker gene recognition module, and an exhausted T cell extraction module), a spatial domain division module, a T cell localization module, and a T cell enrichment analysis module.
[0059] The data acquisition module is configured to acquire single-cell data (single-cell omics data, single-cell sequencing data, or single-cell transcriptome data) and spatial transcriptome data of tumor tissue.
[0060] Single-cell data refers to an expression matrix with cell names as rows and gene expression as columns, where each element represents the expression level of a gene within the cell.
[0061] Single-cell data consists of gene expression matrices composed of cells dissociated from tumor tissue, where rows represent cells and columns represent genes. Using single-cell transcriptome data numbered GSM5354531, the single-cell expression matrix was read from the hard drive using computer commands, yielding 7986 cells, each containing 29733 genes.
[0062] Spatial transcriptome data includes: an expression matrix consisting of gene expression at all sampling points in a tumor tissue slice image (where rows represent sampling points and columns represent genes), and the specific location of each sampling point in the tissue slice image.
[0063] Triple-negative breast cancer tissue sections using 10X Visium sequencing technology were selected as the spatial transcriptome dataset. Their two-dimensional spatial location and histological images are shown below. Figure 1 As shown in the image, the tissue section in the lower left corner is invasive carcinoma tissue, the tissue section in the upper right corner is carcinoma in situ tissue, and the tissue section in the upper left corner is normal breast tissue.
[0064] The preprocessing module is configured to preprocess single-cell data and spatial transcriptome data separately. The preprocessing module includes a first preprocessing module and a second preprocessing module.
[0065] The first preprocessing module is configured to preprocess single-cell data.
[0066] The preprocessing module includes a screening module, a normalization module, and a principal component analysis module.
[0067] The filtering module is configured to filter out genes with excessively low expression levels in all cells, as well as mitochondrial genes, to obtain a filtered expression matrix. For example, filtering out genes with expression levels below 10 reads in all cells, as well as mitochondrial genes, results in a filtered expression matrix containing 18,097 genes.
[0068] The normalization module is configured to normalize and standardize the selected expression matrix and select genes with high expression, for example, the 2000 genes with the highest average expression level.
[0069] The principal component analysis module is configured to perform principal component analysis on the screened expression matrix.
[0070] The first marker gene identification module is configured to: cluster the preprocessed single-cell data and calculate the marker gene for each first cluster.
[0071] The first marker gene recognition module includes: a clustering module, a filtering module, and a testing module.
[0072] The clustering module is configured to embed each cell in the preprocessed expression matrix into the same K-nearest neighbor graph structure and draw edges between cells with similar gene expression patterns (cells with gene expression similarity greater than a threshold); then, the K-nearest neighbor graph is divided into several communities (first subgroups), and the cells within each community are highly interconnected.
[0073] Preferably, the single-cell data is divided into 9 clusters, namely cancerous epithelial cells, ordinary epithelial cells, ordinary endothelial cells, B cells, T cells, NK cells, cancer-associated fibroblasts (CAFs), myeloid cells, plasma cells, and PVL cells.
[0074] The filtering module is configured to: for each first subgroup, calculate the percentage and logarithmic expression (log2FC) of each gene in that subgroup and other subgroups, and filter the genes according to the thresholds of both to obtain the filtered genes.
[0075] The testing module is configured to perform a Wilcoxon test on the filtered genes for each first group, obtain a p-value, then correct the obtained p-value, sort the p-values from smallest to largest, and obtain a series of the most significant marker genes (i.e., select the top-ranked genes as marker genes for the first group).
[0076] The clustering module requires setting a granularity parameter, which is continuously adjusted based on the dataset size and marker gene expression until all clusters have significant and identifiable marker genes. Identifiable marker genes are those that are associated with a known cell type (a cluster).
[0077] The tumor-invasive T-cell extraction module is configured to extract a subset of tumor-invasive T-cells from the dataset (preprocessed single-cell data) based on the marker genes of all clusters and combined with the marker genes of different cells stored in the marker gene database.
[0078] Specifically, T cell marker genes CD4 and CD8 are extracted from the marker gene database; among all marker genes obtained by the verification module, the marker genes with the highest similarity to CD4 and CD8 are calculated to obtain CD4+ and CD8+; in the preprocessed single-cell data, the first cell population with marker genes CD4+ and CD8+ are extracted as the subset of tumor-infiltrating T cells.
[0079] For example, according to CD4 + and CD8 + The T cell clusters were labeled, and a total of 1711 CD8+ markers were obtained. + T cells.
[0080] The marker gene database stores marker genes for several cell types.
[0081] The second marker gene recognition module is configured to further cluster tumor-infiltrating T cells and calculate the marker genes for each second cluster.
[0082] The specific technical approach of clustering tumor-infiltrating T cells again and calculating the marker genes of each second cluster is the same as that of clustering single-cell data and calculating the marker genes of each first cluster. The difference lies in the marker genes and granularity parameters of each cluster.
[0083] Preferably, CD8 + The T-cell subset is divided into 8 subclasses.
[0084] The exhausted T cell extraction module is configured to extract exhausted T cells from tumor-infiltrating T cells based on marker genes of all second subgroups and marker genes of different cells stored in a marker gene database. Specifically, it uses marker genes of various T cell subtypes to label each T cell subset (second subgroup) and screens out exhausted T cells.
[0085] Specifically, marker genes for exhausted T cells were searched from a marker gene database, and cell populations with high expression of these marker genes were extracted and labeled as exhausted T cells; similarly, other marker genes were labeled... T cells, memory T cells, helper T cells, etc., specifically include:
[0086] The marker genes PDCD1, CD200, HAVCR2, TIGIT, and CXCL13 for terminally exhausted T cells were searched from the marker gene database. PDCD1 was extracted. + CD200 + HAVCR2 + TIGIT + The cells were classified into group 4 and labeled as terminally exhausted T cells;
[0087] Searching from marker gene databases T cell marker gene TCF7, TCF7 was extracted. + Cell groups 5 and 7, labeled as T cells;
[0088] The marker genes ZNF683 and CXCR6 for memory T cells were searched in the marker gene database. Cell clusters ZNF683+ and CXCR6+ were extracted and labeled as memory T cells.
[0089] The marker genes GZMB, NKG7, CXCL13, PDCD1, and HAVCR2 for GZMB+ exhausted T cells were searched from the marker gene database. Cell populations of GZMB+, NKG7+, and CXCL13+ were extracted and labeled as GZMB+. + Exhausted T cells;
[0090] The marker genes GZMK, GZMA, PDCD1, and CD74 of GZMK+ exhausted T cells were searched from the marker gene database. Cell clusters of GZMK+, GZMA+, and CD74+ were extracted and labeled as GZMK. + Exhausted T cells;
[0091] The marker genes DUSP2 and DUSP4 for helper T cells were searched from the marker gene database, and DUSP2 was extracted. + DUSP4 + The cells were classified into group 3 and labeled as helper T cells.
[0092] The second preprocessing module is configured to preprocess spatial transcriptome data from tumor tissues.
[0093] Preprocessing of tumor tissue spatial transcriptome data includes: removing sampling points located outside the tissue region, followed by gene filtering, standardization, normalization, and principal component analysis of all sampling points, which is consistent with the preprocessing techniques for single-cell data.
[0094] Spatial transcriptome data refers to spatial transcriptome data obtained by sequencing the breast tumor microenvironment using 10X Visium.
[0095] For example, by preprocessing the expression matrix of the sampling points, we obtained 1162 sampling points in the tissue, with 19237 gene types.
[0096] The spatial domain partitioning module is configured to cluster the preprocessed tumor tissue spatial transcriptome data and calculate the marker genes for each spatial domain (one cluster is one spatial domain, and one spatial domain contains multiple sampling points).
[0097] The spatial clustering operation and the calculation of marker genes for each spatial domain are consistent with the technical approach of clustering single-cell data and calculating marker genes for each cluster, except that the data has been replaced with spatial transcriptome data of the tumor microenvironment.
[0098] The spatial domain partitioning module includes: clustering module, filtering module, and testing module.
[0099] The clustering module is configured to: employ the BayesSpace method to model and cluster each sampling point in the preprocessed spatial representation matrix using a multivariate t-distribution model; and finally update the parameters using the Metropolis-Hastings algorithm, which uses two-dimensional spatial information integrated by the Potts model as its prior distribution. Preferably, in this embodiment, the breast cancer tumor spatial transcription data is divided into 10 spatial domains.
[0100] The filtering module is configured to: for each spatial domain, calculate the percentage and logarithmic expression (log2FC) of each gene in that subgroup and other subgroups, and filter genes based on thresholds of both.
[0101] The testing module is configured to perform a Wilcoxon test on the selected genes for each spatial domain, then correct the obtained p-values, sort the p-values from smallest to largest, and obtain a series of the most significant marker genes.
[0102] The T-cell localization module is configured to use a deconvolution method, taking the single-cell dataset and T-cell data (tumor-invasive T-cell data) segmented by spatial clustering as references, to identify the spatial domain of T cells within the tumor microenvironment. To determine the accuracy of the spatial domain identified by the T cell, the intersection of the T cell's marker genes and the marker genes of that spatial domain is calculated, and the number of marker genes in the intersection is examined. A higher number of marker genes indicates more accurate spatial localization.
[0103] Deconvolution is a method used to infer the proportion of cell types at a sampling point in the spatial transcriptome. Each sampling point contains a mixture of multiple cell types. Deconvolution builds a model based on single-cell data, using gene expression data from each sampling point in the spatial transcriptome as input, and outputs a maximum a posteriori estimate of the cell type distribution in the tissue given the gene expression distribution at a given sampling point.
[0104] The deconvolution method employed is the Cell2Location method, which assumes that single-cell reference expression data follows a negative binomial distribution. Given the provided single-cell data, the method obtains the values of specific parameters for cell type distribution by searching for maximum likelihood estimation (MLE), thereby determining the proportion of cell types. In the results obtained in this embodiment, T cells were found to be significantly enriched in spatial domain 6 obtained by the BayesSpace method, which is distributed in both invasive and in situ carcinoma regions.
[0105] The T cell enrichment analysis module is configured to: again use the deconvolution method, taking the labeled subset of tumor-infiltrating T cells as a reference, identify the spatial domain of exhausted T cells in each tumor-infiltrating T cell spatial domain, and the proportion of cell types (T cells) occupied by exhausted T cells (i.e., obtain the proportion of exhausted T cells in tumor-infiltrating T cells).
[0106] Based on the spatial clustering results of the spatial transcriptome data, the spatial domains with the highest T cell content were identified as T cell enrichment regions. Deconvolution operations were then performed on these regions using the same technique as described earlier, except that the spatial transcriptome data was replaced with T cell enrichment regions, and the single-cell data was replaced with a subset of tumor-infiltrating T cells. The final output represents the proportion of exhausted T cells in different T cell enrichment regions.
[0107] The final output is the proportion of T cells at each sampling point, which, from the perspective of the entire tissue, is also the spatial location of T cells in the tumor microenvironment.
[0108] Using the Cell2Location method, it was found that the proportion of exhausted T cells in the T cell-infiltrating carcinoma enrichment region was significantly higher than that in carcinoma in situ; and The proportion of T cells in the T cell-rich areas of in situ carcinoma was significantly higher than that in invasive carcinoma. This result indirectly confirms the poorer prognosis in invasive carcinoma tissues from a spatial transcriptomics perspective.
[0109] The system described in this embodiment can be applied in fields such as spatial transcriptomics, single-cell transcriptomics, and immunomics. It combines the advantages of spatial transcriptomics data and single-cell transcriptomics data to obtain the proportion of cell types of exhausted T cells in different spatial domains within the same tumor tissue. This greatly improves the difficulty in locating the specific position of T cells in the tumor microenvironment during immunomics analysis, enabling downstream analyses to further study the subtle differences in T cell activity within the tumor microenvironment at a microscopic level. The conclusions drawn are of great significance for exploring the spatial heterogeneity of tumor tissues and the interactions of exhausted T cells in the microenvironment, and form the basis for conducting more downstream immunomics analyses.
[0110] Example 2
[0111] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory to cause the electronic device to perform the following steps:
[0112] Acquire single-cell data and spatial transcriptome data from tumor tissues;
[0113] Single-cell data and spatial transcriptome data were preprocessed separately;
[0114] By performing two clustering and marker gene calculations on the preprocessed single-cell data, tumor-infiltrating T-cell data were extracted from the preprocessed single-cell data, and exhausted T-cell data were extracted from the tumor-infiltrating T-cell data.
[0115] The preprocessed spatial transcriptome data of tumor tissue were clustered to obtain several spatial domains;
[0116] Using deconvolution methods, single-cell data and tumor-infiltrating T-cell data are used as references to identify the spatial domain of tumor-infiltrating T-cells in several spatial domains.
[0117] Using the deconvolution method, tumor-infiltrating T cell data and exhausted T cell data are used as references to identify the spatial domain of exhausted T cells in the spatial domain of tumor-infiltrating T cells, and to obtain the proportion of exhausted T cells in tumor-infiltrating T cells.
[0118] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0119] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0120] In the implementation process, each step of the above method can be completed by the integrated logic circuits in the processor hardware or by software instructions.
[0121] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0122] Those skilled in the art will recognize that the units and algorithm steps described in connection with the various examples of this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.
[0123] Example 3
[0124] This embodiment also provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the following steps:
[0125] Acquire single-cell data and spatial transcriptome data from tumor tissues;
[0126] Single-cell data and spatial transcriptome data were preprocessed separately;
[0127] By performing two clustering and marker gene calculations on the preprocessed single-cell data, tumor-infiltrating T-cell data were extracted from the preprocessed single-cell data, and exhausted T-cell data were extracted from the tumor-infiltrating T-cell data.
[0128] The preprocessed spatial transcriptome data of tumor tissue were clustered to obtain several spatial domains;
[0129] Using deconvolution methods, single-cell data and tumor-infiltrating T-cell data are used as references to identify the spatial domain of tumor-infiltrating T-cells in several spatial domains.
[0130] Using the deconvolution method, tumor-infiltrating T cell data and exhausted T cell data are used as references to identify the spatial domain of exhausted T cells in the spatial domain of tumor-infiltrating T cells, and to obtain the proportion of exhausted T cells in tumor-infiltrating T cells.
[0131] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A system for uncoupling exhausted T cells in a tumor microenvironment, characterized in that, The method comprises the following steps: data acquisition module configured to acquire single cell data and spatial transcriptome data of tumor tissue; a preprocessing module configured to preprocess the single cell data and the spatial transcriptome data respectively; an extraction module configured to extract tumor infiltrating T cell data from the preprocessed single cell data and extract exhausted T cell data from the tumor infiltrating T cell data by performing twice clustering and calculating marker genes on the preprocessed single cell data; a spatial domain division module configured to obtain a plurality of spatial domains by clustering and dividing the preprocessed tumor tissue spatial transcriptome data; a T cell positioning module configured to identify the spatial domain of tumor infiltrating T cells in the plurality of spatial domains by using a deconvolution method and taking the single cell data and the tumor infiltrating T cell data as references; a T cell enrichment analysis module configured to identify the spatial domain of exhausted T cells in the spatial domain of tumor infiltrating T cells by using a deconvolution method and taking the tumor infiltrating T cell data and the exhausted T cell data as references, and obtain the proportion of exhausted T cells in tumor infiltrating T cells; the extraction module comprises a first marker gene identification module; the first marker gene identification module is configured to obtain a plurality of first clusters by clustering and dividing the preprocessed single cell data, and calculate the marker genes of each first cluster; the first marker gene identification module comprises a filtering module; the filtering module is configured to calculate the percentage and logarithmic expression of each gene in a certain first cluster and other first clusters, perform gene filtering, and obtain filtered genes; the first marker gene identification module further comprises a test module; the test module is configured to perform test on the filtered genes for a certain first cluster to obtain p values, correct the obtained p values, sort the p values from small to large, and select a plurality of genes with high ranking as the marker genes of the first cluster; the extraction module comprises a second marker gene identification module; the second marker gene identification module is configured to obtain a plurality of second clusters by clustering and dividing all tumor infiltrating T cell data, and calculate the marker genes of each second cluster; the specific technical route of clustering and dividing tumor infiltrating T cells again and calculating the marker genes of each second cluster is the same as that of clustering and dividing single cell data and calculating the marker genes of each first cluster.
2. The system for decoupling exhausted T cells in a tumor microenvironment as described in claim 1, characterized in that, The preprocessing comprises gene filtering, standardization, normalization and principal component analysis.
3. The system for decoupling exhausted T cells in a tumor microenvironment as described in claim 1, characterized in that, The extraction module further comprises a tumor infiltrating T cell extraction module; the tumor infiltrating T cell extraction module is configured to extract tumor infiltrating T cell data from the preprocessed single cell data based on the marker genes of all first clusters and in combination with the marker genes of different cells stored in a marker gene database.
4. The system for decoupling exhausted T cells in a tumor microenvironment as described in claim 1, characterized in that, the extraction module further comprises an exhausted T cell extraction module; The exhaustion T cell extraction module is configured to extract the exhaustion T cell data in the tumor infiltrating T cell data based on the marker genes of all the second subgroups and the marker genes of different cells stored in the marker gene database.
5. An electronic device implementing a system for uncoupling exhausted T cells in a tumor microenvironment according to any one of claims 1-4, characterized in that, Comprise: a memory for non-transitorily storing computer readable instructions; and a processor for running the computer readable instructions, wherein the computer readable instructions, when run by the processor, perform the following steps: obtaining single cell data and spatial transcriptome data of tumor tissue; preprocessing the single cell data and the spatial transcriptome data respectively; extracting tumor infiltrating T cell data from the preprocessed single cell data by performing twice clustering subgroups and marker gene calculation on the preprocessed single cell data, and extracting exhaustion T cell data from the tumor infiltrating T cell data; performing clustering subgroups on the preprocessed tumor tissue spatial transcriptome data to obtain a plurality of spatial domains; using a deconvolution method, identifying the spatial domain of tumor infiltrating T cells in the plurality of spatial domains by taking the single cell data and the tumor infiltrating T cell data as references; using a deconvolution method, identifying the spatial domain of exhaustion T cells in the spatial domain of tumor infiltrating T cells by taking the tumor infiltrating T cell data and the exhaustion T cell data as references, and obtaining the proportion of exhaustion T cells in tumor infiltrating T cells.
6. A storage medium embodying a system for uncoupling exhausted T cells in a tumor microenvironment as claimed in any one of claims 1 to 4, characterized in that, non-transitorily storing computer readable instructions, wherein when the non-transitory computer readable instructions are executed by a computer, the following steps are performed: obtaining single cell data and spatial transcriptome data of tumor tissue; preprocessing the single cell data and the spatial transcriptome data respectively; extracting tumor infiltrating T cell data from the preprocessed single cell data by performing twice clustering subgroups and marker gene calculation on the preprocessed single cell data, and extracting exhaustion T cell data from the tumor infiltrating T cell data; performing clustering subgroups on the preprocessed tumor tissue spatial transcriptome data to obtain a plurality of spatial domains; using a deconvolution method, identifying the spatial domain of tumor infiltrating T cells in the plurality of spatial domains by taking the single cell data and the tumor infiltrating T cell data as references; using a deconvolution method, identifying the spatial domain of exhaustion T cells in the spatial domain of tumor infiltrating T cells by taking the tumor infiltrating T cell data and the exhaustion T cell data as references, and obtaining the proportion of exhaustion T cells in tumor infiltrating T cells.
Citation Information
Patent Citations
Bladder cancer exhausted T cell subset as well as characteristic genes and application thereof
CN111748627A
Analysis method and system for integrating single cell transcriptome and spatial transcriptome data
CN114944193A