Systems and methods for spatial alignment of cytological specimens and applications thereof

JP2025528016A5Pending Publication Date: 2026-05-26THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
Filing Date
2023-05-18
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing spatial transcriptomics methods struggle to accurately assign cell types to specific regions within histological sections at single-cell resolution, limiting the precision of spatially resolved analyte mapping.

Method used

A method utilizing spatial omics data, combined with reference single-cell omics data, employs a globally optimal solution to assign single cells to spatial coordinates, generating a spatially resolved map of a specimen through convex optimization and shortest augmenting path algorithms, enabling precise cell type allocation.

Benefits of technology

This approach achieves accurate single-cell resolution in spatially resolved mapping, allowing for the detection of genetic and spatial signatures associated with conditions like cancer, predicting treatment responses, and identifying optimal treatment compounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A process is provided for spatially aligning single cells to create a specimen map with single-cell resolution. Various methods can be performed on the specimen to create spatial omics data. The spatial omics data can be used in combination with the single-cell omics data to assign single cells to spatial coordinates to create a resolved specimen map. The systems and methods of the present disclosure render a spatially resolved map of the specimen with single-cell resolution. The spatial omics data can be acquired from the specimen.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 364,935, filed May 18, 2022, entitled "Robust alignment of single-cell and spatial transcriptomes with CytoSPACE," which is incorporated herein by reference in its entirety.

[0002] Federally Sponsored Research and Development Statement This invention was made with government support under contracts CA255450 and CA1871925 awarded by the National Institutes of Health. The government has certain rights in this invention.

[0003] Technical Field The present disclosure provides a description of generating spatially resolved analyte maps at single-cell resolution using spatial omics data. [Background technology]

[0004] background Spatial transcriptomics is a high-throughput methodology for assigning cell types to specific regions within histological sections of tissues or cell cultures, as assessed by the collection of transcriptome profiles from those regions. Generally, this method involves independently analyzing very small regions of histological sections containing a small number of cells (as few as about 5 cells, but typically between 10 and 40 cells) for transcript expression. Transcript expression in individual regions can be assessed using a variety of different methodologies, including fluorescence in situ hybridization (FISH), in situ sequencing, laser capture microdissection followed by transcript analysis, repetitive microdigestion followed by transcript analysis, and in situ capture followed by transcript analysis. Subsequent transcript analysis can be performed using any expression analysis technique, such as quantitative polymerase chain reaction, microarrays, and RNA sequencing. Summary of the Invention

[0005] The disclosed systems and methods render a spatially resolved map of a specimen at single-cell resolution. Spatial omics data can be acquired from the specimen. Reference single-cell omics data can be used to match the spatial omics data. Based on a global optimum solution, single cells from the reference single-cell omics data can be assigned spatial coordinates to generate a spatially resolved map of the specimen.

[0006] In some embodiments, a method comprises analyzing the transcriptomes of a plurality of cells to determine cell types. The method comprises assigning the cells to locations in a tissue sample based on all possible location assignments. The method comprises detecting a genetic and / or spatial signature specific to a condition within the cells assigned to the location in the tissue sample. The method comprises assaying a sample obtained from a subject to detect the signature. The method comprises reporting the presence or severity of the condition in the subject based on the detected signature.

[0007] In some embodiments, the condition is cancer and the spatial signature predicts response to treatment, toxicity of treatment, resistance to treatment, cancer progression, likelihood of metastasis, likelihood of transition from pre-invasive to invasive cancer, or likelihood of recurrence.

[0008] In some embodiments, the method comprises, prior to said assigning step, obtaining an estimate of the fractional abundance of said cell type in said tissue sample and the number of cells at said location.

[0009] In some embodiments, said genetic and / or spatial signature specific to said condition comprises information about the proximity or interactions between different types of cells.

[0010] In some embodiments, the method comprises providing an expression profile for tissue cells at said location within said tissue sample.

[0011] In some embodiments, the tissue sample comprises a section of a solid tumor.

[0012] In some embodiments, the allocating step uses a convex optimization function.

[0013] In some embodiments, the method comprises performing said assaying step on a plurality of test samples, each exposed to one of a plurality of candidate compounds, to identify a compound that treats said condition.

[0014] In some embodiments, said analyzing comprises accessing a database or atlas of said transcriptome of said cell.

[0015] In some embodiments, the allocation step ensures an overall optimal allocation of the cells to the locations.

[0016] In some embodiments, the allocating step uses a shortest augmenting path algorithm.

[0017] In some embodiments, the condition comprises T cell exhaustion.

[0018] In some embodiments, the analyzing step comprises single-cell RNA sequencing (scRNA-Seq) to obtain the transcriptome.

[0019] In some embodiments, a method is for generating a spatially resolved map of a specimen. The method includes using a computer processing system to obtain spatial omics data from a plurality of regions covering the specimen. The specimen is a collection of cells including a plurality of cell types. The method includes using the computer processing system to estimate the number of cells per region of the plurality of regions from the spatial omics data. The method includes using the computer processing system to estimate the proportion of each cell type of the plurality of cell types from the spatial omics data. The method includes using the computer processing system to query reference single-cell omics data to match the number of cells for each cell type of the plurality of cell types to generate a set of single-cell omics data for spatial assignment. The method includes using the computer processing system to assign single cells to spatial coordinates from the set of single-cell omics data to generate a spatially resolved map of the specimen, based on a globally optimal solution.

[0020] In some embodiments, a method is for generating a spatially resolved map of a specimen. The method includes using a computer processing system to obtain spatial omics data from a plurality of regions covering the specimen. The specimen is a collection of cells including a plurality of cell types. The method includes using the computer processing system to estimate the number of cells per region of the plurality of regions. The method includes determining, based on a globally optimal solution but in parallel, the proportion of each cell type in the plurality of cells; using the computer processing system to query single-cell omics data to match the number of cells for each cell type in the plurality of cell types to generate a set of single-cell omics data for spatial assignment; and using the computer processing system to assign single cells to spatial coordinates from the set of single-cell omics data to generate a spatially resolved map of the specimen.

[0021] In some embodiments, a method is provided for generating a spatially resolved map of a specimen for a set of one or more cell types. The method includes using a computer system to obtain spatial omics data from a plurality of regions covering the specimen. The specimen is a collection of cells including a plurality of cell types. The method includes using the computer system to estimate the number of cells per region of the plurality of regions from the spatial omics data. For each region of the plurality of regions, the method includes using the computer system to estimate the proportion of each of the plurality of cell types from the spatial omics data. For each region containing a proportion of a cell type among the set of one or more cell types, the method includes using the computer system to query reference single-cell omics data to match the number of cells for each of the plurality of cell types to generate a set of single-cell omics data for spatial assignment. The method includes, based on a globally optimal solution, but for each region that contains a proportion of a cell type from the set of one or more cell types, using the computing system to assign single cells to spatial coordinates from the set of single-cell omics data to generate a spatially resolved map of the specimen consisting of the cell type from the set of one or more cell types.

[0022] In some embodiments, querying the reference single-cell omics data to match the number of cells for each cell type of the plurality of cell types further comprises, for each region containing a proportion of a cell type in the set of one or more cell types, using the computing system to remove single-cell omics data of one or more single cells within the reference single-cell omics data when the number of single cells within the reference single-cell omics data is greater than the number of cells estimated within the specimen.

[0023] In some embodiments, querying the reference single-cell omics data to match the number of cells for each cell type of the plurality of cell types further comprises, for each region containing a proportion of a cell type in the set of one or more cell types, using the computing system to add single-cell omics data of one or more single cells within the reference single-cell omics data when the number of single cells within the reference single-cell omics data is less than the estimated number of cells within the specimen.

[0024] In some embodiments, each region comprises a number of subregions equal to the estimated number of cells for each region.

[0025] In some embodiments, assigning single cells to spatial coordinates from the set of single-cell omics data comprises using the computer processing system to generate, for each region containing a proportion of a cell type in the set of one or more cell types, a matrix of single-cell omics profiles and a matrix of subregion omics profiles of specimens, and using the computer processing system to utilize the matrices to determine a globally optimal solution.

[0026] In some embodiments, the method further comprises using the computing system to determine the globally optimal solution by a sum of assignments of single cells to subregions that minimizes a linear cost function.

[0027] In some embodiments, the spatial omics is one of spatial transcriptomics, spatial genomics, spatial epigenomics, spatial methylomics, spatial proteomics, or spatial metabolomics.

[0028] In some embodiments, the method further comprises extracting analyte source material from each of the plurality of regions for performing the spatial omics, wherein the analyte source material is extracted via laser capture microdissection, recursive microdigestion, or in situ capture.

[0029] In some embodiments, the spatial omics is spatial transcriptomics. The method further comprises determining the expression of a plurality of transcripts via in situ hybridization.

[0030] In some embodiments, estimating the number of cells per region comprises using the computing system to estimate the number of cells per region of the plurality of regions based on the amount of specimen source material derived from each region as determined by the spatial omics data.

[0031] In some embodiments, estimating the number of cells per region comprises using the computing system to estimate the number of cells per region via cell segmentation.

[0032] In some embodiments, each region of said plurality of regions is examined for segmented nuclear or cell membrane staining, and said estimation of said number of cells is based on a nuclear count or a cell membrane count.

[0033] In some embodiments, estimating the proportion of each cell type of said plurality of cell types is estimated by deconvolution.

[0034] In some embodiments, the deconvolution method is determined from the spatial omics data and an a priori defined reference.

[0035] In some embodiments, the deconvolution method is determined by Spatial Seurat, RCTD, SPOTlight, cell2location, or CIBERSORTx.

[0036] In some embodiments, the step of querying the reference single-cell omics data to match the number of cells for each cell type of the plurality of cell types further comprises using the computing system to remove single-cell omics data of one or more single cells within the reference single-cell omics data when the number of single cells within the reference single-cell omics data is greater than the estimated number of cells within the specimen.

[0037] In some embodiments, querying the reference single-cell omics data to match the number of cells for each cell type of the plurality of cell types further comprises using the computing system to add single-cell omics data of one or more single cells within the reference single-cell omics data when the number of single cells within the reference single-cell omics data is less than the estimated number of cells within the specimen.

[0038] In some embodiments, adding the single-cell omics data of one or more single cells comprises overlapping the single-cell omics data of said single-cell omics data.

[0039] In some embodiments, adding the single-cell omics data of one or more single cells comprises generating single-cell omics data representative of said single-cell omics data.

[0040] In some embodiments, each region comprises a number of subregions equal to the estimated number of cells for each region.

[0041] In some embodiments, assigning a single cell to spatial coordinates from the set of single-cell omics data comprises using the computer processing system to generate a matrix of the single cell in the single-cell omics profile and a matrix of the subregions in the specimen omics profile, and using the computer processing system to utilize the matrices to determine a globally optimal solution.

[0042] In some embodiments, the method further comprises using the computing system to determine the globally optimal solution by a sum of assignments of single cells to subregions that minimizes a linear cost function.

[0043] In some embodiments, the step of determining the globally optimal solution further comprises using the computing system to find the globally optimal solution via a Jonker-Volgenant algorithm based on shortest augmenting paths.

[0044] In some embodiments, determining the globally optimal solution includes using the computing system to find the globally optimal solution via a cost scaling push-relabel method.

[0045] In some embodiments, the spatially resolved map of the specimen has single-cell resolution.

[0046] In some embodiments, a method is for generating a spatially resolved map of a specimen via spatial transcriptomics. The method includes using a computer system to obtain spatial transcriptomics data from a plurality of regions covering the specimen. The specimen is a collection of cells including a plurality of cell types. The method includes using the computer system to estimate the number of cells per region of the plurality of regions from the spatial transcriptomics data. The method includes using the computer system to estimate the proportion of each cell type of the plurality of cell types from the spatial transcriptomics data. The method includes using the computer system to query reference single-cell transcriptomics data to match the number of cells for each cell type of the plurality of cell types to generate a set of single-cell transcriptomics data for spatial assignment. The method includes using the computer system to assign single cells to spatial coordinates from the set of single-cell transcriptomics data to generate a spatially resolved map of the specimen, based on a globally optimal solution.

[0047] In some embodiments, the method includes extracting RNA from each region of the plurality of regions, wherein the RNA is extracted via in situ capture. The method includes sequencing the extracted RNA from each region of the plurality of regions to generate the spatial transcriptomics data from the plurality of regions.

[0048] In some embodiments, the RNA is extracted from each region of the plurality of regions via 10xGenomics Visium, or NanoString GeoMX.

[0049] In some embodiments, the sequencing is performed by one of the following techniques: whole exome sequencing, capture targeted sequencing, amplification-based targeted sequencing, random priming-based sequencing, or end-biased sequencing.

[0050] In some embodiments, the method comprises determining the expression of a plurality of transcripts via in situ hybridization to generate said spatial transcriptomics data from said plurality of regions.

[0051] In some embodiments, said expression of said plurality of transcripts is determined via Vizgen MERSCOPE, NanoString CosMX, 10xGenomics Xenium, or hybridization-based in situ sequencing.

[0052] In some embodiments, estimating the number of cells per region comprises using the computer processing system to estimate the number of cells per region of the plurality of regions based on the number of detectably expressed genes.

[0053] In some embodiments, the number of detectably expressed genes is determined by the number of unique molecular identifiers.

[0054] In some embodiments, estimating the number of cells per region comprises using the computing system to estimate the number of cells per region via cell segmentation.

[0055] In some embodiments, each region of said plurality of regions is examined for segmented nuclear or cell membrane staining, and said estimation of said number of cells is based on a nuclear count or a cell membrane count.

[0056] In some embodiments, estimating the proportion of each cell type of said plurality of cell types is estimated by deconvolution.

[0057] In some embodiments, the deconvolution method is determined from the spatial omics data and an a priori defined reference.

[0058] In some embodiments, the deconvolution method is determined by Spatial Seurat, RCTD, SPOTlight, cell2location, or CIBERSORTx.

[0059] In some embodiments, querying the reference single-cell transcriptomics to match the number of cells for each cell type of the plurality of cell types further comprises using the computing system to remove single-cell transcriptomics data of one or more single cells within the reference single-cell transcriptomics data when the number of single cells within the reference single-cell transcriptomics data is greater than the estimated number of cells within the specimen.

[0060] In some embodiments, querying the reference single-cell transcriptomics to match the number of cells for each cell type of the plurality of cell types further comprises using the computing system to add single-cell transcriptomics data of one or more single cells within the reference single-cell transcriptomics data when the number of single cells within the reference single-cell transcriptomics data is less than the estimated number of cells within the specimen.

[0061] In some embodiments, adding the single cell transcriptomics data of the one or more single cells comprises overlapping the single cell transcriptomics data of said single cell transcriptomics data.

[0062] In some embodiments, adding the single cell transcriptomics data of the one or more single cells comprises generating single cell transcriptomics data representative of said single cell transcriptomics data.

[0063] In some embodiments, each region comprises a number of subregions equal to the estimated number of cells for each region.

[0064] In some embodiments, assigning single cells to spatial coordinates from the set of single-cell transcriptomics data comprises using the computer processing system to generate a matrix of single cells in the single-cell transcriptomics profile and a matrix of subregions in the specimen transcriptomics profile, and using the computer processing system to utilize the matrices to determine a globally optimal solution.

[0065] In some embodiments, the method further comprises using the computing system to determine the globally optimal solution by a sum of assignments of single cells to subregions that minimizes a linear cost function.

[0066] In some embodiments, the step of determining the globally optimal solution further comprises using the computing system to find the globally optimal solution via a Jonker-Volgenant algorithm based on shortest augmenting paths.

[0067] In some embodiments, the determining the globally optimal solution step further comprises using the computing system to find the globally optimal solution via a cost-scaling push-relabel method.

[0068] In some embodiments, the spatially resolved map of the specimen has single-cell resolution.

[0069] In some embodiments, a method is for diagnosing a medical disorder based on a spatial signature. The method includes plotting a spatially resolved map of a tissue specimen extracted from a patient. Plotting the spatially resolved map includes generating spatial omics data from multiple regions covering the tissue specimen. Plotting the spatially resolved map includes querying, using a computing system, reference single-cell omics data to generate a set of single-cell omics data for spatial assignment. Plotting the spatially resolved map includes assigning, based on a globally optimal solution, single cells to spatial coordinates from the set of single-cell omics data to generate the spatially resolved map of the tissue specimen. The method includes evaluating the spatially resolved map to detect the presence of a spatial signature, the spatial signature associated with a characteristic of a medical disorder. The method includes determining that the patient has the characteristic of the medical disorder due to the presence of the spatial signature within the spatially resolved map.

[0070] In some embodiments, evaluating the spatially resolved map to detect the presence of the spatial signature further comprises utilizing the rendered spatially resolved map of the tissue specimen as input in a trained machine learning model to generate a likelihood of the characteristic of a medical disorder, wherein determining that the patient has the characteristic of a medical disorder is determined by the likelihood of the characteristic of a medical disorder.

[0071] In some embodiments, said characteristic of a medical disorder is response to treatment.

[0072] In some embodiments, the method further comprises administering said treatment based on the presence of said spatial signature indicating that said patient will respond to said treatment.

[0073] In some embodiments, the characteristic of a medical disorder is the need for further diagnostic procedures to be performed.

[0074] In some embodiments, the method further comprises performing said diagnostic technique based on the presence of said spatial signature indicating that said patient will require said further diagnostic technique to be performed.

[0075] In some embodiments, the method further comprises performing a spatial omics protocol using said tissue sample extracted from said patient, said spatial omics protocol being utilized to render said spatially resolved map.

[0076] In some embodiments, the method further comprises extracting said tissue sample from said patient for performing said spatial omics protocol.

[0077] In some embodiments, the tissue sample comprises tissue of a tumor, tissue of a multicellular organ, tissue infiltrated by immune cells, tissue infected by a pathogen, or tissue that interacts with a microbiota.

[0078] In some embodiments, the medical disorder is cancer, a pathogenic infection, an organ dysfunction, an inflammatory disorder, an autoimmune disorder, diabetes, liver dysfunction, heart disease, or a neurodegenerative disorder.

[0079] In some embodiments, the characteristic of the medical disorder is a particular pathology, the likelihood of successful or unsuccessful treatment, the severity of the medical disorder, the need for a particular medical intervention, or the likelihood of future medical complications.

[0080] In some embodiments, a method is for diagnosing cancer based on a spatial signature. The method includes: plotting a spatially resolved map of a tumor specimen from a patient; plotting the spatially resolved map includes generating spatial omics data from multiple regions covering the tumor specimen; plotting the spatially resolved map includes querying, using a computer system, reference single-cell omics data to generate a set of single-cell omics data for spatial assignment; plotting the spatially resolved map includes assigning, based on a globally optimal solution, single cells to spatial coordinates from the set of single-cell omics data to generate the spatially resolved map of the tumor specimen, using the computer system; evaluating the spatially resolved map to detect the presence of a spatial signature, the spatial signature being associated with a cancer feature; determining that the patient has the cancer feature based on the presence of the spatial signature within the spatially resolved map.

[0081] In some embodiments, evaluating the spatially resolved map to detect the presence of the spatial signature further comprises utilizing the rendered spatially resolved map of the tumor specimen as input in a trained machine learning model to generate a probability of the cancer signature, wherein determining that the patient has the cancer signature of a medical disorder is determined by the probability of the cancer signature.

[0082] In some embodiments, the cancer characteristic is response to treatment, toxicity of treatment, or resistance to treatment.

[0083] In some embodiments, the cancer characteristic is the response to the treatment, and the method includes administering the treatment based on the presence of the spatial signature indicating that the patient will respond to the treatment.

[0084] In some embodiments, the cancer characteristic is the toxicity of the treatment. The method further comprises administering the treatment based on the presence of the spatial signature indicating that the treatment is not toxic to the patient.

[0085] In some embodiments, the cancer hallmark is resistance to the treatment. The method further comprises administering the treatment based on the presence of the spatial signature indicating that the patient will not be resistant to the treatment.

[0086] In some embodiments, the treatment comprises one of immunotherapy, chemotherapy, radiation therapy, targeted therapy, hormone therapy, or surgical resection.

[0087] In some embodiments, the method includes performing a spatial omics protocol using the tumor specimen extracted from the patient, wherein the spatial omics protocol is utilized to render the spatially resolved map.

[0088] In some embodiments, the method comprises extracting said tumor specimen from said patient for performing said spatial omics protocol.

[0089] In some embodiments, the cancer characteristic is cancer progression, metastatic potential, transition from pre-invasive to invasive cancer, or recurrence potential.

[0090] In some embodiments, a method is provided for training a machine learning model to predict spatial signatures from spatially resolved maps. The method includes plotting spatially resolved maps of a plurality of multicellular specimens, each of which is associated with a biological feature. The method includes plotting spatially resolved maps of a plurality of multicellular control specimens, each of which is not associated with the biological feature. Plotting each spatially resolved map includes generating spatial omics data from a plurality of regions covering the specimen. Plotting each spatially resolved map includes using a computer system to query reference single-cell omics data to generate a set of single-cell omics data for spatial assignment. Plotting each spatially resolved map includes using the computer system to assign single cells to spatial coordinates from the set of single-cell omics data to generate the spatially resolved map of the specimen, based on a globally optimal solution. The method includes training a machine learning model with the spatially resolved maps of each of the plurality of multicellular specimens and the plurality of multicellular control specimens to predict the biological feature from the spatially resolved maps.

[0091] In some embodiments, the biological characteristic comprises a pathology, a medical disorder, a health state, a metabolic state, an organ state, activation of multicellular communication, multicellular migration, or a multicellular response to a stimulus.

[0092] In some embodiments, each multicellular specimen is a tumor specimen, and said biological characteristic is a cancer characteristic selected from response to treatment, toxicity of treatment, or resistance to treatment.

[0093] In some embodiments, the treatment comprises one of immunotherapy, chemotherapy, radiation therapy, targeted therapy, hormone therapy, or surgical resection.

[0094] In some embodiments, each multicellular sample is a tumor sample, and the biological characteristic is a cancer characteristic selected from cancer progression, metastatic potential, transition from pre-invasive to invasive cancer, or recurrence potential.

[0095] In some embodiments, the machine learning model is a classifier.

[0096] In some embodiments, the machine learning model is a regressor.

[0097] In some embodiments, the machine learning model incorporates a deep neural network (DNN), a convolutional neural network (CNN), a graph neural network (GNN), a recurrent neural network, a long short-term memory (LSTM) network, a kernel ridge regression (KRR), or gradient-boosted random forest decision trees.

[0098] In some embodiments, the machine learning model incorporates a spatial encoder. [Brief explanation of the drawings]

[0099] The description and claims will be more fully understood by reference to the following figures and data graphs, which are presented as exemplary embodiments of the invention and should not be construed as a complete description of the scope of the invention.

[0100] [Figure 1] Figure 1 provides an example of a computational method for spatial alignment of cells from spatial omics data.

[0101] [Figure 2] Figure 2 provides a computational system for spatial alignment of cells.

[0102] [Figure 3-1] Figure 3A provides a schematic illustrating CytoSPACE versus conventional methods for decoding the cellular composition of bulk spatial transcriptome data.

[0103] [Figure 3-2] Figure 3B provides an overview of a typical CytoSPACE workflow.

[0104] [Figure 4-1] Figure 4A provides a framework for evaluating CytoSPACE using simulated spatial transcriptome datasets with fully defined single-cell composition and spot resolution.

[0105] [Figure 4-2] Figures 4B-4F provide data depicting the maintenance of gene-level spatial dependency in simulated ST data and the impact of controlled noise on scRNA-seq query data. 4B: Pearson correlation analysis of (i) log2 expression levels in scRNA-seq mapped to Slide-seq beads (as part of the simulated ST dataset construction) against (ii) the original Slide-seq beads. The resulting p-values ​​were Benjamini-Hochberg adjusted separately for each brain region and are shown as q-values. *Q<0.05; ***Q<0.001; ****Q<0.0001; ns, not significant. Subiculum. 4C and 4D: Boxplots showing the impact of adding noise to the scRNA-seq query dataset used in the simulation experiments. 4E and 4F: UMAPs of scRNA-seq after adding noise for the mouse cerebellum dataset (4E) and the mouse hippocampus dataset (4F). [Figure 4-3] Same as above. [Figure 4-4] Same as above.

[0106] [Figure 5-1]Figures 5A-5E provide estimates of alignment uncertainty in simulated ST datasets. 6A: Confidence scores for all mapped cells. Figure 6B: Boxplots showing confidence scores stratified by brain region and correct / incorrect assignment. For a given cell of type i, "correct" was defined as a spot containing at least one cell of type i. Statistical significance was determined by a two-tailed Wilcoxon test. ****P<2e-16. 6C: Boxplots showing the area under the curve (AUC) for distinguishing correct from incorrect spots by cell type (n=11, cerebellum; n=17, hippocampus). 6D and 6E: The effect of imposing a 10% confidence score threshold (>0.1) on the proportion (6D) and absolute number (6E) of retained cells. The box centerline, box boundary, and whiskers in panels b and c represent the median, first and third quartiles, and minimum and maximum values, respectively. [Figure 5-2] Same as above.

[0107] [Figure 6-1]Figures 6A-6E provide estimates of cell type proportions and the number of cells per spot in bulk spatial transcriptome data. 6A: Application of Spatial Seurat to infer cell type proportions in simulated ST datasets. Scatter plots show ground truth cell type proportions (x-axis) against estimated proportions (y-axis) for simulated ST data of mouse cerebellar slices (top row) and mouse hippocampal slices (bottom row) with different spot resolutions. For single-cell RNA sequencing data, perturbations were first introduced by adding noise to 5% of the transcriptome. 6B: Scatter plots showing the number of cells per spot (y-axis) estimated by CytoSPACE in simulated ST datasets against the ground truth (x-axis), with an average of 5 cells per spot for mouse cerebellar and hippocampal slices. Relative density is depicted by point size. Agreement and significance were assessed by Pearson r or Spearman rho, and two-tailed t-tests, respectively. 6C: Same as 6B, except correlation coefficients (Pearson and Spearman) are shown for all analyzed spot resolutions. All correlations are significant (P<10-20). 6D: Paired analysis showing the difference in performance between log2-adjusted and non-log-linear scales for predicting the number of cells per spot for all six combinations of spot resolutions (means of 5, 15, and 30) in the simulated ST dataset for Pearson and Spearman correlation coefficients. Statistical significance was calculated by two-tailed paired Wilcoxon test. 6E: Agreement between the number of cells per spot (y-axis) imputed by the default RNA-based approach implemented in CytoSPACE and the cell segmentation algorithm (VistoSeg) applied to paired gene expression data and histological images of adult mouse brain coronal samples profiled with 10x Visium, respectively.The box centerline, box boundaries, and whiskers indicate the median, the first and third quartiles, and the minimum and maximum values ​​within 1.5 times the interquartile range of the box limits, respectively. Linear regression (indicated by 95% confidence intervals) was applied to the medians of the box plots. In 6A and 6B, agreement was assessed by Pearson correlation (r), Spearman correlation (ρ), and / or linear regression (dashed lines). A two-tailed t-test was used to assess whether each correlation result was significantly non-zero. No adjustment for multiple comparisons was made. [Figure 6-2] Same as above.

[0108] [Figure 7-1] Figure 7A provides a heatmap depicting CytoSPACE performance for aligning scRNA-seq data (with 5% added noise) to spatial locations in an ST dataset simulated with an average of 5 cells per spot.

[0109] [Figure 7-2] Figure 7B provides performance across a variety of different methods, mouse brain regions, and noise levels for assigning individual cells to the correct spots in the simulated ST dataset. Each point represents one single cell type (mouse cerebellum, n = 11; mouse hippocampus, n = 17). The box centerline, box boundary, and whiskers indicate the median, first and third quartiles, and minimum and maximum values ​​within 1.5 times the interquartile range of the box limits, respectively. Statistical significance was assessed for CytoSPACE using a two-tailed paired Wilcoxon test. The resulting p-values ​​were Benjamini-Hochberg adjusted for each noise level and tissue type combination and reported as the maximum Q value (*Q<0.05, ***Q<0.001).

[0110] [Figure 7-3]Figure 7C provides an extended benchmark analysis of the simulated ST data (related to Figure 7C). Boxplots depicting the percentage of all single-cell transcriptomes assigned to the correct ST spot are shown for 13 broad-based methods, with different spot resolutions (average of 5, 15, and 30 cells per spot) and scRNA-seq noise levels (perturbations applied to 5%, 10%, and 25% of the transcriptome). Statistical significance was determined using a two-tailed paired Wilcoxon test for CytoSPACE. P values ​​were corrected using the Benjamini-Hochberg method. P values ​​are expressed as q values ​​(**Q<0.01). The box centerline, box boundaries, and whiskers indicate the median, first and third quartiles, and the minimum and maximum values ​​within 1.5 times the interquartile range of the box limits, respectively.

[0111] [Figure 7-4] Figures 7D and 7E provide CytoSPACE, Tangram, and CellTrek alignments for all cell types analyzed in the simulated ST datasets (related to Figure 7A). Heatmaps depicting single-cell mapping accuracy, defined as the percentage of single cells that map correctly to ground truth spots, are shown for the three methods and for all evaluated cell types mapped to the simulated mouse cerebellum ST dataset (n = 11 cell types) (7D) and mouse hippocampus ST dataset (n = 17 cell types) (7E), with an average of 5 cells per spot.

[0112] [Figure 8]Figures 8A-8C provide the performance of CytoSPACE with RCTD. 8A and 8B: Comparison of cell type proportions estimated by Spatial Seurat and RCTD for (8A) a simulated dataset with an average of 5 cells per spot and 5% noise added to the scRNA-seq data, and (8B) a simulated dataset across all analyzed spot resolutions and noise levels. Agreement was assessed by Pearson correlation (r), Spearman correlation (ρ), and linear regression (dashed lines). A two-tailed t-test was used to assess whether each correlation result was significantly non-zero. 8C: Same as Figure 7C, except showing the application of CytoSPACE with RCTD (rather than Spatial Seurat) for cell type proportion estimation versus the selected comparator method. The box centerline, box boundary, and whiskers in b and c indicate the median, first and third quartiles, and the minimum and maximum values ​​within 1.5 times the interquartile range of the box limits, respectively. Statistical significance was determined using a two-tailed paired Wilcoxon test for CytoSPACE. P values ​​were corrected using the Benjamini-Hochberg method. P values ​​are expressed as q values ​​(**Q<0.01).

[0113] [Figure 9]Figures 9A and 9B provide data showing the association between CytoSPACE performance and inferred overall cell type abundance in simulated spatial transcriptomics datasets. 9A: Scatterplot comparing single-cell mapping accuracy in simulated ST datasets (with an average of 5 cells per spot) to the average cell type abundance fraction inferred by Spatial Seurat for all cell types and noise levels. Linearity was determined by Pearson correlation. 9B: Same as 9A, except Pearson correlation significance values ​​are summarized across all evaluated simulated ST datasets, spot resolutions, and noise levels. The box centerline, box boundary, and whiskers indicate the median, first and third quartiles, and the minimum and maximum values ​​within 1.5 times the interquartile range of the box limits, respectively. A two-tailed t-test was used to assess whether each correlation coefficient was significantly non-zero. P values ​​were corrected using the Benjamini-Hochberg method and expressed as q values.

[0114] [Figure 10-1]Figures 10A-10E provide data showing the effect of perturbations on estimates of cell type abundance fractions and the number of cells per spot. 10A: Boxplots showing the effect of perturbations on cell type abundance fraction estimates across five separate trials. Boxplots are expressed relative to the original estimates (left) and absolute units (right) for a mouse cerebellum (top) and hippocampus (bottom) datasets with an average of 5 cells per spot and 5% noise added to the scRNA-seq query dataset. 10B: Boxplots showing CytoSPACE performance on simulated ST datasets before and after perturbations to cell type fractions, for all spot resolutions and scRNA-seq noise levels. 10C: Scatterplots showing the effect of suppressed perturbations on the estimated number of cells per spot for a representative simulated ST dataset (mouse hippocampus with an average of 5 cells per spot). 10D: Boxplot showing the Pearson correlation between perturbed and original estimates of the number of cells per spot for all evaluated simulated ST datasets across five runs. 10E: Boxplot showing CytoSPACE performance for all simulated ST datasets before and after perturbing the estimates of the number of cells per spot (related to 10D). P-values ​​were corrected using the Benjamini-Hochberg method. P-values ​​are expressed as q-values ​​(*Q<0.05; **Q<0.01). The box centerline, box boundaries, and whiskers in 10A, 10B, 10D, and 10E indicate the median, first and third quartiles, and the minimum and maximum values ​​within 1.5 times the interquartile range of the box limits, respectively. [Figure 10-2] Same as above. [Figure 10-3] Same as above.

[0115] [Figure 11]Figures 11A and 11B provide data showing the stability of CytoSPACE's pairwise cell assignment across multiple seeds and distance metrics. 11A: Same as Figure 7C, except that CytoSPACE performance is shown for 10 different random samplings of each scRNA-seq query dataset. Statistical significance was calculated using one-way repeated measures ANOVA. ns, not significant. 11B: Same as Figure 7C, except that CytoSPACE performance on the simulated ST dataset is shown using Pearson correlation, Spearman correlation, or Euclidean distance to calculate the CytoSPACE cost matrix. P values ​​were corrected using the Benjamini-Hochberg method. P values ​​are expressed as q values ​​(**Q<0.01).

[0116] [Figure 12-1] Figures 12A-C provide single-cell RNA-seq data mapped to ST profiles of a wide variety of human tumor specimens. Grey boxes represent cell types without author-provided annotation in the corresponding scRNA-seq atlas. [Figure 12-2] Same as above. [Figure 12-3] Same as above.

[0117] [Figure 13] Figure 13A provides a workflow for assessing spatial enrichment in the center or periphery of tumors. DEGs, differentially expressed genes.

[0118] Figure 13B provides spatial enrichment of T cell exhaustion genes in the T cell transcriptome mapped to melanoma samples by CytoSPACE (first row, panel a). NES, normalized enrichment score.

[0119] Figure 13C provides the same data as Figure 13B, except that NES is shown for six scRNA-seq / ST pairs (n=12 values / box) and three methods.

[0120] [Figure 14-1] Figures 14A and 14D provide spatial enrichment of tumor-associated cell states across various methods and datasets. 14A: Left: Bubble plot (related to Figure 13C) showing spatial enrichment of exhaustion genes in the transcriptomes of CD4 and CD8 T cells, mapped to ST spots by CytoSPACE, Tangram, and CellTrek. Right: Same as Figure 14B, except that performance is shown for Tangram and CellTrek. Single-cell RNA-seq datasets without annotated plasma cells are indicated by gray boxes ("N / A"). Bubbles represent normalized enrichment scores calculated by pre-ranked GSEA. 14D: Percentage of datasets per cell type for which the predicted spatial enrichment direction was correctly inferred by each of the 13 evaluable methods for each of the gene sets analyzed in this study (n=11 different gene sets with 12 data points per method, since canonical exhaustion genes were analyzed for CD4 T cells and CD8 T cells). The box centerline, box boundaries, and whiskers indicate the median, first and third quartiles, and minimum and maximum values ​​within 1.5 times the interquartile range of the box limits, respectively. Statistical significance was determined using a two-tailed paired Wilcoxon test for CytoSPACE, and p-values ​​were corrected using the Benjamini-Hochberg method.

[0121] [Figure 14-2] Figure 14B provides spatial enrichment of CE9- and CE10-specific cell states in data mapped by CytoSPACE and analyzed by pre-ranked GSEA. Datasets without annotations are shown in gray.

[0122] Figure 14B provides spatial enrichment of CE9- and CE10-specific cell states in the data across 13 methods and 66 combinations of dataset pairs and cell states. To unify the expected enrichment direction of cell states, the NES value for CE10 was multiplied by -1.

[0123] [Figure 15-1] Figures 15A-15F provide data demonstrating the robustness of CytoSPACE applied to tumor spatial transcriptomics datasets. 15A: Same as Figure 9A, except that inferred cell type abundances are analyzed against the average CytoSPACE performance across six tumor ST datasets. In this case, performance is defined as cell state enrichment, as estimated by the normalized enrichment score (NES). Notably, the NES value for CE10 was multiplied by -1 to unify the expected enrichment direction. 15B: Same as Figure 10A, except that cell type proportion perturbations are shown for a representative CRC ST dataset. 15C: Same as Figures 13C and 14C, except that the effect of perturbing cell type proportions on CytoSPACE performance is shown. 15D: Boxplots showing the Pearson correlation between perturbed and original estimates of the number of cells per spot for all six tumor ST datasets across five runs. 15E: CytoSPACE performance for all six tumor scRNA-seq / ST dataset pairs before and after perturbing the estimates of the number of cells per spot across five trials, along with "flattening" the number of cells per spot, in which spots were assigned to the same number of cells. P-values ​​were corrected using the Benjamini-Hochberg method and expressed as q-values. *Q<0.05; **Q<0.01. 15F: Same as Figures 13C and 14C, except NES values ​​are compared for cell state enrichment between the default seed and nine additional random samplings of the scRNA-seq query dataset. [Figure 15-2] Same as above.

[0124] [Figure 16] Figures 16A and 16B provide single-cell spatial analysis of TREM2+ and FOLR2+ macrophage status across various datasets and methods. 16A: Predicted spatial localization of TREM2+ and FOLR2+ macrophages in human tumors (Nalio Ramos et al.). 16B: Boxplots comparing the log2 fold change in TREM2 and FOLR2 expression in single macrophage / monocyte transcriptomes grouped into "near" (Euclidean distance to tumor = 0) and "distant" (Euclidean distance to tumor > 0) categories. Each point represents the scRNA-seq / ST pair analyzed in Figure 14B. The box centerline, box border, and whiskers represent the median, first and third quartiles, and minimum and maximum values, respectively. Comparisons of two groups were performed using a two-tailed paired Wilcoxon test (indicated by the horizontal lines above each pair of TREM2+ and FOLR2+ boxes). ns, not significant.

[0125] [Figure 17-1] Figures 17A and 17B provide UMAP projections of the scRNA-seq tumor atlas labeled by predicted spatial location. UMAP embedding showing all single-cell transcriptomes mapped by CytoSPACE to ST samples. Cells are colored by lineage (17A) and relative distance to tumor cells (17B). [Figure 17-2] Same as above.

[0126] [Figure 18-1] Figure 18A provides a schematic of the mouse nephron / collecting duct system. Known locations of epithelial tissue are represented by numbers.

[0127] [Figure 18-2]FIG. 18B provides the epithelial cell transcriptome from the mouse kidney scRNA-seq atlas mapped by CytoSPACE to a 10x Visium sample of normal mouse kidney, shown with jitter within the assigned spots.

[0128] Figures 18C and 18D provide single-cell mapping of a normal mouse kidney using CytoSPACE. 18C: Mouse kidney scRNA-seq atlas mapped to 10x Visium samples of normal mouse kidney, showing epithelial cell transcriptomes mapped by CytoSPACE and colored by the known zonal area of ​​each cell (as in Figure 18A) overlaid on the Visium histological image. The zonal colors of individual epithelial cells mapped by CytoSPACE are averaged per spot. 18D: Scatter plot showing the statistical significance of co-association (x-axis) between podocytes (epithelial state 1) and all other cell types / states, and parietal cells (epithelial state 2) mapped by CytoSPACE. A spot was scored as "present" if at least one cell of a given cell type was mapped by CytoSPACE; otherwise, it was scored as "absent." Subsequently, the significance of co-association was calculated using two-tailed Fisher's exact test, −log 10 Expressed as p-values. Self-comparisons are represented by NA (not applicable).

[0129] [Figure 19-1] Figures 19A-C provide epithelial cell transcriptomes from a mouse kidney scRNA-seq atlas (each cell is colored by its known distance to the inner medulla) mapped by CytoSPACE, Tangram, and CellTrek to a 10x Visium sample of normal mouse kidney.

[0130] [Figure 19-2]FIG. 19D provides the degree of agreement between the predicted distance of each epithelial condition and the known distance to the base of the inner medulla.

[0131] [Figure 20-1]Figures 20A-20F provide a CytoSPACE-guided reconstruction of the nephron / collecting duct system. 20A: Similar to Figure 18A, except showing epithelial cell states colored by physically adjacent phenotypes. The corresponding cell state ontology is provided in Table 5. 20B: UMAP embedding of the normal mouse kidney scRNA-seq atlas (mapped by CytoSPACE) and colored as in 20A. 20C: Left: Heatmap (related to Figure 19A) showing pairwise spatial overlap between all renal epithelial cell states mapped by CytoSPACE to a 10x Visium sample of normal mouse kidney. Overlap was determined by Jaccard coefficient and normalized to the maximum value per row. Right: Heatmap showing known neighboring states (as in 20A). 20D: Spring layout of the data in 20C, where each cell state is plotted along with its four nearest neighbors (in rank space) as inferred by CytoSPACE. Selected kidney structures are shown. Ridge thickness is proportional to the degree of overlap in rank space. Statistical significance was calculated by a one-sided permutation test. 20E: Scatterplot comparing (i) the distance (y-axis) between each state i and its nth nearest neighbor (state j) inferred by CytoSPACE (median rank across all evaluable states) with (ii) the distance (x-axis) between state i and its ground truth nth nearest neighbor. The distance between states was calculated as the number of known continuous states between i and j. 1 to 10 nearest neighbors were evaluated. Agreement was assessed by Pearson correlation and Lin's concordance correlation coefficient (CCC). A two-tailed t-test was applied to determine whether the correlation coefficient was significantly non-zero. 20F: Same analysis as in panel 20E, except for all evaluated methods, comparing performance using the CCC. DEEPsc assigned all cells to the same spot and was omitted. [Figure 20-2] Same as above.

[0132] [Figure 21-1]Figure 21A provides the MERSCOPE profile of a breast cancer specimen (left), colored by cell type. Right: scRNA-seq data mapped to the MERSCOPE profile by CytoSPACE, where previously annotated cell types from this scRNA-seq atlas are color-coded.

[0133] [Figure 21-2] Figure 21B provides enrichment (pre-ranked GSEA) of CD4 T cell status within the tumor region comparing scRNA-seq data (CytoSPACE) mapped to MERSCOPE with MERSCOPE alone.

[0134] [Figure 22-1]Figures 22A-22I provide a technical evaluation of CytoSPACE applied to single-cell ST data. 22A: Workflow for the analysis in Figures 22B-22E. 22B: Left: MERSCOPE reference profile of a breast cancer specimen, with major cell types differentiated by color. Right: MERSCOPE query dataset mapped to the reference profile by CytoSPACE, with query cell types differentiated by color. 22C: Phenotypic concordance between reference and query cells after alignment. 22D: Mapping accuracy analysis showing the significance of Pearson correlation between the log2 GEPs of (i) the reference cells and (ii) the query cells mapped to the reference cells, stratified by cell type. The matrix diagonal captures the comparison between query cell GEPs and their corresponding reference cell assignments. Unmatched pairwise combinations (off-diagonal terms) represent cell-type-specific controls. 22E: Analysis of the retention of pairwise distances between cells after mapping with CytoSPACE. For each cell type, a scatter plot shows the retention index, defined as the Pearson correlation between matrix Q and matrix R, versus the variance in matrix R (panel a). The significance of the linear regression line was assessed by a two-tailed t-test. 22F: Analysis of concordance at the gene level, showing the significance of the Pearson correlation between (i) the scRNA-seq data mapped to MERSCOPE (Wu et al.) and (ii) the log2 expression levels of the original MERSCOPE data, analyzed separately for each gene (commonly n=497) and cell type. As a control, disconfirmed pairwise combinations of the same 497 genes were also evaluated (off-diagonal terms in the correlation matrix). 22G: Concordance of cell type labels between MERSCOPE and the aligned scRNA-seq. 22H: Left: tumor and adjacent normal region. Right: FOLR2 expression in single-cell transcriptomes annotated as "macrophage / monocyte" and mapped by CytoSPACE (Wu et al.) showing elevated levels in adjacent normal regions, consistent with expectations. 22I, Same as Figure 21B, except for CD8 T cells.In d and f, up to 1,000 cells per cell type and 1,000 off-diagonal correlations were randomly sampled for analysis. For each cell type, p-values ​​were Benjamini-Hochberg adjusted and expressed as -log10 q-values ​​(multiplied by -1 for negative correlations). Group comparisons in d and f were assessed using a one-sided Wilcoxon test on the matrix diagonal, and p-values ​​were Benjamini-Hochberg adjusted. ****Q<0.0001. In d and f, the center line, borders, and whiskers indicate the median, first and third quartiles, and the minimum and maximum values ​​within 1.5 times the interquartile range of the box limits, respectively. GEP, gene expression profile. [Figure 22-2] Same as above. [Figure 22-3] Same as above. [Figure 22-4] Same as above. DETAILED DESCRIPTION OF THE INVENTION

[0135] Detailed Description Turning now to the figures and data, systems and methods are provided for spatially aligning cells within a population of cells. In various embodiments of the systems and methods of the present invention, spatial omics analysis is performed to assign cell types to specific locations within a spatially defined population, resulting in a mapping of cells within the population. The systems and methods of the present invention can be performed on a variety of multicellular networks that include multiple cell types within an analysis region. The systems and methods of the present invention can determine the spatial relationships between cell types within an analysis region, resulting in single-cell resolution within the region. The results of the systems and methods of the present invention can be mapped, annotated, and visualized, thereby elucidating the spatial interactions of each cell within the multicellular network being evaluated.

[0136] The various systems and methods of the present invention can be applied to various omics. The term "omics" is understood to refer to any of a variety of substantially comprehensive cellular analyses. In some embodiments, "omics" refers to transcriptomics, genomics, epigenomics, methylomics, proteomics, and metabolomics. Furthermore, as understood in the art, any and all of these omics can be utilized for spatial analysis, and thus the systems and methods of the present invention can be adapted to specific parameters for performing such analyses. In general, the systems and methods described herein can be applied when any of a specific set of omics can describe one cell as distinct from another cell within a population. For example, genomics can be utilized to distinguish cells within a population of cells with mixed genomics, such as in mixed-species environments (e.g., biofilms, microbiota), and environments of high genomic heterogeneity (e.g., tumors, neural tissue). For more information about spatial transcriptomics, see, for example, M. Asp. J. Bergenstrahle, and J. Lundeberg, Bioessays. 2020 Oct; 42 (10): e1900221; and P. L. Stahl et al., Science. 2016 Jul. 1; 353 (6294): 78-82; the disclosures of each of which are incorporated herein by reference. For more information about spatial genomics, see, for example, T. Zhao et al., Nature. 2022 Jan; 601 (7891): 85-91; and Rusheth et al., Nat Biotechnol. 2019 Aug; 37 (8): 877-883; the disclosures of which are incorporated herein by reference. For more information regarding spatial epigenomics, see, e.g., T. Lu et al., Cell. 2022 Nov 10;185(23):4448-4464.e17; the disclosure of which is incorporated herein by reference.For more details regarding spatial methylomics, see, for example, N. Loyfer et al., Nature. 2023 Jan;613(7943):355-364; the disclosure of which is incorporated herein by reference. For more details regarding spatial proteomics, see, for example, E. Lundberg and G.H.Borner, Nat Rev Mol Cell Biol. 2019 May;20(5):285-302; the disclosure of which is incorporated herein by reference. For more details regarding spatial metabolomics, see, for example, L.R. Conroy et al., Nat Commun. 2023 May 13;14(1):2759; the disclosure of which is incorporated herein by reference. Furthermore, it should be understood that various omics can be combined for spatial analysis, as described, for example, in D. Zhang et al., Nature. 2023 Apr;616(7955):113-122 (the disclosure of which is incorporated herein by reference).

[0137] Various systems and methods of the present invention refer to the spatial alignment of cell types. The term "cell type" refers to a specific label of a cell that can be distinguished from other cells based on its omics profile. Various contributions can affect the omics profile, and therefore, cell type is broadly interpreted to potentially include subtle variations in its detectable omics profile. In some instances, cell types refer to cells with specific functions. For example, cell types may refer to various immune cells (e.g., macrophages, CD4 T cells, CD8 T cells, B cells, etc.) or various cells of an organ system (e.g., cardiomyocytes, pericytes, myeloid cells, fibroblasts, adipocytes, endothelial cells, etc.). In some instances, cell types refer to levels of developmental maturation or stemness. For example, cell types may refer to various cells of hematopoietic development (e.g., hematopoietic stem cells, myeloid progenitor cells, myeloblasts, monocytes, macrophages). In some instances, cell types refer to genetic heterogeneity. For example, tumor cells may have different amounts of somatic mutations that can be distinguished. In some instances, cell types represent cells that have responded in a particular way to one or more stimuli. For example, T cells that are naive to cancer and T cells that have infiltrated cancer. For example, epithelial cells that have contact with pathogens and epithelial cells that are naive to pathogen contact. In some instances, cell types represent cells of various species or lineages. For example, cells in a microbiota may contain various types of bacteria. In some instances, cell types represent a mixture of various definitions (e.g., cells with a specific function, a specific developmental maturity, a specific somatic genetic makeup, and / or a specific response to a stimulus).

[0138] Spatial omics, particularly spatial transcriptomics, has become a powerful tool for delineating spatial differences (e.g., spatial expression patterns) in spatially organized specimens (e.g., primary tissue specimens). Commonly used platforms remain limited to bulk omics measurements, where each spatially resolved expression profile is derived from a region with as many as 10, 20, or 40 cells (or more). To compensate for this, several computational methods have been developed to infer the cellular composition in a given bulk omics sample representing a region. Most such methods use reference profiles derived from representative single-cell omics data from specific cell types to deconvolve these into a matrix of cell type proportions (e.g., a region containing X% cell type 1, Y% cell type 2, and Z% cell type 3). These methods lack granularity, which hinders the discovery of spatially defined various cellular states, their interaction patterns, and the communities surrounding them.

[0139] Alternatively, spatial omics can be performed in situ, which means that biomolecular evaluation is performed and visualized inside the specimen. The advantage of in situ spatial omics is that it provides subcellular resolution. However, the improved resolution comes with the drawback that evaluation is limited to a small number of biomolecules (up to about 1000 probes) and lacks complex analysis of those molecules (e.g., somatic mutations within genes cannot be evaluated). Therefore, these methods lack the depth and complexity of omics that would be desired at single-cell resolution.

[0140] To address these limitations, the systems and methods described herein were developed to provide single-cell spatial organization. The systems and methods of the present invention utilize efficient computational approaches to align individual cells from cell-type references to precise spatial locations within spatially organized regions of a specimen. Unlike other methods, the solution described herein formulates single-cell spatial allocation as a convex optimization problem, which is solved using a global approach to find the optimum or minimum error. This system and method produces optimal spatial allocation results and has greater noise tolerance than other common methods. The output is a reconstructed spatial alignment of cells that can be visualized down to single-cell resolution, thereby enabling a better understanding of multicellular ecosystems. As illustrative examples, tumor microenvironments, sites of immune cell infiltration, various multicellular organ systems, host-pathogen interactions, and microbiota ecosystems can be evaluated to depict the spatial organization and communication between various cells.

[0141] The systems and methods of the present invention can spatially align cells using spatial omics and single-cell references for cell types as inputs. For example, in the field of spatial transcriptomics, a set of reference single-cell RNA-seq results classified by cell type can be used. The systems and methods of the present invention can use the inputs to determine the abundance fraction of each cell type in the spatial omics sample and the number of cells per spot. In some embodiments, the abundance fraction can be determined using a deconvolution tool such as (for example) Spatial Seurat, RCTD, SPOTlight, cell2location, or CIBERSORTx. In some embodiments, the abundance fraction can be determined iteratively as cells are mapped. In some embodiments, the number of cells is inferred by estimating RNA abundance. In some embodiments, the number of cells is determined by cell segmentation. The systems and methods of the present invention further randomly sample the single-cell reference to match the predicted number of cells per cell for a cell type. The systems and methods of the present invention further assign each cell to a spatial coordinate determined by a convex optimization method. In some embodiments, the optimization method minimizes a correlation-based cost function, constrained by the estimated number of cells per region, via a shortest augmentation path optimization algorithm.

[0142] The innovative systems and methods described herein convert spatial omics data (at a resolution of approximately 5 to 20 cells per region) into a map of the spatial arrangement of cells at single-cell resolution. These systems and methods provide a dramatic improvement to computational spatial mapping of cells that has not yet been achieved in the art. This improvement can be easily understood by the results of implementing the methods of the present invention, which provide high-accuracy single-cell resolution output that can be visualized in a color-coded map. The examples described herein compare the innovative methods of the present invention with conventional state-of-the-art methods, and the results of the comparison clearly show a dramatic improvement. Specimen spatial alignment

[0143] Some embodiments are directed to assigning single cells to spatial alignment from spatial omics data. In many embodiments, spatial omics data is collected from multiple regions and compared to single-cell omics data. In some embodiments, a global optimization solution is utilized to align single cells to yield a spatial arrangement.

[0144] FIG. 1 provides a computational method for generating a spatial arrangement of a single cell based on spatial omics data. Method 100 can begin by obtaining spatial omics data from multiple regions of a specimen. The specimen is a collection of cells with multiple cell types defined by a spatial arrangement. In some embodiments, the specimen is derived from an in vivo specimen source. In some embodiments, the specimen is derived from an ex vivo specimen source. In some embodiments, the specimen is derived from an environmental specimen source. The spatially defined entity can be a primary tissue specimen, a biofilm or other organized cell growth, a cell culture, an organoid, or any other specimen that can be defined by multiple cell types in a defined spatial arrangement. In various examples, the specimen is a tumor, a multicellular organ specimen, a multicellular organoid specimen, a specimen containing tissue infiltrated by immune cells, a specimen containing host tissue and pathogens, or a specimen containing host tissue and microbiota. The omics data can be derived from a biological specimen or a fixed specimen, as appropriate for the methodology used to perform spatial omics assessments.

[0145] Any spatial omics data can be utilized as long as the spatial omics data can identify the cell type of the specimen. Spatial omics that can be evaluated include, but are not limited to, spatial transcriptomics, spatial genomics, spatial epigenomics, spatial methylomics, spatial proteomics, or spatial metabolomics. Depending on the omics type, various biomolecules can be collected from the cell type and processed to perform omics analysis.

[0146] To perform spatial transcriptomics, RNA can be extracted from a specimen, processed, and evaluated. Any suitable method for evaluating the transcriptome can be used, including, but not limited to, in situ hybridization, in situ sequencing, microarrays, and RNA sequencing. RNA sequencing can be whole-exome sequencing, capture-targeted sequencing, amplification-based targeted sequencing, random-priming-based sequencing, or end-biased sequencing, with or without unique molecular identifiers (UMIs). When deciding how to evaluate the transcriptome, there is a balance between the depth of the genes analyzed and the spatial resolution. Illustratively, in situ methods have subcellular resolution but cannot assess a large depth of genes, whereas sequencing methods have lower resolution (approximately 5 to 20 cells per region) but can provide near-complete transcriptome depth. Numerous platforms have been developed to perform spatial transcriptomics. Examples for in situ hybridization transcriptomics include, but are not limited to, Vizgen MERSCOPE, NanoString CosMX, 10xGenomics Xenium, and hybridization-based in situ sequencing (HybISS) (for more on MERSCOPE, see J. Liu et al., Life Sci Alliance. 2022 Dec 16;6(1):e202201701; for more on CosMX, see S. He et al., Nat Biotechnol. 2022 Dec;40(12):1794-1806; for more on Xenium, see S. M Salas et al., bioRxiv 2023.02.13.528102; for more on HybISS, see D. Gyllborg et al., Nucleic Acids Res. 2020 Nov 2020). 4;48(19):e112, the disclosures of each of which are incorporated herein by reference).Examples for RNA-seq transcriptomics include (but are not limited to) 10xGenomics Visium and NanoString GeoMX, each of which can be combined with a high-throughput sequencer (e.g., Illumina HT series) (for further information regarding Visium, see PL Stahl et al., Science. 2016 Jul 1;353(6294):78-82; for further information regarding GeoMX, see K. Roberts, bioRxiv 2021.03.20.436265; the disclosures of each of which are incorporated herein by reference).

[0147] To perform spatial transcriptomics, genomic DNA can be extracted from a specimen, processed, and evaluated. Any suitable method for evaluating the genome can be used, including but not limited to microarrays and DNA sequencing. DNA sequencing can be whole-genome sequencing, whole-exome sequencing, capture-targeted sequencing, or amplification-based targeted sequencing.

[0148] To perform spatial transcriptomics, DNA or RNA can be extracted from specimens, processed, and evaluated. Any suitable method for evaluating epigenome can be used, including but not limited to, chromatin immunoprecipitation sequencing, chromatin accessibility assessment, and those inferred from RNA sequencing. Chromatin accessibility assessment can be performed (for example) using the assay for transposase-accessible chromatin with sequencing (ATAC-Seq).

[0149] To perform spatial methylomics, DNA or RNA can be extracted from a specimen, processed, and evaluated. Any suitable method for evaluating the methylome can be used, including, but not limited to, methylation assessment and methods inferred from RNA sequencing. Methylation assessment can be performed using (for example) bisulfite conversion sequencing or enzymatic methyl sequencing (EM-Seq).

[0150] To perform spatial genomics, proteinaceous species can be extracted from a specimen, processed, and evaluated. Any suitable method for evaluating the proteome can be utilized, including, but not limited to, mass spectrometry and protein microarrays.

[0151] To perform spatial metabolomics, metabolite species can be extracted from a specimen, processed, and evaluated. Any suitable method for evaluating the metabolome can be used, including, but not limited to, mass spectrometry and nuclear magnetic resonance spectroscopy.

[0152] In many embodiments, analyte source material for performing spatial omics is captured in multiple regions. Generally, the multiple regions encompass the analytes to be evaluated, or at least a portion thereof. Various methods can be used to capture analyte source material from the multiple regions, and the analyte source material may depend on various protocols and the particular type of omics to be evaluated. In some embodiments, analyte source material is extracted from the multiple regions using laser capture microdissection. In some embodiments, analyte source material is extracted from the multiple regions using iterative microdigestion. In some embodiments, analyte source material is extracted from the multiple regions using in situ capture.

[0153] In many embodiments, spatial omics is performed in situ, which means that omics analysis is performed directly on intact specimens.Generally, fixed specimens (for example, formalin-fixed paraffin-embedded tissues) or fresh frozen specimens are permeabilized, and the detection of biomolecules for omics analysis is performed therein.Because in situ omics is performed directly on specimens and provides subcellular resolution, multiple regions can be defined as desired by users, and can be as granular as single cells.

[0154] After evaluating multiple regions, spatial omics data can be searched.Spatial omics data can be further processed to ensure high data quality for further downstream evaluation.For example, the reads that map poorly in sequencing results can be discarded.Many other processing steps can be carried out as is routine when evaluating omics data.

[0155] Method 100 determines (103) an estimate of the number of cells per region. The number of cells per region provides an estimate of the average region size and the number of subregions within each region. Several different techniques can be utilized to estimate the number of cells per region. In some embodiments, the number of cells per region is estimated based on omics analysis. In some embodiments, the number of cells per region is estimated via cell segmentation.

[0156] To estimate cell number from omics analysis, the assumption is that sample source material derived from multiple cells can be used to estimate cell number. For example, the number of detectably expressed genes per cell corresponds well to the total captured mRNA content, which can be used to determine the number of cells. For example, when single-cell RNA-seq is performed, the number of detectably expressed genes is used to determine when the results contain more than a single cell (e.g., a duplicate result). When transcriptome analysis is performed, the number of unique molecular identifiers provides a surrogate for the number of detectably expressed genes and can therefore provide an estimate of the number of cells per region. Similar analyses can be performed for other omics using DNA, proteinaceous species, and metabolite inputs.

[0157] To estimate cell number via cell segmentation, regions of the specimen are examined for segmented nuclear and / or cell membrane staining. Based on the nuclear or cell membrane count, a cell count per region is estimated. Various image processing methods can be utilized to perform cell segmentation, such as VistoSeg and CellPose (M. Tippani et al., bioRxiv, 2021.2008.2004.452489 (2022); and C. Stringer et al., Nature Methods 18, 100-106 (2021); the disclosures of which are incorporated herein by reference).

[0158] Method 100 estimates the proportions of multiple cell types in the spatial omics data (105). Various techniques can be used to estimate the cell proportions. In some embodiments, the cell proportions are estimated by deconvolution. In some embodiments, the cell proportions are calculated as part of an optimization solution for assigning single cells to spatial coordinates, as discussed in more detail in step 111.

[0159] Numerous cellular deconvolution methods for estimating cellular proportions for omics data from multiple regions are available as computational applications. Generally, a global determination of cell type proportions within a specimen is determined from bulk omics profiles using an a priori defined reference (typically guided single-cell analysis). Various methods for cell deconvolution that can be utilized include (but are not limited to) Spatial Seurat, RCTD, SPOTlight, cell2location, and CIBERSORTx (for Spatial Seurat, see T. Stuart et al., Cell 177, 1888-1902 e1821 (2019); for RCTD, see D. M. Cable et al., Nature Biotechnology 40, 517-526 (2022); for SPOTlight, see M. Elosua-Bayes et al., Nucleic Acids Res 49, e50 (2021); for cell2location, see V. Kleshchevnikov et al., Nature Biotechnology 40, 661-671 (2022); for CIBERSORTx, see A. M. Newman, Nat Biotechnol 37, 773-782 (2019); the disclosures of each of which are incorporated herein by reference.

[0160] In some embodiments, cell deconvolution is performed on individual regions (rather than globally) to generate cell proportions for each region. Local cell deconvolution can be performed on each region of a specimen, or on a specific set of regions. An advantage of performing local cell deconvolution is that if only a specific cell type (or set of cell types) needs to be evaluated for spatial location, regions lacking that cell type (or set of cell types) can be ignored when assigning cells to spatial coordinates.

[0161] The method 100 obtains reference single-cell omics data (107). The single-cell omics data is utilized to infer single-cell omics data for a particular cell type. The reference single-cell data can be obtained via a database, a publicly available (or otherwise available) data set, or can be determined experimentally. To determine experimentally, cells of a particular cell type can be isolated (e.g., via flow cytometry) and their single-cell omics data can be determined.

[0162] Method 100 queries the reference single-cell omics data to match the number of cells for each cell type of a plurality of cell types to create a set of single-cell omics for spatial assignment (109). This step aligns the queried reference single-cell omics data with the specimen's omics data. The alignment is repeated for each cell type. In some embodiments, specimen cell types that are underrepresented or unrepresented may be excluded from the analysis because their contribution may not be important to the final spatial mapping alignment (e.g., cell types with a proportion below a threshold may be excluded from the analysis).

[0163] If the queried single-cell omics data has sequencing data for a number of cells greater than the estimated number of cells in the specimen, single-cell omics data for one or more single cells are removed so that the single-cell omics data matches the estimated number of cells in the specimen. If the queried single-cell omics data has sequencing data for a number of cells less than the estimated number of cells in the specimen, single-cell omics data for one or more single cells are added so that the single-cell omics data matches the estimated number of cells in the specimen. Any method for adding single-omics data can be used. In some embodiments, adding single-omics data is achieved by overlapping the single-cell data of the single-cell omics data. In some embodiments, adding single-omics data is achieved by generating single-cell data to be added to the single-cell omics data, which can be generated to represent the single-cell omics data.

[0164] Method 100 assigns single cells from the single-cell omics dataset to spatial coordinates based on a globally optimal solution (111). In some embodiments, a global convex optimization is performed to assign the single cells. In some embodiments, the optimization is linear. In some embodiments, the optimization is nonlinear. To perform the optimization, each region may include a set of subregions corresponding to the estimated number of cells for each region. A matrix of single cells in the single-cell omics profile and a matrix of subregions in the analyte omics profile. Single cells may be assigned to subregions such that the sum of optimal cell / subregion assignments provides the global optimization. In some embodiments, the global optimization is determined by the sum of cell / subregion assignments that minimizes a linear cost function.

[0165] Various solvers can be used to determine the globally optimal solution. In some embodiments, the Jonker-Volgenant algorithm based on shortest augmenting paths is used to determine the globally optimal solution (R. Jonker and A.A. Volgenant, Computing 38, 325-340 (1987), the disclosure of which is incorporated herein by reference). In some embodiments, the cost-scaling push-relabel method is used to determine the globally optimal solution (AV Goldberg and R. Kennedy, Math. Program. 71, 153-177 (1995), the disclosure of which is incorporated herein by reference).

[0166] In some embodiments, instead of predetermining the proportion of each cell type in the spatial omics data (as is done in step 105), the proportion of each cell type is determined as part of the global optimization. Thus, the summation of optimal cell / subregion assignments also evaluates variations in cell type numbers to yield the global optimization.

[0167] In some embodiments, when local cell deconvolution is performed to determine the cell proportions in each region, global optimization is only performed for regions containing a set of one or more cell types to which coordinates will be assigned. Other regions can be ignored.

[0168] Based on the global optimization of assigning single cells to subregions, a spatially aligned map of cells can be generated.When the global optimization is only performed on a region containing a set of one or more cell types, a focused spatially aligned map of the cells of the set of one or more cell types can be generated.Various examples of generated maps are provided in the examples below.

[0169] While specific examples of various methods for generating spatial arrangements of single cells using spatial omics data are described above, those skilled in the art can understand that various steps of the process may be performed in different orders and that certain steps may be optional according to some embodiments of the present invention. As such, it should be apparent that various steps of the process may be used as appropriate for the requirements of a particular application. Furthermore, any of a variety of processes for generating spatial arrangements of single cells using spatial omics data that are appropriate for the requirements of a given application may be utilized according to various embodiments of the present invention.

[0170] The spatial alignment of single cells to generate a map of a specimen can be utilized in numerous downstream applications. In some embodiments, the reconstructed spatial alignment of cells can be visualized to single-cell resolution. In some embodiments, the spatial alignment of single cells can provide detailed information about the ecosystem of the microenvironment. For example, the evaluation of a tumor specimen can provide details about tumor growth, cancer progression, and / or response to treatment. Any of a number of multicellular ecosystems can be evaluated; for example, the tumor microenvironment, sites of immune cell infiltration, multicellular organ systems, host-pathogen interactions, and microbiota can be evaluated.

[0171] The results of the spatial alignment can be used to determine various signatures associated with the spatial context. For example, when evaluating cancer specimens, signatures associated with treatment response, treatment resistance, cancer progression, and cancer recurrence can be determined. These signatures can then be used to formulate a diagnosis.

[0172] Spatial signatures can further be delineated by training computational machine learning models to provide predictions. For example, multiple tissue samples associated with a particular biological feature can each be evaluated for a spatial signature. The particular biological feature can be any feature, such as a pathology, a medical disorder, a health status, a metabolic status, an organ status, activation of multicellular communication, multicellular migration, a multicellular response to a stimulus, or any other feature that can be associated with a particular spatial arrangement of cells. In the field of cancer diagnosis, cancer features are evaluated to evaluate (for example) treatment response, treatment resistance, cancer progression, and cancer recurrence. Machine models can be trained to predict particular biological features based on the rendered spatially resolved maps of single cells. Various machine models can essentially detect spatial signatures from the spatially resolved maps even in situations where a trained clinician would not be able to detect the spatial signature. Models can be trained with multicellular specimens known to have an association with the particular feature. Models can also be trained with multicellular control specimens known not to be associated with the particular biological feature. For example, spatial alignments derived from tumor samples from multiple patients who were resistant to a particular treatment and from multiple patients who were responsive to a particular treatment can be used to train a model to predict the likelihood that a tumor sample will resist that particular treatment. In various embodiments, the training can be supervised, partially supervised, or unsupervised. In some embodiments, the machine learning model is a classifier. In some embodiments, the machine learning model is a regressor. The model can incorporate one or more of any suitable architecture, such as (e.g.,) a deep neural network (DNN), a convolutional neural network (CNN), a graph neural network (GNN), a recurrent neural network, a long-short-term memory (LSTM) network, a kernel ridge regression (KRR), or a gradient-boosted random forest decision tree.In some embodiments, the model incorporates a spatial encoder.

[0173] Diagnostic procedures can be developed using the spatial signatures. These diagnostic procedures can include the following steps: -Draw spatially resolved maps of patient-derived tissue samples Evaluating the spatially resolved map to detect spatial signatures. determining a diagnosis based on said spatial signature;

[0174] In some embodiments, the diagnostic procedure may include determining a treatment based on the spatial signature. In some embodiments, the diagnostic procedure may include administering a treatment determined based on the spatial signature. In some embodiments, the diagnostic procedure may include determining a further diagnostic technique to be performed. In some embodiments, the diagnostic procedure may include performing a further diagnostic technique determined based on the spatial signature.

[0175] In some embodiments, the diagnostic procedure involves performing a spatial omics protocol using a tissue sample from a patient, where the spatial omics protocol is utilized to develop spatial alignment. In some embodiments, the diagnostic procedure involves obtaining a tissue sample from a patient. In some embodiments, the diagnostic procedure involves obtaining a tissue sample from a patient. The tissue sample can include a tumor, a multicellular organ sample, a sample containing tissue infiltrated by immune cells, a sample containing host tissue and pathogens, or a sample containing host tissue and microbiota. In some instances, the patient has a disease or medical disorder, and the tissue sample contains or is affected by the disease or medical disorder. The medical disorder can include, but is not limited to, cancer, pathogenic infection, organ dysfunction, inflammatory disorder, autoimmune disorder, diabetes, liver dysfunction, heart disease, or neurodegenerative disorder. In some embodiments, the diagnostic procedure can predict the characteristics of the medical disorder. Characteristics may include (but are not limited to) a particular pathology, the likelihood of successful or unsuccessful treatment, the severity of a medical disorder, the need for a particular medical intervention, or the likelihood of a future medical complication.

[0176] In the field of cancer diagnosis, various cancer-related characteristics can be diagnosed. In some embodiments, the diagnostic procedure can predict response to treatment. In some embodiments, the diagnostic procedure can predict toxicity of treatment. In some embodiments, the diagnostic procedure can predict resistance to treatment. Treatments can include (but are not limited to) immunotherapy, chemotherapy, radiation therapy, targeted therapy, hormonal therapy, and surgical resection. In some embodiments, the diagnostic procedure can predict cancer progression. In some embodiments, the diagnostic procedure can predict the likelihood of metastasis. In some embodiments, the diagnostic procedure can predict the likelihood of transition from pre-invasive cancer to invasive cancer. In some embodiments, the diagnostic procedure can predict the likelihood of recurrence. Spatial Alignment System

[0177] Turning now to FIG. 2 , a computing system for spatial alignment of cells according to various embodiments of the present disclosure typically utilizes a processing system including one or more of a CPU, a GPU, and / or a neural processing engine. In numerous embodiments, spatial omics input data is processed via the computing system to spatially align cells with single-cell omics data. In some embodiments, the computing system is housed within a computing device that interfaces directly with the system for capturing spatial omics data. In some embodiments, the computing system is housed separately from and receives the acquired spatial omics data. In certain embodiments, the computing system is in communication with the system for capturing spatial omics data. In various embodiments, the processing system communicates with the system for capturing spatial omics data by any suitable means (e.g., wireless connection). In certain embodiments, the computing system is implemented as a software application on a computing device such as, but not limited to, a remote processor, CPU, mobile phone, tablet computer, and / or portable computer.

[0178] A computing system according to various embodiments of the present disclosure is illustrated in FIG. 2. The computing system 201 includes a processor system 203, an I / O interface 205, and a memory system 207. As can be readily appreciated, the processor system 203, the I / O interface 205, and the memory system 207 can be implemented using any of a variety of components appropriate to the requirements of a particular application, including, but not limited to, a CPU, a GPU, an ISP, a DSP, a wireless modem (e.g., a WiFi modem, a Bluetooth modem), a serial interface, volatile memory (e.g., DRAM), and / or non-volatile memory (e.g., SRAM and / or NAND flash). In the illustrated embodiment, the memory system can store numerous applications and / or data. The applications can include, but are not limited to, an application 209 for determining cell number (e.g., number of cells in a spot), an application 211 for determining cell type proportions (e.g., cell types within a spot), an application 213 for matching single-cell omics data (e.g., using single-cell sequencing data references to match cell types to spots), and an application for assigning spatial coordinates of cells (e.g., assigning cells to particular spots). The various applications can be downloaded and / or stored in non-volatile memory. When executed, the various applications can configure the processing system to perform computational processes, each including, but not limited to, the computational methods described above and / or combinations and / or modified versions of the computational methods described above.In some embodiments, various applications utilize input data 217, generate and / or utilize intermediate data 219, and generate output data 221, each of which can be stored in a memory system, either temporarily to perform computational methods or for a longer period so that the data can be retrieved at a later time. Input data can include, but is not limited to, spatial omics data and single-cell sequencing data. Intermediate data can include, but is not limited to, cell counts per spot, cell type ratios, and the likelihood that single-cell sequencing results match the spatial omics data. Output data can include, but is not limited to, assignment of cells to spatial coordinates and visualization of the spatial alignment of cells. It will be understood that input data 217, intermediate data 219, and output data 221 can be utilized in many different ways and therefore should not be limited in any particular way. By way of illustration, any data can be utilized as output to an output interface (e.g., a monitor or other computing system) or as input for any other process.

[0179] While a specific computing system is described above with reference to FIG. 2, it should be readily understood that the computational and / or other processes utilized in providing spatial cell alignment according to various embodiments of the present disclosure may be implemented in any of a variety of processing devices, including combinations of various processing devices. Accordingly, it should be understood that various computing devices according to embodiments of the present disclosure are not limited to specific computing systems and / or spatial cell alignment applications. Various computing devices may be implemented using any of various combinations of the systems described herein and / or modified versions of the systems described herein to perform the processes, combinations of processes, and / or modified versions of the processes described herein. [Example]

[0180] The embodiments of the present disclosure will be better understood through the several examples provided therein. Many exemplary results of various methods for generating spatial alignment of individual cells from scRNA-seq are described. As can be easily seen from one particular embodiment, namely, CytoSPACE, the method as described herein is superior to several other methods currently in practice. High-resolution single-cell and spatial transcriptome alignment with CytoSPACE

[0181] Single-cell spatial organization is a key determinant of cellular state and function. For example, in human tumors, various local signaling networks differentially influence individual cells and their surrounding microenvironments, with various implications for tumor growth, progression, and response to therapy. While spatial transcriptomics (ST) has become a powerful tool for characterizing spatial gene expression in primary tissue specimens, commonly used platforms, such as 10x Visium, remain limited to bulk gene expression measurements, where each spatially resolved expression profile is derived from as many as 10 or more cells (J. Hu et al., Comput Struct Biotechnol J 19, 3829-3841 (2021); the disclosure of which is incorporated herein by reference).

[0182] Therefore, several computational methods have been developed to infer the cellular composition in a given bulk ST sample. Most such methods use reference profiles derived from single-cell RNA sequencing (scRNA-seq) data to deconvolve ST spots into a matrix of cell type proportions. However, these methods lack single-cell resolution, thus hindering the discovery of spatially defined various cellular states, their interaction patterns, and the communities surrounding them (Figure 3A).

[0183] To address this challenge, Cyto Spatial Positioning Analysis via Constrained Expression Alignment (CytoSPACE) was developed as an example to provide single-cell spatial organization. CytoSPACE is an efficient computational approach for mapping individual cells from a reference scRNA-seq atlas to their precise spatial locations in bulk or single-cell ST datasets (Figure 3A and Figure 3B). Unlike other methods (see T. Biancalani et al., Nature Methods, 18, 1352-1362 (2021); and R. Wei, Nature Biotechnology, 40, 1190-1199 (2022)), the solution described herein formulates single-cell / spot assignment as a convex optimization problem and solves it using the Jonker-Volgenant shortest augmentation path algorithm (R. Jonker and A.A. Volgenant, Computing, 38, 325-340 (1987); the disclosure of which is incorporated herein by reference). This approach exhibits improved noise tolerance while ensuring optimal mapping results. The output is a reconstructed tissue specimen with both high gene coverage and spatially resolved scRNA-seq data suitable for downstream analysis, including the discovery of context-dependent cellular states. On both simulated and real ST datasets, CytoSPACE was found to substantially outperform related methods for elucidating single-cell spatial composition.

[0184] CytoSPACE proceeds in four major steps (Figure 3B). First, to explain the discrepancy between the scRNA-seq and ST data in the number of cells per cell type, two parameters are required: (i) the abundance fraction of each cell type within the ST sample, and (ii) the number of cells per spot. Abundance fractions are determined using external deconvolution tools, such as Spatial Seurat, RCTD, SPOTlight, cell2location, or CIBERSORTx (for Spatial Seurat, see T. Stuart et al., Cell 177, 1888-1902 e1821 (2019)); for RCTD, see D. M. Cable et al., Nature Biotechnology 40, 517-526 (2022)); for SPOTlight, see M. Elosua-Bayes et al., Nucleic Acids Res 49, e50 (2021); for cell2location, see V. Kleshchevnikov et al., Nature Biotechnology 40, 661-671 (2022)); for CIBERSORTx, see A. M. Newman, Nat Biotechnol 37, 773-782 (2019); the disclosures of each of which are incorporated herein by reference. By default, the number of cells is directly inferred by CytoSPACE using an approach to estimate RNA abundance, although alternative methods involving cell segmentation approaches can also be used (see, e.g., M. Tippani et al., bioRxiv, 2021.2008.2004.452489 (2022); and C. Stringer et al., Nature Methods 18, 100-106 (2021); the disclosures of which are each incorporated herein by reference). Once both parameters are estimated, the scRNA-seq dataset is randomly sampled to match the predicted number of cells per cell type in the ST dataset.Upsampling is performed for cell types with poor representation, either by extraction with restoration or by introducing placeholder cells. Finally, CytoSPACE assigns each cell to a spatial coordinate in a manner that minimizes a correlation-based cost function, constrained by the estimated number of cells per spot, via a shortest augmented path optimization algorithm. An efficient integer programming approximation method that produces comparable results is also provided (AV Goldberg and R. Kennedy, Math. Program. 71, 153-177 (1995); the disclosure of which is incorporated herein by reference).

[0185] To validate the performance of CytoSPACE, the ST dataset was simulated with a fully defined single-cell composition. To this end, we leveraged previously published mouse cerebellum (n = 11 major cell types) and hippocampus (n = 17 major cell types) data. These data were generated using Slide-seq, a platform with high spatial resolution (approximately single cells) but limited gene coverage (Figure 4A). (For further information on Slide-seq, see S.G. Rodriques et al., Science 363, 1463-1467 (2019); the disclosure of which is incorporated herein by reference.) To increase transcriptome representation while maintaining spatial dependency, each Slide-seq bead was replaced with the most correlated single-cell expression profile of the same cell type from an scRNA-seq atlas of the same brain region (Figure 4B) (for further information regarding the atlas, see A. Saunders, Cell 174, 1015-1030 e1016 (2018) ; the disclosure of which is incorporated herein by reference). A spatial grid with adjustable dimensions was then overlaid to pool the single-cell transcriptomes into a pseudo-bulk transcriptome. This was done across a range of realistic spot resolutions (averages of 5, 15, and 30 cells per spot). To ensure unique spatial addresses for all cells in the scRNA-seq query dataset, a paired scRNA-seq atlas was created from the cells underlying each pseudo-bulk ST array. Finally, varying amounts of noise were added to the scRNA-seq data to emulate technical and platform-specific variations between the scRNA-seq and ST datasets (Figure 4C-F). Together, these datasets allow for rigorous assessment of pairwise cell alignments, including orthogonal approaches to study alignment quality (Figure 5A-E).

[0186] Next, various methods for CytoSPACE parameter estimation were evaluated. For cell type enumeration, Spatial Seurat was used, which showed strong agreement with known overall proportions in the simulated ST dataset (Figure 6A). To approximate the number of cells per spot, a simple approach was implemented based on RNA abundance estimation. This approach was correlated with ground truth predictions in the simulated ST data and cell segmentation analysis of matched histological images from real ST data (Figures 6B-E).

[0187] CytoSPACE was benchmarked against 12 previous methods, including two recently described algorithms for scRNA-seq and ST alignment: Tangram, which integrates scRNA-seq and ST data via maximization of a spatial correlation function using nonconvex optimization, and CellTrek, which uses Spatial Seurat to identify shared embeddings between scRNA-seq and ST data and then applies random forest modeling to predict spatial coordinates. A few naive approaches, including Pearson correlation and Euclidean distance, were also evaluated. To compare outputs, each cell was assigned to the spot with the highest score (all approaches except CellTrek) or the spot with the closest Euclidean distance to the cell's predicted spatial location (CellTrek only).

[0188] Across multiple assessed noise levels and cell types, CytoSPACE achieved substantially higher accuracy than other methods for mapping single cells to their known locations in simulated ST datasets (Figures 7A-7E and Table 1). This was true for multiple spatial resolutions, independent of brain region, both for individual cell types and across all assessable cells (Figures 7B and 7C). We also obtained similar results with an independent method (RCTD) for determining cell type abundance in ST data (Figures 8A-8C).

[0189] The robustness of CytoSPACE to variations in key input parameters was assessed (steps 1–3 in Figure 3B). First, estimated cell type abundances were examined: in this case, estimated cell type abundances ranged from an average of 0.025% to an average of 32% in the simulated ST datasets (Figures 9A and 9B). Despite this range, no significant correlation with mapping accuracy was observed (Figures 9A and 9B). Next, experiments were performed in which estimates of (i) cell type abundance and (ii) the number of cells per spot were systematically perturbed. In all cases, CytoSPACE continued to outperform previous methods (Figures 10A-10E). Finally, output stability was verified when sampling scRNA-seq query datasets with different seeds (step 3 in Figure 3B) and when different distance metrics were used to calculate the CytoSPACE cost function. Results remained consistent across multiple runs and distance metrics (Figures 11A and 11B). Collectively, these data highlight the robustness of CytoSPACE and its potential for improved spatial mapping of scRNA-seq data.

[0190] To evaluate performance on real ST datasets, primary tumor samples were examined. Primary tumor samples were from three types of solid malignancies: melanoma, breast cancer, and colon cancer. One HER2 + A total of six scRNA-seq / ST combinations were analyzed, encompassing six bulk ST samples (n = 4, Visium; n = 2, legacy ST), including formalin-fixed, paraffin-embedded (FFPE) breast tumor specimens and three scRNA-seq datasets from matched tumor subtypes (Table 2). All cell types in each scRNA-seq dataset were aligned using CytoSPACE and compared against Tangram and CellTrek (Figures 12A-C). CytoSPACE was highly efficient, processing a Visium-scale dataset in approximately 5 minutes on a single CPU core (Table 3). This was true regardless of whether a shortest augmentation path approach or an integer programming approximation approach was applied; both achieved comparable results (Table 4). To quantitatively compare cell state recovery with respect to spatial localization patterns in the tumor microenvironment (TME), assigned cells were dichotomized within each cell type according to their proximity to tumor cells. Gene sets that mark TME cell states by known localization were then assessed for whether they were skewed in the expected orientation ( Fig. 13A ).

[0191] T cell exhaustion, a canonical state of dysfunction resulting from prolonged antigen exposure in tumor-infiltrating T cells, was first investigated. Consistent with expectations, CytoSPACE recovered spatial enrichment of T cell exhaustion genes in CD4 and CD8 T cells, which map closest to cancer cells in all six combinations of scRNA-seq and ST datasets (Figures 13B, 13C, and 14A). In contrast, Tangram and CellTrek produced single-cell mapping with substantially lower enrichment of T cell exhaustion genes in the expected direction, with 25%–33% of cases showing enrichment in the opposite direction, i.e., away from the tumor center (Figures 13C and 14A).

[0192] To demonstrate applicability to other spatially biased cell states, the analysis was extended to various TME lineages, thereby identifying cell type-specific genes whose expression varied as a function of distance from the tumor cells. To validate the results, the inventors analyzed two recently defined cell ecosystem subtypes in human carcinomas: CE9 and CE10 (see BALuca et al., Cell 184, 5482-5496.e5428 (2021) for further information on CE9 and CE10; the disclosure of which is incorporated herein by reference). These "ecotypes" are also observed in melanoma, and each encompasses B cells, plasma cells, CD8 T cells, CD4 T cells, and monocytes / macrophages with stereotypic spatial localization. The CE9 cell state preferentially localizes to the tumor center, whereas the CE10 state preferentially localizes to the tumor periphery. Using marker genes specific to each state, we asked whether the single cells mapped by each method were consistent with the CE9- and CE10-specific patterns of spatial localization. Indeed, as observed for T cell exhaustion factors, CytoSPACE successfully recaptured the expected spatial bias in the CE9 and CE10 cell states across lymphoid and myeloid lineages (Figure 14B), thereby outperforming 12 previous methods in both magnitude and direction of marker gene enrichment (Figures 14A, 14C, and 14D). Furthermore, consistent with simulation experiments, CytoSPACE results remained robust to perturbations of its input parameters (Figures 15A-F). As a further confirmation of validity, TREM2 + Macrophages and FOLR2 +The predicted spatial localization patterns of macrophages were evaluated, as these macrophages have recently been shown to localize to the tumor stroma and tumor mass, respectively, across diverse cancer types (Figure 16A). When compared against Tangram and CellTrek, only CytoSPACE reproduced these previous findings with statistical significance (Figure 16B). Moreover, when inferred spatial locations (near tumor versus far from tumor) were projected against the UMAP embedding of scRNA-seq data, single cells generally could not be clustered based on their distance from tumor cells (Figures 17A and 17B). These data highlight the ability of CytoSPACE to correctly identify spatially resolved cell states, including those not discernible from scRNA-seq or ST data alone.

[0193] To further demonstrate how CytoSPACE can reveal spatial biology, two additional scenarios were explored. First, it was asked whether CytoSPACE could discover densely packed cellular substructures in bulk ST data. For this purpose, normal mouse kidneys were chosen because they have a highly granular spatial architecture. After mapping a well-annotated scRNA-seq atlas with over 30 spatially resolved subtypes of renal epithelium onto a 10x Visium profile (55 μm diameter per spot) of normal mouse kidney (Figure 18A and Table 5), it was assessed whether CytoSPACE could recapitulate known patterns of spatial organization. Indeed, CytoSPACE (i) reconstructed known zonal regions (Figures 18B and 18C), (ii) identified cell types (approximately 70 μm diameter) that preferentially colocalize with glomeruli (Figure 18D), and (iii) sequenced nearly 30 epithelial states in spots consistent with their known locations in the nephron epithelium and collecting duct system, thus outperforming previous methods (Figures 19A and 19B and Figures 20A-F).

[0194] Finally, we asked whether CytoSPACE could enhance single-cell ST datasets with low gene throughput. To do so, a breast cancer specimen was analyzed. This specimen contained over 550k annotatable cells and 500 preselected genes profiled by MERSCOPE (Vizgen). We first confirmed that CytoSPACE could correctly map single cells profiled by MERSCOPE and reproduce their spatial dependencies (Figures 22A-22E). Next, the scRNA-seq breast cancer atlas was mapped to the same MERSCOPE dataset. In addition to observing strong inter-platform agreement for most annotated cell types (Figure 21A, and Figures 22F and 22G), we observed a significant bias in cancer-associated T cell signatures enriched in tumor tissue or adjacent normal tissue (Figure 21B, and Figures 22H and 22I, and Table 6). Such enrichment was significantly greater than that calculated from MERSCOPE data alone and correlated with expected enrichment (Figures 21B and 22I, and Table 6). Collectively, these data highlight the versatility of CytoSPACE for complex tissue reconstruction at the single-cell level.

[0195] CytoSPACE is a tool for aligning single-cell and spatial transcriptomes via global optimization. Unlike related methods, CytoSPACE ensures globally optimal single-cell / spot alignments conditional on a correlation-based cost function and the number of cells per spot. Furthermore, CytoSPACE can be easily extended to accept additional constraints, such as the compositional proportion of each cell type per spot (e.g., as inferred by RCTD or cell2location). In contrast, CellTrek relies on co-embedding learned by Spatial Seurat, which can erase subtle yet important biological signals (e.g., differences in cell state). While Tangram is robust in idealized situations, it cannot guarantee a globally optimal solution. While CytoSPACE requires two input parameters, both parameters can be reasonably well estimated using standard approaches, suggesting that they are unlikely to pose significant barriers in practice. Furthermore, CytoSPACE was substantially more accurate than related methods on both simulated and real datasets. As such, CytoSPACE is useful for deciphering the spatial variation and community structure of single cells in a variety of physiological and pathological situations. Example of the method CytoSPACE analytical framework

[0196] CytoSPACE utilizes linear optimization to efficiently reconstruct ST data using single-cell transcriptomes from a reference scRNA-seq atlas. To formulate the assignment problem of mapping individual cells in scRNA-seq data to spatial coordinates in ST data, let N × C matrix A represent the single-cell gene expression profile with N genes and C cells; let M × S matrix B represent the gene expression profile of the spatial transcriptomics (ST) data with M genes and S spots; and let G be a vector of length g containing a desired subset of genes shared by both datasets. For both gene expression profile matrices, values ​​are first normalized to counts per million (or transcripts per million for platforms that cover the entire gene body) and then transcribed into log2 space. Thus, in its default implementation, CytoSPACE uses all genes as input and does not involve a dimensionality reduction step. Next, (by default) we calculate the sth s- ... th Number of cells giving the RNA content in the spot

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0197] Two possible solvers were provided in CytoSPACE, both of which would return a globally optimal solution to the above problem as formulated. The first of these implementations is the Jonker-Volgenant algorithm based on shortest augmenting paths, where the dual problem of the above formulation is defined as:

number

number

number

[0198] An alternative solver was based on a cost-scaling push-relabel method using the Google OR-Tools software package in Python 3 (AV Goldberg and R. Kennedy, Math. Program. 71, 153-177 (1995); the disclosure of which is incorporated herein by reference). This solver converts exact costs to integers with some loss of numerical precision and has time complexity O(L 2log(LC)), where C represents the maximum magnitude of edge costs. In practice, this solver is nearly as fast as a solver based on Jonker-Volgenant. However, when a very large number of cells are to be mapped, this solver may provide faster run times. Furthermore, this solver has more widespread support across a variety of operating systems and may therefore be useful for users working on systems that do not support AVX2 intrinsics, as required by the lapjv solver. For users who want to obtain lapjv-accurate results on operating systems that do not support the lapjv package, an equivalent, but significantly slower, broadly compatible solver that implements the Jonker-Volgenant algorithm is provided via the lap package (version 0.4.0). Estimating cell type proportions

[0199] To overcome variability in cell type abundance ratios between a given ST sample and a reference scRNA-seq dataset, the first step of CytoSPACE requires estimating cell type ratios in the ST sample (Figure 3B). Notably, only global estimates for the entire ST array are required; these may be obtained by combining spot-level ratios by cell type. While an interesting future extension of CytoSPACE would be to estimate cell type ratios as part of an optimization routine, many deconvolution methods have been proposed to determine cell type composition from ST spots, and any such method could be deployed for this purpose. In this example, Spatial Seurat from Seurat version 3.2.3 was used for the primary analysis. Spatial Seurat demonstrates high correlation between the estimated and true ratios of different cell types in the simulated data (Figure 6A). After loading the raw count matrix, SCTransform() and RunPCA() were performed with default parameters, followed by FindTransferAnchors(), with the preprocessed scRNA-seq data and ST data serving as reference and query, respectively. Spot-level predictions were obtained by TransferData(), and global predictions were obtained by summing the prediction scores per cell type across all spots and scaling the sum of cell type scores to 1.

[0200] In addition to Spatial Seurat, RCTD performance was validated to estimate overall cell type proportions as input to CytoSPACE (Figures 8A-8C). RCTD version 2.0.0 (package spacexr in R) was used with doublet_mode='full' and otherwise default parameters to obtain cell type proportion estimates per spot, then sum the spot-normalized result weights per cell type across all spots and rescale the sum to 1. Estimating the number of cells per spot

[0201] The number of detectably expressed genes per cell ("gene count") corresponds closely to the total captured mRNA content, as estimated by the summation of unique molecular identifiers (UMIs) per cell 45Because gene counts are routinely used as a surrogate for duplicates or multiplexes in scRNA-seq experiments, it was hypothesized that the sum of UMIs per ST spot might reasonably approximate the number of cells per spot, as required for the second step of CytoSPACE (Figure 3B). To test this hypothesis, UMIs were normalized to counts per million per spot, while blunting the effects of outliers, technical variation, and cell volume, and then a log2 adjustment was performed. The number of cells per ST spot was then estimated by fitting a linear function through two points: for the first point, it was assumed that the minimum number of cells per spot was 1, and that this minimum value in cell number corresponded to the minimum sum of UMIs in log2 space. For the second point, it was assumed that the average number of cells per spot corresponded to the average sum of UMIs in log2 space, and this value was set according to user input. For 10x Visium samples, where spots generally contain 1–10+ cells per spot, an average of 5 cells per spot was used throughout this study. For legacy ST samples with larger spot dimensions, an average of 20 cells per spot was selected. The number of cells for each spot was calculated from this fitted function. Supporting this hypothesis, we found that for simulated ST datasets, the Pearson correlation between the estimated and actual number of cells ranged between 0.80 and 0.93, depending on the dataset and spot resolution evaluated, indicating that log2 modulation outperforms the summation of UMIs on the original linear scale (i.e., without CPM) (Figure 6B-6D). The same was true when comparing the number of cells per spot analyzed by cell segmentation (VistoSeg) applied to previously analyzed imaging data from a mouse brain Visium sample (Figure 6E), further validating this approach.While this estimate component is provided by default, the user may also provide their own estimates for this step, including estimates generated by cell segmentation methods (e.g., VistoSeg, CellPose). Harmonize the number of cells per cell type

[0202] The third step of CytoSPACE is to align the number of cells per cell type between the query scRNA-seq dataset and the target ST dataset (Figure 3B). This is accomplished by sampling the former to match the predicted abundance in the latter using one of the following methods: ·Duplicate. num sc,k and num ST,k Let denote the actual and estimated numbers of cells per cell type k in the scRNA-seq and ST data, respectively. For cell type k,

number

number

number

[0203] To evaluate the accuracy and robustness of CytoSPACE (Figure 4A), ST datasets with known single-cell composition were simulated using previously annotated Slide-seq datasets of mouse cerebellar and hippocampal slices. Let S be the M × B gene expression matrix of the Slide-seq pack with M genes and B beads. To create a higher gene coverage version of S (denoted as Sc), a previously annotated scRNA-seq dataset of the same brain region was used to replace S beads with single-cell transcriptomes. After quality control, in which outlier cells with more than 1,500 genes were removed, each bead in the Slide-seq dataset was matched to the nearest cell of the same cell type in the scRNA-seq dataset by Pearson correlation. This was performed separately for each mouse brain region. Because a single cell may be matched to more than one bead, genes were permuted among cells of the same cell type to obtain unique single-cell transcriptomes. For each cell, 20% of its transcriptome of randomly selected genes was replaced per cell by the transcriptome of another randomly selected cell of the same cell type, ensuring that the latter was not a duplicate of the former. For simplicity, the number of beads present in these two tissues was matched by randomly sampling beads from the hippocampus data until the number present in the cerebellum data was reached.

[0204] Having created the Sc matrix for each brain region, it was next desired to generate an ST dataset with a defined spot resolution. To this end, an m × n spatial grid was imposed over the entire pack. Each grid spot

number

number

[0205] Finally, noise was added to the scRNA-seq data in a defined amount to (i) utilize the underlying scRNA-seq data for each Sc matrix as a query dataset, and (ii) emulate technical variations between platforms. To this end, a percentage of genes p to be perturbed was selected, and then a corresponding subset of genes from each cell was randomly selected, for which the noise was distributed according to an exponential Gaussian distribution. N(0,1) Noise perturbations were investigated for the following values ​​of p: 5%, 10%, and 25%. Despite the addition of noise, the UMAP plot of the perturbed transcriptome remained similar to the original data, demonstrating the maintenance of a biologically realistic data structure (Figures 4C-F). Quality control studies for paired spot cell alignment

[0206] There are two important scenarios in which discrepancies between scRNA-seq and ST data can occur. In the first scenario, a cell type is detectable in the scRNA-seq dataset but not in the spatial dataset. CytoSPACE addresses this issue by requiring cell type abundance estimates as input (e.g., using Seurat, RCTD, or cell2location). In doing so, cell types missing from the ST dataset will generally be omitted from the spatial mapping (if imputed with an abundance fraction of zero) or inferred with an abundance fraction, thereby minimizing their impact on performance.

[0207] In the second scenario, cell types are detectable in the spatial dataset but not in the scRNA-seq dataset, causing incorrect mapping. This scenario is rare because droplet sequencing can easily dissect all major cell types in a given tissue sample, excluding cell types for which dissociation-induced loss is either rare or prone to dissociation-induced loss. Other methods for spatial spot resolution, including Seurat, RCTD, and cell2location, have the same limitation, but this is usually negligible in practice.

[0208] While the Jonker-Volgenant algorithm is guaranteed to optimally solve the assignment problem given its cost function, it lacks an underlying probabilistic framework for estimating mapping uncertainty. An alternative is to determine whether a given cell type belongs to a given spatial spot after mapping, i.e., whether the spot contains at least one cell of the same cell type. Notably, this definition is significantly less demanding than the metrics described in "Performance Assessment." Nevertheless, to explore this possibility, the following procedure was implemented: First, to identify the top marker genes for each cell type mapped by CytoSPACE, NormalizeData(), ScaleData(), and FindAllMarkers() from Seurat v4.0.1 were sequentially applied to the scRNA-seq query dataset using default parameters. The ST dataset was then normalized and rescaled using the same workflow. The -log 2 fold change was calculated with a log2 fold change greater than 0. 10For each cell type i with at least five and up to 50 marker genes (represented by m) identified by the adjusted p-value, 50 spatial spots to which CytoSPACE assigned at least one cell of cell type i were randomly selected, and 50 spatial spots that did not contain at least one cell of cell type i were also randomly selected. If fewer than 50 spots met the given condition, the 50 spots were resampled. Next, pairwise cell assignments were used to reconstruct a pseudo-bulk transcriptome from the normalized and scaled scRNA-seq dataset by averaging each selected spot across its assigned cells. A support vector machine (e1071 v1.7.8 in R) was then trained to distinguish between the two groups of pseudo-bulk transcriptomes from the previous step using the top m marker genes of cell type i. This model was used to calculate the probability, referred to as a confidence score, that cell type i belonged to each spot in the normalized and scaled ST dataset. Finally, for each mapped cell of type i, its spot-specific confidence score was retrieved.

[0209] This approach was evaluated on simulated ST data for which the ground truth was known (Figure 5A). Despite the already low (<5%) proportion of incorrectly mapped cells (defined above) prior to applying this filter, the approach successfully distinguished correctly mapped cells from incorrectly mapped cells with high statistical significance, with nearly all AUCs exceeding 0.8 for classifying individual cell types (Figures 5B and 5C). Moreover, at a confidence threshold above 10%, virtually every correctly mapped cell was retained, whereas over 75% of incorrectly mapped cells were removed (Figures 5D and 5E). Therefore, although available via the CytoSPACE GitHub repository, this procedure may be used as an optional post-processing step to explore alignment quality. Benchmark analysis using simulated datasets

[0210] To fully evaluate CytoSPACE's performance, an expanded benchmark analysis was conducted that included Tangram, CellTrek, and 10 additional methods that may be adapted (Figure 7C). Methods were included if they (i) were applicable to a single-cell query dataset and a spatial reference dataset, including bulk ST data, (ii) produced an output or involved intermediate steps (e.g., scRNA-seq integration techniques, some gene imputation methods, naive distance metrics) that aligned these two datasets, thereby enabling the imputation of single-cell spatial coordinates in the query dataset, and (iii) were cross-reviewed by publicly available software implementations.

[0211] Most previous methods failed to meet these requirements, including those designed for spot-level resolution (e.g., cell2location, RCTD), those designed for spatial clustering (e.g., BayesSpace), and those designed for spatial coordinate prediction without spatial reference (e.g., novoSpaRc). Therefore, the benchmark analysis consisted of three dedicated pairwise cell mapping methods (CytoSPACE, Tangram, and CellTrek), three single-cell integration methods (Harmony, LIGER, and Seurat V3), four methods from which pairwise cell assignments can be derived (DistMap, SpaGE, DEEPsc, and SpaOTsc), and three naive methods (Pearson correlation, Spearman correlation, and Euclidean distance). The application of each approach is described below.

[0212] CytoSPACE. For each ST resolution and scRNA-seq noise level, the abundance fraction of known cell types in the ST samples was estimated via Spatial Seurat as described in "Estimating cell type fractions." CytoSPACE was run with the "generated cells" option and with the lapjv solver implemented in Python (package lapjv, version 1.3.14).

[0213] Tangram. Similar to CytoSPACE, and in contrast to the other methods discussed here, Tangram attempts to optimally position input cells across spots, ensuring that the pairwise cell mapping for each input cell is not strongly separable from the pairwise cell mappings of other cells. Therefore, to ensure a fair comparison with CytoSPACE, Tangram (version 1.0.2) was run using the same input cells mapped by CytoSPACE, including newly generated cells after resampling to match the predicted cell type counts. This also included providing a normalized vector of CytoSPACE's cell number per spot estimate as the density prior (density_prior argument). Tangram was trained on CPM-normalized scRNA-seq data in two ways: (i) using all available genes per cell, and (ii) using top marker genes stratified by cell type. To identify marker genes using Seurat (version 4.1.0), NormalizeData() was applied with default parameters, and FindAllMarkers() was applied with only.pos=TRUE, min.pct=0.1, and logfc.threshold=0.25. The top 100 genes by mean log2 fold change were then selected for each cell type.

[0214] CellTrek. CellTrek (version 0.0.0.9000) was given all cells present in each simulated ST dataset (without newly generated cells mapped by CytoSPACE and Tangram), assuming that CellTrek (by default) largely overlaps input cells and also filters input cells based on whether mutual nearest neighbors are identified between cells and spots. After single cells were assigned spatial coordinates, the nearest ST spot for each cell was selected via Euclidean distance. Because the CellTrek wrapper does not handle ST input without associated h5 and image files, the code was modified to accept ST datasets from other sample sources. CellTrek was run with default parameters, except for (i) limiting the repel functionality (repel_r = 0.0001) (because this parameter forces imputed spatial coordinates to deviate arbitrarily from their original predictions) and (ii) setting spot_n to twice the average number of cells per spot for each spatial resolution tested.

[0215] DistMap. DistMap attempts to computationally reconstruct ST data from paired scRNA-seq at single-cell resolution. DistMap uses marker genes and a binarization approach that calculates Matthews correlation coefficients to obtain a distributed location assignment for each cell. 50 .

[0216] For benchmarking, DistMap (v0.1.1) was given all input cells and spots, which restricted genes to marker genes expressed in at least five cells and five spots (selected as described for benchmarking Tangrams by top genes). Count matrices were CPM normalized and log2-adjusted. After creating a DistMap object with the normalized ST data given in the insitu argument, the scRNA-seq data were binarized via binarizeSingleCellData(dm, seq(0.15, 0.5, 0.01)). A binarized version of the ST data matrix was prepared by setting all non-zero counts to 1, and then the insitu.matrix member variable of the DistMap object was replaced with this binarized version. Cell-to-spot mapping was performed with mapCells(), and each cell was assigned to the spot with the highest score, as returned by the mcc.scores member variable.

[0217] SpaOTsc. SpaOTsc is a method for inferring spatial properties of scRNA-seq data, primarily designed for investigating spatial cell-cell communication. As the first step in this process, SpaOTsc computes a map between single cells and spatial datasets using an optimal transmission approach for marker genes.

[0218] For benchmarking, SpaOTsc (v0.2) was given all input cells and spots, which restricted genes to marker genes expressed in at least five cells and five spots (selected as described for benchmarking Tangrams by top genes). Following the tutorial instructions, SpaOTsc was implemented as follows: First, counts were normalized to sum to 10,000 per cell or spot, respectively, and then the resulting scRNA-seq matrix (df_sc) and ST matrix (df_is) were log2-transformed. From the normalized scRNA-seq data, principal component analysis (PCA) was performed with prcomp in R, and then the Pearson correlation coefficient matrix (sc_pcc) was calculated between single cells from the top 40 principal components. To obtain the Matthews correlation coefficient matrix (mcc) between cells and spots, each normalized data matrix was binarized with a quantile threshold of 0.7 (resulting in df_sc_bin and df_is_bin for the scRNA-seq and ST matrices, respectively), and then the Pearson correlation coefficient was calculated across all cell-spot pairs. SpaOTsc was then run with the following set of commands: C = np.exp(1-mcc), issc = SpaOTsc.spatial_sc(sc_data = df_sc, sc_data_bin = df_sc_bin, is_data = df_is, is_data_bin = df_is_bin, sc_dmat = np.exp(1-sc_pcc), is_dmat = is_dmat), out = issc.transport_plan(C ** 2, alpha=0.1, rho=100.0, epsilon=1.0, cor_matrix=mcc, scaling=False). Each cell was then assigned to the spot with the highest score as returned in the output of issc.transport_plan().

[0219] DEEPsc. DEEPsc is a deep learning-based method for imputing spatial information to scRNA-seq data given a spatial reference atlas. DEEPsc first transfers the spatial reference atlas data to a space of reduced dimensionality via PCA, and then performs network training on it. The scRNA-seq data is projected into the same PCA space and fed into the DEEPsc network, which then outputs a matrix of the likelihood that each cell originated from each spot in the ST tissue.

[0220] To benchmark, DEEPsc (version number not available; last GitHub commit at time of clone: ​​June 5, 2022) was given all input cells and spots, the respective input matrices CPM normalized and then log-transformed via log1p, and genes restricted to genes present in both matrices. DEEPsc was run with 50,000 iterations in parallel mode for training and with default parameters otherwise.

[0221] SpaGE. SpaGE, or Spatial Gene Enhancement using scRNA-seq, is a method for increasing gene coverage in ST measurements by integrating spatial data with higher-coverage scRNA-seq datasets. SpaGE uses the domain-fitting algorithm PRECISE to project datasets into a shared space, where gene expression predictions are then calculated via a k-nearest neighbor approach. Although SpaGE was designed for gene expression prediction rather than mapping cells to spots, it includes an integration step, allowing this integrated space to be used for spot-to-spot cell mapping.

[0222] To do so while making maximum use of the SpaGE framework (version number not available; last GitHub commit at time of clone: ​​July 20, 2021), a command was added to the source code to return a single nearest spot neighbor for each cell in the SpaGE integrated space. The modified SpaGE code was then given all input cells and spots. Per tutorial advice, genes not expressed in at least 10 cells were filtered out, and the scRNA-seq matrix was then CPM-normalized and log2-transformed, while the ST matrix was normalized to median counts per spot and subsequently log2-transformed. SpaGE was run, again per tutorial advice, with n_pv=30 and otherwise default parameters.

[0223] Spatial Seurat. Seurat is a widely known method for integrating single-cell expression datasets that works by identifying "anchors" between datasets and can be used with spatial data as well. Spatial Seurat integration was validated for assigning cells to spots using Seurat v3. After loading the scRNA-seq and ST count matrices into Seurat objects, the scRNA-seq and ST count matrices were preprocessed with SCTranform() and then evaluated using the standard integration protocol of FindTransferAnchors(normalization.method="SCT") followed by TransferData(). Pair-to-spot cell assignments were determined by the predicted IDs returned from the resulting prediction assays.

[0224] Harmony. Harmony is a method for integrating multiple scRNA-seq datasets into a joint embedding space that uses various clustering methods across a principal component representation of the data to obtain linear correction factors for integration. As a dataset integration method, Harmony does not provide direct spot-to-spot cell mapping results. Therefore, to benchmark the method, we first integrated the complete single-cell and corresponding spatial datasets, and then assigned each cell to its nearest spot in the integrated space by selecting the spot with the smallest Euclidean distance to the cell.

[0225] To obtain an integrated spatial representation, the standard Harmony protocol was followed. First, the Seurat objects created from the scRNA-seq and ST count matrices were merged, and then the standard Seurat processing pipeline of NormalizeData(), FindVariableFeatures(), ScaleData(), and RunPCA() was applied, each with default parameters. The resulting Seurat objects were used to run Harmony v0.1 with group.by.vars="orig.ident" and default parameters otherwise.

[0226] LIGER. Like Harmony, LIGER is another method designed for single-cell expression dataset integration; however, LIGER instead relies on an integrated non-negative matrix factorization approach to embed features into a low-dimensional space, thereby incorporating both dataset-specific and shared factors. As described above for Harmony, LIGER was used to obtain a shared embedding space between the scRNA-seq and ST datasets, and then cells were assigned to spots according to the minimum Euclidean distance.

[0227] To run LIGER (v1.0.0), a LIGER object was created and processed with the package functions normalize(), selectGenes(var.thresh=0.2), and scaleNotCenter() for normalization, gene selection, and scaling, respectively, and then applied to align the dataset using online_iNMF() and quantile_norm(). All parameters not specified here were set to default values. An embedding was extracted from the LIGER object member - variable H.norm.

[0228] In addition to the above methods, Euclidean distance (calculated using the spatial.distance.cdist function in scipy v1.8.0), Pearson correlation, and Spearman correlation were evaluated. Each cell was assigned to either the spot that minimized the distance (Euclidean distance) or the spot that maximized the correlation (Pearson correlation and Spearman correlation). All ground truth cells were evaluated without resampling, and the input dataset was CPM normalized and log2 adjusted prior to analysis.

[0229] Performance evaluation. To determine the accuracy of single-cell mapping, assigned locations that matched exactly with ground truth spots were classified as correct. sc where represents the number of correct assignments, and single-cell precision (Pr sc ) was defined as:

number

[0230] To be widely useful, any computational method, such as CytoSPACE, must exhibit robustness to reasonable variations or errors in the inputs. With this in mind, the consistency and robustness to variations of CytoSPACE was verified across a range of input parameters.

[0231] Robustness to cell fraction estimation errors. To mimic realistic technical errors in estimating cell type fractions, where proportionally larger errors can be expected for rarer cell types, multiplicative noise was introduced within a 4-fold range, with noise depending inversely on the original fraction estimate. First, for each cell type i in the sample, y i has a mean of 0 and the original proportion estimates for cell types x i randomly sampled from a Gaussian distribution with standard deviation and which depend inversely on:

number

number

number

number

[0232] CytoSPACE was validated with this noise model in simulations with five replicate experiments for each simulated test case ("Simulation Framework"), and the results were evaluated via single-cell assignment accuracy as described in "Performance Evaluation" (Figures 10A and 10B).

[0233] Robustness to errors in estimating cell number per spot. Noise was introduced into the estimate of the number of cells per spot by a protocol similar to that described above for perturbing cell type proportion estimates. First, for each spot in the sample, y i has mean 0 and the original estimate n for cell type i i randomly sampled from a Gaussian distribution with standard deviation and which depend inversely on:

number

number

[0234] To restrict the range of values ​​to a feasible region, a minimum number of cells per spot of 1 and a maximum number of cells per spot of 110% of the original maximum M were imposed. Thus, the perturbed values ​​were

number

[0235] Robustness to sampling variations. While most steps in the algorithm are deterministic, CytoSPACE requires that the input scRNA-seq dataset be resampled to create a pool of cells that match the expected cells in the ST dataset; this sampling is done randomly. To verify the consistency of results across different samples, CytoSPACE was run 10 times with different seeds for each simulation case described in "Simulation Framework." Single-cell accuracy of assignment was calculated as described above ("Performance Evaluation"). Results for this analysis are shown in Figure 11A.

[0236] Robustness to distance metrics. In addition to Pearson correlation, the default distance metric for CytoSPACE was implemented, and CytoSPACE performance was validated with the alternative distance metrics Spearman correlation and Euclidean distance as shown in Figure 11B. For each ST resolution and scRNA-seq noise level in the simulated data (simulated data as described in "Simulation Framework"), CytoSPACE was run with Spearman correlation and Euclidean distance used instead of the distance metric. ST dataset for TME community analysis

[0237] Melanoma ST data generated by Thrane et al. 35were downloaded from spatialresearch.org / resources-published-datasets / doi-10-1158-0008-5472-can-18-0747 / (K. Thrane et al., Cancer Research 78, 5970-5979 (2018); the disclosure of which is incorporated herein by reference). Preprocessed spatial transcriptomics datasets of breast cancer specimens (Visium fresh-frozen and FFPE) and colorectal cancer (fresh-frozen) were downloaded from 10x Genomics (www.10xgenomics.com / spatial-transcriptomics / ). Annotations of tumor cell-containing regions were downloaded from 10x Genomics for the Visium FFPE breast cancer samples and shared by 10x Genomics upon request for the Visium fresh-frozen breast cancer samples analyzed in this study. Preprocessed Visium arrays of fresh / frozen TNBC specimens (1160920F) along with tumor borders (SZWu et al., Nature Genetics 53, 1334-1347 (2021); the disclosure of which is incorporated herein by reference) were obtained from Wu et al. 37 . scRNA-seq tumor atlas

[0238] All analyzed tumor scRNA-seq data (downloaded as preprocessed count matrices (UMI-based) or transcript matrices (non-UMI-based)) were selected and curated to clinically match the ST specimens analyzed in this study (see "Molecular Classification of Breast Cancer Specimens"). In addition, author-provided annotations were used for all scRNA-seq reference datasets, with the following modifications: For the melanoma dataset generated by Tirosh et al., we excluded normal melanocytes and separated T cells into CD4 and CD8 subsets according to the expression of CD8A / CD8B and CD4 / IL7R (I. Tirosh et al., Science 352, 189-196 (2016); the disclosure of which is incorporated herein by reference). For the breast cancer dataset from Wu et al., and in the colorectal cancer dataset from Lee et al. (HOLee et al., Nature Genetics 52, 594-603 (2020); the disclosure of which is incorporated herein by reference), author annotations were mapped to cell types according to the scheme in Table 2. Notably, T cells that could not be confidently classified as CD8 or CD4 T cells and myeloid cells that could not be confidently classified as monocytes / macrophages or dendritic cells were excluded. Molecular Classification of Breast Cancer Specimens

[0239] When available, author annotations were used to determine the enrichment status of estrogen receptor (ER) and human epidermal growth factor receptor 2 (HER2) for each scRNA-seq and ST tissue breast cancer sample. FFPE breast cancer specimens from 10xGenomics without receptor status annotation were examined for expression of the ESR1 (ER) and ERBB2 (HER2) genes. FFPE breast cancer ST specimens as HER2+ / ER- were reclassified based on high expression of ERBB2 without appreciable ESR1 expression. Mapping single-cell transcriptomes to tumor ST samples

[0240] For various analyses herein, CytoSPACE and other benchmarking methods described in "Benchmark analyses using simulated datasets" were applied as follows: CytoSPACE. Cell type proportions were calculated using Spatial Seurat ("Estimate cell type proportions"). CytoSPACE was run with the "duplicated cells" option and the lapjv solver implemented with the lapjv Python package on a single CPU core. For all Visium samples, the average number of cells per spot was set to 5, while for legacy ST samples (melanoma ST data), this parameter was set to 20. Tangram. As input, the same single-cell transcriptomes mapped by CytoSPACE were analyzed, including duplicates, with a density prior (density_prior argument) determined by the number of cells per spot estimated by CytoSPACE. Because Tangram worked best with all genes when used on simulated ST datasets, Tangram (version 1.0.2) was run on the CPM-normalized scRNA-seq data using 24 CPU cores for all available genes. Other parameters were set to default. CellTrek. Assuming CellTrek's internal filtering mechanism (see "Benchmark Analysis Using Simulated Datasets"), all cells in the corresponding scRNA-seq atlas were provided as input (without duplication or downsampling). For the Visium samples, CellTrek (version 0.0.0.9000) was run with 24 CPU cores with default parameters (reduction='pca', intp=T, intp_pnt=10000, intp_lin=F, nPC=30, ntree=1000, dist_thresh=0.4, top_spot=10, spot_n=10, repel_r=5, repel_iter=10, keep_model=T). Cells were then assigned from their raw output coordinates to their nearest spots by Euclidean distance. For the legacy ST samples (melanoma), the code was modified to handle input without h5 and image files, as detailed above. To fit the larger spot resolution in the legacy ST dataset, spot_n was fixed at 40. Other parameters were the same as above. Other methods. Other benchmark methods (DistMap, SpaOTsc, DEEPsc, SpaGE, Spatial Seurat, Harmony, LIGER, Euclidean distance, Pearson correlation, and Spearman correlation) were implemented according to the details described in their corresponding sections in "Benchmark analysis using simulated datasets," with the following exception: for computational feasibility across particularly large scRNA-seq datasets, SpaOTsc was run for two scRNA-seq / ST pairs (CRC and TNBC) with the protocol described above for "Tangram," which provided cells mapped by CytoSPACE, rather than the entire scRNA-seq dataset. Execution Time Analysis

[0241] To truly evaluate the efficiency of CytoSPACE and benchmark it against recent dedicated pairwise cell-mapping methods, run times were recorded for CytoSPACE, Tangram, and CellTrek across all scRNA-seq tumor atlas / ST pairs tested (n = 4 pairs for Visium ST data; n = 2 pairs for low-resolution legacy ST data) with the parameter details described above (Table 3). For CytoSPACE, run times were reported for both the exact (shortest growth path via lapjv solver) and integer approximation solvers, and with and without the Spatial Seurat preprocessing step to obtain input cell type abundance fractions. Data loading and file writing steps were excluded from run times for all methods. Various methods were tested on comparable, but not identical, systems: CytoSPACE, the Spatial Seurat preprocessing step, and Tangram were tested on a compute cluster equipped with an Intel E5-2640v4 processor (2.4 GHz base frequency and 3.4 GHz maximum frequency with 128 GB of associated RAM), an Intel 5118 processor (2.3 GHz base frequency and 3.2 GHz maximum frequency with 191 GB of associated RAM), and an AMD 7502 processor (2.5 GHz base frequency and 3.35 GHz maximum frequency with 256 GB of associated RAM), while CellTrek was tested on a server with an Intel E5-2680v3 processor and 230 GB of associated RAM. All methods were equipped with 24 cores except for CytoSPACE, whose mapping function uses only a single core. Validating alternative solvers

[0242] To demonstrate that the integer approximation solver is a fast alternative to the recommended exact solver (lapjv) and yields comparable results, the proportion of single cells mapping to the same location was estimated across the two solver methods. For each scRNA-seq tumor atlas / ST pair tested, the same single cell was mapped by the exact solver and the integer approximation solver after preprocessing for duplication and downsampling to match the estimated cell type proportion in the tissue via CytoSPACE, and the percentage of cells mapping to the same spot for each method was reported (Table 4). For duplicated cells, no distinction was made between copies. Spatial enrichment analysis

[0243] To determine whether single cells mapping to ST spots exhibited enrichment of known spatially resolved gene expression programs, cells were first divided into two groups ("near" and "far") based on their distance from cancer cells. For breast cancer ST samples, all of them were profiled by 10x Visium, and tumor boundary annotations determined by a pathologist were used to group cells. For the melanoma and CRC datasets, the average Euclidean distance of each TME cell to the five nearest tumor cells (mapped by the respective alignment method) was determined. For the melanoma dataset, melanoma cells were considered tumor cells, while for the CRC dataset, tumor epithelial cells were considered tumor cells for the purpose of identifying tumor location in the tissue. For each TME cell type, the resulting distances were stratified by median into "near" and "far" groups. This was done for two main reasons. First, the CRC samples lacked tumor boundary annotations. Second, while the melanoma dataset contained such annotations, the low spatial resolution of the legacy ST platform prevented accurate co-registration with spatial spots at the tumor / stroma boundary.

[0244] To quantify spatial enrichment, pre-ranked gene set enrichment analysis (GSEA) was performed in fgsea (v1.14.0) with nperm = 10,000. As input, all spatially mapped single-cell transcriptomes were loaded into Seurat v4.1.0 (min.cells = 5) by cell type and normalized with NormalizeData(). For each method and cell type, gene lists ranked by log2 fold change were generated for "near" and "far" identity classes using FoldChange(). If fewer than 10 cells of a cell type were assigned to a spot within a partition by at least one method, that cell type was excluded from the enrichment analysis. Notably, some methods (SpaOTsc, DEEPsc, Seurat, Hamony, and Euclidean distance) were unable to map all assessed cell types to regions both closer and farther from tumor cells, precluding the use of GSEA (as described below in "Spatial Enrichment Analysis") for affected cell types. In such cases, statistical comparisons to CytoSPACE were performed ignoring NAs. Because CytoSPACE and Tangram were each run with the same scRNA-seq input, prior to running Seurat and fgsea, a random sampling of cells was performed to match the number of cells per cell type mapped by CytoSPACE and Tangram and to ensure a fair comparison between the methods. This was done as described in "Matching the Number of Cells Per Cell Type—Overlap." Gene sets for T cell exhaustion and CE9 / CE10-related cell states were derived by Zheng et al. and Luca et al., respectively (C. Zheng et al., Cell 169, 1342-1356.e1316 (2017); and BALuca, Cell 184, 5577-5592.e5518 (2021); the disclosures of which are incorporated herein by reference). Estimating the robustness of CytoSPACE using real data

[0245] The robustness validation described above in "Estimating CytoSPACE robustness with simulated data" was repeated using real data. CytoSPACE was applied to the task of spatial enrichment analysis in TME samples under various perturbations, and performance was quantified according to the recovery of the expected spatial enrichment of gene sets in the TME, as described in "Spatial enrichment analysis" (Figures 15A-15F). The perturbation analysis was performed in the same manner as with simulated data, except for the analysis of robustness to estimation errors in the number of cells per spot. For the scRNA-seq / ST dataset pair, the tuning parameter p was set as follows: 1.4 (Visium data), 1.9 (legacy ST data, melanoma slide 2), and 2.3 (legacy ST data, melanoma slide 1). Spatially resolved macrophage states

[0246] TREM2 + Macrophages and FOLR2 +To assess the spatial localization of macrophages (Figures 16A and 16B), single-cell transcriptomes annotated as "macrophage / monocyte" were mapped to ST spots as described above ("Mapping single-cell transcriptomes to tumor ST samples") and ordered based on their spatial distance (Euclidean) from tumor cells. All cells were processed by Seurat as described in "Spatial enrichment analysis." To calculate distance, the same metrics described for the melanoma and CRC datasets were used ("Spatial enrichment analysis"). For cells mapping within the tumor boundary annotated by the pathologist (breast cancer dataset), distance was set to zero. Cells were then divided into "close" (distance = 0) and "distant" (distance > 0) groups, and the log2 fold change of each gene was calculated using FoldChange() in Seurat (Figure 16B). Integrated single-cell spatial analysis of the healthy mouse kidney

[0247] For analysis of healthy mouse kidneys, the following were downloaded: (i) a fully annotated scRNA-seq atlas encompassing over 30 spatially resolved subtypes of immune cells, stromal elements, and renal epithelium, and (ii) 10x Visium samples of normal mouse kidneys. Renal epithelial cell states lacking a numeric identifier (as in Figure 18A) were omitted, and states corresponding to the same phenotype were merged (3 and 4, 5 and 6, 7 and 8). The datasets were then aligned using CytoSPACE as described in "Mapping Single-Cell Transcriptomes to Tumor ST Samples," except with the average number of cells per spot set to 10. Using epithelial cells with ground truth locations in the scRNA-seq atlas, the following stripes were analyzed: cortex (outermost region), outer medulla (central region), and inner medulla (innermost region), where the outer medulla was further subdivided into an outer stripe (proximal to the cortex) and an inner stripe (proximal to the inner medulla) (Figures 18B and 18C).

[0248] A ground truth rank was established for each epithelial cell state, reflecting its relative distance to epithelial state 32 ("pelvic deep medullary epithelium"), which corresponds to the base of the ureteral epithelium (UE) in the inner medulla as previously reported (Figure 18A and Table 5). The single-cell spatial coordinates determined by CytoSPACE were then used to calculate the average Euclidean distance of each epithelial cell state to the centroid of the epithelial cells mapped to the epithelial state. Regardless of whether the nephron or UE was examined, the correlation between the predicted distance and the ground truth distance was high, demonstrating the potential of CytoSPACE for granular mapping (Figures 19A-D).

[0249] For the analysis in Figures 20A–20E, we verified whether CytoSPACE could resolve known structures of the nephron and UE assembly system (Figure 20A), which are not recognizable from the scRNA-seq atlas (Figure 20B) or the ST dataset alone. To this end, a spatial spot was scored as 1 if at least one cell of a given cell type was mapped by CytoSPACE, and 0 otherwise. The resulting binary square matrix (cell type as rows and cell type as columns) was then converted to a Jaccard similarity matrix J, which quantifies the spatial overlap between epithelial states (Figure 20C, left). After filtering all but the four nearest neighbors of each epithelial state in J, each row was converted to rank space, and an undirected graph was created from the data using igraph v1.2.6 in R. The graph was then visualized using layout_with_fr(), the force-directed layout algorithm of Fruchterman and Reingold implemented in igraph (Figure 20D). To determine statistical significance (Figure 20D), the nearest neighbors N of each epithelial state i in J i A sorting approach was devised in which is first determined. Then, N iand the ground truth nearest neighbor (1 or more than 1) i The minimum number of physically adjacent epithelial states (represented by x) was calculated (Figure 20C, right). i was calculated for all evaluable epithelial conditions, and the results were then averaged. [ka] After this, each row of J is randomly permuted and the average distance [ka] was recalculated. This means that [ka] A total of 100,000 iterations were performed to calculate the empirical p-value of . To generate the UMAP plot in Figure 20B, the following Seurat v4.0.1 commands were applied sequentially to the log-normalized scRNA-seq data of epithelial cell states from Ransick et al.: FindVariableFeatures() with selection.method="vst" and nfeatures=2000, ScaleData(), RunPCA(), FindNeighbors() with dims=1:10, and RunUMAP() with dims=1:30. Application to single cell ST data

[0250] While CytoSPACE's primary goal is the reconstruction of bulk ST data at the single-cell level, it can also be directly applied to single-cell ST data. To do this efficiently for very large single-cell ST datasets, a sampling routine was implemented to uniformly distribute single-cell ST datasets into bins (by default) of up to 10,000 cells each, without replacement; this balances considerations of cell diversity and mapping efficiency. Specifically, the single-cell ST dataset is first randomly distributed into n bins of 10,000 ST cells each, without replacement. Next, for each bin (1,...,n), 10,000 single-cell transcriptomes are sampled from the scRNA-seq query dataset (by default) according to the procedure described above in "Matching the Number of Cells Per Cell Type—Duplicate." While the entire procedure is reproducible and fixed to a specific seed in the initial setup, the scRNA-seq dataset is resampled anew for each bin (1,...,n) to enhance robustness. Finally, CytoSPACE is run on each bin and the results are combined to produce a single aggregated output.

[0251] For the analyses in Figures 21A and 21B and the expanded data in Figures 22A–22I, the preprocessed MERSCOPE profile of an FFPE human breast cancer sample (HumanBreastCancerPatient1) was downloaded from Vizgen (info.vizgen.com / merscope-ffpe-access). Cells with fewer than 100 detected transcripts and cells with fewer than 10 detected genes were excluded from the analysis, resulting in 560,655 cells with an average of 149 detected genes per cell. The gene by cell count matrix was normalized by downsampling, which eliminated potential confounding factors (e.g., cell volume) by normalizing the total transcripts per cell to be the same (300 transcripts per cell). Seurat v4.1.1 was used to analyze the normalized data. The top 100 variable genes were identified using FindVariableFeatures(), and cells were clustered using FindCluster() with resolution = 0.8. Using canonical marker genes, clusters were annotated as fibroblasts (high COL1A1 or COL5A1), endothelial cells (high PECAM1 or VWF), macrophages (high FCGR3A or C1QC), dendritic cells (high CD1C or CD207), lymphocytes (high CD3E, TRAC, ZAP70, MS4A1, GNLY, or MZB1), and epithelium (remaining). Lymphocytes were further clustered using the top 300 variable genes with a resolution of 1.2 and annotated as CD4 T cells (high CD3E, TRAC, ZAP70, or FOXP3 and absent CD8A), CD8 T cells (high CD3E, TRAC, or ZAP70 and high CD8A), NK cells (high GNLY and absent CD3E), B cells (high MS4A1), and plasma cells (high MZB1). Clusters that did not meet these criteria but showed strong expression of non-lymphoid markers were annotated accordingly using the epithelial and stromal markers described above.

[0252] To account for errors in transcript assignment arising from overlapping cells in the z-series, gene expression in the central z-plane (z = 3) was compared with expression in the peripheral z-plane (z = 0) for each segmented cell. Transcripts detected in either of these z-planes were first isolated as individual genes by cell count matrix. All genes whose expression significantly differed between these two z-planes for one or more cell types were then identified using a two-tailed Wilcoxon test (nominal P < 0.05). For each of these genes, if expression was significantly higher in the central z-plane for one cell type but significantly higher in the z = 0 plane for another cell type, the gene was considered a potential contaminant and set to 0 in all cells of the latter cell type.

[0253] For the analyses in Figures 22A-22E, the MERSCOPE dataset was randomly split (50:50) into an "scRNA-seq" query dataset and an ST reference dataset (Figure 22A). Query cells were then mapped to the reference as described above, running CytoSPACE with 5 CPU cores, the number of cells per spot set to 1, and the overall abundance fraction of each cell type set to its ratio in the reference dataset (Figure 22B). Strong agreement was observed for various cell type labels (Figure 22C), and for each cell type, the gene expression profiles (GEPs) of mapped cells correlated with their assigned reference cells to a greater extent than with other reference cells of the same cell type (Figure 22D). We then asked whether pairwise transcriptome distances between single cells were preserved (Figure 22A). To do so for each evaluable cell type, the pairwise correlation matrix Q of single-cell GEPs (in log2 space) in the scRNA-seq query dataset was calculated. This was done after assigning the query cell to a spatial location in the reference. The same was then done for the reference dataset, resulting in matrix R. Both matrices were ordered identically according to the same single-cell spatial coordinates, which allowed us to determine whether the spatial correlation structure was reproduced between the mapped cells. Indeed, by calculating the retention index for each cell type, defined as the Pearson correlation between these two matrices, highly significant retention of pairwise distances was observed for each cell type (P < 2.2e-16; Figure 22E). To ensure unbiased evaluation, equal numbers of cells per cell type were sampled (without restoration) based on the lowest common denominator (n = 150 cells) in the reference dataset prior to constructing each matrix.The degree of retention was found to be proportional to the variance between GEPs in the reference dataset: i.e., cell types with lower transcriptome heterogeneity in the reference (i.e., more uniform GEPs) had less spatial structure and lower retention of pairwise distances, consistent with expectations (Figure 22E).

[0254] Because the MERSCOPE dataset lacked ESR1 (estrogen receptor) and PGR (progesterone receptor) among its 500 target genes but showed elevated expression of ERBB2 (encoding HER2), HER2+ breast tumors profiled by scRNA-seq were selected as the query dataset in Figures 21A-B (Table 2). To ensure sufficient overlap in co-detected genes, cells from the scRNA-seq dataset with fewer than 50 expressed genes (CPM > 0) overlapping with the MERSCOPE panel were removed. The scRNA-seq atlas was then mapped to MERSCOPE by running CytoSPACE on five CPU cores, with the number of cells per spot set to 1, and the overall abundance fraction of each cell type set to its ratio as determined above.

[0255] To assess spatial enrichment of cell states in Figures 21A and 21B and 22F-22I, individual cells were first partitioned into two regions based on their Euclidean distance to epithelial cells. If an epithelial cell was located within 100 μm of more than 50 epithelial cells, it was assigned to the tumor region. This threshold was selected based on a density-based analysis, which revealed two major distributions of epithelial cell density, with approximately 50 epithelial cells per 100 μm radius representing a local minimum between these two distributions. Of the remaining cells, a cell was then assigned to the tumor region if it was located within 100 μm of a tumor epithelial cell; otherwise, it was assigned to the adjacent normal region (i.e., stroma) (Figure 22H). For the analyses shown in Figures 21B and 22I, the log2 fold change of each gene in the tumor region relative to the stromal region was determined for CD4 and CD8 T cells using raw MERSCOPE data (500 genes) or scRNA-seq data (whole transcriptome) mapped to MERSCOPE. Pre-ranked gene set enrichment analysis (GSEA) was applied as described in "Spatial Enrichment Analysis" for the top 200 signature genes of each pan-cancer T-cell state defined by Zheng et al., except for 'CD4T_IL7R-Tn', which lacked signature genes in the MERSCOPE dataset. For this analysis, the fgsea package version 1.20.0 was used. Ground truth was determined as the rank of the log2 fold change between the tumor odds ratio and normal odds ratio for each assessed T-cell state. statistics

[0256] All statistical tests were two-sided unless otherwise specified. The Wilcoxon test was used to evaluate statistical differences between two groups. Adjustments for multiple hypothesis testing were made via Benjamini-Hochberg where applicable. Linear agreement was determined by Pearson (r) correlation or Spearman correlation (ρ), and two-sided t-tests were used to assess whether results were significantly non-zero. All statistical analyses were performed using R v3.5.1 and 4.0.2+, Python 3.8, MATLAB_R2019a®, and Prism9+ (Graphpad Software, La Jolla, CA). Table 1. Benchmark evaluation results for the simulated ST dataset. a-b, Proportion of cells per cell type mapping to the correct spot in ST data across various noise levels and spatial resolutions by method for mouse cerebellum samples (a) and mouse hippocampus samples (b). c-d, Overall assignment accuracy across various noise levels and spatial resolutions by method for mouse cerebellum samples (c) and mouse hippocampus samples (d). Single-cell assignment accuracy is the proportion of cells that map correctly to the corresponding ground truth spot in the ST data, and cell type accuracy is the proportion of cells that map to spots in the ST data that contain at least one cell of the same cell type in the ground truth. e-h, Same as a-d, respectively, except for the selected method with RCTD used for cell type proportion estimation for CytoSPACE inputs instead of Spatial Seurat. [Table 1a-1] [Table 1a-2] [Table 1a-3]

Table 1a-4

Table 1b-1

Table 1b-2

Table 1b-3

Table 1b-4

Table 1b-5

Table 1b-6

Table 1c-1

Table 1c-2

Table 1d

Table 1e-1

Table 1e-2

Table 1e-3

Table 1f-1

Table 1f-2

Table 1f-3

Table 1g

Claims

[Claim 1] The invention described in the specification.