Systems and methods for comparative analysis of single cell RNA sequencing data

The scCompare pipeline addresses the limitations of existing scRNA-seq data analysis methods by enabling comprehensive comparative analysis and identifying novel cell types through correlation-based mapping, achieving higher precision and sensitivity than current tools.

WO2025137517A1PCT designated stage expired Publication Date: 2025-06-26BLUEROCK THERAPEUTICS LP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/061387
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current methods for analyzing single cell RNA sequencing (scRNA-seq) data are inadequate for comparative analysis, particularly in ensuring consistency and quality of cell therapies across different batches and conditions.

Method used

The development of a computational pipeline, referred to as scCompare, which uses correlation-based mapping approaches to determine phenotypic identities across distinct scRNA-seq datasets, allowing for comprehensive and nuanced comparisons of cellular phenotypes.

Benefits of technology

scCompare effectively identifies phenotypically equivalent cell populations and detects novel cell types by applying statistically determined lower thresholds for phenotype mapping, outperforming existing tools like Single-cell Variational Inference (scVI) in precision and sensitivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024061387_26062025_PF_FP_ABST
    Figure US2024061387_26062025_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure provides computational methods and systems for analyzing and comparing single-cell RNA sequencing data. In particular, this disclosure provides methods and systems employing correlation-based mapping approaches to determine phenotypic identities across distinct datasets.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR COMPARATIVE ANALYSIS OF SINGLE CELL RNA SEQUENCING DATA CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application Serial No. 63 / 614,097, entitled “SYSTEMS AND METHODS FOR COMPARATIVE ANALYSIS OF SINGLE CELL RNA SEQUENCING DATA,” filed on December 22, 2023, which is hereby incorporated by reference in its entirety. FIELD

[0002] This disclosure relates to the field of computational biology and genomics. This disclosure includes methods and systems for analyzing and comparing single cell RNA sequencing (scRNA-seq) data sets. BACKGROUND

[0003] Cell therapy generally involves administering living cells to patients for the treatment of a disease or condition. The development of induced pluripotent stem cell (iPSC) technology has revolutionized this field by enabling the generation of diverse cell types, such as dopaminergic neurons and cardiomyocytes, for therapeutic purposes. However, one challenge in the development of cell therapies is ensuring consistency, safety, and efficacy of these cellular products from complex bioprocesses.

[0004] One approach to evaluating cell populations for therapeutic purposes is through single cell RNA sequencing (scRNA-seq). Common practices in analyzing data from scRNA- seq approaches primarily focus on analyzing single datasets. Such practices include clustering (grouping similar cells), dimensionality reduction (simplifying complex data for better understanding and visualization), and differential gene expression analysis (identifying active genes in different cell types). And while these approaches are valuable for studying individual cell populations, they are less equipped for the comparative analysis needed in cell therapy development, where consistency across different cell batches and conditions is important. Page 1 of 62 ACTIVE 692268512v5

[0005] The lack of robust comparative analysis tools in scRNA-seq data interpretation presents a challenge in the development of cell therapies (e.g., particularly in quality control, process optimization, and / or preclinical research). Existing tools do not adequately address the nuances and complexities involved in comparing cellular profiles across different conditions and developmental stages, which is essential for developing safe and effective cell therapies. The present disclosure satisfies this need and provides related advantages. SUMMARY

[0006] This disclosure relates to computational methods and systems for analyzing single cell RNA sequencing (scRNA-seq) data. In some embodiments, this disclosure provides novel computational pipeline, se.g. that use correlation-based mapping approaches to determine phenotypic identities across distinct scRNA-seq datasets (e.g., leveraging average transcriptomic signatures from individual phenotypes). These approaches allow for comprehensive and nuanced comparisons of cellular phenotypes between different scRNA- seq datasets. One such pipeline, referred to herein as scCompare, possesses the capacity to identify potentially novel cell types by applying statistically determined lower thresholds for phenotype mapping, thereby marking unique, unmapped cells for additional exploration. This disclosure is accompanied by experimental evidence showing scCompare outperforms existing tools, such as Single-cell Variational Inference (scVI), in precision and sensitivity, and in differentiating between cell types of complex datasets.

[0007] The methods and systems described herein are applicable to a variety of applications. In some examples, the methods and system may be used for optimizing cell bioprocesses, ensuring the purity of therapeutic cell populations, maintaining consistent cell quality across various production batches, and identifying novel cell types.

[0008] In some aspects, the techniques described herein relate to a method for using single-cell RNA sequencing (scRNAseq) data to determine whether a first set of cells, derived using a first bioprocess, is phenotypically equivalent to a second set of cells, derived using a second bioprocess. The method may include using at least one computer hardware processor to perform obtaining a plurality of clusters of the first set of cells, each of the plurality of clusters corresponding to a respective phenotype of a plurality of phenotypes, obtaining scRNAseq expression data for the second set of cells, the scRNAseq expression data including, for each of multiple cells in the second set of cells, a plurality of expression Page 2 of 62 ACTIVE 692268512v5levels for a respective plurality of genes, determining, using the plurality of clusters and the scRNAseq expression data, a phenotype of the plurality of phenotypes for each of at least some of the cells in the second set of cells, determining, for each of at least some phenotypes of the plurality of phenotypes, a proportion of a number of the first set of cells of the particular phenotype to a total number of the first set of cells, thereby obtaining a first plurality of phenotype proportions, determining, for each of the at least some phenotypes of the plurality of phenotypes, a proportion of a number of the second set of cells determined to be of the particular phenotype to a total number of the second set of cells, thereby obtaining a second plurality of phenotype proportions, and determining whether the first set of cells and the second set of cells are phenotypically equivalent using the first plurality of phenotype proportions and the second plurality of phenotype proportions.

[0009] In some aspects, the method may further include, in response to determining that the first set of cells and the second set of cells are not phenotypically equivalent, determining whether the second set of cells includes pluripotent stem cells or progenitor cells, and in response to determining that the second set of cells includes pluripotent stem cells, adjusting an aspect of the second bioprocess.

[0010] In some aspects, adjusting the aspect of the second bioprocess may include adding one or more agents to the second bioprocess, adjusting a concentration of one or more agents in the second bioprocess, or adjusting an initial concentration of pluripotent stem cells used in the second bioprocess. In some aspects, the method may further include, in response to determining that the first set of cells and the second set of cells are not phenotypically equivalent, identifying cells of the second set of cells that are of an unknown phenotype, and determining a gene expression signature for the identified cells.

[0011] In some aspects, the method may further include using the gene expression signature to develop an assay for quantifying cells of the unknown phenotype in one or more other samples.

[0012] In some aspects, the method may further include determining a phenotype for the identified cells using the gene expression signature. In some aspects, determining the phenotype for the identified cells using the gene expression signature includes identifying one or more marker genes included in the gene expression signature, and determining the phenotype based on the identified marker genes. Page 3 of 62 ACTIVE 692268512v5

[0013] In some aspects, the method may further include determining whether to administer a cell therapy based on a result of determining whether the first set of cells and the second set of cells are phenotypically equivalent, the cell therapy including one or more products of cells derived using the second bioprocess. In some aspects, the method may further include determining whether to discard the second set of cells based on a result of determining whether the first set of cells and the second set of cells are phenotypically equivalent, and in response to determining to discard the second set of cells based on the result, discarding the second set of cells.

[0014] In some aspects, determining whether the first set of cells and the second set of cells are phenotypically equivalent using the first plurality of phenotype proportions and the second plurality of phenotype proportions includes determining a coefficient of determination (R2) between the first plurality of phenotype proportions and the second plurality of phenotype proportions, the coefficient of determination (R2) representing a degree to which the first plurality of phenotype proportions and the second plurality of phenotype proportions are phenotypically equivalent, and determining that the first set of cells and the second set of cells are phenotypically equivalent if the coefficient of determination (R2) satisfies at least one criterion.

[0015] In some aspects, the method may further include obtaining a first scRNAseq expression data for the first set of cells, the first scRNAseq expression data including, for each of multiple cells in the first set of cells, a plurality of expression levels for a respective first plurality of genes, wherein obtaining the plurality of clusters of the first set of cells includes clustering the first set of cells into the plurality of clusters using the first scRNAseq expression data.

[0016] In some aspects, clustering the first scRNAseq expression data includes clustering the first scRNAseq expression data using a hierarchical clustering algorithm. In some aspects, clustering the first scRNAseq expression data using the hierarchical clustering algorithm includes clustering the scRNAseq expression data using Leiden clustering.

[0017] In some aspects, obtaining the first scRNAseq expression data for the first set of cells includes obtaining, for each cell in the first set of cells, expression levels for between 5,000 and 10,000 genes, between 10,000 and 15,000 genes, between 15,000 and 20,000 genes, between 20,000 and 25,000 genes, or between 25,000 and 30,000 genes. Page 4 of 62 ACTIVE 692268512v5

[0018] In some aspects, the method may further include determining a gene expression signature for each particular cluster in the plurality of clusters using the first scRNAseq expression data, the determining including, for each particular gene in a subset of the first plurality of genes, determining an average expression level of the particular gene across cells included in the particular cluster. In these aspects, determining the phenotype for each of at least some of the cells in the second set of cells using the plurality of clusters may include determining the phenotype using the gene expression signatures determined for the plurality of clusters.

[0019] In some aspects, the method may further include, for a particular cluster of the plurality of clusters, generating a distribution of correlation coefficients (e.g., by determining, for each cell included in the particular cluster, a correlation coefficient between a gene expression signature of the cell and the gene expression signature determined for the particular cluster, and determining a statistical cut-off value for the particular cluster). In one such aspects, the statistical cut-off may represent a mean absolute deviation from a median of the generated distribution of correlation coefficients.

[0020] In some aspects, determining the phenotype of the plurality of phenotypes for each of at least some of the cells in the second set of cells includes, for a particular cell of the at least some cells in the second set of cells determining a correlation coefficient between scRNAseq expression data obtained for the particular cell and the gene expression signature determined for each cluster, and determining the phenotype for the particular cell based on the correlation coefficients and the statistical cut-off value for each cluster.

[0021] In some aspects, the method may further include identifying the particular cell as an outlier if none of the determined correlation coefficients exceed the respective statistical cut-off value. In some aspects, the first bioprocess is different from the second bioprocess. For example, in some instances the first bioprocess comprises a differentiation protocol that is different than the differentiation protocol employed in the second bioprocess. In some instances, the first bioprocess and the second bioprocess comprise one or more different agents for differentiating cells.

[0022] In some aspects, the method is used for quality control purposes, including purity assessments, to verify the consistency of induced pluripotent stem cell (iPSC)-derived cell populations. In some aspects, the method is used to refine one or more cell manufacturing Page 5 of 62 ACTIVE 692268512v5processes, including the comparison of cell batches produced under varying conditions or protocols. In some aspects, the method is used in research and development of cell therapies.

[0023] In some aspects, the techniques described herein relate to a system, including at least one computer hardware processor, and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method for using single-cell RNA sequencing (scRNAseq) data to determine whether a first set of cells, derived using a first bioprocess, is phenotypically equivalent to a second set of cells, derived using a second bioprocess, the method including obtaining a plurality of clusters of the first set of cells, each of the plurality of clusters corresponding to a respective phenotype of a plurality of phenotypes, obtaining scRNAseq expression data for the second set of cells, the scRNAseq expression data including, for each of multiple cells in the second set of cells, a plurality of expression levels for a respective plurality of genes, determining, using the plurality of clusters and the scRNAseq expression data, a phenotype of the plurality of phenotypes for each of at least some of the cells in the second set of cells, determining, for each of at least some phenotypes of the plurality of phenotypes, a proportion of a number of the first set of cells of the particular phenotype to a total number of the first set of cells, thereby obtaining a first plurality of phenotype proportions, determining, for each of the at least some phenotypes of the plurality of phenotypes, a proportion of a number of the second set of cells determined to be of the particular phenotype to a total number of the second set of cells, thereby obtaining a second plurality of phenotype proportions, and determining whether the first set of cells and the second set of cells are phenotypically equivalent using the first plurality of phenotype proportions and the second plurality of phenotype proportions. In some aspects, the first bioprocess in the system is different from the second bioprocess. For example, in some instances the first bioprocess comprises a differentiation protocol that is different than the differentiation protocol employed in the second bioprocess. In some instances, the first bioprocess and the second bioprocess comprise one or more different agents for differentiating cells.

[0024] In some aspects, the techniques described herein relate to at least one non- transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform a method for using single-cell RNA sequencing (scRNAseq) Page 6 of 62 ACTIVE 692268512v5data to determine whether a first set of cells, derived using a first bioprocess, is phenotypically equivalent to a second set of cells, derived using a second bioprocess, the method including obtaining a plurality of clusters of the first set of cells, each of the plurality of clusters corresponding to a respective phenotype of a plurality of phenotypes, obtaining scRNAseq expression data for the second set of cells, the scRNAseq expression data including, for each of multiple cells in the second set of cells, a plurality of expression levels for a respective plurality of genes, determining, using the plurality of clusters and the scRNAseq expression data, a phenotype of the plurality of phenotypes for each of at least some of the cells in the second set of cells, determining, for each of least some phenotypes of the plurality of phenotypes, a proportion of a number of the first set of cells of the particular phenotype to a total number of the first set of cells, thereby obtaining a first plurality of phenotype proportions, determining, for each of the at least some phenotypes of the plurality of phenotypes, a proportion of a number of the second set of cells determined to be of the particular phenotype to a total number of the second set of cells, thereby obtaining a second plurality of phenotype proportions, and determining whether the first set of cells and the second set of cells are phenotypically equivalent using the first plurality of phenotype proportions and the second plurality of phenotype proportions. BRIEF DESCRIPTION OF DRAWINGS

[0025] FIG. 1 illustrates an exemplary scCompare workflow. The workflow includes a semi-supervised clustering that is performed to group a mapping data set into phenotypically relevant subsets (e.g., clusters). Bulk signatures derived from the prototype gene expression signature for each cluster are generated and distributions (e.g., of within-cluster Pearson correlations of each member cell to the cluster’s prototype signature) are formed. These provide statistical cutoffs for cluster inclusivity. Then, cells from a test data set are correlated with each prototype signature and a resulting Pearson correlation is compared to the statistical cutoffs for cluster inclusivity. Cells within the test data set that pass the threshold (e.g., the statistical cutoffs to be included in a cluster) are considered mapped and are labeled with the cluster’s phenotype, otherwise they are labeled as unmapped. Finally, fractions of mapped cells are compared between the mapping data set and the testing data set (e.g., to facilitate comparability of phenotypic representations). Page 7 of 62 ACTIVE 692268512v5

[0026] FIGS. 2A-2C show exemplary results of a comparison of scCompare and scVI models, both initially derived from the 3k PBMC data set when applied to the 68k PBMC data set. This comparison includes the identification of an unannotated cell type, plasmacytoid dendritic cells. FIG. 2A shows the 3k PBMC data set. (i) UMAP of 3k PBMC data set. (ii) Correlation dendrogram of 3k PBMC phenotypes (iii) 3k PBMC gene expression dot plot showing expression of key phenotypic markers. FIG. 2B shows the results of mapping 3k PBMC identities onto the 68k PBMC data set using scCompare and scVI. (i-ii) UMAPs showing phenotypes assigned by scCompare (i) and scVI (ii). (iii) Gene expression dotplots showing consistency of key gene expression between the 3k PBMC data set and the assigned phenotypes in the 68k data set (top – scCompare, bottom – scVI). scVI does not predict any megakaryocytes in the 68k PBMC dataset despite evidence for their presence (small PPBP-expressing subcluster (B iii). The intensity scale represents the mean expression of a gene within a cell type. The size of the dots represents the fraction of the cells within a cell type. FIG. 2C shows statistical cutoff-assisted discovery of plasmacytoid dendritic cells. Cells not meeting statistical cutoff were assessed for novel phenotype status using differential gene expression list analysis and UMAP Euclidean distance metrics. After assessment and remapping, a cluster of plasmacytoid dendritic cells were discovered. This phenotype is confirmed by expression of MZB1.

[0027] FIGS. 3A-3E shows exemplary results of experiments validating scCompare. In particular, FIGS. 3A-3E shows representative outcomes of scCompare successfully reproducing published mapping of phenotypes between two directed differentiation protocols. FIG. 3A shows a sample schematic indicating the types of differentiation protocols and timepoints (dashed box) analyzed. FIG. 3B shows UMAP of Leiden clustering of “mapping” data set, Protocol 1 Days 12 and 24, and violin plots displaying marker gene expression profiles for each Leiden cluster, with cell type indicated above. CM = cardiomyocyte, PCM = progenitor CM, STR= stromal-like, EC = endothelial-like, SM = smooth muscle-like, END = endodermal, ECT = ectodermal. FIG. 3C shows a correlation dendrogram of highly variable genes and Leiden clusters in the “mapping” data set, Protocol 1, and UMAP with shading indicating cell type. FIG. 3D shows a UMAP of “testing” data set, Protocol 2 Days 14 and 26, after mapping using the annotated Protocol 1 (training) cell types, and UMAP of Pearson correlation coefficients of each labeling. FIG. 3E show a scatter plot showing the relative ratios (natural log) of annotated cell types between Protocol 1 and Protocol 2. Compared to Protocol 1, Protocol 2 contains relatively fewer endodermal, endothelial-like, ectodermal, and Page 8 of 62 ACTIVE 692268512v5smooth muscle-like cells, while having relatively more stromal and cardiomyocyte cells. An x=y line is plotted to aid in the visual inspection of the ratio comparisons. Dots falling to the lower right of the line represent cell types enriched in Protocol 1 whereas dots falling to the upper left of the line represent cell types enriched in Protocol 2. Dots that fall directly on the line indicate an equal fraction of the cell types is shared between Protocol 1 and 2.

[0028] FIGS. 4A-4B show dotplot visualizations of marker genes between different protocols (Protocol 1 and Protocol 2) for making cardiomyocytes. Differentially expressed genes between clusters of cells are shown in FIG. 4A for Protocol 1 vs Protocol 2 for timepoint 1 (D12 / D14) and FIG. 4B shows the same for Protocol 1 vs Protocol 2 for timepoint 2 (D24 / D26). The legend denotes the size of the dots as the percentage of the cells in the groups expressing the genes and the shaded bar represents the level of mean expression of the genes within the groups.

[0029] FIGS. 5A-5D shows an scCompare grid search experiment exploring the effects of parameter optimization on results using the Human Protein Atlas scRNA-seq data set. Within FIG. 5A, each point (a subset of 10,000 cells) represents an individual run of scCompare, with its highly variable genes assigning position in X and Leiden clustering resolution assigning its position in Y. The shading of each point is assigned by its classification performance as measured by class-size weighted F1 score. FIG. 5B and FIG. 5C present a different view of the data in FIG. 5A, where the points are positioned by map and test misclassification rates in X and Y respectively. The identify line acts as a visual guide, where points deviating from this line have unequal percentages of map and test set misclassification rates. The plots within FIG. 5B are shaded by the number of Leiden clusters and the plots within FIG. 5C are shaded by the number of highly variable genes. FIG. 5D shares the same axis’ as FIG. 5A, but points are instead shaded by the absolute difference in misclassification rate between map and test sets.

[0030] FIGS. 6A-6B shows exemplary results of an assessment of unmapped cells using list subset and UMAP centroid distance metrics. FIG. 6A shows a table of subset similarity and UMAP centroid distance metrics for each unmapped cell category. FIG. 6B shows a dotplot showing a unique gene expression signature of the plasmacytoid dendritic cells after assigning the unmapped_Dendritic and B cells to this new phenotypic category. Of note is the expression of LRRC26, a gene uniquely expressed by plasmacytoid dendritic cells in the data set and a known marker of this cell type. Page 9 of 62 ACTIVE 692268512v5

[0031] FIGS. 7A-7C show exemplary UMAPs and corresponding dotplots for different categories of unmapped cells. For each unmapped cell type, the differentially expressed genes post filtering are plotted. Note that for each example, the cells in the unmapped category show a reduction in the level of expression for the gene signature. In general, this is likely due to an increased level of sparsity in the unmapped cells due to high dropout. Furthermore, for all excluding ‘unmapped_B’ and ‘unmapped_dendritic’, there is a strong overlap in UMAP space between the confidently mapped and unmapped cells in each category.

[0032] FIGS. 8A-8D shows exemplary visualization of scCompare's phenotypic mapping results between different tissue and cell type classifications in the Tabula Sapiens (TS) dataset. FIG. 8A shows a circular visualization (circos plot) of endothelial cells by tissue distribution, with colored segments representing different tissue origins (Heart, Liver, Lung, Muscle, Prostate, Salivary Gland, Skin, Thymus, and Vasculature) and connecting bands showing relationships between correctly classified (CDR) and misclassified (MIS) cells. FIG. 8B shows mapping classification results for lung tissue by cell type, displaying the distribution and relationships between different cell populations (Alveolar cells type 2, Club cells, Endothelial cells, Fibroblasts, Macrophages, Respiratory ciliated cells, and T-cells). FIG. 8C shows T-cell subtype distribution across tissues, with colored segments representing different T-cell subtypes (CD4+ helper T cell, CD8+ cytotoxic T cell, and various memory and regulatory T cell populations) and their mapping relationships. FIG. 8D shows intestinal tissue cell type distribution and classification accuracy, with segments representing different cell populations specific to intestinal tissue (CD4+ positive T cells, CD8+ positive T cells, enterocytes of epithelium of large / small intestine, enteroendocrine cells, goblet cells, and intestinal goblet cells).

[0033] FIG. 9 is a block diagram of an example system for performing comparative analysis of scRNA-seq data, according to some embodiments of the technology described herein.

[0034] FIG. 10 is a schematic diagram of an illustrative computing device with which aspects described herein may be implemented. DETAILED DESCRIPTION Page 10 of 62 ACTIVE 692268512v5

[0035] Aspects of this disclosure relate to computational methods and systems for analyzing genomics data. In some aspects, this disclosure provides computational methods and systems for analyzing and comparing genomics data, such as single-cell RNA sequencing (scRNA-seq) data. In certain embodiments, these computational methods and systems are enabled by a series of one or more computer algorithms and statistical models. The algorithms and statistical models may facilitate the comparison of cell populations from different sources and identify unique cell types (e.g., by setting specific thresholds for phenotype mapping). The methods and systems described herein provide a wide range of utility, enabling analyses with unparalleled precision and sensitivity. In some embodiments, the disclosed computational methods and systems can discern differences between different cell types and / or differences between cell populations prepared using different bioprocesses. Such capabilities make them well suited for optimizing cell bioprocessing workflows, ensuring the purity and consistency of cell populations for therapeutic applications, and / or maintaining quality across different cell production batches.

[0036] scRNA-seq allows for an unbiased assessment of cellular phenotype by enabling the extraction of transcriptomic data (e.g., data delineating a set of RNA transcripts present in a cell or population of cells at a specific point in time) from a sample. One question that arises in downstream analysis is how to evaluate biological similarities and differences between samples in high dimensional space. This becomes especially complex when there is cellular heterogeneity within the samples. In certain embodiments, this disclosure provides a novel computational system, referred to herein as scCompare, which allows for these types of comparisons of scRNA-seq data sets. Using scCompare, phenotypic identities from a known data set can be transferred onto another data set (e.g., using correlation-based mapping to average transcriptomic signatures from each phenotype identified in the known data set). In some examples, scCompare may use statistically derived lower thresholds (e.g., cut-offs) for phenotype inclusivity. Using statistically derived lower thresholds allows for cells to be unmapped if they are distinct from the known phenotypes, facilitating potential novel cell type detection.

[0037] One advantage of scCompare, as described herein, is that the pipeline makes use of computational strategies that are straightforward and more easily understood as in comparison to more complex artificial learning systems. scCompare harnesses well- established computer algorithms, which not only ensures reliability but also makes the Page 11 of 62 ACTIVE 692268512v5process and outcomes more transparent and interpretable. This simplicity and clarity in data interpretation are important factors, especially in contexts where understanding the nuances of data is important, such as in supporting drug products for regulatory approval. Furthermore, the user-friendly nature of scCompare's analytics can reduce the need for extensive computational expertise, broadening its accessibility and applicability in various therapeutic development settings.

[0038] The present disclosure is accompanied with results of experiments from a head-to- head comparison of scCompare and Single-cell Variational Inference (scVI) for the analysis of scRNA-seq data sets from human Peripheral Blood Mononuclear cells (PBMCs). These experimental results show that scCompare outperforms Single-cell Variational Inference (scVI) in terms of higher precision and sensitivity for most of the cell types. Furthermore, experimental results presented herein demonstrate the ability of scCompare to detect novel or unique populations of cells. For example, as will described in the Examples, scCompare was used on a cardiomyocyte differentiation data set where it discovered a novel cluster of cells that might be related to the addition of a cytokine in the differentiation bioprocess. Further experimental evidence is provided to demonstrate the use of scCompare on cell atlas data sets, which revealed insights into the cellular heterogeneity underpinning biological diversity between samples. In addition, the cell atlas was used to better understand the effect of key parameters used in the scCompare pipeline.

[0039] While existing tools such as Scanpy provide certain preprocessing capabilities for scRNA-seq data analysis, the computational methods and systems described herein extend beyond such functionalities to provide enhanced phenotypic mapping strategies. In some embodiments, while the methods and systems may utilize Scanpy or similar tools for initial data filtering, preprocessing, and clustering, one core insight of this disclosure lies in the subsequent analysis pipeline that enables precise phenotypic comparison between datasets and identification of novel cell types. This approach differs fundamentally from existing methods, such as Scanpy's methods, by implementing correlation-based mapping with statistically derived thresholds that can identify unmapped cell populations.

[0040] Furthermore, methods and systems described herein provide several technical advantages over existing approaches. For example, the correlation-based mapping strategy employs statistical thresholds that are empirically derived from the data itself, allowing for more robust phenotype matching between datasets. In addition, unlike existing tools that Page 12 of 62 ACTIVE 692268512v5force all cells to map to known categories, the methods and systems described herein can identify cells that do not correspond to any known phenotype, thereby enabling the discovery of novel cell types. Also, exemplary computational pipelines described herein can quantitatively assess phenotypic shifts in response to experimental conditions or bioprocess modifications, providing a powerful tool for optimizing cell therapy manufacturing processes. These technical advantages make the methods and systems particularly well-suited for quality control and process optimization in therapeutic applications where precise phenotypic comparison is important.

[0041] Moreover, in some embodiments, the methods and systems described herein are designed for compatibility with widely-used data formats and analysis workflows. For example, the methods and systems can process data in the AnnData format commonly used in the Python ecosystem or accept converted Seurat objects from the R environment via the h5Seurat file format specification. This compatibility enables seamless integration with existing bioinformatics pipelines while providing novel analytical capabilities not present in current tools. The methods and systems described herein can be particularly advantageous in therapeutic development settings where understanding cellular phenotype changes across different manufacturing conditions is important for optimizing production processes and ensuring product consistency. I. Definitions

[0042] The following definitions supplement those in the art and are directed to the present disclosure only. The following definitions are not to be imputed to any related or unrelated case, e.g., to any commonly owned patent or patent application. Although some methods and materials similar or equivalent to those described herein can be used to practice features of the disclosure, some preferred materials and methods are described herein. Accordingly, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.

[0043] Unless otherwise defined herein, scientific and technical terms used in connection with the present application shall have the meanings that are commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Page 13 of 62 ACTIVE 692268512v5

[0044] It should be understood that this invention is not limited to the particular methodology, protocols, and reagents, etc., described herein and as such may vary. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the present invention, which is defined solely by the claims.

[0045] As used herein, the articles “a,” “an,” and “the” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.

[0046] The use of the alternative (e.g., “or”) should be understood to mean either one, both, or any combination thereof of the alternatives.

[0047] As used herein, the term “about” or “approximately” refers to a quantity, level, value, number, frequency, percentage, dimension, size, amount, weight or length that varies by as much as 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or 1% compared to a reference quantity, level, value, number, frequency, percentage, dimension, size, amount, weight or length. In some instances, the term “about” or “approximately” refers a range of quantity, level, value, number, frequency, percentage, dimension, size, amount, weight or length ± 15%, ± 10%, ± 9%, ± 8%, ± 7%, ± 6%, ± 5%, ± 4%, ± 3%, ± 2%, or ± 1% of a reference quantity, level, value, number, frequency, percentage, dimension, size, amount, weight or length.

[0048] As used herein, the term “bioprocess” refers to any biological or biochemical procedure, typically involving the use of living organisms or their components, employed in the production, modification, or manipulation of biological materials. For example, a bioprocess can involve the controlled growth and differentiation of cells in vitro. This may include various stages, such as the initial culturing of cells, their expansion under specific conditions, induction of differentiation (as in the case of iPSC-derived cell types like dopaminergic neurons and cardiomyocytes), and any subsequent treatments or modifications.

[0049] As used herein, the term “cells” can refer to one or more types of cells noted in this or any other paragraph of the disclosure, or substitutes and / or equivalents thereof. For instance, the cells can include neural cells, myeloid cells, T cell (e.g., regulatory T cells, microglial cells, or cardiac cells. The cells can also include stem cells, such as embryonic stem cells or induced pluripotent stem cells. Exemplary cells include, but are not limited to, mesenchymal stem cells, hematopoietic stem cells, embryonic stem cells or induced Page 14 of 62 ACTIVE 692268512v5pluripotent stem cells, red blood cells, platelets, chondrocytes, skin cells, immune cells (e.g. tumor infiltrating lymphocytes, viral reconstitution T cells, dendritic cells, regulatory T cells, macrophages), neural crest stem cells, neurons, glia, smooth muscle, cardiac tissue, chondrocytes, osteocytes, glial restricted progenitors, astrocytes, oligodendrocytes, neuroblast cells, megakaryoblasts, megakaryocytes, monoblasts, monocytes, macrophages, myeloid cells, myeloid dendritic cells, microglial cells, differentiated microglial cells, microglial progenitor cells, proerythroblasts, erythroblasts, normoblasts, reticulocytes, thrombocytes, myeloblasts, progranulocytes, neutrophilic myelocytes, neutrophilic band cells, neutrophils, eosinophilic myelocytes, eosinophilic band cells, eosinophils, basophilic myelocytes, basophilic band cells, basophils, committed lymphoid progenitors, pre-NK cells, NK lymphoblasts, NK cells, thymocytes, T-lymphoblasts, T-cells, plasmacytoid dendritic cells, pre-B cells, B-lymphoblasts, B cells, plasma cells, osteoblasts, chondrocytes, myoblasts, myotubes, fibroblasts, adipocytes, mesoderm, ectoderms, cardiomyocytes, fibroblasts, endothelial cells, pericytes, smooth muscle cells, mesothelial cells, primordial germ cells, sperm, eggs, or any other suitable type of cell. In some instances, the cells are dopaminergic neurons. Exemplary dopaminergic neurons include midbrain dopaminergic neurons, authentic midbrain dopaminergic neurons, midbrain dopaminergic neuron progenitor cells, dopaminergic neuron progenitor cells, and dopaminergic neuron precursor cells.

[0050] As used herein, the term “correlation-based mapping technique” and any grammatical equivalents can refer to a computational method that can be used to analyze and compare gene expression profiles of single cells across different scRNA-seq datasets. This technique can involve assessing the degree of similarity (e.g., correlation) between the gene expression data of individual cells in one dataset (e.g., test dataset) with predefined reference profiles or prototypes in another dataset (e.g., mapping dataset). As described herein, this technique can include calculating the correlation coefficients between gene expression signatures of single cells and prototype signatures of identified cell types or clusters in the reference dataset. These coefficients can quantify how closely the gene expression of each test cell matches that of each reference cell type. Based on the correlation coefficients, cells in the test dataset can be assigned phenotypic labels. Cells with high correlation to a particular prototype can be labeled as belonging to the corresponding cell type. Cells in a dataset, e.g., the test dataset, that do not show a high correlation to any of the prototypes in the reference dataset can be classified as ‘unmapped.’ In some instances, this can indicate that these cells represent novel or significantly different cell types not present in the mapping Page 15 of 62 ACTIVE 692268512v5dataset. This technique can be useful in comparative analyses conducted by methods and systems described herein (e.g., allowing for the assessment of phenotypic similarities and differences between cell populations derived from different conditions, development stages, and / or treatments). Methods described herein can include setting statistical thresholds for determining the correlation levels necessary for a cell to be considered as belonging to a specific phenotype.

[0051] As used herein, the term “correlation coefficient” and its grammatical equivalents can refer to a measure of linear fit. A “Pearson product-moment correlation coefficient” or “Pearson’s R” can refer to a measure of the strength and direction of the linear relationship between two variables and is defined as the covariance of the variables divided by the product of their standard deviations.

[0052] As used herein, the term “clusters” or “clustering” and any grammatical equivalents can refer to groups of single cells organized based on a common characterization. In some instances, a cluster may refer to a group of cells organized based on the cells’ gene expression profiles, e.g., as analyzed from single-cell RNA sequencing (scRNA-seq) data. The formation of clusters can be driven by the objective of categorizing cells according to similarities, densities, intervals, or specific statistical measures within the scRNA-seq data. This can be achieved using algorithms (e.g., algorithms such as the Leiden clustering algorithm) for accurate identification and clustering based on complex gene expression networks. These clusters can be visualized in a two-dimensional space (e.g., using techniques such as Uniform Manifold Approximation and Projection (UMAP)), facilitating the interpretation of cellular relationships and types within each dataset. In some examples, each cluster can correspond to a unique cell phenotype (e.g., ascertained from shared gene expression profiles) and / or can be annotated using known cell type identities from a mapping dataset. In one instance, one or more prototype signatures can be created for a cluster. The prototype signatures may represent an average gene expression of the cluster’s cells, serving as a model for the gene expression characteristics of the cells’ cell type’. Thus, clusters can be important for organizing and interpreting scRNA-seq data, enabling comprehensive analyses, including the comparison of cell types across datasets and the identification of novel cell types.

[0053] As used herein, the term “dataset” and its grammatical equivalents can refer to a collection of data specifically organized for analysis, e.g., scRNA-seq data. In some Page 16 of 62 ACTIVE 692268512v5instances, a dataset can include gene expression information of a multitude of single cells (e.g., characterizing each single cell by its levels of expression across thousands of genes). For example, a dataset in scRNA-seq can include quantitative data representing the expression levels of various genes in individual cells. This data can be derived from sequencing the RNA of each cell in a sample, thereby capturing the cell’s unique gene expression profile. One purpose of a dataset in the context of this disclosure can be to serve as input for analytical processes. For example, a dataset can be used as a basis for computational tasks such as clustering, mapping, and comparative analysis. Within the context of this disclosure, datasets can include a mapping dataset and a testing dataset. A mapping dataset can refer to a dataset with known cell type identities, which can be used as a reference in a comparative analysis. For example, it can serve as the standard against which test datasets are compared. A test dataset can thus be analyzed and compared against a mapping dataset. The test dataset can be derived from a different sample or under different experimental conditions, including using a different bioprocess. The dataset can be obtained from various sources (e.g., experimental data and / or publicly available repositories) and / or can be generated from specific biological samples.

[0054] As used herein, the term “phenotype” and its grammatical equivalents, can refer to observable characteristics and / or traits of a single cell, or groups of cells, as determined by its gene expression profile. These characteristics can include physical attributes and / or functional aspects, such as cellular behavior and response to stimuli, which are directly influenced by the cell’s gene expression. In some instances, phenotype can indicate a cell’s type (e.g., neuron, immune cell, epithelial cell, etc.) and / or its functional state (e.g., active, dormant, differentiating, etc.). As described herein, the phenotype of a cell can be identified through the analysis of gene expression data. Different patterns of gene activation or repression can give rise to distinct phenotypes. In some instances, phenotypes can be involved in the clustering process. For example, cells can be grouped into clusters based on similarities in their phenotypes, as reflected in their gene expression profiles. Furthermore, phenotypes can be used to compare cell populations between different datasets. By analyzing how phenotypes are represented in different groups of cells (e.g., different samples, different conditions, etc.) the methods and systems described herein can assess similarities and / or differences between cell groups and / or identify potential novel cell types. Page 17 of 62 ACTIVE 692268512v5

[0055] As used herein, the term “progenitor cell” and its grammatical equivalents can refer to a descendant of a stem cell that can further differentiate into specialized cell types within a particular cell lineage.

[0056] As used herein, a statement that a cell or population of cells is “positive” for a particular marker (e.g., a molecule and / or structure), or “expresses” a particular marker, refers to the detectable presence (e.g., presence detected above a threshold level) of the particular marker on or in the cell or group of cells. By contrast, a statement that a cell or population of cells is “negative” for a particular marker, or fails to express a particular marker or gene, refers to an absence of substantial detectable presence (e.g., presence detected above a threshold level) on or in the cell of a particular marker. The disclosed systems and processes may test for presence of a variety of types of markers. Exemplary markers include, for example, surface markers or intracellular markers. A surface marker may refer to a marker that is present on the surface of a cell. An intracellular marker may refer to a marker located within a cell. A transcription factor (or group of transcription factors) is one example of an intracellular marker. In some examples, a cell or population of cells is determined to be “positive” or “negative” for a marker by flow cytometry.

[0057] As used herein, the term “flow cytometry” refers to a process in which a sample of a cell (or group of cells) is marked with a probe, such as an antibody (e.g., a fluorochrome- conjugated probe such as a fluorochrome-conjugated antibody), that specifically binds to a particular marker. In instances in which the particular marker is an intracellular marker, the cell, or group of cells, may have been permeabilized (e.g., with a methanol or detergent) to allow the probe to access the marker within the intracellular space. In some instances, multiple fluorescently distinct probes may be used (e.g., each binding to a different marker). After marking with the probe, the presence or absence of the marker may be determined based on a detection of the probe. In one example, the probe may be detected by injecting the sample into a flow cytometer instrument where the cells of the sample flow (e.g., one cell at a time) through a laser beam. As the cell intersects the laser beam, the laser may excite any markers bound to the probe, emitting fluorescence that may be measured and analyzed to detect and quantify the presence of markers. In some instances, the level of staining identified in the sample may be compared with a level of marking identified (under similar or identical conditions) in a control (e.g., an isotype-matched control or fluorescence minus one (FMO) gating control). The presence of the marker may be determined if the level of marking Page 18 of 62 ACTIVE 692268512v5identified in the sample is substantially above the level of marking in the control, at a level substantially similar to that for cell known to be positive for the marker, and / or at a level substantially higher than that for a cell known to be negative for the marker. The absence of the marker may be determined if the level of marking identified in the sample is not detected at a level substantially above the level of marking detected in the control and / or at a level substantially lower than that for cell known to be positive for the marker, and / or at a level substantially similar as compared to that for a cell known to be negative for the marker.

[0058] As used herein, the term “progenitor cell,” and its grammatical equivalents, refers to a descendant of a stem cell that can further differentiate into specialized cell types within a particular cell lineage.

[0059] As used herein, the term “stem cell” refers to a cell with the ability to divide (e.g., for indefinite periods) in culture and to give rise to specialized cells.

[0060] As used herein, the term “substantially” or “essentially” refers to a quantity, level, value, number, frequency, percentage, dimension, size, amount, weight or length that is about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or higher compared to a reference quantity, level, value, number, frequency, percentage, dimension, size, amount, weight or length. In some instances, the terms “essentially the same” or “substantially the same” refer a range of quantity, level, value, number, frequency, percentage, dimension, size, amount, weight or length that is about the same as a reference quantity, level, value, number, frequency, percentage, dimension, size, amount, weight or length.

[0061] Throughout this disclosure, various aspects of the invention can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range. Page 19 of 62 ACTIVE 692268512v5II. scCompare

[0062] Advances in next generation sequencing (NGS) allow for high resolution gene expression. Single cell RNA sequencing (scRNA-seq) has provided unprecedented power to profile transcriptome-wide at the single cell level. Coupled with advances in computational biology to analyze scRNA-seq data, (see Traag VA, Waltman L, van Eck NJ. From Louvain to Leiden: guaranteeing well-connected communities. Sci Rep. 2019;9(1):5233; and McInnes L, Healy J, Melville J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. Published online September 17, 2020, each of which are incorporated herein by reference) biological samples can be resolved to identify heterogeneity among the single cells in a pooled sample (Choi YH, Kim JK. Dissecting Cellular Heterogeneity Using Single-Cell RNA Sequencing. Mol Cells. 2019;42(3):189-199. doi:10.14348 / molcells.2019.2446, which is incorporated herein by reference). Analyzing an scRNA-seq data set or integrating several data sets can be challenging because the data tend to be extremely high dimensional and highly sparse. However, several methods have been recently developed to aid in the computation (Ryu Y, Han GH, Jung E, Hwang D. Integration of Single-Cell RNA-Seq Datasets: A Review of Computational Methods. Mol Cells. 2023;46(2):106-119. doi:10.14348 / molcells.2023.0009, which is incorporated herein by reference). For example, batch correction and harmonization of scRNA-seq data sets help to minimize systematic differences so that the data can be compared on the gene expression level (Tran HTN, Ang KS, Chevrier M, et al. A benchmark of batch-effect correction methods for single-cell RNA sequencing data. Genome Biol. 2020;21(1):12. doi:10.1186 / s13059-019- 1850-9, which is incorporated herein by reference). However, there is a lack of computational approaches that serve to compare scRNA-seq data sets at the phenotypic level (e.g., according to similarity and differences in phenotypic heterogeneity).

[0063] This disclosure addresses that need by providing a novel computational pipeline, such as those described herein for scCompare, which is useful for comparing scRNA-seq data sets (e.g., according to similarities and / or differences in phenotypic heterogeneity). In some embodiments, this computational pipeline may enable comparing the phenotypic characterization of single cells.

[0064] Accordingly, certain aspects of the disclosure relate to methods for using single- cell RNA sequencing (scRNAseq) data to determine whether a first set of cells, derived using a first bioprocess, is phenotypically equivalent to a second set of cells, derived using a second Page 20 of 62 ACTIVE 692268512v5bioprocess. In some embodiments, the method includes using at least one computer hardware processor to perform obtaining a plurality of clusters of the first set of cells, each of the plurality of clusters corresponding to a respective phenotype of a plurality of phenotypes; obtaining scRNAseq expression data for the second set of cells; determining, using the plurality of clusters and the scRNAseq expression data, a phenotype of the plurality of phenotypes for each of at least some of the cells in the second set of cells; determining, for each of at least some phenotypes of the plurality of phenotypes, a proportion of a number of the first set of cells of the particular phenotype to a total number of the first set of cells, thereby obtaining a first plurality of phenotype proportions; determining, for each of the at least some phenotypes of the plurality of phenotypes, a proportion of a number of the second set of cells determined to be of the particular phenotype to a total number of the second set of cells, thereby obtaining a second plurality of phenotype proportions; and determining whether the first set of cells and the second set of cells are phenotypically equivalent using the first plurality of phenotype proportions and the second plurality of phenotype proportions. In some embodiments, the scRNAseq expression data includes, for each of multiple cells in the second set of cells, a plurality of expression levels for a respective plurality of genes.

[0065] In some embodiments, the method further includes, in response to determining that the first set of cells and the second set of cells are not phenotypically equivalent, determining whether the second set of cells includes pluripotent stem cells or progenitor cells.

[0066] In some embodiments, the method further includes, in response to determining that the second set of cells includes pluripotent stem cells, adjusting an aspect of the second bioprocess. In some embodiments, adjusting the aspect of the second bioprocess includes adding one or more agents to the second bioprocess. In some embodiments, adjusting the aspect of the second bioprocess includes adjusting a concentration of one or more agents in the second bioprocess. In some embodiments, adjusting the aspect of the second bioprocess includes adjusting an initial concentration of pluripotent stem cells used in the second bioprocess.

[0067] In some embodiments, the method further includes, in response to determining that the first set of cells and the second set of cells are not phenotypically equivalent, identifying cells of the second set of cells that are of an unknown phenotype, and determining a gene expression signature for the identified cells. Page 21 of 62 ACTIVE 692268512v5

[0068] In some embodiments, the method further includes include using the gene expression signature to develop an assay for quantifying cells of the unknown phenotype in one or more other samples.

[0069] In some embodiments, the method further includes determining a phenotype for the identified cells using the gene expression signature. In some embodiments, determining the phenotype for the identified cells using the gene expression signature includes identifying one or more marker genes included in the gene expression signature, and determining the phenotype based on the identified marker genes.

[0070] In some embodiments, the method further includes determining whether to administer a cell therapy based on a result of determining whether the first set of cells and the second set of cells are phenotypically equivalent. In some embodiments, the cell therapy includes one or more products of cells derived using the second bioprocess.

[0071] In some embodiments, the method further includes determining whether to discard the second set of cells based on a result of determining whether the first set of cells and the second set of cells are phenotypically equivalent, and in response to determining to discard the second set of cells based on the result, discarding the second set of cells.

[0072] In some embodiments, determining whether the first set of cells and the second set of cells are phenotypically equivalent using the first plurality of phenotype proportions and the second plurality of phenotype proportions includes determining a coefficient of determination (R2) between the first plurality of phenotype proportions and the second plurality of phenotype proportions, and determining that the first set of cells and the second set of cells are phenotypically equivalent if the coefficient of determination (R2) satisfies at least one criterion. In some embodiments, a leave one out analysis is performed on multiplicity of exemplary mapping datasets to determine a distribution of R2 scores, and the testing dataset R2 score is then compared to the distribution to determine testing dataset inclusivity. In some embodiments, the coefficient of determination (R2) represents a degree to which the first plurality of phenotype proportions and the second plurality of phenotype proportions are phenotypically equivalent. In some embodiments, determining whether the first set of cells and the second set of cells are phenotypically equivalent using the first plurality of phenotype proportions and the second plurality of phenotype proportions includes determining a Spearman Rank correlation coefficient between the first plurality of phenotype Page 22 of 62 ACTIVE 692268512v5proportions and the second plurality of phenotype proportions, and determining that the first set of cells and the second set of cells are phenotypically equivalent if the Spearman Rank correlation coefficient satisfies at least one criterion. In some embodiments, the Spearman Rank correlation coefficient represents a degree to which the first plurality of phenotype proportions and the second plurality of phenotype proportions are phenotypically equivalent.

[0073] In some embodiments, the method further includes obtaining a first scRNAseq expression data for the first set of cells. In some embodiments, the first scRNAseq expression data includes, for each of multiple cells in the first set of cells, a plurality of expression levels for a respective first plurality of genes. In some embodiments, obtaining the plurality of clusters of the first set of cells includes clustering the first set of cells into the plurality of clusters using the first scRNAseq expression data. In some embodiments, clustering the first scRNAseq expression data includes clustering the first scRNAseq expression data using a hierarchical clustering algorithm. In some embodiments, clustering the first scRNAseq expression data using the hierarchical clustering algorithm includes clustering the scRNAseq expression data using Leiden clustering. In some embodiments, the clusters of the first set of cells include a projection into a Uniform Manifold Approximation and Projection (UMAP) 2-dimensional space.

[0074] In some embodiments, obtaining the first scRNAseq expression data for the first set of cells includes obtaining, for each cell in the first set of cells, expression levels for between 5,000 and 30,000 genes. In some embodiments, obtaining the first scRNAseq expression data for the first set of cells includes obtaining, for each cell in the first set of cells, expression levels for between 5,000 and 10,000 genes, between 5,000 and 15,000 genes between 10,000 and 15,000 genes, between 10,000 and 20,000 genes, between 15,000 and 20,000 genes, between 15,000 and 25,000 genes between 20,000 and 25,000 genes, between 20,000 and 30,000 genes or between 25,000 and 30,000 genes. In some embodiments, obtaining the first scRNAseq expression data for the first set of cells includes obtaining, for each cell in the first set of cells, expression levels for between 5,000 and 10,000 genes, between 10,000 and 15,000 genes, between 15,000 and 20,000 genes, between 20,000 and 25,000 genes, or between 25,000 and 30,000 genes. In some embodiments, obtaining the first scRNAseq expression data for the first set of cells includes obtaining, for each cell in the first set of cells, expression levels for between 5,000 and 10,000 genes. In some embodiments, obtaining the first scRNAseq expression data for the first set of cells includes obtaining, for Page 23 of 62 ACTIVE 692268512v5each cell in the first set of cells, expression levels for between 10,000 and 15,000 genes. In some embodiments, obtaining the first scRNAseq expression data for the first set of cells includes obtaining, for each cell in the first set of cells, expression levels for between 15,000 and 20,000 genes. In some embodiments, obtaining the first scRNAseq expression data for the first set of cells includes obtaining, for each cell in the first set of cells, expression levels for between 20,000 and 25,000 genes. In some embodiments, obtaining the first scRNAseq expression data for the first set of cells includes obtaining, for each cell in the first set of cells, expression levels for between 25,000 and 30,000 genes.

[0075] In some embodiments, the method further includes determining a gene expression signature for each particular cluster in the plurality of clusters using the first scRNAseq expression data. In some embodiments, determining a gene expression signature for each particular cluster in the plurality of clusters using the first scRNAseq expression data includes, for each particular gene in a subset of the first plurality of genes, determining an average expression level of the particular gene across cells included in the particular cluster. In some embodiments, determining the phenotype for each of at least some of the cells in the second set of cells using the plurality of clusters includes determining the phenotype using the gene expression signatures determined for the plurality of clusters.

[0076] In some embodiments, the method further includes, for a particular cluster of the plurality of clusters, generating a distribution of correlation coefficients. In some embodiments, generating a distribution of correlation coefficients includes determining, for each cell included in the particular cluster, a correlation coefficient between a gene expression signature of the cell and the gene expression signature determined for the particular cluster, and determining a statistical cut-off value for the particular cluster. In some embodiments, the statistical cut-off represents a mean absolute deviation from a median of the generated distribution of correlation coefficients.

[0077] In some embodiments, determining the phenotype of the plurality of phenotypes for each of at least some of the cells in the second set of cells includes, for a particular cell of the at least some cells in the second set of cells, determining a correlation coefficient between scRNAseq expression data obtained for the particular cell and the gene expression signature determined for each cluster, and determining the phenotype for the particular cell based on the correlation coefficients and the statistical cut-off value for each cluster. Page 24 of 62 ACTIVE 692268512v5

[0078] In some embodiments, the method further includes identifying the particular cell as an outlier if none of the determined correlation coefficients exceed the respective statistical cut-off value. In some embodiments, the first bioprocess is different from the second bioprocess For example, in some instances the first bioprocess comprises a differentiation protocol that is different than the differentiation protocol employed in the second bioprocess. In some instances, the first bioprocess and the second bioprocess comprise one or more different agents for differentiating cells. Accordingly, in some embodiments the method can be used to optimize a bioprocess used to make therapeutic cells

[0079] In some embodiments, the method is used for quality control purposes. In some embodiments, the quality control purposes include purity assessments, or verifying the consistency of induced pluripotent stem cell (iPSC)-derived cell populations. In some embodiments, the method is used to refine one or more cell manufacturing processes. In some embodiments refining one or more cell manufacturing processes includes comparing cell batches produced under varying conditions or protocols. In some embodiments, the method is used in the research and development of cell therapies.

[0080] In some embodiments, the method as described herein for using scRNAseq data to determine whether a first set of cells, derived using a first bioprocess, is phenotypically equivalent to a second set of cells, derived using a second bioprocess relates to a system, wherein that system includes at least one computer hardware processor, and at least one non- transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method for using scRNAseq data to determine whether a first set of cells, derived using a first bioprocess, is phenotypically equivalent to a second set of cells, derived using a second bioprocess as described herein.

[0081] In some embodiments, the system includes at least one computer hardware processor, and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method for using scRNAseq data to determine whether a first set of cells, derived using a first bioprocess, is phenotypically equivalent to a second set of cells, derived using a second bioprocess, the method including obtaining a plurality of clusters of the first set of cells, each of the plurality of clusters corresponding to a respective phenotype of a plurality of phenotypes; obtaining Page 25 of 62 ACTIVE 692268512v5scRNAseq expression data for the second set of cells; determining, using the plurality of clusters and the scRNAseq expression data, a phenotype of the plurality of phenotypes for each of at least some of the cells in the second set of cells; determining, for each of at least some phenotypes of the plurality of phenotypes, a proportion of a number of the first set of cells of the particular phenotype to a total number of the first set of cells, thereby obtaining a first plurality of phenotype proportions; determining, for each of the at least some phenotypes of the plurality of phenotypes, a proportion of a number of the second set of cells determined to be of the particular phenotype to a total number of the second set of cells, thereby obtaining a second plurality of phenotype proportions; and determining whether the first set of cells and the second set of cells are phenotypically equivalent using the first plurality of phenotype proportions and the second plurality of phenotype proportions. As one skilled in the art can appreciate, the features of the methods disclosed herein can be similarly applied to such a system. For example, in some embodiments, the first bioprocess in the system is different from the second bioprocess. In some embodiments, the scRNAseq expression data includes, for each of multiple cells in the second set of cells, a plurality of expression levels for a respective plurality of genes.

[0082] In some aspects, the method as described herein for using scRNAseq data to determine whether a first set of cells, derived using a first bioprocess, is phenotypically equivalent to a second set of cells, derived using a second bioprocess relates to at least one non-transitory computer-readable storage medium storing processor-executable instructions, wherein that at least one non-transitory computer-readable storage medium, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform a method for using scRNAseq data to determine whether a first set of cells, derived using a first bioprocess, is phenotypically equivalent to a second set of cells, derived using a second bioprocess as described herein.

[0083] In some embodiments, the at least one non-transitory computer-readable storage medium stores processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform a method for using scRNAseq data to determine whether a first set of cells, derived using a first bioprocess, is phenotypically equivalent to a second set of cells, derived using a second bioprocess, the method including obtaining a plurality of clusters of the first set of cells, each of the plurality of clusters corresponding to a respective phenotype of a plurality of Page 26 of 62 ACTIVE 692268512v5phenotypes, obtaining scRNAseq expression data for the second set of cells; determining, using the plurality of clusters and the scRNAseq expression data, a phenotype of the plurality of phenotypes for each of at least some of the cells in the second set of cells; determining, for each of least some phenotypes of the plurality of phenotypes, a proportion of a number of the first set of cells of the particular phenotype to a total number of the first set of cells, thereby obtaining a first plurality of phenotype proportions; determining, for each of the at least some phenotypes of the plurality of phenotypes, a proportion of a number of the second set of cells determined to be of the particular phenotype to a total number of the second set of cells, thereby obtaining a second plurality of phenotype proportions; and determining whether the first set of cells and the second set of cells are phenotypically equivalent using the first plurality of phenotype proportions and the second plurality of phenotype proportions. As one skilled in the art can appreciate, the features of the methods disclosed herein can be similarly applied to such a non-transitory computer-readable storage medium. For example, in some embodiments, the scRNAseq expression data includes, for each of multiple cells in the second set of cells, a plurality of expression levels for a respective plurality of genes.

[0084] FIG. 1 illustrates an exemplary scCompare workflow in accordance with the methods described herein.

[0085] In one exemplary scCompare workflow, a first data set (e.g., a mapping data set with known cell type identities) can be processed (e.g., using a Leiden clustering algorithm) to cluster single cells into multiple clusters (e.g., for projection into Uniform Manifold Approximation and Projection (UMAP) 2-dimensional space). Given the phenotypic annotations of the single cells in the clusters, cell-type-specific prototype signatures (e.g., cell-type-specific gene expression signatures) can be generated for each cluster by calculating, for each cluster, the average gene expression of all cells within the cluster. Then, a distribution (e.g., a correlation distribution of gene signatures) can be generated for the cells of each cluster by comparing each cell’s gene expression signature to the prototype signature of its respective cluster. Next, a gene expression signature of each single cell in a second data set (e.g., the test data set) can be evaluated against statistical thresholds (e.g., inclusion or exclusion thresholds derived from the distributions generated for the clusters of the first data set), to assign a phenotypic label to single cells within the second data set that fall within a cluster distribution. Single cells of the second data set that fall outside of the cluster Page 27 of 62 ACTIVE 692268512v5distributions of the first data set can be labeled as “unmapped.” This feature (enabling the identification of unmapped single cells) enables the discovery of unique cell types.

[0086] In some embodiments, a comparison of the proportion of assigned clusters between data sets (e.g., the mapping data set and the test data set) may be visualized (e.g., in a scatter plot). In one embodiment, a coefficient of determination and / or a Spearman Rank correlation coefficient may be calculated for the data sets, as statistics measuring the similarity between the phenotypic composition of the two data sets.

[0087] FIG. 1 also illustrates an scCompare workflow (e.g., corresponding to the workflow just described) in which known phenotypes of clusters in an scRNA-seq data set (e.g., the cluster map dataset in FIG. 1) are used to generate bulk signatures for each cluster within the data set (e.g., by correlating highly variable genes within each cluster to a mean expression prototype). Based on statistical cut-offs (e.g., lower thresholds) generated for each distribution of bulk signatures, identities from the cluster map data set are mapped onto clusters from a test scRNA-seq data set (e.g., by comparing the clusters from the test data set to the statistical cut-offs for each distribution of the clusters’ bulk signatures) to assign phenotypic characterization to the cells of the test data set or label cells as unmapped. The unmapped cells provide opportunity for further investigation of similarities and differences in representation of cell identities and / or biological discovery. Additionally, as shown in FIG. 1, the ratio of phenotypes between the two data sets may be compared.

[0088] As mentioned previously, this disclosure provides results of experiments comparing the performance of the scCompare pipeline, on benchmark scRNA-seq data sets from peripheral blood mononuclear cells (PBMCs), with the performance of Single-cell Variational Inference (scVI), a state-of-the-art computational tool for general scRNA-seq analyses including clustering single cells (Lopez R, Regier J, Cole MB, Jordan MI, Yosef N. Deep generative modeling for single-cell transcriptomics. Nat Methods. 2018;15(12):1053- 1058. Doi:10.1038 / s41592-018-0229-2, which is incorporated herein by reference). In addition, this disclosure provides results of using publicly available scRNA-seq data sets, from atlases and experimental protocols that differentiate human induced pluripotent stems cells (hiPSCs) into cardiomyocytes (CMs), to test the utility of scCompare to detect similarities and differences in single cell populations. Page 28 of 62 ACTIVE 692268512v5

[0089] In comparison to scVI, a deep learning probabilistic model for perusing unexplored biological diversity for scRNA-seq data, scCompare matched or outperformed the tool in terms of higher accuracy and specificity for most of the cell types (see Table 1 and FIGS. 2A-2C). More importantly, in contrast with scVI, scCompare has a flexible utility. For example, scCompare can compare atlases of scRNA-seq data (Table 2) and / or (e.g., in the case of cardiomyocyte differentiation protocols) can reveal an unmapped cluster of single cells (see FIGS. 3A-3D). scCompare is also an improvement from scVI in that it allows the identification of “unmapped” cells based on their dissimilarity to computed signatures (e.g., by setting statistical cutoffs). Demonstrated herein is the utility of this feature in the comparison of cardiomyocyte differentiation protocols, which revealed an unmapped group of cells which may have some biological relevance (see FIGS. 3A-3D and FIG. 4A-4B). A comparison of scCompare and scVI is provided in greater detail in connection with the Examples section of this application.

[0090] scRNA-seq data is large, complex and a fair amount of bioinformatics expertise, computational savviness, and biological intuition is required to mine the data effectively. scCompare provides a straightforward analysis pipeline for researchers and clinicians to use. In some examples, an assessment of key parameters used in the scCompare pipeline can give users guidance on how to optimize analysis results (e.g., see FIGS. 5A-5D).

[0091] Accordingly, in some aspects, this disclosure provides a method for analyzing single-cell RNA sequencing (scRNAseq) data to compare cell populations, the method comprising: using at least one computer hardware processor to perform: obtaining a first scRNAseq dataset comprising expression data from a first population of cells and phenotype annotations for the first population of cells; generating, for each phenotype annotation, a prototype gene expression signature by determining an average expression level of highly variable genes across cells having that phenotype annotation; determining, for each phenotype annotation, a statistical distribution of correlation coefficients between the prototype gene expression signature for that phenotype and individual gene expression profiles of cells having that phenotype annotation; determining a statistical cutoff value for each phenotype annotation based on its distribution of correlation coefficients; obtaining a second scRNAseq dataset comprising expression data from a second population of cells; for each cell in the second population: determining correlation coefficients between that cell's gene expression profile and each prototype gene expression signature, and assigning that cell Page 29 of 62 ACTIVE 692268512v5to a phenotype annotation only if its correlation coefficient with the corresponding prototype signature exceeds the statistical cutoff value for that phenotype annotation, otherwise marking the cell as unmapped; determining proportions of cells assigned to each phenotype annotation in both populations; and identifying potentially novel cell types by analyzing gene expression patterns of unmapped cells.

[0092] In some embodiments, the method includes selecting highly variable genes for generating prototype signatures. The selection of highly variable genes may include calculating, for each gene, a mean expression and variance across all cells in the first dataset. In some embodiments, the method normalizes variance by mean expression to account for the relationship between variance and mean in RNA sequencing data. The method may then fit a curve to this mean-variance relationship and select genes that deviate significantly from this curve as highly variable genes. In some embodiments, the number of highly variable genes selected may be between 1,000 and 5,000 genes. The method may additionally filter the selected genes based on their expression frequency, requiring each selected gene to be expressed in at least 1% of cells to reduce the impact of technical artifacts.

[0093] In some embodiments, the method generates statistical distributions of correlation coefficients to determine phenotype-specific thresholds. For each phenotype annotation, the method calculates Pearson correlation coefficients between each cell assigned to that phenotype and the prototype signature for that phenotype. The resulting distribution of correlation coefficients may be analyzed for normality using standard statistical tests. In some embodiments, where the distribution is approximately normal, the method uses parametric statistics such as mean and standard deviation to set thresholds. In some embodiments, where the distribution is non-normal, the method uses non-parametric statistics such as median absolute deviation or employ the Fisher Transformation to normalize the distribution. In some embodiments, the method stores these distributions and their associated statistics for use in subsequent analyses.

[0094] In some embodiments, the method performs a systematic analysis of unmapped cells to identify potentially novel cell types. In some embodiments, this analysis includes clustering unmapped cells based on their gene expression profiles using the same clustering algorithm applied to the original dataset. For each cluster of unmapped cells, the method may perform differential expression analysis to identify genes that are specifically expressed in that cluster compared to all other cells. In some embodiments, the method calculates both the Page 30 of 62 ACTIVE 692268512v5statistical significance and the magnitude of differential expression for each gene, requiring a minimum fold-change of 2 and an adjusted p-value less than 0.05. The method may then compare the resulting gene signatures against known cell type markers from published literature and databases. Additionally, the method may analyze the position of unmapped cell clusters in reduced dimensionality space (e.g., UMAP coordinates) relative to mapped cells to assess their relationship to known cell types. Clusters of unmapped cells that show both distinct gene expression signatures and spatial separation from known cell types may be classified as potentially novel cell types.

[0095] In some embodiments, the method employs various validation steps for identified novel cell types. In some embodiments, the method performs pathway enrichment analysis on the differentially expressed genes to identify biological processes associated with the potential novel cell type. In some embodiments, the method calculates various quality metrics, such as the number of detected genes per cell and the proportion of mitochondrial reads, to ensure that the novel cell type classification is not driven by technical artifacts. In some embodiments, the method may require a minimum number of cells (e.g., 50 cells) showing similar expression patterns before designating a group of unmapped cells as a potentially novel cell type. III. Exemplary embodiments of scCompare applications Safety and Efficacy Testing Embodiment

[0096] In some embodiments, the scCompare pipeline can be used to assess safety and efficacy of therapeutic cell populations. For example, scCompare can be used in safety and efficacy assessments to verify the safety and efficacy of iPSC-derived cell populations, such as dopaminergic neurons and cardiomyocytes (e.g., using the scCompare comparison analysis described herein to align their gene expression profiles with established benchmarks). As another example, the scCompare pipeline can be used to identify any contaminants or undesired cell types, ensuring the production of safe and effective therapeutic cells. Quality control and cell Line characterization Page 31 of 62 ACTIVE 692268512v5

[0097] In some embodiments, the scCompare pipeline can be used to improve quality control. For example, scCompare can be used in purity assessments to verify the consistency of iPSC-derived cell populations, such as dopaminergic neurons and cardiomyocytes, across different batches. In this example, the comparison analysis (enabled by the scCompare pipeline as described herein) can be used to verify the consistency of iPSC-derived cell populations by aligning their gene expression profiles with established benchmarks. As another example, the scCompare pipeline can be used to identify any contaminants or undesired cell types. In some embodiments, scCompare can be used to monitor cell lineages and analyze differentiation processes. Cell Manufacturing Process Optimization Embodiment

[0098] In some embodiments, scCompare can be used to refine cell manufacturing processes. In some embodiments, the comparison analysis (enabled by the scCompare pipeline as described herein) can be used to compare cell batches produced under varying conditions or protocols. In some embodiments, scCompare can be used to identify more effective methods for generating cells with desired attributes by the comparison of at least two different protocols or under at least two different conditions. In some embodiments, scCompare can be used to maintain the quality and characteristics of cells during scale-up processes, ensuring consistent cell quality irrespective of production volume. Research and Development Advancement Embodiment

[0099] In some embodiments, scCompare can be used in the research and development of cell therapies and / or to discover or identify novel cell types. In some embodiments, the novel cell types can be selected or expanded for use in a therapeutic cell product. IV. Computer Implementation

[0100] FIG. 9 is a block diagram of an example system 900 for performing comparative analysis of scRNA-seq data. System 900 includes computing device(s) 910 configured to have software 950 execute thereon to perform various functions in connection with performing comparative analysis of scRNA-seq data. As shown in FIG. 9, software 950 includes a plurality of modules. A module may include processor-executable instructions Page 32 of 62 ACTIVE 692268512v5that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform function(s) of the module. Such modules are sometimes referred to herein as “software modules.” A module may include processor- executable instructions that, when executed, cause performance one or more acts of one or more processes, such as one or more acts of the illustrative scCompare process shown in FIG. 1.

[0101] In some embodiments, computing device(s) 910 obtain scRNA-seq data from sequencing platform 930. The sequencing platform 930 may include a next generation sequencing platform (e.g., Illumina®, Roche®, Ion Torrent®, etc.), any high-throughput or massively parallel sequencing platform, and / or a platform configured to perform sequencing techniques other than next generation sequencing (e.g., microarrays, etc.). For example, the sequencing platform 930 may include the Chromium® single cell gene expression platform from 10X Genomics (e.g., version 2 or 3). The computing device(s) 910 may directly or indirectly obtain the scRNA-seq data from the sequencing platform 930 via at least one communication network, such as the Internet or any other suitable communication network(s), as aspects of the technology described herein are not limited in this respect.

[0102] Additionally or alternatively, the computing device(s) 910 may obtain the scRNA- seq data from user(s) 920. For example, the user(s) 920 may provide the scRNA-seq data as input to the computing device(s) 190. Additionally or alternatively, user(s) 920 may provide input specifying processing or other methods to be performed on scRNA-seq data. User(s) 920 may provide input by uploading one or more files and / or interacting with a user interface (e.g., user interface module 970) of the computing device(s) 910.

[0103] Software 950 may include one or more modules configured to analyze the scRNA-seq data. For example, the software 950 may include a cluster generation module 955, a gene expression signature generation module 960, a phenotype cut-off determination module 965, a phenotype determination module 980, and / or a novel cell type identification module 975.

[0104] The cluster generation module 955 may be configured to cluster cells based on the cells’ expression profiles to obtain a plurality of clusters of cells. For example, the cluster generation module 955 may be used to cluster scRNA-seq data obtained for cells with known phenotypes. In some embodiments, the clustering may be performed using hierarchical clustering (e.g., Leiden clustering). Leiden clustering is described by Traag, V., et al. ("From Louvain to Leiden: guaranteeing well-connected communities." Scientific reports 9.1 (2019): 1-12.), which is incorporated by reference herein in its entirety. Page 33 of 62 ACTIVE 692268512v5

[0105] The gene expression signature generation module 960 may be configured to determine a gene expression signature for each of a plurality of clusters of cells. For example, the gene expression signature generation module 960 may obtain the clusters from the cluster generation module 955. In some embodiments, a gene expression signature is generated using scRNA-seq data (e.g., the scRNA-seq data used to generate the clusters). Generating a gene expression signature for a particular cluster may involve, for each of multiple genes included in a subset of genes, determining an average expression level of the particular gene across the cells included in the particular cluster. In some embodiments, the gene expression signature for the particular cluster includes the average expression levels determined for the multiple genes.

[0106] The phenotype cut-off determination module 965 may be configured to determine, for a particular cluster (e.g., the clusters generated by the cluster generation module 955), a statistical cut-off value for the cluster. As described herein, a statistical cut-off value for a cluster may be used to determine whether or not to identify a particular cell as having the phenotype associated with the cluster. In some embodiments, statistical cut-off values may be determined using (i) the scRNA-seq data for the cells included in the clusters, and (ii) gene expression signatures generated for the clusters (e.g., obtained from the gene expression signature generation module 960). For example, this may involve determining, for each of multiple cells in a particular cluster, a correlation coefficient between the scRNA-seq data for the particular cell and the gene expression signature generated for the cluster. The resulting distribution of correlation coefficients may be used to determine the statistical cut-off value for the particular cluster. For example, the statistical cut-off value may be the mean absolute deviation from a median of the distribution.

[0107] The phenotype determination module 980 may be configured to determine a phenotype for a cell based on the cell’s scRNA-seq data and data associated with the clusters (e.g., generated by cluster generation module 955). For example, the cell’s phenotype may be determined based on gene expression signatures (e.g., generated by gene expression signature generation module 960) and / or statistical cut-off values (e.g., determined by the phenotype cut-off determination module 965) associated with the clusters. In some embodiments, this involves (i) determining a respective correlation coefficient between the cell’s scRNA-seq data and the gene expression signature generated for each cluster, and (ii) identifying the phenotype for the cell based on the resulting correlation coefficients and / or the statistical cut- off values. For example, the phenotype determination module 980 may be configured to identify, for the cell, the phenotype of the cluster associated with the highest correlation Page 34 of 62 ACTIVE 692268512v5coefficient. Additionally or alternatively, the phenotype determination module 980 may be configured to identify, for the cell, the phenotype of the cluster associated with a correlation coefficient exceeding the cluster’s statistical cut-off value. In some embodiments, if none of the correlation coefficients exceed a respective statistical cut-off value, then the cell may be identified as an outlier or unmapped cell.

[0108] The phenotype determination module 980 may additionally or alternatively be configured to determine whether different sets of cells are phenotypically equivalent. For example, the phenotype determination module 980 may be configured to determine (i) for each phenotype of cells included in a first set of cells (e.g., the cells for which clustering was performed by cluster generation module 955), a respective proportion of the cells of the particular phenotype to the total number of cells in the first set of cells, and (ii) for each phenotype of cells included in second set of cells (e.g., cells for which phenotypes were determined by phenotype determination module 980), a respective proportion of the cells of the particular phenotype to the total number of cells in the second set of cells. The phenotype determination module 980 may use the phenotype proportions for the first set of cells and the phenotype proportions for the second set of cells to determine whether the two sets of cells are phenotypically equivalent. For example, this may involve determining a coefficient of determination (R2) between the phenotype proportions, and using the coefficient of determination to determine whether the two sets of cells are phenotypically equivalent.

[0109] The novel cell type identification module 975 may be configured to identify cell types for cells identified as being outliers (e.g., by phenotype determination module 980). In some embodiments, identifying the cell types includes clustering the gene expression profiles of cell types identified as being outliers. For example, the clustering may be the same clustering performed by cluster generation module 955. For a particular cluster, the novel cell type identification module 975 may perform differential analysis to identify genes that are specifically expressed in that cluster compared to all cells. The resulting gene signatures may be compared against known cell type markers from published literature. Additionally or alternatively, the novel cell type identification module 975 may analyze the position of the clusters of outlier cells in reduced dimensionality space relative to clusters for which phenotypes are known (e.g., the clusters generated by cluster generation module 955). In some embodiments, clusters of outlier cells that show distinct gene expression signatures and / or spatial separation from known cell types may be classified as potentially novel cell types. Page 35 of 62 ACTIVE 692268512v5

[0110] As shown in FIG. 9, software 950 also includes a user interface module 970. User interface module 970 may be configured to generate a graphical user interface (GUI) through which user(s) 920 may provide input and / or view information generated by software 950. For example, in some embodiments, the user interface module 970 may be a webpage or web application accessed through an Internet browser. In some embodiments, the user interface module 970 may generate a GUI of an app executing on a user’s mobile device. In some embodiments, the user interface module 970 may generate a number of selectable elements through which a user may interact. For example, the user interface module 970 may generate dropdown lists, checkboxes, text fields, or any other suitable element, as aspects of the technology described herein are not limited in this respect.

[0111] In some embodiments, scRNA-seq data store(s) 940 stores scRNA-seq data, the output of one or more software modules, or any other suitable information, as aspects of the technology described herein are not limited in this respect. For example, the scRNA-seq data store(s) 940 may store scRNA-seq data, cluster(s) generated by cluster generation module 955, gene expression signature(s) generated by gene expression signature generation module 960, phenotype cut-off value(s) determined by phenotype cut-off determination module 965, phenotype(s) of cell(s) determined by phenotype determination module 980, novel cell type(s) identified by novel cell type identification module 975, and / or user input from the user interface module 970. In some embodiments, the scRNA-seq data store(s) 940 include any suitable type of data store (e.g., a flat file, a database system, a multi-file, etc.) and may store data in any suitable format, as aspects of the technology described herein are not limited in this respect. The scRNA-seq data store(s) may be part of software 950 (not shown) or excluded from software 950, as shown in FIG. 9.

[0112] An illustrative implementation of a computer system 1000 that may be used in connection with any of the embodiments of the technology described herein (e.g., such as the scCompare workflow shown in FIG. 1) is shown in FIG. 10. The computer system 1000 includes one or more processors 1010 and one or more articles of manufacture that comprise non-transitory computer-readable storage media (e.g., memory 1020 and one or more non- volatile storage media 1030). The processor 1010 may control writing data to and reading data from the memory 1020 and the non-volatile storage media 1030 in any suitable manner, as the aspects of the technology described herein are not limited to any particular techniques for writing or reading data. To perform any of the functionality described herein, the processor 1010 may execute one or more processor-executable instructions stored in one or more non-transitory computer-readable storage media (e.g., the memory 1020), which may Page 36 of 62 ACTIVE 692268512v5serve as non-transitory computer-readable storage media storing processor-executable instructions for execution by the processor 1010.

[0113] Computing system 1000 may include a network input / output (I / O) interface 1040 via which the computing device may communicate with other computing devices. Such computing devices may be interconnected by one or more networks in any suitable form, including a local area network or a wide area network, such as an enterprise network, and intelligent network (IN) or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks, wired networks or fiber optic networks.

[0114] Computing system 1000 may also include one or more user I / O interfaces 1050, via which the computing device may provide output to and receive input from a user. The user I / O interfaces may include devices such as a keyboard, a mouse, a microphone, a display device (e.g., a monitor or touch screen), speakers, a camera, and / or various other types of I / O devices.

[0115] Further, it should be appreciated that a computer may be embodied in any of a number of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer, as examples. Additionally, a computer may be embedded in a device not generally regarded as a computer but with suitable processing capabilities, including a Personal Digital Assistant (PDA), a smartphone, a tablet, or any other suitable portable or fixed electronic device.

[0116] The above-described embodiments can be implemented in any of numerous ways. For example, the embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor (e.g., a microprocessor) or collection of processors, whether provided in a single computing device or distributed among multiple computing devices. It should be appreciated that any component or collection of components that perform the functions described above can be generically considered as one or more controllers that control the above-described functions. The one or more controllers can be implemented in numerous ways, such as with dedicated hardware, or with general purpose hardware (e.g., one or more processors) that is programmed using microcode or software to perform the functions recited above.

[0117] In this respect, it should be appreciated that one implementation of the embodiments described herein comprises at least one computer-readable storage medium (e.g., RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital Page 37 of 62 ACTIVE 692268512v5versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or other tangible, non-transitory computer-readable storage medium) encoded with a computer program (i.e., a plurality of executable instructions) that, when executed on one or more processors, performs the above- described functions of one or more embodiments. The computer-readable medium may be transportable such that the program stored thereon can be loaded onto any computing device to implement aspects of the techniques described herein. In addition, it should be appreciated that the reference to a computer program which, when executed, performs any of the above- described functions, is not limited to an application program running on a host computer. Rather, the terms computer program and software are used herein in a generic sense to reference any type of computer code (e.g., application software, firmware, microcode, or any other form of computer instruction) that can be employed to program one or more processors to implement aspects of the techniques described herein.

[0118] The terms “program” or “software” are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects as described above. Additionally, it should be appreciated that according to one aspect, one or more computer programs that when executed perform methods of the present disclosure need not reside on a single computer or processor but may be distributed in a modular fashion among a number of different computers or processors to implement various aspects of the present disclosure.

[0119] Computer-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.

[0120] Also, data structures may be stored in computer-readable media in any suitable form. For simplicity of illustration, data structures may be shown to have fields that are related through location in the data structure. Such relationships may likewise be achieved by assigning storage for the fields with locations in a computer-readable medium that convey relationship between the fields. However, any suitable mechanism may be used to establish a relationship between information in fields of a data structure, including through the use of pointers, tags or other mechanisms that establish relationship between data elements. Page 38 of 62 ACTIVE 692268512v5

[0121] When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers.

[0122] The foregoing description of implementations provides illustration and description but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications and variations are possible in light of the above teachings or may be acquired from practice of the implementations. In other implementations the methods depicted in these figures may include fewer operations, different operations, differently ordered operations, and / or additional operations. Further, non-dependent blocks may be performed in parallel.

[0123] It will be apparent that example aspects, as described above, may be implemented in many different forms of software, firmware, and hardware in the implementations illustrated in the figures. EXAMPLES

[0124] The Examples below describe work that was performed to develop and evaluate scCompare. Example 1: Comparison of scCompare to Single-cell Variational Inference (scVI) to map phenotypic labels

[0125] In this example, we detail an experimental comparison conducted between scCompare and an established computational method, Single-cell Variational Inference (scVI), for the analysis of single cell RNA sequencing (scRNA-seq) data. The objective was to evaluate and demonstrate the efficacy of scCompare in relation to scVI.

[0126] As described elsewhere, scVI is an scRNA-seq analysis tool that uses an optimization routine and deep neural networks to join information across cells and genes that are similar, and estimate the distributions that represent the observed expression values (see Lopez R, Regier J, Cole MB, Jordan MI, Yosef N. Deep generative modeling for single-cell transcriptomics, Nat Methods, 2018;15(12):1053-1058. Doi:10.1038 / s41592-018-0229-2, which is incorporated herein by reference). scVI uses a neural network-based approach to construct distributions from the matrix of read counts and batch information to account for experimental and technical variation in the data, and then estimate biological differences Page 39 of 62 ACTIVE 692268512v5between cells. A couple of disadvantages of scVI is that it uses a non-deterministic algorithm leading to alternative results with different initializations, and for genes with few cells, the prior and the inductive bias of the neural network may not fit the data optimally. The expected read count data (batch-corrected and normalized) is used in scVI for comparison to scCompare.

[0127] To assess the ability of scCompare and scVI to align biological annotation between scRNA-seq data sets, the 3K PBMCs scRNA-seq data was randomly sampled 50 times in training and testing splits (80:20) and the top 2,000 highly variable genes were used. Performance was based on the accuracy, precision, and sensitivity of scCompare and scVI mapping the phenotypic labels to the test data sets. As shown in Table 1, the accuracy of predictions was similar, but scCompare matches or outperforms scVI in terms of precision and sensitivity for most cell types, most strikingly with the dendritic cells and megakaryocytes. (However, scVI scored slightly higher for CD4 T cells, FCGR3A monocytes and natural killer (NK) cells in precision for the former and sensitivity for the latter two). Table 1. Performance comparison between scCompare and scVI using the 3k PBMC scRNA-seq data set* Accuracy Precision Recall Cell Types scCompare scVI scCompare scVI scCompare scVI B 0.998 0.995 0.990 0.969 0.996 0.996 CD14 Monocytes 0.982 0.976 0.948 0.943 0.943 0.905 CD4 T 0.975 0.955 0.981 0.986 0.962 0.911 CD8 T 0.967 0.949 0.829 0.721 0.877 0.861 Dendritic 0.998 0.992 0.957 0.628 0.925 0.665 FCGR3A Monocytes 0.991 0.984 0.915 0.830 0.938 0.948 Megakaryocytes 1.000 0.995 1.000 0.160 0.993 0.133 NK 0.984 0.979 0.897 0.822 0.903 0.947 *Performance parameters based on the average of 50 iterations of 80%:20% splits of the data training to test and using 2,000 highly variable genes. Page 40 of 62 ACTIVE 692268512v5

[0128] The ability of scCompare and scVI to map known transcriptomic signatures from the 3K PBMCs scRNA-seq data set to the 68K PBMCs scRNA-seq data set was also evaluated. As shown in FIG. 2B, the cell type identities from the 3K data set were mapped equivalently to the 68K data set except for the Megakaryocytes in the case of scVI where they were labeled as NK cells. The gene representations of the clusters of single cells for the phenotypic labels were comparably similar except for the megakaryocytes (FIG. 2B). However, due to the misclassification of the megakaryocytes to NK cells by scVI, the gene expression for the NK cluster of cells was more elevated for the Granulysin (GNLY) and Natural Killer Cell Granule Protein 7 (NKG7) genes than the mapping with scCompare. Furthermore, scCompare identified a previously unannotated cell type, plasmacytoid dendritic cell. (Kapoor T, Corrado M, Pearce EL, Pearce EJ, Grosschedl R. MZB1 enables efficient interferon α secretion in stimulated plasmacytoid dendritic cells. Sci Rep. 2020;10(1):21626. doi:10.1038 / s41598-020-78293-3, which is incorporated herein by reference). The application of a Median Absolute Deviation (MAD)-based statistical cutoff allowed for discovery of novel cell types, though it was also considered that not all unmapped cells would be novel (not accounted for in the mapping dataset). In order to determine if a group of unmapped cells was novel, two parameters were assessed: 1) a differentially expressed gene-based metric and 2) a UMAP Euclidean distance-based metric (FIGS. 2A-C). The cell cluster identified as plasmacytoid dendritic cells was observed to have a gene expression signature divergent from that of dendritic cells (the pre-cutoff assigned phenotype) and was separate from the mapped dendritic cells in UMAP space. Furthermore, differential gene expression analysis revealed the expression of MZB1 in this cluster, a marker gene for plasmacytoid dendritic cells. This not only demonstrated scCompare’s usefulness in mapping phenotypes between data sets, but it also provided the added ability to discover novel cell types not accounted for in the mapping data. Example 2: Mapping between scRNA-seq cell atlases

[0129] To demonstrate a broader applicability of the scCompare pipeline, cell atlas mapping was performed using the Human Protein Atlas (HPA; proteinatlas.org) as the training data set and the Tabula Sapiens (TS) as the test data set. These atlases contain samples obtained from primary human adult tissue from multiple individuals across many tissues. These samples have been extensively annotated with various levels of descriptive labels identifying the breadth and depth of heterogeneity present within each data set. At the Page 41 of 62 ACTIVE 692268512v5broadest level of categorization available in the HPA, bulk signatures were obtained corresponding to each of the 12 annotation groups using the top 2,000 highly variable genes. The bulk signatures and associated statistical cutoffs were then applied to the TS data set. Because the two data sets did not have perfectly overlapping levels of annotation, a categorization featuring five classifications was used as the ground truth label in the TS data set. Table 2 displays the results, with assigned labels and ground truth labels along the vertical and horizontal axes, respectively. As can be surmised, the vast majority of cells fall into the correct corresponding cluster grouping, with 92.5% of immune cells, 95% of stromal cells, 85% of epithelial cells and 78% of endothelial cells falling into corresponding categories. Table 2. scCompare comparability of cell types in the Human Protein Atlas to those in Tabular Sapiens Actual Class (from TS) Predicted Class (Using HPA) Endothelial Epithelial Germ line Immune Stromal Adipocytes 0.194 0 0 0 0.005 Blood & immune cells 0.01 0.004 0 0.925 0.001Neuronal cells 0 0 0 0 0 Pigment cells 0 0 0 0.001 0.003 Undifferentiated cells 0 0.093 0 0.001 0 Unmapped 0.007 0.025 0 0.059 0.014 Footnote: HPA cell type labels are from the canon_label_assignment metadata. Bolded numbers represent prediction accuracy for true positives of the actual classes (Column-wise). Example 3: Comparison of cardiomyocyte differentiation protocols

[0130] To demonstrate the utility of scCompare to compare scRNA-seq data sets from a study design, data was leveraged from an experiment that evaluated two different protocols to differentiate human induced pluripotent stem cells (hiPSCs) to cardiomyocytes (CMs) (see Grancharova T, Gerbin KA, Rosenberg AB, et al. A comprehensive analysis of gene Page 42 of 62 ACTIVE 692268512v5expression changes in a high replicate and open-source dataset of differentiating hiPSC- derived cardiomyocytes. Sci Rep. 2021;11(1):15845. doi:10.1038 / s41598-021-94732-1, which is incorporated herein by reference). The two protocols included Differentiation Protocol 1 and Differentiation Protocol 2. Differentiation Protocol 1 used small molecules whereas Protocol 2 used a cytokine in addition to the small molecules. The differentiation of the hiPSCs was carried out for 90 days. As represented in FIG. 3A, Protocol 1 samples were taken for scRNA-seq at day 0 (D0), day 12 (D12) and day 24 (D24), whereas Protocol 2 samples were taken for scRNA-seq at D0, day 14 (D14) and day 26 (D26). The difference in the days of collection is due to the lag in the differentiation initiation day between the two protocols. The scRNA-seq data was selected from Protocol 1 as the training (mapping) data set and the scRNA-seq data from Protocol 2 as the testing data set. scRNA-seq analysis of the training data set generated 13 Leiden clusters (FIG. 3B). The UMAP representation of the clusters showed well-formed clusters and an abundance of heterogeneity in the data which was expected in a “developmental” (differentiation) setting compared to “adult” / mature tissues. Differential expression analysis of the Leiden clusters 0 - 9 revealed marker genes TNNT2, MYH6 and MYH7 representative of CMs. Cluster 11 includes the CM marker genes in addition to MKI67 and FN1 indicative of progenitor CMs (PCMs). Clusters 12, 7, and 5 all include marker genes FN1, but also GRHL2 and AFP for the latter two clusters respectively, and they represent stromal-like cells, endodermal cells and ectodermal cells individually. The two remaining clusters 10 and 13 have marker genes TRPM3 and CTNNA2, and EGFL7 respectively, and correspondingly they suggest smooth muscle-like and endothelial-like cell types. FIG. 3C depicts a correlation heat map of the highly variable genes expressed by the single cells in the Leiden clusters and revealed a high degree of similarity between the CM cell types (clusters 0 – 9, and progenitor CM cluster 11). The other clusters have moderate to low correlation to each other. The cell type annotations in the Protocol 1 (mapping) data showed good uniformity of the CM clusters and relatively good individual clustering of the other cell types. However, after scCompare mapping of the phenotypic identities to the Protocol 2 (test) data, there appeared to be heterogeneity in the CM cluster, very few ectodermal, endothelial-like and smooth muscle-like cells, and less representation of endodermal cells (FIG. 3D). The aforementioned reduction in cell type mapping in Protocol 2 is represented in FIG. 3E numerically as proportions and graphically as a scatter plot of the cluster fraction of cells with the annotated cell type in Protocol 1. The coefficient of determination and the Spearman Rank Correlation of the cell type fraction between Protocol 1 and Protocol 2 was moderate at 0.441 and 0.393, respectively; markedly Page 43 of 62 ACTIVE 692268512v5low due to the underrepresentation of END, SM, ECT and EC cell type fractions. Single cells in Protocol 2 that were not assigned a cell type were labeled as unmapped. From the higher proportion of CM cells in Protocol 2 (~19%) and less off-target cell types. Differential expression analysis for D12 / 14 timepoint shows a set of marker genes CNTN5, SPHKAP, MECOM and TENM3, previously associated with iPSC-derived CM identity, that is differentially expressed in Protocol 1 vs Protocol 2 samples (FIG. 4A). Interestingly, the same analysis for later timepoint (D24 / 26) suggested a completely different set of marker genes implicated in cardiac morphogenesis (ROBO2), maturation (KCNH7), hypoxic response (NEAT1) (FIG. 4B). Example 4: scCompare number of genes and cluster resolution parameters assessment

[0131] This example describes work that was conducted to evaluate the sensitivity of the scCompare pipeline to two critical user-defined parameters: the selection of highly variable genes and clustering resolution. For this work, a subset of 10,000 cells was obtained from the HPA data set and separated into 5,000 mapping and test cells. The scCompare pipeline was repeatedly performed on this data set using an increasing size selection of highly variable genes and clustering resolutions (FIGS. 5A-5D). Highly variable gene selection was bounded between 50 and 10,000. Clustering resolution was bounded between 0.05 and 2.5, which corresponded to 2-40 distinct clusters. FIG. 5A shows shades for each individual iteration by weighted-F1, a statistical measure of prediction accuracy bounded by 0 and 1. With respect to clustering resolution, the results demonstrate the impact of over clustering on performance, as greater resolutions tend to produce ambiguity between classes resulting in misclassification. When assessing the impact of gene signature length, the results suggest that beyond a certain minimum threshold, most signature length selections tend to perform well. FIGS. 5B and 5C illustrate the same data set looking at each individual run by map and test misclassification rates, with shades provided by input variable. These plots best capture the trend towards higher input values of clustering resolutions and signature lengths tending to produce higher test misclassification than map misclassification, a situation that could be considered as overfitting. FIG. 5D provides another view of overfitting, where points are shaded by the absolute difference in map and test misclassification rates. Higher values tend to occur at the highest clustering resolutions and gene signature lengths. As a note, FIG. 5D masks cases in which the misclassification rate is equally poor in test and mapping data sets, as is the case in the low gene signature length (FIG. 5A). In general, our Page 44 of 62 ACTIVE 692268512v5results suggest that aside from excessive clustering resolutions and gene signature lengths, our pipeline produces robust results. Example 5: Identifying novel cell types

[0132] This example describes work that was conducted to explore scCompare’s capacity to identify novel cell types by using the unmapped functionality. In this example, unmapped cells were those whose correlation coefficients were outside the prescribed range. It is likely that some cells, whose ground truth phenotype was correctly assigned by the phenotype assignment step, would fall outside the range and be labeled unmapped. This could be due to differences in cell sparsity, a result of dropout and a well-known technical limitation of scRNAseq. Therefore, further analysis would be necessary to parse unmapped cells that fell out of the range for technical reasons from those that fell out of the range due their having a novel phenotype. The main criteria for novel cell type status assessment was the identification of a divergent gene expression pattern in a set of unmapped cells that did not conform to both that of the cell type the cell had been assigned to and any of the other annotated phenotypes. This could have been strengthened if the unmapped cells in question were also tightly clustered and fell outside of the UMAP coordinates of the cells that had been confidently assigned to the same phenotype. In order to facilitate the assessment of these criteria, unmapped cells were parsed into categories where each category was simply the concatenation of ‘unmapped’ and the assigned phenotype prior to the statistical cutoff.

[0133] In order to address this need, two metrics were devised to facilitate the unmapped assignment. The first metric used differential gene expression and filtering. For each category of unmapped cells, their filtered, differential gene expression list was compared to the filtered, differential gene expression list of the cells confidently assigned to the same phenotype. A list subset similarity metric (subset_similarity= (|list1∩list2|) / (|list2|), where |list2| and |list1∩list2| represent the sizes of ‘list2’ and the intersection of ‘list1’ and ‘list2’ respectively) was designed to compare these gene lists. If the unmapped cells of a specific phenotype had expression signature that is a perfect subset of the confidently mapped cells, then the subset score was identically 1. This score fell to 0 if the list of genes in the unmapped cells did not intersect at all with the list of genes produced by the confidently mapped corresponding phenotype. If the unmapped cells of a given phenotype had a subset similarity equal to, or near 1, then it was likely that these cells had fallen out of the statistical Page 45 of 62 ACTIVE 692268512v5range for technical reasons and their ground truth phenotype was in fact the one assigned prior to the statistical cutoff. A low subset similarity score indicated that the unmapped cells had a divergent gene expression pattern and should be considered for novel cell type status. The secondary metric used the Euclidean distance between centroids in UMAP space. This metric was calculated between the centroids of the unmapped cells of each phenotypic category and the cells that were confidentially mapped to the corresponding phenotype, normalized to the largest distance between any two centroids in UMAP space. As UMAP coordinates provide a rough estimate of cell-cell similarity, cells whose unmapped status were due to a technical reason should not diverge (in UMAP space) from the cells confidently mapped. Unmapped cells that represented a novel cell type diverged from their assigned phenotype in UMAP space. Therefore, these two metrics were used to assess unmapped cell status and indicated cells that had a high likelihood of being a novel phenotype (FIGS. 6A- 6B and FIGS. 7A-7C).

[0134] Using these metrics, both ‘unmapped_B’ and ‘unmapped_Dendritic’ cells were identified as having divergent gene expression patterns and increased UMAP centroid distances and warrant assessment for novel cell type status. Further gene expression analysis of these cells indicates that they are likely plasmacytoid dendritic cells18 (FIG. 6A-6B). Example 6: Fine-Resolution Phenotypic Mapping Between Cell Atlases

[0135] This example describes experimental work that was conducted to evaluate scCompare's ability to perform high-resolution phenotypic mapping between complex cell atlases containing diverse cell types. While previous examples described herein demonstrated mapping of broad cell type categories, this example explores the exemplary pipelines’ capacity to maintain phenotypic fidelity at multiple levels of cellular identity.

[0136] The Human Protein Atlas (HPA) and Tabula Sapiens (TS) datasets were selected for this analysis due to their comprehensive annotation of human cell types across multiple tissues. Prior to mapping, a phenotype label alignment strategy was developed to create corresponding categories between the two atlases. This alignment considered three hierarchical levels: broad cell type, sub-type, and organ-specific annotations.

[0137] To assess mapping fidelity without bias from statistical thresholding, scCompare was initially run without applying statistical cutoffs, allowing for direct evaluation of the 1:1 mapping accuracy. The pipeline achieved 88% unweighted accuracy across all phenotypic Page 46 of 62 ACTIVE 692268512v5categories, demonstrating robust performance even with fine-grained cell type classifications (Table 3).

[0138] Differential gene expression analysis was performed between matched cell types across the two atlases to validate the biological relevance of the mapping results. This analysis confirmed the preservation of cell-type-specific marker gene expression patterns between the datasets. The results were visualized using circos plots (FIGS. 8A-8D) to represent the relationships between phenotype labels at different hierarchical levels.

[0139] The mapping results revealed several insights about the performance of scCompare at high phenotypic resolution:

[0140] Hierarchical Preservation: The pipeline maintained the hierarchical relationships of cell types, successfully mapping cells not only to broad categories but also to appropriate sub-categories and organ-specific annotations.

[0141] Tissue-Specific Variation: Analysis of organ-level mapping demonstrated how the same cell types can exhibit tissue-specific variations while maintaining core phenotypic characteristics.

[0142] Marker Gene Consistency: Differential expression analysis confirmed that cells mapped to the same phenotypic categories shared consistent marker gene expression patterns across atlases, validating the biological accuracy of the mapping.

[0143] These results demonstrate that scCompare can effectively handle complex phenotypic hierarchies and maintain mapping accuracy even when dealing with datasets containing numerous distinct cell types Table 3. Exemplary results of scCompare performance Classification # Total Misclassified Misclassified Count Accuracy Phenotypes (>5%) (%) Hepatocytes 0 300 100 N / A Basal keratinocytes 0 300 100 N / A Exocrine glandular 3cells Muller glia cells 3 300 99 N / A Alveolar cells type 2 4 300 98.67 N / A Page 47 of 62 ACTIVE 692268512v5Endothelial cells 69 4200 98.36 N / A Fibroblasts 58 1800 96.78 N / A Macrophages 131 3600 96.36 N / A Intestinal goblet cells 27 600 95.5 N / A Respiratory ciliatedcells Smooth muscle cells 30 600 95 N / A Cardiomyocytes 17 300 94.33 N / A Basal prostatic cells 17 300 94.33 N / A Nk-cells 47 600 92.17 T-cells: 6.5 B-cells 370 3300 88.79 T-cells: 9.85 Intestinal goblet cells: Distal enterocytes 35 300 88.33 8.67 T-cells 1480 11400 87.02 Nk-cells: 6.97 Plasma cells 119 900 86.78 B-cells: 6.11 Distal enterocytes: Proximal enterocytes 71 300 76.33 10.33, B-cells: 7.67 T-cells: 32.0, Salivary duct cells 116 300 61.33 Macrophages: 5.67 Alveolar cells type 2: Club cells 256 600 57.33 39.5 Exocrine glandular Ductal cells 190 300 36.67 cells: 62.0 Weighted Accuracy 3058 31200 90.2 N / A Unweighted Accuracy N / A N / A 88.05 N / A N / A: Not applicable. Materials and Methods for Examples Data Sets

[0144] Human Protein Atlas (proteinatlas.org). The scRNA-seq data was collected as previously described (Uhlén M, Fagerberg L, Hallström BM, et al. Proteomics. Tissue-based map of the human proteome. Science. 2015;347(6220):1260419. doi:10.1126 / science.1260419, which is incorporated herein by reference). Briefly, scRNA- seq read count data from 81 cell types from 31 different data sets was downloaded (HPA atlas) (Table 4). The scRNA-seq data is based largely on the Chromium® single cell gene expression platform from 10X Genomics (version 2 or 3), constituting a single cell suspension from tissues without pre-enrichment of cell types. This includes only studies with Page 48 of 62 ACTIVE 692268512v5> 4,000 cells and 20 million read counts, and only data sets whose pseudo-bulk transcriptomic expression profile is highly correlated with the transcriptomic expression profile of the corresponding HPA tissue bulk samples. An exception was made for the eye (~12.6 million reads allowed), the rectum (2,638 cells allowed) and the heart muscle (plate- based scRNA-seq platform) in order to include these additional cell types in the analysis. Table 4. Exemplary cell type classification scheme

[0145] Tabula Sapiens. The scRNA-seq data was collected as previously described (see, Tabula Sapiens Consortium*, Jones RC, Karkanias J, et al. The Tabula Sapiens: A multiple- organ, single-cell transcriptomic atlas of humans. Science. 2022;376(6594):eabl4896. Page 49 of 62 ACTIVE 692268512v5doi:10.1126 / science.abl4896, which is incorporated herein by reference). Briefly, 24 tissues in total were collected from two cohorts of donors. This allowed biological replicates for almost all tissues. More details of the samples are available from the metadata posted on figshare (figshare.com / articles / dataset / Tabula_Sapiens_release_1_0 / 14267219). The scRNA- seq raw read count data from the 10X Genomics pipeline was downloaded (TS cell atlas) for analysis.

[0146] Peripheral blood mononuclear cells (PBMCs). scRNA-seq data from approximately 3,000 single PBMCs (3K data set) from a donor was contributed by 10X Genomics as a public resource. The raw read count data was downloaded for analysis (see, 10X Genomics 3k PBMCs from a Healthy Donor. s3-us-west- 2.amazonaws.com / 10x.files / samples / cell / pbmc3k / pbmc3k_filtered_gene_bc_matrices.tar.gz). In addition, scRNA-seq data from 68,000 single PBMCs (68K data set) from a donor was generated using the 10X Genomics platform (see, Zheng GXY, Terry JM, Belgrader P, et al. Massively parallel digital transcriptional profiling of single cells. Nat Commun. 2017;8:14049. doi:10.1038 / ncomms14049, which is incorporated herein by reference). The raw read count data was used for analysis.

[0147] Cardiomyocytes. Human induced pluripotent stem cells (hiPSCs) were differentiated to cardiomyocytes in 90 days using two different protocols as previously described (see, Grancharova T, Gerbin KA, Rosenberg AB, et al. A comprehensive analysis of gene expression changes in a high replicate and open-source dataset of differentiating hiPSC-derived cardiomyocytes. Sci Rep. 2021;11(1):15845. doi:10.1038 / s41598-021-94732- 1, which is incorporated herein by reference). Briefly, in Protocol 1, cardiomyocytes were differentiated using small molecules with CHIR99021 and IWP2. In Protocol 2, cardiomyocytes were differentiated using cytokines Activin A, BMP4 and XAV939 plus small molecules with CHIR99021. Samples were collected on days D12 and D24 for Protocol 1 and days D14 and D26 for protocol 2. The difference in the days of collection is due to the lag in the differentiation initiation day between the two protocols. scRNA-seq raw read count data was generated using the SPLiT-seq (see, Rosenberg AB, Roco CM, Muscat RA, et al. Single-cell profiling of the developing mouse brain and spinal cord with split-pool barcoding. Science. 2018;360(6385):176-182. doi:10.1126 / science.aam8999, which is incorporated herein by reference) (split-pool barcoding) library preparation methodology and from sequencing on an Illumina NextSeq. Page 50 of 62 ACTIVE 692268512v5Data Filtering, Preprocessing and Clustering

[0148] Single-Cell Analysis in Python (scanpy) v1.9.2 was used for filtering, preprocessing, and clustering the data prior to running scCompare (see, Wolf FA, Angerer P, Theis FJ. SCANPY: large-scale single-cell gene expression data analysis. Genome Biol. 2018;19(1):15. Doi:10.1186 / s13059-017-1382-0, which is incorporated herein by reference). The unique molecular identifiers (UMIs) in each data set were annotated using the human Ensembl gene model (see, Cunningham F, Allen JE, Allen J, et al. Ensembl 2022. Nucleic Acids Res. 2022;50(D1):D988-D995. doi:10.1093 / nar / gkab1049, which is incorporated herein by reference). The data was filtered such that transcripts not present in at least 3 cells were excluded from further analysis. Cell barcode IDs were filtered to only include those with a minimum of 2,000 non-zero transcripts, those having >= 5% mitochondrial transcripts, and those having <= 10% ribosomal transcripts. The count data was normalized to counts per million (CPM) for each cell by dividing the expression of each gene by the total count of its respective cell, multiplying the result by a million, and then applying a log base 2 transformation with an offset of 1. The normalized data was scaled across single cells to a mean expression = 0 and variance = 1. Highly variable genes were selected using the variance-stabilizing transformation option in scanpy followed by principal component analysis to reduce the dimension of the data. Using the Kneedle heuristic16to determine number of principal components (PCs) according to the point of maximum curvature of the explained variance and k = 100 nearest neighbors, single cells were embedded in a graph with the edges represented as distances drawn between cells. Using a resolution of 0.8, the Leiden algorithm was applied to group the single cells into clusters. Finally, the clusters of cells were visualized in uniform manifold approximation and projection (UMAP) space. Generate Bulk Signatures, statistical thresholding and mapping phenotypic labels

[0149] Count measurements and cluster identities from the training (e.g., mapping) data were used to build cell-cluster-specific prototype signatures that are based on the average expression of each cluster using only highly variable genes. For each cluster, distributions of the correlations of each cell’s highly variable genes to the cluster prototype were generated. Page 51 of 62 ACTIVE 692268512v5The correlation of each cell’s gene signature to each prototype signature is calculated, and the cell is initially assigned the cluster to which it has highest correlation.

[0150] The median absolute deviation (MAD), a measure of dispersion in data, uses the median as a statistic of central tendency and is robust against outliers (see, Iglewicz B, Hoaglin DC. How to Detect and Handle Outliers. ASQ Quality Press; 1993, which is incorporated herein by reference). If x1, x2, …, xn represent a set of n Pearson correlation coefficients between the scRNA-seq expression data of signature genes for each cell in a specific cluster from the mapping data set, and if ^^is the median, then

[0151] 5*MAD below the median was used as the statistical cut-off for the distribution of the Pearson correlation coefficients for a given cluster to exclude phenotypic identity to the single cells in the testing data set. In addition, the training data set was mapped back to the bulk signatures. Single cells that fell below the statistical cut-offs of the clusters were labeled “unmapped.” The coefficient of determination (R2), Spearman Rank correlation coefficient (Rs) and scatter plots were used to assess similarity and compare the proportion of phenotypes between the training and test data sets. A leave one out analysis can be performed on multiplicity of exemplary mapping datasets to determine a distribution of R2 scores. The testing dataset R2 score can then be compared to the distribution to determine testing dataset inclusivity. As an alternative approach to determining the statistical cut-off for distributions that might be highly skewed, Fisher Transformation to convert the Pearson correlation coefficients to z-score can be employed. Two or more standard deviations below the mean as the statistical cut-off demarcates ≥ 97.5 % of the correlations.

[0152] The Fisher Transformation may be calculated as: z = (1 / 2)ln((1 + r) / (1 - r)) wherein ln is the natural logarithm and r is the Pearson correlation coefficient. The resulting z-scores follow an approximately normal distribution with a standard deviation equal to 1 / √(N - 3), where N represents the sample size. This approach may be particularly advantageous when dealing with non-normally distributed correlation coefficients, as the Fisher Transformation helps normalize the distribution. In some embodiments, the method may automatically select between the MAD-based approach and the Fisher Transformation based on the skewness of the correlation coefficient distribution. For example, if the absolute Page 52 of 62 ACTIVE 692268512v5skewness of the correlation coefficient distribution exceeds a threshold value (e.g., 1.0), the Fisher Transformation approach may be automatically selected. Marker Genes Identification

[0153] Marker genes for each protocol (small molecule and cytokine-driven) were identified at matched timepoints. Specifically, D12 and D24 of the small molecules in Protocol 1 were compared to D14 and D26 of the small molecules and a cytokine in Protocol 2, respectively. Marker genes were identified using scanpy’s rank_genes_groups method. A t-test with overestimated variance was used and Benjamini-Hochberg p-value correction was applied. The resulting genes were filtered requiring a minimum in group fraction of 0.5, maximum out group fraction of 0.5, and a minimum fold change of 2. Performance Metrics for the Phenotypic Mapping

[0154] Accuracy may be measured as the proportion of the mapping of the correct phenotypic labels to the test data set over all levels. True Positives + True NegativesAccuracy =True Positives + True Negatives + False Positives + False Negatives

[0155] Specificity may refer to the probability of not falsely mapping the correct phenotypic label to the test data set.Specificity =True Negatives False Positives + True Negatives

[0156] Recall (sensitivity) may represent the probability of mapping the correct phenotypic label to the test data set.Recall =True Positives True Positives + False NegativesPage 53 of 62 ACTIVE 692268512v5Data and Programming Code Availability Statements

[0157] The HPA data is publicly available at www.proteinatlas.org / about / download. The TS gene count data is publicly available at the Gene Expression Omnibus under accession GSE201333. The 3K PBMC data is publicly available at s3-us-west- 2.amazonaws.com / 10x.files / samples / cell / pbmc3k / pbmc3k_filtered_gene_bc_matrices.tar.gz. The 68K PBMC data is publicly available at www.10xgenomics.com / resources / data sets / fresh-68-k-pbm-cs-donor-a-1-standard-1-1-0. The CM data is publicly available at open.quiltdata.com / b / allencell / packages / aics / wtc11_hipsc_cardiomyocyte_scrnaseq_d0_to_d 90. The scCompare repository is github.com / bluerocktx / bfx-scCompare EQUIVALENTS AND SCOPE, INCORPORATION BY REFERENCE

[0158] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. It is understood that modifications which do not substantially affect the activity of the various embodiments of this disclosure are also provided within the description of the disclosure provided herein. The scope of the present disclosure is not intended to be limited to the above description, but rather is as set forth in the appended claims.

[0159] In the claims articles such as “a,” “an,” and “the” may mean one or more than one unless indicated to the contrary or otherwise evident from the context. Claims or descriptions that include “or” between one or more members of a group are considered satisfied if one, more than one, or all of the group members are present in, employed in, or otherwise relevant to a given product or process unless indicated to the contrary or otherwise evident from the context. The disclosure includes embodiments in which exactly one member of the group is present in, employed in, or otherwise relevant to a given product or process. The disclosure also includes embodiments in which more than one, or all of the group members, are present in, employed in, or otherwise relevant to a given product or process.

[0160] Furthermore, it is to be understood that the disclosure encompasses all variations, combinations, and permutations in which one or more limitations, elements, clauses, descriptive terms, etc., from one or more of the claims or from relevant portions of the description is introduced into another claim. For example, any claim that is dependent on another claim can be modified to include one or more limitations found in any other claim that is dependent on the same base claim. Furthermore, where the claims recite a Page 54 of 62 ACTIVE 692268512v5composition, it is to be understood that methods of using the composition for any of the purposes disclosed herein are included, and methods of making the composition according to any of the methods of making disclosed herein or other methods known in the art are included, unless otherwise indicated or unless it would be evident to one of ordinary skill in the art that a contradiction or inconsistency would arise.

[0161] Where elements are presented as lists, e.g., in Markush group format, it is to be understood that each subgroup of the elements is also disclosed, and any element(s) can be removed from the group. It should be understood that, in general, where the disclosure, or aspects of the embodiments, is / are referred to as comprising particular elements, features, steps, etc., certain embodiments of the disclosure or aspects of the embodiments consist, or consist essentially of, such elements, features, steps, etc. Thus, for each embodiment of the disclosure that comprises one or more elements, features, steps, etc., the disclosure also provides embodiments that consist or consist essentially of those elements, features, steps, etc.

[0162] Where ranges are given, endpoints are included. Furthermore, it is to be understood that unless otherwise indicated or otherwise evident from the context and / or the understanding of one of ordinary skill in the art, values that are expressed as ranges can assume any specific value within the stated ranges in different embodiments of the disclosure, to the tenth of the unit of the lower limit of the range, unless the context clearly dictates otherwise. It is also to be understood that unless otherwise indicated or otherwise evident from the context and / or the understanding of one of ordinary skill in the art, values expressed as ranges can assume any subrange within the given range, wherein the endpoints of the subrange are expressed to the same degree of accuracy as the tenth of the unit of the lower limit of the range.

[0163] In addition, it is to be understood that any particular embodiment of the present disclosure may be explicitly excluded from any one or more of the claims. Where ranges are given, any value within the range may explicitly be excluded from any one or more of the claims. Any embodiment, element, feature, application, or aspect of the compositions and / or methods of the disclosure, can be excluded from any one or more claims. For purposes of brevity, all of the embodiments in which one or more elements, features, purposes, or aspects is excluded are not set forth explicitly herein. Page 55 of 62 ACTIVE 692268512v5

[0164] Throughout this disclosure various publications, patents, and sequence database entries are mentioned. The disclosures of these publications, patents, and sequence database entries, including those items listed above, are hereby incorporated by reference in their entirety as if each individual publication or patent was specifically and individually indicated to be incorporated by reference. In case of conflict, the present application, including any definitions herein, will control.

[0165] Although the disclosure has been described with reference to the examples provided above, it should be understood that various modifications can be made without departing from the scope of the disclosure. Accordingly, the above examples are intended to illustrate but not limit the present disclosure. Page 56 of 62 ACTIVE 692268512v5

Claims

CLAIMS 1. A method for using single-cell RNA sequencing (scRNAseq) data to determine whether a first set of cells, derived using a first bioprocess, is phenotypically equivalent to a second set of cells, derived using a second bioprocess, the method comprising: using at least one computer hardware processor to perform: obtaining a plurality of clusters of the first set of cells, each of the plurality of clusters corresponding to a respective phenotype of a plurality of phenotypes; obtaining scRNAseq expression data for the second set of cells, the scRNAseq expression data comprising, for each of multiple cells in the second set of cells, a plurality of expression levels for a respective plurality of genes; determining, using the plurality of clusters and the scRNAseq expression data, a phenotype of the plurality of phenotypes for each of at least some of the cells in the second set of cells; determining, for each of at least some phenotypes of the plurality of phenotypes, a proportion of a number of the first set of cells of the particular phenotype to a total number of the first set of cells, thereby obtaining a first plurality of phenotype proportions; determining, for each of the at least some phenotypes of the plurality of phenotypes, a proportion of a number of the second set of cells determined to be of the particular phenotype to a total number of the second set of cells, thereby obtaining a second plurality of phenotype proportions; and determining whether the first set of cells and the second set of cells are phenotypically equivalent using the first plurality of phenotype proportions and the second plurality of phenotype proportions.

2. The method of claim 1, further comprising: in response to determining that the first set of cells and the second set of cells are not phenotypically equivalent, determining whether the second set of cells comprises pluripotent stem cells or progenitor cells; and in response to determining that the second set of cells comprises pluripotent stem cells, adjusting an aspect of the second bioprocess. Page 57 of 62 ACTIVE 692268512v53. The method of claim 2, wherein adjusting the aspect of the second bioprocess comprises: adding one or more agents to the second bioprocess; adjusting a concentration of one or more agents in the second bioprocess; or adjusting an initial concentration of pluripotent stem cells used in the second bioprocess.

4. The method of any of claims 1-3, further comprising: in response to determining that the first set of cells and the second set of cells are not phenotypically equivalent, identifying cells of the second set of cells that are of an unknown phenotype; and determining a gene expression signature for the identified cells.

5. The method of claim 4, further comprising: using the gene expression signature to develop an assay for quantifying cells of the unknown phenotype in one or more other samples.

6. The method of any of claims 4-5, further comprising: determining a phenotype for the identified cells using the gene expression signature.

7. The method of claim 6, wherein determining the phenotype for the identified cells using the gene expression signature comprises: identifying one or more marker genes comprised in the gene expression signature; and determining the phenotype based on the identified marker genes.

8. The method of any of claims 1-7, further comprising: determining whether to administer a cell therapy based on a result of determining whether the first set of cells and the second set of cells are phenotypically equivalent, the cell therapy comprising one or more products of cells derived using the second bioprocess.

9. The method of any of claims 1-8, further comprising: determining whether to discard the second set of cells based on a result of determining whether the first set of cells and the second set of cells are phenotypically equivalent; and Page 58 of 62 ACTIVE 692268512v5in response to determining to discard the second set of cells based on the result, discarding the second set of cells.

10. The method of any of claims 1-9, wherein determining whether the first set of cells and the second set of cells are phenotypically equivalent using the first plurality of phenotype proportions and the second plurality of phenotype proportions comprises: determining a coefficient of determination (R2) between the first plurality of phenotype proportions and the second plurality of phenotype proportions, the coefficient of determination (R2) representing a degree to which the first plurality of phenotype proportions and the second plurality of phenotype proportions are phenotypically equivalent; and determining that the first set of cells and the second set of cells are phenotypically equivalent if the coefficient of determination (R2) satisfies at least one criterion.

11. The method of any of claims 1-10, further comprising: obtaining a first scRNAseq expression data for the first set of cells, the first scRNAseq expression data comprising, for each of multiple cells in the first set of cells, a plurality of expression levels for a respective first plurality of genes, wherein obtaining the plurality of clusters of the first set of cells comprises clustering the first set of cells into the plurality of clusters using the first scRNAseq expression data.

12. The method of claim 11, wherein clustering the first scRNAseq expression data comprises clustering the first scRNAseq expression data using a hierarchical clustering algorithm.

13. The method of claim 12, wherein clustering the first scRNAseq expression data using the hierarchical clustering algorithm comprises clustering the scRNAseq expression data using Leiden clustering.

14. The method of any of claims 11-13, wherein obtaining the first scRNAseq expression data for the first set of cells comprises obtaining, for each cell in the first set of cells, expression levels for between 5,000 and 10,000 genes, between 10,000 and 15,000 genes, between 15,000 and 20,000 genes, between 20,000 and 25,000 genes, or between 25,000 and 30,000 genes. Page 59 of 62 ACTIVE 692268512v515. The method of any of claims claim 11-14, further comprising: determining a gene expression signature for each particular cluster in the plurality of clusters using the first scRNAseq expression data, the determining comprising: for each particular gene in a subset of the first plurality of genes, determining an average expression level of the particular gene across cells comprised in the particular cluster, wherein determining the phenotype for each of at least some of the cells in the second set of cells using the plurality of clusters comprises: determining the phenotype using the gene expression signatures determined for the plurality of clusters.

16. The method of claim 15, further comprising: for a particular cluster of the plurality of clusters: generating a distribution of correlation coefficients, the generating comprising: determining, for each cell comprised in the particular cluster, a correlation coefficient between a gene expression signature of the cell and the gene expression signature determined for the particular cluster; and determining a statistical cut-off value for the particular cluster, wherein the statistical cut-off is a mean absolute deviation from a median of the generated distribution of correlation coefficients.

17. The method of claim 16, wherein determining the phenotype of the plurality of phenotypes for each of at least some of the cells in the second set of cells comprises, for a particular cell of the at least some cells in the second set of cells: determining a correlation coefficient between scRNAseq expression data obtained for the particular cell and the gene expression signature determined for each cluster; and determining the phenotype for the particular cell based on the correlation coefficients and the statistical cut-off value for each cluster.

18. The method of claim 17, further comprising: identifying the particular cell as an outlier if none of the determined correlation coefficients exceed the respective statistical cut-off value. Page 60 of 62 ACTIVE 692268512v519. The method of any of claims 1-18, wherein the first bioprocess is different from the second bioprocess.

20. The method of any of claims 1-19, wherein the method is used to verify consistency of induced pluripotent stem cell (iPSC)-derived cell populations.

21. The method of any of claims 1-19, wherein the method is used to compare cell batches produced under at least one of varying conditions or varying protocols.

22. The method of any of claims 1-19, wherein the method is used in research and development of cell therapies.

23. A system, comprising: at least one computer hardware processor; and at least one non-transitory computer-readable storage medium storing processor- executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform the method of any of claims 1- 22.

24. The system of claim 23, wherein the first bioprocess is different from the second bioprocess.

25. At least one non-transitory computer-readable storage medium storing processor- executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform the method of any of claims 1- 22. Page 61 of 62 ACTIVE 692268512v5

Citation Information

Patent Citations

  • Methods for predicting outcomes and treating colorectal cancer using a cell atlas

    US20210047694A1

  • RNA sequencing method for the analysis of b and t cell transcriptome in phenotypically defined b and t cell subsets

    US20220333194A1

  • Method for Identifying Functional Disease-Specific Regulatory T Cells

    US20230107291A1