Systems and methods for characterization of cell phenotype and applications thereof
Patent Information
- Application Number
- PCT/US2025/018429
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-04
- Filing Date
- 2025-03-04
- Publication Date
- 2026-02-26
AI Technical Summary
Existing computational methods struggle to accurately and reproducibly determine the absolute potency status of cells across different datasets, making it difficult to compare and contextualize cellular potency in diverse physiological and pathological processes, particularly in cancer.
A novel interpretable deep learning framework utilizing binary neural networks is employed to predict single-cell potency categories and absolute developmental potential from omic data, enabling the identification of molecular programs associated with cell phenotypes.
The framework provides remarkably accurate and interpretable predictions of cellular potency, facilitating insights into cancer diagnostics, therapy response, and clinical outcomes by leveraging omic data from single cells.
Smart Images

Figure US2025018429_26022026_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR CHARACTERIZATION OF CELL PHENOTYPE ANDAPPLICATIONS THEREOF CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Appl. No.63 / 561,252, entitled “Systems and Methods for Characterization of Cell Phenotype andApplications Thereof,” filed March 4, 2024, the disclosure of which is hereby incorporatedby reference in their entirety. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with Government support under contract CA255450awarded by the National Institutes of Health. The Government has certain rights in the invention. TECHNICAL FIELD
[0003] The disclosure provides description of an interpretable computational networkthat utilizes omic data from single cells or other entities for characterizing cell or entitywith a phenotype status, including interpretable diagnostic applications applicable tocancer and other medical disorders. BACKGROUND
[0004] Cell differentiation is a biological process can be described as linear set ofphenotypes from which less specialized cells become more specialized in a maturationprocess. In mammals, the zygote (i.e., the initial cell post conception) is the least differentiated cell, which provides genetic material for development of all cell types via cell differentiation. During development, a multitude of stem cells are formed which havethe main function of dividing into more cells that will populate throughout the body. Uponvarious internal and external cues, stem cells differentiate into mature cells to provide aparticular function, forming a complex system of tissues and cell types.
[0005] Cell potency is the ability of a cell to divide and give rise to various differentiatedcell types. Embryonic stem cells are pluripotent and have the potential to differentiate into any cell within the body. As cells mature, their potency is further limited. For example, hematopoietic stem cells are limited to the potential to differentiate into immune and bloodcells, and likewise, neural stem cells are limited to the potential to differentiate into cellsof the nervous system. Accordingly, an embryonic stem cell can divide and mature into aneural stem cell or a hematopoietic stem cell. A neural stem cell can then divide and further mature into a neuron or astrocyte (cells that provide brain function). Likewise, a hematopoietic stem cell can divide and further mature into a B-lymphocyte or T- lymphocyte (immune cells that respond to pathogen infection).
[0006] Cell potency phenotypes include the cellular properties of stemness, celllineage, and differentiation status. Stemness is a characteristic that marks the ability to self-renew and mature into a variety of cells. Accordingly, the stemness of an embryoniccell is greater than that of a neural stem cell or hematopoietic stem cell, and these cellshave greater stemness than neurons, astrocytes, B-lymphocytes and T-lymphocytes. Cell lineage is the developmental pathway that a particular mature cell underwent from an originator cell. For instance, a B-lymphocyte’s cell lineage starting at fertilization is as follows: zygote embryonic stem cells hematopoietic stem cells common lymphoid progenitors pro-B lymphocytes pre-B lymphocytes immature B lymphocytes B- Differentiation state is the status of cell within its developmental pathway. For example, a B-lymphocyte has a more mature differentiation state than a hematopoietic stem cell.
[0007] Differentiation and potency status is also important in cancer pathology.Cancerous tissues and tumors comprise mixed populations of cells of varieddifferentiation and potency status. Cancer stem cells are lowly differentiated, highly potentcells that possess the ability to self-renew, driving cancer growth. The cancer stem cellscan have great influence on whether a cancer is fast growing, able to metastasize, and likely to relapse. Presence of these cells can dictate a need for more potent treatments.SUMMARY OF THE INVENTION
[0008] Many embodiments are directed to methods of diagnosing cellular phenotype(e.g., potency) status of individual cells or entities based upon omic data (e.g., single-cell RNA sequencing (scRNA-seq), cell-free DNA methyl sequencing (cfDNA methyl-seq), etc.). In many of these embodiments, omic features signatures (e.g., gene expression signatures, methylation signatures, etc.) are utilized to determine a cell or entity phenotype. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The description and claims will be more fully understood with reference to thefollowing figures and data graphs, which are presented as exemplary embodiments of the invention and should not be construed as a complete recitation of the scope of the invention.
[0010] Figure 1 provides an example of an overview of a computational process forcharacterizing phenotypes from omic data.
[0011] Figure 2 provides an example of a computational process for standardizingomic data.
[0012] Figure 3 provides an example of a computational process for characterizing aphenotype.
[0013] Figure 4 provides an example of computing systems for characterizingphenotypes.
[0014] Figure 5 provides an example of a module for characterizing a phenotype.
[0015] Figures 6A-19M provide experimentation and data results of a computationalframework for predicting potency of single cells.DETAILED DESCRIPTION
[0016] The description provides various computational processes to characterizecellular and biological phenotypes from omic data and various applications based onphenotype characterization. Many embodiments of computational processes utilize omicdata results to infer phenotypes of a cell or an individual. In one example, single cell RNAsequencing (scRNA-seq) results infer absolute potency status of the sequenced cell. Anadvantage of determining absolute potency is that cells of different origins and lineages,sequences via different platforms, at different locations, or by different methods can be compared. The determination of absolute potency is a great enhancement over prior ascomputational methods that rely on relative potency within a population. Phenotype statusof an biological entity (e.g., a cell, a tissue, an individual) can be inferred by a computational framework comprising one or more phenotype modules, where each module yields an assessment score for each entity for a particular phenotype. Phenotype assessment scores from two or more modules can be combined to yield a probability of whether a particular entity has a certain phenotype.
[0017] In several embodiments, a phenotype module can comprise learning of a geneset that describes a biological phenotype via binarization of genes based on omic data results having assignable phenotypes. In one example, a gene set is learned for cellular potency from scRNA-seq data. Binary neural networks have been historically employed when the amount of computation was desired to be reduced, often sacrificing model accuracy. Herein, however, binary neural networks are utilized to achieve interpretable sets of omic features for describing a phenotype. These interpretable sets of omic features are utilized within the module to yield an assessment score for each entity for a particular phenotype. Furthermore, these sets of omic features can be utilized as biomarkers for a phenotype. Accordingly, in some implementations, a binary neural network layer is trained to classify an entity to an assigned phenotype using omic features as input. As it is trained, the binary neural network learns interpretable sets of omic features for describing a phenotype. The set of omic features can then be utilized as biomarkers for that phenotype (within the neural network or via a biological assay).
[0018] Phenotype scores from the one or more phenotype modules can be utilized toyield a medical diagnosis. For example, it has been found that potency phenotype modules are useful in diagnosis of cancer. In regards to cancer diagnostics, omic data of cancer cells can be assessed to characterize the cancer by the presence and / or percentage of cells having particular potency phenotypes. Characterization of the cancer in this manner can inform survival outcomes, cancer aggressiveness, risk of metastasis, risk of recurrence, and therapy response. Further clinical diagnostics and / or treatments can be determined accordingly.
[0019] Various embodiments of the computational methods described herein providean ability to accurately and reproducibly determine multiple phenotypes, such as absolutepotency status of a cell, which has not been achieved previously. While lineage tracing,functional transplantation assays, and single-cell genomics have revealed key insights into developmental hierarchies, the molecular programs underlying potency hadremained unclear. For example, totipotent, multipotent, and unipotent cells each havevastly different developmental capacities, yet little was known about the omic profiles thatdistinguish them. An improved understanding of cell potency would facilitate new insights into diverse physiological and pathological processes, including cancer, where altered differentiation states and plasticity programs influence clinical outcomes.
[0020] Prior computational models for characterizing single cell potency yieldedpredictions of single-cell orderings in a manner that is relative among the cells within adataset, but not across datasets. This has made it difficult to compare predictions ofdifferent datasets and contextualize the resulting cellular potency of different experiments.To overcome these challenges, the systems and methods described herein utilize aninterpretable deep learning framework for jointly predicting single-cell potency categoriesand absolute developmental potential from single-cell omic data. Unlike most deeplearning methods, which are not inherently explainable, the current systems and methods implement a novel architecture that learns multivariate gene expression programs for each phenotype of interest. Such gene expression programs can be seamlessly extracted from the model, are readily interpretable, and deliver remarkably accurate predictions of developmental potential. To demonstrate the utility of systems and methods described, human and mouse scRNA-seq datasets with gold standard potency levels anddifferentiation trajectories were curated and the used for training the computationalframeworks. The systems and methods were further validated and benchmarked. Theresults from such assessments illuminate pan-tissue determinants of cell potency andhighlight the value of interpretable deep learning for characterizing single-cell phenotypes in health and disease.Overview of methods to characterize class or phenotype status of a cell
[0021] A number of embodiments are directed to characterizing a phenotype status ofan entity from omic data. Herein, an entity is a single cell or a sample, or an assembly ofcells that form a biological entity that can be characterized. Examples of assemblies of cells that form a biological entity include a cell, a cellular network, a tissue, a tissue system, an organ, an organ system, an individual, or any other assembly of one or more cells having a collective class and / or phenotype. In one example, an entity is a cell, and the phenotype is potency phenotype. Omic data such as RNA sequencing (scRNA-seq) or methyl-sequencing can be used as features within a computational framework topredict the likelihood of a phenotype category or predict a status on a regressioncontinuum. In various embodiments, an assembly of cells are characterized, such as apopulation of single cells within a tissue or a cancer growth. Any biological phenotype orbiological class can be characterized, such as a potency status, a metabolic status, asenescence status, a spatial status, a status related to cell-to-cell contact, a status related to a signaling pathway, a status related to a stimulation, a status related to an immune response, a status related to a pathogen, a status related to an injury, a status related to a medical condition, a status related to a treatment, etc. As an example of a status can be a differentiation status. Examples of categorical assignments of differentiation statusphenotypes include (but are not limited to) totipotent, pluripotent, multipotent, oligopotent,unipotent, and differentiated. Other status categorical assignments, class assignments,and continuums would be readily apparent to the phenotype being characterized.
[0022] Several particular embodiments are also directed towards identifying otherpotency phenotypes such as quiescent stem cells, and especially adult quiescent stem cells, from mitotic stem cells, and senescent cells within a population of cells. Quiescentstem cells (also referred to long-term stem cells) are traditionally difficult to infer due totheir hibernation-like state. In many embodiments, quiescent stem cells are non-mitotic stem cells that are capable of reentering into a mitotic phase and further differentiating into a more differentiated state. Accordingly, quiescent stem cells have a similar differentiation and stemness status as mitotic stem cells but will have substantialdifferences in character and gene expression profiles as related to its quiescent state.
[0023] Provided in Fig. 1 is a flow chart of an example of an overview computationalprocess to characterize a phenotype or class status of entities based on its omic data. Inone example, the process can be utilized to characterize and compare absolute potency of biological cells from different populations or different sequencing reactions. Notably, potency is based on intrinsic gene expression associated with potency. Thecomputational process can incorporate one or more binary neural networks that are bothdeep-learning and interpretable, resulting in a learning of the molecular programsassociated with a class or a phenotype. The results of the process (e.g., determination ofcell’s absolute potency) and the interpretability of the process (e.g., molecular programs associated with potency) can be utilized in a variety of downstream processes, such as (for example) biomarker identification, development of diagnostics, diagnosis of medical conditions, and indication of therapy response as related to the phenotype.
[0024] Computational process 100 can begin by standardizing (101) omic data.Standardization can comprise assigning a standardized rank of relative omic datafeatures within an entity (e.g., a gene’s relative expression level within a cell), such thatfeature levels can be compared among the various sources. In some implementations in which multiple species are utilized, orthologous features (e.g., orthologous genes) are harmonized. It is to be understood that any omic method for determining classes or phenotypes can be utilized, including (but not limited to) RNA-seq, spatial transcriptomics, methyl-seq, ChIP-seq, ATAC-seq, proteome analysis, etc.
[0025] In one example, omic data features are gene expression levels from single cellRNA sequencing (scRNA-seq), which is a method of sequencing RNA transcripts on theindividual cell level (and not a bulk collection of cells). Any of a variety of single cellsequencing techniques and technologies can be utilized for assessment, such as droplet- based sequencing and microwell-plate-based sequencing (see, e.g., Y. Kashima, et al., Exp Mol Med. 2020 Sep;52(9):1419-1427, the disclosure of which is incorporated by method). Further, because computational process 100 can determine a composite class or composit phenotype status such as absolute potency, scRNA-seq data derived from several distinct reactions can be combined into the process, resulting in an ability to compare single cells from any of a variety of sources, experimental protocols, and even a variety of species. Standardization of the scRNA-seq data allows for assessment ofabsolute potency among virtually any scRNA source. Similar techniques can be utilized for single-cell methy-seq data, single cell ChIP-seq data, single cell ATAC-seq data, orsingle-cell proteome data.
[0026] Steps 103 and 105 of computational process 100 can be performed to assessa class status or phenotype status. Accordingly, a computational process can performthese steps as part of a particular phenotype module, and the modules may be combined to yield an ensemble computational framework. In some implementations, a computational framework comprises one or more modules, each module for classificationof a status, which can be combined to yield a composite class status. Or a module cancomprise a module that utilizes regression analysis to yield a status on a continuum,which can optionally be combined with other modules to yield a composite response. Inone example, a potency phenotype is a phenotype that characterizes a cell’s potency, such as a differentiation status selected from: totipotent, pluripotent, multipotent, oligopotent, unipotent, and differentiated. In particular implementations, a computational framework for assessing potency phenotype comprises a module for each of the following potency phenotypes: totipotent, pluripotent, multipotent, oligopotent, unipotent, and differentiated.
[0027] At step 103, computational process 100 determines omic features associatedwith a phenotype. An omic feature depends on the omic analysis performed. For example,omic features can be gene expression (as determinable by RNA-seq or proteomeanalysis), methylation patterns (as determinable by methyl-seq), histone marks (asdeterminable by ChIP-seq), chromatin accessibility (as determinable by ATAC-seq), and protein expression (as determined by proteomics). To determine omic feature association, training data can comprise omic data having an assignable status. Utilizing the training data, a computational model can be used to learn which omic features are associated with a class status or phenotype. In some implementations, a binarized neuralnetwork is utilized, which can provide deep learning and interpretability by assessing afull set of omic features and discriminating which features (1= yes; 0=no) best describe the phenotype.
[0028] The results of step 103 are utilized downstream in the computational processbut can also be utilized in other downstream applications. Unlike most deep learningmodels (e.g., non-binarized neural networks) in which the process for determining association remains hidden, step 103 reveals a set of omic features associated with a phenotype. The presence of these omic features can be utilized as biomarkers for phenotype assessment and within various diagnostics.
[0029] At step 105, computational process 100 determines a class status or phenotypeenrichment score for one or more entities. This step can be performed utilizing omic datautilized to train the model or omic data that was not utilized for training, such as utilizing omic data that does not yet have an assignable potency phenotype. This step assesses omic data for expression of a set of omic features associated with the phenotype to determine if the omic data set is enriched for presence of the set of omic features associated with the phenotype. A phenotype enrichment score can further be based on genotype feature count number.
[0030] Steps 103 and 105 can be repeated for a plurality of phenotypecharacterizations. As mentioned, a computational framework can comprise one or moremodules, each module for determining the likelihood of a class or phenotype status, which can be categorical. The resulting enrichment scores for each phenotype status can be integrated (107) to yield a composite. In some implementations, potency phenotype enrichment scores are integrated to yield a probability that cell has a certain potency phenotype. In some implementations, potency phenotype enrichment scores are integrated to yield a continuous potency score, e.g., when a computational framework comprises modules of a number differentiation statuses. In some implementations, a continuous potency score can be interpreted as an absolute potency scale, especially when the computational framework comprises modules for differentiation statuses ranging from totipotent to differentiated. Standardization of omic data for phenotype analysis
[0031] Several embodiments of the disclosure are directed to standardizing omic data(e.g., scRNA-seq data) among different omic data acquisitions. Different omic data acquisitions is to mean capture of omic data of different entities, performed by differentprotocols, captured by different laboratories, captured at different times, utilization ofdifferent machines, etc. Standardization improves computational model training andcomparison of class or phenotype characterization results among different entities. A goalof standardization is to enable comparison of omic features from one entity to another, regardless of how the omic data was acquired. For example, certain methods can compare gene expression from one cell to another, regardless of cell tissue source and across sequencing platforms.
[0032] Provided in Fig. 2 is a computational process 200 to standardize omic dataresults, such as sequencing results of single cells. Process 200 obtains omic data resultsfrom a collection of entities, such as sequencing data results from a collection of singlecells. scRNA-seq is a method of sequencing RNA transcripts on the individual cell level(and not a bulk collection of cells). Accordingly, a collection of cells are sorted into individual cells and have each cell’s RNA transcripts individually isolated and processedfor sequencing, resulting in a collection of RNA transcript data of single cells. Similarmethods can be employed for other omic data reactions.
[0033] In numerous embodiments, omic material (e.g., transcriptomic RNA) can bederived from a cellular source, such as a tissue (e.g., biopsy), an in vitro cell line, orisolated single cells (e.g., circulating tumor cells isolated from blood). Tissue sources may include a biopsy of an organ (e.g., blood, brain, lymph node, thymus, bone marrow, spleen, skeletal muscle, heart, colon, stomach, small intestine, kidney, liver, lung, etc.). In some embodiments, the biological sample may be a tumor biopsy. A cancer biopsy refers to any tissue sample containing cancer cells that is obtained (e.g., by excision, needle aspiration, etc.) from a subject. The appropriate cellular source will often depend on the process, the features to be examined, and resultant measurements to be revealed.
[0034] For analysis of cellular sources, cells can be sorted into single cells, omiccomponents can be isolated, purified, or concentrated by a number of methods known inthe field. Cellular sources can be sorted into single cells by various methods, as understood in the field, including (but not limited to) flow cytometry, laser microdissection, manual cell picking, and microfluidics. For scRNA-seq analysis, RNA transcript molecules can be amplified and prepared into a library and sequenced by a number of known methods, such as droplet-based sequencing and microwell-plate-based sequencing.
[0035] Several embodiments derive omic data from publicly or privately availabledatabases. These databases store a number of experimental omic analysis results fromnumerous biological sources. Databases that have a large amount of chromatin derived nucleic acid sequence data include (but are not limited to) National Center for Biotechnology Information (NCBI) and scRNAseqDB.
[0036] In various implementations, class orphenotype analysis is performed across anumber of species. Many biological processes are fundamental to biology often shared across species, such as development. For example, a number of research models can yield insight into human health and can be utilized to train computational models for understanding classes or phenotypes within human systems. Accordingly, computationalmethod 200 optionally harmonizes (203) identifiers in omic data results when resultsinclude multiple species.
[0037] In one example, gene identifiers between species are harmonized. Toharmonize gene identifiers, all gene symbols of a species can be mapped to the closest orthologs of another species, as determined by gene sequence similarity. In some implementations, when multiple genes from one species has maximum similarity to agene in another species, only the gene of the multiple genes from the one species havinghighest sequence similarity is utilized for within the computational model and the other genes are discarded. In some implementations, unique genes of a species (e.g., genes having no ortholog with sequence similarity over a threshold), is kept. Upon harmonization, a harmonized gene set reflective of the species from which the scRNA- seq data is derived, is utilized within the computational framework. Similar techniques can be utilized for methylation patterns (e.g., sequence similarity of and proximal to CpG sites), histone patterns (e.g., sequence similarity of and proximal to histone sites), chromatin accessibility (e.g., sequence similarity of open / closed chromatin), and proteome (e.g., orthologous proteins, or sequence similarity of proteins).
[0038] In some implementations, the harmonized data set trims away omiccomponents that are not often present within omic acquisition data. In the example of RNA-seq analysis, if a gene is not present within at least X% of training datasets, it is discarded. X can be any percentage, and may be dependent on the number datasets utilized for training. In some implementations, X is between 50% and 90%, which removes many uncommonly expressed genes. For methyl seq analysis, certain CpG sites can be trimmed away; for histone analysis or chromatin accessibility analysis, certain genomiclocations can be trimmed away; and for protein analysis, certain gene products can be trimmed away.
[0039] For each entity, computational process 200 further assigns (205) each omicfeature a relative integer rank based on level within the entity. Accordingly, the most present omic feature in an entity (e.g., highest expressed gene in a cell) is assigned a rank of 1 and the least present omic feature in an entity (e.g., lowest expressed gene in a cell) is assigned a rank equivalent to the total number of omic features in the learned omic feature set. It should be understood that an inverse ranking can be utilized instead.Upon assignment of omic feature rank, a matrix can be built of omic features x entities, inwhich omic feature according to rank is aligned with its entity. Integer ranking of omicfeatures circumvents batch effects, mitigates the influence of extreme values and outliers,and reduces the risk of model overfitting. In some implementations, a matrix of omicfeatures x entities is utilized as an input layer within a computational module to yield a a class status or phenotype score.
[0040] While specific examples of processes for standardizing omic data aredescribed above, one of ordinary skill in the art can appreciate that various steps of the process can be performed in different orders and that certain steps may be optional. As such, it should be clear that the various steps of the process could be used as appropriate to the requirements of specific applications. Furthermore, any of a variety of processesfor standardizing omic data appropriate to the requirements of a given application can beutilized. Computational modules for yielding a class sutatus or phenotype score
[0041] Several embodiments are directed to training and utilizing a computationalmodule to yield a class status or phenotype score of an entity based on its omics (e.g.,transcriptome, genome, histone marks, chromatin, proteome, etc.). A computationalmodule can be trained to determine a likelihood that a particular cell has a phenotype status, such as (for example) a particular differentiation status. Accordingly, in that example, a computational module can be trained to determine a likelihood a cell is one of: totipotent, pluripotent, multipotent, oligopotent, unipotent, or differentiated. In some implementations, a computational module is part of a computational framework in whichthe framework comprises a plurality of computational modules, each computationalmodule for predicting a class or phenotype status. The computational framework canintegrate phenotype scores of each the modules to yield an aggregated assessment of phenotypes. In some particular implementations, a computational framework for assessing differentiation status comprises a computational module for each of the following: totipotent, pluripotent, multipotent, oligopotent, unipotent, and differentiated. Other examples include (but not limited to) a medical condition status and a treatment status. A medical condition status can be classified as affected or healthy, or classified as subtype of two or more subtypes of a medical condition. A treatment condition status can be classified as responsive or unresponsive, or classified as an optimal treatment type selected from two more treatment types. Furthermore, status of multiple phenotypesor class status can be combined, such as (for example) affected / healthy and medicalcondition subtype; responsive / unresponsive and optimal treatment type; and medical condition subtype and optimal treatment type. Further, a status can be yield on acontinuum utilizing regression. Selection of module for class or phenotype statusdetermined by the task to be completed and the assigned status of omic data utilized for training.
[0042] Provided in Fig. 3 is a computational method 300 to be performed by acomputational module to yield a phenotype score. Computational method 300 can beginby obtaining (301) a matrix of omic features x entities derived from omic data results. Forexample, in single-cell gene expression data, omic features are genes and entities arecells. Accordingly, a matrix for a collection of cells can be genes X cells, where genes isexpression level of a particular gene and cell refers to a particular single-cell geneexpression result. As noted previously, omic data can be methylation patterns, histonepatterns, chromatin accessibility, and proteome expression. Further, examples of entities include (but are not limited to) cell systems, tissues, tissue systems, organs, organ systems, or individuals. It should be further noted that data can be acquired in one manner (e.g., single-cell sequencing) and the acquired data can be combined to yield or is utilized to infer a higher order entity.
[0043] In several implementations, each omic feature is a relative integer rank basedon level within the entity. Accordingly, the most present omic feature in an entity (e.g.,highest expressed gene in a cell) is assigned a rank of 1 and the least present omicfeature in an entity (e.g., lowest expressed gene in cell) is assigned a rank equivalent tothe total number of omic entities in the set. It should be understood that an inverse rankingcan be utilized instead. In some implementations, the omic features x entities matrix isprepared in accordance with computational method 200 as shown in Fig. 2, but anymethodology to generate a omic features x entities matrix can be utilized. In someimplementations, a omic features x entities matrix comprises data for training a module, and thus the cell component has an assignable class or phenotype status. In someimplementations, a omic features x entities matrix comprises data not utilized in training,and the computational module is utilized for assessment to classify an entity’s phenotype or class sstatus.
[0044] During training, a computational module utilizes a omic features x entitiesmatrix with entities having an assignable phenotype status to learn omic features associated with the phenotype status via a binary neural network. When validating a module or when utilizing a module for assessment, the learned omic features and binary neural network are utilized to determine a phenotype enrichment score for each entity being assessed or used for validation. Accordingly, the next set of steps of computational method 300 can be utilized for training, validation, and / or assessment, as appropriate to how the module is being utilized.
[0045] Computational process 300 optionally trims (303) the omic features x entitiesmatrix to a learnable maximum rank, which allows for enrichment assessment. For numeric stability, the cutoff rank for trimming input the omic features x entities matrix canbe computed as a function of learnable parameter instead of being learned directly. Insome implementations, the cutoff rank for trimming can be computed to ensure that amatrix preserved ranks of more omic features beyond the maximum omic feature set sizeof the module following trimming, which ensures that a omic feature set pool is larger thanthe omic feature set itself when computing a omic feature set enrichment scorecalculation, aiding in comparison of enrichment scores. Unranked processes can beutilized as well. Enrichment scores can be computed by aggregation of omic features
[0046] During training (whether initiating or updating), computational process 300learns (305) omic sets for each module based on binarization binarization of omic featureswithin a set of features. This step determines which omic features are highly associatedwith the module’s assigned phenotype status. Model predictions can be assessed at eachiteration against ground truth, with a loss function and its gradient computed and used tobackpropagate updates to binary network weights. Backpropagation for the binary neuralnetwork component of each potency status can be performed by any appropriate means.In some implementations, Straight-Through Estimator and hardtanh activation function isutilized for backpropagation (see, e.g., M. Courbariaux, et al., (2015). BinaryConnect: training deep neural networks with binary weights during propagations. Proceedings ofthe 28th International Conference on Neural Information Processing Systems - Volume2. MIT Press, the disclosure of which is incorporated herein by reference).
[0047] The number of omic features in a learned omic feature set associated with aphenotype or class status can vary and is dependent on the model training. Accordingly,in some embodiments the number of omic features in a learned omic feature setassociated with a module includes any number of omic features between 5 and 1000,and in some specific embodiments the number of genes in a learned gene set associatedwith a module includes at least: 5 to 15 genes, 10 to 20 genes, 15 to 25 genes, 20 to 30genes, 25 to 35 genes, 30 to 40 genes, 35 to 45 genes, 40 to 50 genes, 45 to 55 genes,50 to 60 genes, 55 to 65 genes, 60 to 70 genes, 65 to 70 genes, 70 to 80 genes, 75 to 85genes, 80 to 90 genes, 85 to 95 genes, 90 to 100 genes, 50 to 150 genes, 100 to 200genes, 150 to 250 genes, 200 to 300 genes, 250 to 350 genes, 300 to 400 genes, 350 to450 genes, 400 to 500 genes, 450 to 550 genes, 500 to 600 genes, 550 to 650 genes,600 to 700 genes, 650 to 750 genes, 700 to 800 genes, 750 to 850 genes, 800 to 900genes, 850 to 950 genes, or 900 to 1000 genes. As an example, provided in Table 1 isa set of learned genes from scRNA-seq data associated with modules of the following potency phenotype status: totipotent, pluripotent, multipotent, oligopotent, unipotent, and differentiated.
[0048] Computational process 300 also computes (307) one or more enrichmentscores for each entity. For example, in phenotype module based on scRNA, a gene expression enrichment score and a gene count enrichment score for each cell wascomputed (see Examples and Data Results). A gene expression enrichment scoreaggregates overall expression activity of the learned set of genes of the module in rankedspace. Whereas a gene count enrichment score compares the number of genes expressed within the learned gene set versus background expectation. In some implementations, the enrichment scores are subject to the learnable maximum rank. In some implementations, enrichment scores are normalized across entities to facilitate comparability of the scores.
[0049] Computational process 300 optimally further integrates (309) two or moreenrichment scores to yield a module phenotype or class status score (when two or moreenrichment scores are present). In some implementations, the one or more enrichment scores are passed through a feedforward layer with a gene set enrichment weight to yield a score vector, indicating a likelihood that the entity is of has this phenotype or class status.
[0050] As described previously, a module can be part of larger computationalframework, which can comprise several modules. Each module can yield a module scorefor a particular phenotype or class status. These module phenotype scores can becombined to yield an overall phenotype composite or class assessment.
[0051] In some implementations, the computational framework is for predicting cellpotency status. In some implementations, the computational framework predicts the most likely potency phenotype of a cell. In some implementations, the framework applies asoftmax activation function, yielding entity X module omic feature set matrix representingthe likelihood of each entity belonging to each of the combined phenotypes. Thecomputational framework can then predict phenotype class status by assigning thephenotype status with highest likelihood for each entity via a maximization function (e.g.,argmax function). Phenotypes or class status can be predicted on a continuum usingregression.
[0052] In some particular implementations, the computational framework computes anaggregated phenotype status score, which can be a continuous phenotype score. For instance, when a computational framework comprises modules for multiple potency phenotypes along a differentiation timeline, a continuum of potency can be computed. In some implementations, when a computational framework incorporates modules from totipotent to differentiated, the continuum of potency is absolute, and thus an absolute potency score can be determined. In some implementations, an absolute potency scaleis created (e.g., from 0 to 1), and the likelihood of a cell having a certain potency phenotype status is combined with this scale to yield an absolute potency score.Alternatively, thresholds for categorical status can be utilized such that the aggregatedphenotype status score yields a categorical status. The aggregated phenotype categorical status can be further combined with the likelihood of a cell having a certain potency phenotype status to yield an overall categorical status classfication.
[0053] In some implementations, an aggregated phenotype score (e.g., an absolutepotency score) is smoothened by denoising across similar entities. To prepare fordenoising, a similarity matrix of the omic data is generated to identify similar entities basedon their omics data. To generate a similarity matrix (e.g., Markov matrix), an original omicfeature matrix is filtered to a set of omic features present in at least a particular percentof entities (e.g., 5%). Then a dispersion index is calculated for each omic feature by takingthe ration of its variance to its mean across all entities. A number of most dispersed omicfeatures (e.g., top 1000 most dispersed genes in scRNA-seq data) are used to computea similarity matrix that is converted into a similarity matrix by setting all diagonal elementsto 0 and all correlations less than the null similarity coefficient to 0, then normalizing thesum of each row to 1. The null similarity coefficient can be estimated by bootstrapping10,000 random Pearson product-moment correlation coefficients between pairs ofentities. The elements of the resulting similarity matrix indicate the neighborhoods ofentities across which denoising will be applied, where 0 indicates no similarity between apair of entities and values between 0 and 1 indicate different degrees of similarity.
[0054] Once a similarity matrix is prepared, a number of methods of denoising can beused, including non-negative least squares (NNLS) regression, diffusion, binning, kNNsmoothing, or any combination thereof. Denoising phenotype scores may reduce intra-phenotypic variance, resulting in more precise inference of phenotype.
[0055] While specific examples of processes for yielding a phenotype score aredescribed above, one of ordinary skill in the art can appreciate that various steps of the process can be performed in different orders and that certain steps may be optional. As such, it should be clear that the various steps of the process could be used as appropriate to the requirements of specific applications. Furthermore, any of a variety of processesfor yielding a phenotype score appropriate to the requirements of a given application canbe utilized.Applications of phenotype status and phenotype omic sets
[0056] Various embodiments are also directed towards utilizing phenotype status oromic features sets of a phenotype learned within a module. A number of embodimentsare also directed to identification of biomarkers and / or key features of a phenotype status.In some instances, unique biomarkers (e.g., omic features as identified within aphenotype status module) can be used to identify phenotype of an entity or an omic dataset. In some instances, biomarkers may also have an important biological role that canbe exploited. For example, provided in Table 1 is a list of biomarkers associated with potency phenotypes (i.e., totipotent, pluripotent, multipotent, oligopotent, unipotent, and differentiated) as identified using the computational process described herein. In another example, provided in Table 2 is a list of potency phenotype biomarkers associated particular cancers as identified using the computational processes described herein. The biomarkers can be utilized for identifying particular presence of cells within developing tissues or within particular cancers. Some of these biomarkers may be important for maintaining stemness and / or promoting differentiation and thus manipulation ofexpression of these biomarkers can alter a cell’s ability to maintain or progress from itsdifferentiation status. Likewise, biomarker expression (or lack thereof) may also beimportant for differentiation into a particular cell lineage and thus manipulation of expression of these biomarkers can alter a cell’s maturation and ultimately which type of mature cell is developed. And in some instances, biomarkers can be identified that are specific to particular differentiation trajectories of diseased tissues, which can be exploited to further study, diagnose, and / or treat a particular disease. For example, identifyingbiomarkers specific to highly potent cells within tumors can be exploited to further studyand understand the cells and how to better diagnose and treat various cancers.
[0057] Biomarker identification can be applied to any entity and any phenotype status.Examples of phenotype status include (but are not limited to) potency status, metabolic status, senescence status, spatial status, status related to cell-to-cell contact, status related to a signaling pathway, status related to a stimulation, status related to an immuneresponse, status related to a pathogen, status related to an injury, status related to a medical condition, and status related to a treatment. Utilizing methods described herein, biomarkers can be identified for presence of medical condition, identification of medical condition subtype, responsiveness to a therapy, identification of an optimal therapy, and combinations thereof.
[0058] A binary model as described herein not only identifies biomarkers but alsoprovides predictive power of the biomarker for a phenotype status. The predictive power is indicative of a combination of: the uniqueness of the biomarker to other status categories, commonality of the biomarker shared by entities of the same status, and biomarker signal power. Th predictive power thus facilitates selection of biomarkers for a phenotype status. This greatly improves biomarker identification for several phenotypesthat are difficult to ascertain robust biomarkers. For example, it is often to difficult toidentify biomarkers of a phenotype status across different cell types, such as unipotent status. There are hundreds of cell types that have unipotent status, so finding a shared biomarker is an arduous task. However, the computational method described herein can (and has) found biomarkers for unipotent status and several other potency status (See table 1). Another arduous task is to identify of biomarkers for medical condition subtype or therapy optimization of a medical condition because the affected tissue is seemingly homogenous by classical methods of comparison. The feature importance provided by a binary model as provided herein can select discriminative biomarkers between similar tissues, greatly improving the ability to develop diagnostic assays for identifying medical disorder subtypes, responsive therapies, and optimal therapies. In various embodiments, one or more biomarkers are selected from: a top 0.1% of omic features ranked by feature importance, a top 0.5% of omic features ranked by feature importance, a top 1% of omic features ranked by feature importance, a top 2% of omic features ranked by feature importance, a top 3% of omic features ranked by feature importance, a top 4% of omic features ranked by feature importance, a top 5% of omic features ranked by feature importance, or a top 10% of omic features ranked by feature importance.
[0059] In some embodiments, a diagnostic assay to detect an identified biomarker isdeveloped. Detection assays, in accordance with various embodiments, include detectionof the biomarker in a sample. In some embodiments, RNA expression biomarkers aredetected by hybridization, polymerase chain reaction, or sequencing techniques. Nucleic acid oligomers designed hybridize to sequences of the biomarker for capture, amplification, or sequencing can be synthesized to yield a diagnostic assay. In someembodiments, protein expression biomarkers are detected by immunodetectiontechniques utilizing antibodies that recognize the biomarker, such as western blot,immunostaining, flow cytometry, enzyme-linked immunosorbent assay (ELISA), or dotblot assays. Antigen-binding proteins (e.g., antibodies or antigen binding CDRs derivedfrom an antibody) can be generated techniques to generate polyclonal antibodies (e.g., immunization of antigen within an animal), monoclonal antibodies (e.g., hybridoma cell line generation), and recombination molecular biology (e.g., cloning a CDR sequence). Methylation biomarkers can be detected by assessing certain genomic loci via qPCR or sequencing with bisulfite or enzymatic conversion. Histone biomarkers can be detected at certain genomic loci via chromatin-immunoprecipitation and PCR or sequencing of the pulled down DNA. Chromatin accessibility can be assessed by ATAC-seq, and assessing certain genomic loci within the sequencing data.
[0060] In some embodiments, expression of a gene indicative of a phenotype statusis further examined by modulating its expression. A gene indicative of a phenotype canbe selected based on predictive power within a phenotype status computational moduel. Accordingly, various embodiments are directed to knocking down, knocking out, and / or inhibiting expression of various genes to determine their function in relation to thephenotype. In some particular embodiments, short-hairpin RNAs (shRNAs) or CRISPRconstructs are developed to modulate expression of a particular gene.
[0061] Various embodiments are also directed towards evaluating a disease in anindividual, and / or predicting a clinical outcome of a disease therapy, based on potencyphenotype of a collection of cells. This can be especially useful in assessing cancers anddevelopmental disorders. In regards to cancer, potency phenotype of cancerous cells can inform survival outcomes, cancer aggressiveness, risk of metastasis, risk of recurrence, and therapy response. Diagnostics and treatments can be determined accordingly.Accordingly, potency phenotype can yield insight into disease pathology and / ortherapeutic options to assist a clinician in treatment of an individual.
[0062] A number of embodiments are directed toward methods that include obtaininga biological sample from an individual having a medical condition, and determiningphenotype status of a collection of a biological sample. The biological sample can be bulkor single cell derived from a tissue source or cell-free nucleic acids from a cell-free source. Cell-free sources include (but are not limited to) blood / plasma, cerebral spinal fluid,lymph, saliva, urine, and stool. The phenotype status can be utilized to determine a clinicalphenotype status, such as presence of medical condition, severity or grade of medical condition, subtype of medical condition, prognosis of a medical condition, therapyresponse, optimal therapy, or any combination thereof. Computational modules can betrained to directly learn a clinical phenotype status. Alternatively, some phenotype status (e.g., differentiation status in cancer cells) are clinical indicators when correlative with aclinical phenotype status.
[0063] In some cases, the medical condition is a cancer, which may be any suitablecancer, such as, but not limited to, human sarcomas and carcinomas, e.g., fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, colon carcinoma, pancreatic cancer, breast cancer, ovarian cancer, prostate cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, cystadenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, cervical cancer, testicular tumor, lung carcinoma, small cell lung carcinoma, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, melanoma, neuroblastoma, retinoblastoma; leukemias, e.g., acute lymphocytic leukemia and acute myelocytic leukemia (myeloblastic, promyelocytic, myelomonocytic, monocytic and erythroleukemia); chronic leukemia (chronic myelocytic (granulocytic) leukemia and chronic lymphocytic leukemia); and polycythemia vera, lymphoma (Hodgkin's diseaseand non-Hodgkin's disease), multiple myeloma, Waldenstrom macroglobulinemia, follicular lymphoma and heavy chain disease.
[0064] As described in the Examples and Data Results, it has been found that thecomputational processes described herein were able to identify the differentiation states of cells on absolute scale in several human malignancies. It was found that cancers comprising high number of cells with immature potency (e.g., more multipotent and oligopotent) were correlated with shorter survival time. It was further found that cancers comprising high number of cells with immature potency (e.g., more multipotent and oligopotent) were correlated with resistance to immune checkpoint inhibitors.
[0065] A number of embodiments of the disclosure is directed towards diagnosis andtreatment of cancers. An example of a diagnostic method for informing a treatment is provided below: Assess cancer cells for potency Determine number of cells with immature potency (e.g., multipotent and oligopotent) – which can be determined via assessment of biomarkers or using anabsolute potency learned model If immature potency is above a threshold, treat with cancer therapeutic other than immune checkpoint inhibitors (ICIs) (e.g., chemotherapy, targeted therapy, hormone therapy) If immature potency is below a threshold, treat with ICIs (which can be combined with other therapies) This diagnostic method and treatment can be useful for a number of cancers in which ICIs provide great treatment response in some individuals but not in others. For example,the use of ICIs is common to treat triple negative breast cancer (TNBC) but for a numberof recipients, resistance is high and treatment ultimately does not provide benefit. Identification of these individuals prior to determining treatment may improve their outcome, or at least forgo administration of ICIs that will have little benefit but can have various side effects. Examples of immune checkpoint inhibitors include anti-PD1 antibodies, anti-PDL1 antibodies, anti-CTLA-4 antibodies, and anti-LAG-3 antibodies.
[0066] A number of embodiments of the disclosure is directed towards diagnosistreatment of cancer metastasis. An example of a diagnostic method for metastasis treatment is provided below: Assess cancer cells for potency Determine number of cells with immature potency (e.g., multipotent and oligopotent) – which can be determined via assessment of biomarkers or using anabsolute potency learned model If cells with immature potency is above a threshold, treat cancer for metastasis, even when metastasis is not detected by traditional methods AND / OR – perform further biopsies and scans to identify metastasisThis diagnostic method and treatment can be useful for a number of cancers in which risk of metastasis is high. Cancer cells with immature potency are associated with a greater likelihood of metastasis. Thus, even when traditional methods (e.g., lymph biopsies) do not detect metastasis, the treatment provided is at a level for metastatic cancer because of the increased risk. Generally, metastatic cancer can be treated aggressively with chemotherapy combined with immunotherapy and / or targeted therapy. Nonmetastatic cancer can be treated with chemotherapy or targeted therapy, which can be combined with immunotherapy.
[0067] A number of embodiments of the disclosure is directed towards diagnosis andsurveillance during treatment of cancers. An example of a diagnostic method for surveillance during treatment is provided below: Provide cancer treatment Assess cancer cells for potency over time– which can be determined viaassessment of biomarkers or using an absolute potency learned model Determine number of cells with immature potency (e.g., multipotent and oligopotent) If relative immature potency is increases, increase treatment dosage This diagnostic method and treatment can be useful for a number of cancers in which risk of recurrence is high. Cancer cells with immature potency are associated with a greater likelihood of recurrence. As treatment progresses, overall cancer cells may be decreasingbut if cancer cells with immature potency are remaining, there may be desire to increase dosage amount (as compared to the standard of care based on traditional assessments).
[0068] A number of embodiments of the disclosure is directed surveillance of minimalresidual disease (MRD). An example of a diagnostic method for surveillance during treatment is provided below: Provide cancer treatment Assess cancer cells for potency prior and / or during treatment Determine number of cells with immature potency (e.g., multipotent and oligopotent) – which can be determined via assessment of biomarkers or using anabsolute potency learned model If immature potency is over a threshold, increase the number of diagnostic examinations for minimal residual disease
[0069] This diagnostic method and treatment can be useful for a number of cancers inwhich there is a risk of recurrence. Cancer cells with immature potency are associated with a greater likelihood of recurrence. If immature potency is over a threshold, there is a higher risk of recurrence and thus more diagnostic examinations for minimal residual disease should be performed (as compared to the standard of care).
[0070] A number of embodiments of the disclosure are directed towards diagnosis ofpresence of a medical condition, severity of medical condition, or identification of a medical condition subtype. An example of a diagnostic method for treatment of a medical condition is provided below: Assess a biological sample for biomarkers identified and selected based on predictive power within a binary neural network moduleDetermine the biological sample indicates presence of a medical condition, severity of medical condition, or identification of a medical condition subtype Administer treatment for the medical condition, medical condition severity, ormedical condition subtype
[0071] This diagnostic method and treatment can be useful for variety of medicaldisorders. Generally, the method can train a set of binary neural network modules thatpredict the phenotype status of medical condition presence, medical condition severity,or medical condition subtype utilizing omic data from cohort of individuals having thedisorder. From the binary network, the omic features for classifying the medical disorder can be learned. Biomarkers can be selected from the omic features based on feature importance. Selected biomarkers can then be utilized to assess a biological sample of an individual to determine whether the biological sample indicates a medical condition, amedical condition severity, or a medical condition subtype. Based on the indication, theindividual can be treated with a therapy, such as administration of a therapeutic for the disorder.
[0072] A number of embodiments of the disclosure are directed towards determiningresponse of a treatment or an optimal treatment. An example of a diagnostic method for treatment of a medical condition is provided below: Assess a biological sample for biomarkers identified and selected based on predictive power within a binary neural network module Determine the biological sample indicates response to a treatment or an optimal treatment Administer a treatment in accordance with treatment response or optimal treatment
[0073] This diagnostic method and treatment can be useful for variety of treatments ofmedical disorders. Generally, the method can train a set of binary neural network modules that predict the phenotype status of responsiveness to a treatment or an optimal treatment among 2 or more treatments utilizing omic data from cohort of individuals having been treated. From the binary network, the omic features for classifying the treatment response status or treatment response can be learned. Biomarkers can be selected from the omic features based on feature importance. Selected biomarkers can then be utilized to assess a biological sample of an individual to determine whether the biological sample indicates the individual will respond to a treatment or which treatment is optimal for the individual. Based on the indication, the individual can be treated with the indicated treatment, such as administration of a therapeutic for a medical disorder.Systems of potency phenotype status
[0074] Figure 4 provides an example of a computational system 401, which may beimplemented on a single or a plurality of computing device(s). Computational system 401may be a personal computer, a laptop computer, any other computing device withsufficient processing power, or any plurality and / or combination of computing devices forthe processes described herein. Computational system 401 can include a processor 403,which may refer to one or more devices within the computing systems 401 that can beconfigured to perform computations via machine readable instructions stored within amemory 407 of computational system 401. The processor may include one or moremicroprocessors (CPUs), one or more graphics processing units (GPUs), and / or one or more digital signal processors (DSPs).
[0075] Memory 407 may contain one or more applications for performing the variouscomputational process described herein. Accordingly, in some implementations, memory 407 comprises a omic Data Standardization 409, and in some implementations, memory407 comprises and a Phenotype Application 411. As an example, processor 403 mayperform a process as described herein, during which memory 407 may be used to storevarious input data, intermediate datea, and / or output data such as gomic data 409a,harmonized omic feature lists 409b, ranked omic features X entities matrices 409c,phenotype modules 411a, denoising data 411b, learned feature sets 411c and phenotype scores 411d.
[0076] Computational system 401 may include an input / output interface 405 that canbe utilized to communicate with a variety of devices, including but not limited to other computing systems, a projector, and / or other display devices. As can be readily appreciated, a variety of software architectures can be utilized to implement a computer system as appropriate to the requirements of specific applications.
[0077] Provided in Fig. 5 is an example of module for yielding a phenotype likelihoodscore from ranked space omic data. Phenotype module 502 comprises binary neuralnetwork (BNN) layers 504 and a feed forward layers 506. As input to the phenotypemodule 502, ranked space omic data 501 with an assignable phenotype status is utilized by BNN layers 504 to learn omic feature sets and a binary weights to yield a learned omicfeature binary weight matrix 505. Ranked space omic data 501 can be trimmed to reducethe number of features to be assessed to yield trimmed rank space omic data 503, removing unnecessary features, as may be determinable by the learned omic featuresets. To perform assessment, ranked omic data 501 or trimmed ranked space omic data503 is applied to the learned omic feature binary matrix to yield binarized omic feature setenrichment scores 507, which is utilized as input for the feedforward layers 506. Binarizedomic feature set enrichment scores are applied to omic feature set enrichment scoreweights to yield a phenotype likelihood score 511.
[0078] A computational framework can yield an overall phenotype assessment bycombining two or more phenotype modules 502 to. Accordingly, a phenotype likelihood score 511 of two or more modules are combined. Any combination can be performed. In some implementations, each phenotype likelihood score 511 is a vector and thus each vector can be combined to yield an overall vector for the overall phenotype assessment. In scenarios in which the two or more phenotypes of the modules yields a scalar phenotype (e.g., any phenotype on a spectrum), an overall phenotype score can be placed on a scale to yield a phenotype assessment. Further, an entity can be classified as having a particular phenotype. EXAMPLES AND DATA RESULTS
[0079] The embodiments of the disclosure will be better understood with the severalexamples and data results provided in the attached addendum, which describes specificimplementations of systems and computational methods for assessing cell potency fromscRNA-seq data. Described therein is implementations of a computational frameworkreferred to as CytoTRACE 2 that yielded insight into the molecular programs forassociated with several potency phenotypes and further outperformed several othercomputational processes for assessing cell potency. Mapping single-cell developmental potential in health and disease with interpretable deep learning
[0080] Several computational models have been developed to map developmentalpotential in a populations of cells. CytoTRACE 1 (G. S. Gulati, et al., Science. 2020 Jan24;367(6476):405-411, the disclosure of which is incorporated herein by reference)–likeother methods for trajectory inference, including scVelo (V. Bergen, et al., Nat Biotechnol.2020 Dec;38(12):1408-1414), CellRank (M. Lange, et al., Nat Methods. 2022Feb;19(2):159-170), and Monocle 3 (J. Cao, et al., Nature. 2019 Feb;566(7745):496-502)–predicts single-cell orderings in a manner that is relative to each dataset. This has made it difficult to unify such predictions across datasets and contextualize them against the backdrop of cellular potency.
[0081] To overcome these challenges, we developed CytoTRACE 2, a specificimplementation of an interpretable deep learning framework for jointly determiningphenotype categories from omic data. CytoTRACE 2 was developed to determine single-cell potency categories and absolute developmental potential from scRNA-seq data.Unlike most deep learning methods, which are not inherently explainable, CytoTRACE 2implements a novel architecture that learns multivariate omic signatures (e.g., geneexpression programs) for each phenotype status of interest. Such programs can beseamlessly extracted from the model, are readily interpretable, and deliver remarkably accurate predictions of developmental potential. To demonstrate the utility of CytoTRACE 2, we curated human and mouse scRNA-seq datasets with gold standard potency levels and differentiation trajectories. We then trained, validated, and benchmarked the performance of our approach, explored its potential for reconstructing mouse developmental ontogeny, and used it to uncover candidate differentiation programs in cancer. Our results illuminate pan-tissue determinants of cell potency and highlight the value of interpretable deep learning for characterizing single-cell developmental states in health and disease. ResultsA single-cell atlas of developmental potency
[0082] CytoTRACE 2 was designed to provide a unique view of developmentalpotential. Unlike other methods, it predicts single-cell potency categories on an absolute scale using classically defined developmental stages as “anchor points”. It then producesa continuous potency score ranging from the fertilized egg to the least primitive cells,enabling high-resolution assessment of cellular differentiation status from scRNA-seq data.
[0083] To achieve these goals, we first established an annotation scheme consistingof six broad potency categories spanning the full range of cellular ontogeny: totipotentcells capable of generating an entire multicellular organism; pluripotent cells with thecapacity to differentiate into all adult cell types; immature multipotent, oligopotent, and unipotent cells, each generally capable of producing >3, 2 or 3, or 1 downstream cell type(s), respectively, and differentiated cells, ranging from mitotic to post-mitotic phenotypes (Figure 6A).
[0084] We then screened online repositories for scRNA-seq data with assignablepotency levels from humans and mice, for which extensive genomic data are available. Following curation, we identified 33 gold standard datasets spanning nine scRNA-seq platforms, 406k cells, and 125 standardized cell phenotypes, each with experimentallyconfirmed potency levels (Figure 7A). To distinguish fine-grained developmentalintermediates (e.g., hematopoietic stem cells and their immediate progeny, both of whichare multipotent), we also subdivided the six categories into 24 granular potency levelsusing evidence from lineage tracing, transplantation, colony formation, and related experiments (Figure 7A). From these data, we selected a training set consisting of 93 cell phenotypes (78 mouse and 32 human), 16 tissue types, 13 studies, and six scRNA-seq platforms (Figure 6B). The remaining datasets were used to assess performance. Modeling potency with interpretable AI
[0085] Like its predecessor, CytoTRACE 2 was designed with an emphasis ondecoding the molecular determinants of developmental potential. However, since most deep learning methods require dedicated procedures to ascertain feature importance, we engineered CytoTRACE 2 with a novel, fully explainable architecture for single-cell classification tasks. Drawing inspiration from binarized neural networks, where weightstake values of 1 or –1, our approach, termed a Gene Set Binary Network (GSBN),automatically activates (=1) or deactivates (=0) individual genes, with the goal of identifying highly discriminative gene sets that best describe each phenotype of interest(e.g., potency levels). Multiple gene sets can be learned per phenotype, enhancingflexibility, and informative genes can be easily extracted from the model. As such, unlike most deep learning architectures, GSBNs can yield immediate insight into the molecular programs underpinning performance.
[0086] We leveraged this framework to model the six broad potency categories directly(Figures 6C and 7B). We then devised an approach to integrate predictions acrossGSBNs to produce a continuous output. In doing so, we hypothesized that we could exploit the natural ordering of potency levels to impute fine-grained developmental intermediates. As such, CytoTRACE 2 yields two key outputs for each single-cell transcriptome: (i) the potency category with maximum likelihood and (ii) a continuous ‘potency score’ calibrated to cellular ontogeny that ranges from 1 (totipotent) to 0(differentiated) (Figures 6C and 7C). Based on the assumption that transcriptionallysimilar cells occupy related differentiation states, CytoTRACE 2 also leverages a nearestneighbor approach to smooth individual potency scores (Figures 7C and 7D). With thisdesign, CytoTRACE 2 blends features of cell type classification and trajectory inferencewhile precisely annotating cells and pinpointing the genes underlying its predictions.Technical assessment of CytoTRACE 2 in gold standard datasets
[0087] Having established a large compendium of gold standard data, we nextcharacterized the performance of CytoTRACE 2. To this end, we calculated both the correctness of potency predictions and the ordering of known trajectories. For the latter, we implemented two definitions of developmental ordering: ‘absolute order’, in which predictions were analyzed in aggregate across datasets with respect to known potencylevels, and ‘relative order’, in which cells were evaluated on a relative scale – from theleast to most differentiated within each dataset (Figure 7E). In all cases, we quantified the agreement between known and predicted developmental orderings using weighted Kendall correlation. This allowed us to assess concordance between ranks while balancing key variables to minimize bias.
[0088] We began by evaluating model hyperparameters using a cross-validationscheme. Across a wide range of values, we observed only minimal variation inperformance (Figures 7F and 7G). As such, we selected robust hyperparameter valuesto retrain the model and employed leave-one-out cross-validation as a first step toward evaluating it.
[0089] On the training set, CytoTRACE 2 delivered remarkably high accuracy fordiscriminating potency (Figure 6D, left), immediately distinguishing it from pseudotime methods that infer relative orderings. This was true regardless of whether we analyzedbroad (n = 6) or granular potency levels (n = 23 evaluable stages) (Figure 6D), the latterof which far exceeded the resolution on which the model was trained. In line with this,when evaluating a canonical multi-branching process – mouse bone marrowhematopoiesis – potency scores not only tracked with ground truth differentiation statesin a continuous manner but also distinguished multilineage precursors from their unipotent and mature progeny (Figure 6E). Thus, CytoTRACE 2 can impute fine-grained developmental hierarchies on an absolute scale.
[0090] To validate our approach, we next extended our analysis to fully unseen data.We first examined a test cohort comprised of 14 held-out datasets encompassing nine tissue systems, seven platforms, and 94k evaluable cells (Figure 8A). Across diverse metrics, CytoTRACE 2 exhibited nearly identical performance to the training set and was not significantly affected by differences in species, tissue type, or platform (Figures 8B,9A and 9B). Moreover, as before, CytoTRACE 2 successfully resolved highly granularpotency levels (n = 17 evaluable stages) (Figure 8B, right), with no discernible loss ofperformance on phenotypes that were absent during training (Figure 9C). Generalizability of learned potency representations
[0091] To more deeply study generalizability, we next performed three additionalexperiments. First, we re-trained CytoTRACE 2 on alternative subsets of our potency atlas, including scenarios in which distinct cell-type clades were serially held out fromtraining (e.g., immune cells, neural cells, endothelial cells, bone cells, etc.). In all cases,results were well-correlated with ground truth, implying that potency-related biology isstrikingly conserved (Figures 9D, 9E, and 9F). Next, we assessed single-cell chromatinaccessibility data as an orthogonal readout of cell potency. This allowed us to explore theuniversality of potency representations without relying on our annotation scheme. On multiome datasets spanning human hematopoiesis, human fetal pancreas development,and mouse intestinal epithelium, CytoTRACE 2 was significantly correlated with thedegree of open chromatin in diverse developmental phenotypes (Figure 9G). Moreover, it was ~4-fold more consistent, on average, with known hierarchies, highlighting thebenefit of modeling cell potency through deep learning (Figure 9H; Table 9F).
[0092] We next analyzed a time series experiment of mouse bone marrowhematopoiesis profiled by the lineage and RNA recovery (LARRY) heritable clonalbarcoding system (Figure 8C). This allowed us to determine whether CytoTRACE 2 can forecast cell fate potential in a setting where cell fate is known. Focusing on cells with shared barcodes across timepoints, CytoTRACE 2 predictions were highly aligned with expectation, with cells predicted as multipotent, oligopotent, or unipotent each linked to amedian of 4, 2, or 1 downstream phenotype(s), respectively (Figures 8C and 8D). Cellspredicted as differentiated were also associated with 1 downstream phenotype, on average, consistent with cells arising from a common progenitor (Figure 8D). Thus, CytoTRACE 2 can effectively model developmental state in a lineage tracing system. Robustness of CytoTRACE 2
[0093] Having established key performance characteristics, we next sought todetermine the impact of technical factors on CytoTRACE 2. We first asked whether CytoTRACE 2 is robust to potential annotation errors in the training set. Using a strategyphenotypes were incorrectly annotated (Figures 10A, 10B, and 10C). We next wonderedwhether CytoTRACE 2 is robust to sparsity and variation in mRNA content. In evaluable test datasets, CytoTRACE 2 predictions were virtually identical down to (i) 750 detectably expressed genes per cell (Figure 10D), (ii) 2,000 unique molecular identifiers per cell (Figure 10E), and (iii) ~5 cells per phenotype (Figure 10F). Thus, CytoTRACE 2 is resistant to moderate annotation errors and broadly applicable when practically obtainable count thresholds are met. Benchmarking of CytoTRACE 2
[0094] To systematically benchmark our approach, we next evaluated cell potencyclassification performance. We hypothesized that CytoTRACE 2 would be especially advantageous in this setting owing to dedicated features that mitigate overfitting, including the use of dropout during training and binary components that encourage model sparsity. In support of our hypothesis, CytoTRACE 2 achieved markedly higher accuracy (median multi-class F1 score of nearly 0.7) and lower mean absolute error than eight competingmachine learning methods, each trained and tested on the same data as CytoTRACE 2(Figures 11A and 11B).
[0095] We then considered eight broadly applicable methods for inferring high-resolution developmental hierarchies, but not potency labels, from scRNA-seq data. To extend our analysis beyond our test cohort, we included Tabula Sapiens, a multi-donor scRNA-seq atlas covering nearly 500k postmortem cells from 24 human tissue types. Following curation and ground truth potency annotation, we quantified performance at the single-cell level using relative and absolute order as described above.
[0096] Overall, CytoTRACE 2 outperformed all competing measures, in most casesby a substantial margin. Indeed, when applied to reconstruct relative orderings of 57 systems in which ground truth is known, CytoTRACE 2 achieved a median correlationover 60% higher than the second leading approach – CytoTRACE 1 (Figure 8E). Whenapplied to reconstruct absolute developmental orderings, the boost in performance was similarly pronounced (Figure 8F). Comparable findings were obtained when comparing CytoTRACE 2 against 18,706 annotated gene sets (Figure 8F) and scVelo, a generalizedRNA velocity model for predicting future cell states (Figures 11C and 11D).
[0097] Thus, these data further validate CytoTRACE 2, establish its extensibility tonew data, and emphasize its promise for revealing differentiation states and biological insights that are unattainable by existing in silico methods and gene signatures.Advantages of absolute developmental potential from the zygote to birth
[0098] The ability to profile absolute developmental potential has immediateadvantages over conventional single-cell analysis strategies. For example, cell potency can vary dramatically within and across biospecimens, datasets, and developmental systems of interest, confounding their interpretation. CytoTRACE 2 is the first method capable of meaningfully extracting this key parameter from single-cell data, setting it apart from other methods, including its predecessor. For example, when applied to a diverse array of held-out datasets (Figure 12A), CytoTRACE 2 readily (i) corroborated upregulation of a pluripotency program in cranial neural crest cell precursors; (ii) delineated multilineage branchpoints and their developmental transitions in human and mouse tissues; and (iii) correctly distinguished datasets with and without immaturepopulations (Figure 12A, top and center). In contrast, CytoTRACE 1 entirely failed toappreciate these differences, leading to misleading results (Figure 12A, bottom).
[0099] CytoTRACE 2 can also interrogate potency dynamics in vitro, including in time-resolved cellular reprogramming and differentiation experiments. For example, it successfully traced the path to induced pluripotent stem cell (iPSC) generation from mouse embryonic fibroblasts (Figure 12B), verified stark differences in iPSC reprogramming efficiency between protocols (Figure 12C), and monitored the loss ofpotency during directed differentiation of iPSCs into multipotent endoderm (Figure 11E).CytoTRACE 2 can thus serve as a “scorecard” to optimize somatic cell reprogramming, directed differentiation, and cellular engineering experiments where the generation of cells with specific potency levels is required. Large-scale integrative analysis of cell potency
[0100] In recent years, tissue and organ-level cell atlases have generated millions ofsingle-cell expression profiles; however, many biological systems are only partiallycovered by individual datasets. Given that CytoTRACE 2 is anchored to a predefineddevelopmental range, we next tested whether it can perform large-scale integrativeanalysis of cell potency across datasets – without co-embedding or batch correction.
[0101] For this purpose, we focused on mouse embryogenesis, with the goal ofcharting single-cell potency over the full prenatal period. Notably, no existing atlas covers this entire process at daily intervals. To achieve continuity, we assembled six timelapse datasets consisting of >11M mouse single-cell transcriptomes, 62 developmental time points, and five platforms, including traditional plate-based methods, droplet sequencing, and combinatorial index sequencing of single-cell nuclei. Together, these longitudinal data encompass every embryonic day from E0.5 (zygote) through P0 (birth) (Figure 13A, bottom). We then applied CytoTRACE 2 to each dataset individually and aggregated the resulting potency scores, both by author-supplied phenotypes and major embryonic time points (Figure 13A). Mean potency scores produced by CytoTRACE 2 were highlycorrelated with developmental time (absolute Kendall correlation = 0.95, P = 1.1 × 10–24,Figure 13B). Moreover, as expected, overall potency levels trackedin a graded fashion, with CytoTRACE 2 documenting the transition from totipotent to pluripotent cells during pre-gastrulation; from pluripotent to multipotent cells duringgastrulation; and from multipotent to differentiated states during organogenesis (Figure 13A).
[0102] To benchmark this result, we applied eight previous methods to each of thesedatasets. CytoTRACE 2 showed performance gains over competing methods for concordance with developmental time (Figure 13C). This was evident before and duringorganogenesis where both dramatic and subtler changes in median potency levels wererespectively observed (Figure 13C).
[0103] As another form of validation, we analyzed a data-driven lineage tree of mouseembryogenesis. Notably, this tree, which covers 258 evaluable phenotypes arising from the zygote (root), was inferred using a heuristic strategy based on single-cell transcriptional covariance. In contrast to CytoTRACE 1, which exemplifies methods that produce a relative ordering in each dataset (e.g., Monocle, RNA velocity), CytoTRACE 2effectively captured the progressive drop in potency across developmental lineagesindependent of source dataset (Figures 13D, center and right, and 13E). Moreover, despite the inclusion of several immature cell types as terminal nodes (e.g., owing to sampling bias), CytoTRACE 2 was remarkably well-correlated with the number of downstream lineages per phenotype, further validating its output (Figure 13F). Potency-related programs determined by interpretable AI
[0104] Previous genomic surveys of stemness have largely focused on molecularfeatures of pluripotent stem cells arising from the inner cell mass or somatic cellreprogramming. In comparison, relatively few studies have examined conserved gene expression signatures of other potency levels, and none have characterized such programs at the granularity analyzed in this work. Thus, given the inherent interpretability of our modular GSBN design, we sought to bridge this gap and nominate new hypotheses.
[0105] We began by visualizing CytoTRACE 2’s internal representation of geneexpression signatures (Figure 14A, left). We first compiled results from our earlierbenchmarking analysis of 33 gold standard datasets (Figure 7A). We then converted gene set activity levels from the GSBN modules into a low dimensional embedding (Figure14B). Despite analyzing cells from 27 studies and 18 tissue types, predicted differentiationstates formed a cohesive gradient spanning the breadth of cellular ontogeny (Figure 14B).This gradient was faithful to ground truth potency categories (Figure 5B), independent of training or test sets, and devoid of obvious batch effects (Figure 15A). Moreover, within each potency category, gene sets were highly discriminatory in training and test sets (Figure 14C) and showed notable overlap in constituent genes after accounting for the polarity of enrichment scores (Figure 15B).
[0106] To dissect the biology underlying these programs, we next created six rankedgene lists, one for each broad potency level, summarizing the relative importance of eachgene to the model (Figure 14A, right; Table 10A). This allowed us to aggregate the top-ranking genes and interrogate their specificity and functional themes (Figure 15C). Remarkably, the resulting signatures were highly conserved across species, platforms,and developmental clades, revealing both positive and negative correlates of eachpotency level (Figure 14D).
[0107] Given this strong degree of conservation, we hypothesized that CytoTRACE 2might enrich for factors that are essential for potency maintenance. Indeed, the coretranscription factors Pou5f1 (encoding OCT4) and Nanog were each ranked within thetop 0.2% of pluripotency genes prioritized by the model. To further explore this hypothesis, we selected multipotency as a key developmental state with poorly understood molecular programs. We then queried our results against a large-scale CRISPR screen, in which ~7,000 genes in mouse hematopoietic stem cells (HSCs) were individually knocked out and assessed for developmental consequences in vivo (Figure14E). After intersecting these genes with the CytoTRACE 2 feature space, we rank-ordered the resulting knockout genes (n = 5,757) by their ability to either promote or inhibitHSC differentiation into mature myeloid or lymphoid progeny.
[0108] Strikingly, the top 100 positive markers of multipotency identified byCytoTRACE 2 were significantly skewed toward factors that promote HSC differentiationwhen knocked out (Q = 0.04; Figure 14F, top left). Conversely, the top 100 negativemarkers were skewed toward factors that inhibit differentiation when knocked out (Q =0.04; Figure 14F, bottom left). These results, which support our hypothesis, were both robust to variation in the number of evaluated markers (50, 100, 200, 500) and specific togenes associated with multipotency by the model (Figure 15D). As a control, we repeatedthis experiment using markers of multipotent cells obtained by standard differentialexpression analysis. Whether applied to expression profiles of mouse hematopoietic subsets or gold standard training datasets from our curated compendium, results were indistinguishable from random chance, both when considering the top 100 control markers and median enrichment scores across various gene set sizes (Figures 14F, right,and 15E). Thus, given the interpretability, expressiveness, and performancecharacteristics of CytoTRACE 2, we could readily identify critical potency-related genes that are otherwise latent or challenging to extract. CytoTRACE 2 implicates unsaturated fatty acid synthesis in multipotency
[0109] The development of CytoTRACE 2 also provided an opportunity to investigategeneralizable biomarkers of developmental stage. While such factors have been largelylimited to pluripotency in prior literature (e.g., Nanog, Pou5f1), with CytoTRACE 2, it wasstraightforward to identify candidate markers for all evaluable potency categories (Figures15F and 15G). This was particularly notable for multipotent and differentiated cells, forwhich over 25% of the top 100 genes per category were shared by at least threedevelopmental systems, including unseen tissues, in held-out data (Figures 15F and15G).
[0110] Among these genes, Fads1 – a factor involved in unsaturated fatty acid (UFA)synthesis – emerged as the top candidate determinant of multipotency in our atlas,showing preferential expression in 7 of 8 evaluated tissues. Consistent with this, cholesterol metabolism was a leading pathway among multipotency-associated genesidentified by CytoTRACE 2 (Q = 0.001; Figure 16A). Previous studies have establishednumerous connections between fatty acid metabolism and stem cell biology, with theformer playing important roles in energy balance, signal transduction, niche interactions,and differentiation cues. However, no studies have unambiguously attributed lipidmetabolism genes to individual potency levels. As such, we decided to further investigateFads1 along with two other players in UFA biosynthesis and cholesterol metabolism thatwere enriched in multipotent cells: Fads2 and Scd2 (Figure 16B).
[0111] We began by verifying the specificity and cross-tissue representation of thesethree genes by single-cell transcriptomics. Indeed, by quantifying their joint enrichmentacross 125 phenotypes in our potency atlas, Fads1, Fads2, and Scd2 were significantlyhigher in multipotent cells (Figure 16C). This was true across training and test datasets, with area under the curve (AUC) values of 0.87 and 0.92, respectively, when balancing by tissue type (Figure 16C).
[0112] Next, to rule out potential artifacts, such as those arising from tissuedissociation or single-cell sequencing, we characterized their expression experimentally. We did this in two ways: (i) by quantitative PCR of mouse hematopoietic cells sorted intomultipotent, oligopotent, and differentiated subsets, and (ii) by multiplexed in situ mRNAimaging of mouse intestinal epithelium. In both cases, we observed preferentialexpression of Fads1, Fads2, and Scd2 in multipotent cells (Figures 16D, 16E, and 17A-17E). Moreover, within the intestinal epithelium, we found that mRNA transcripts of these genes not only colocalized with canonical multipotent Lgr5+cells in the lower crypt but also colocalized with cells in the upper crypt expressing Fgfbp1, a recently reportedmarker of multipotent intestinal stem cells (Figures 16E, 17C, 17D, and 17E).
[0113] Thus, these data strongly corroborate our in silico predictions and nominatespecific UFA factors as cross-tissue biomarkers of multipotent cells. Future studies willbe needed to verify these findings and elucidate the specific roles of Fads1, Fads2, andScd2 in multipotent biology, including whether they underpin – or are simply a byproductof – developmental state. Moreover, functional enrichment data, candidate biomarkers,and putative transcription factors enriched in potency are provided in the supplement(Figures 15C, 15F, and 15G).Single-cell developmental states across human malignancies
[0114] In cancer, stem cell regulatory programs are strongly implicated intumorigenesis, therapy resistance, immune evasion, and metastasis. However, deciphering cancer cell differentiation states with single-cell genomics has remained challenging, especially when combining results across different samples, tumor types, and platforms.
[0115] To determine whether CytoTRACE 2 is applicable in this setting, we began byanalyzing malignant cells from a lineage tracing system in which evolutionary relationships are known. To this end, we leveraged published data from the Kras;Trp53(KP)-driven lung adenocarcinoma mouse model (termed KP-tracer), in whichevolving lineage barcodes and transcriptomes were profiled from the same cells. As partof prior analysis, distinct cell populations were identified in each tumor sample with eitherlow or high clonal diversity – defined as cells that have, or have not, preferentiallyundergone clonal expansion, respectively. Based on these data, the authors proposed a model in which a population with high clonal diversity generally appears earlier in tumor evolution and gives rise to populations with low clonal diversity through clonal expansions.
[0116] Given these data, and notwithstanding the possibility that de-differentiation canoccur, we hypothesized that cancer cells with higher clonal diversity in this system preferentially exhibit higher potency. Indeed, after applying stringent selection criteria to identify evaluable tumors with both high and low clonal diversity populations (Figure 19A), we found that CytoTRACE 2 scores were significantly elevated in cells from high clonaldiversity populations (Figures 18A and 18B). Moreover, this result was independent ofmutation load as inferred by copy number analysis (Figure 19B), implying that potency scores capture an orthogonal axis of cellular diversity.
[0117] To extend our analysis to human malignancies, we next applied CytoTRACE 2to 430,332 previously annotated malignant cells. These data span 17 cancer types and43 scRNA-seq datasets covering 555 patients with carcinoma, melanoma, sarcoma, blood cancer, or brain cancer (Figure 18C, left). When median-aggregated across cancer types, over half of malignant cells were predicted to be differentiated (55%), with unipotent-like, oligopotent-like, and multipotent-like cells each encompassing 15%, 11%, and 3% of malignant cells, respectively. Pluripotent-like and totipotent-like phenotypes were either extremely rare (0.008% on average) or not detected, respectively, consistent with the absence of teratomas in these samples (Figure 19C). As observed for the KP- tracer model, potency classifications were largely independent of inferred copy number profiles, in line with epigenetic alterations as a major source of phenotypic variation (Figure 19D).
[0118] To check our results against baseline malignancies in which cellular hierarchieshave been established, we focused on two cancer types, acute myeloid leukemia (AML)and oligodendroglioma. In AML, markers of predicted potency levels were well-validated by published signatures of leukemic stem cells and their progeny (Figure 19E). In oligodendroglioma, stem-like cells, but not other cells, were predicted to have multi-lineage potential (oligopotent with a minor multipotent-like population) (Figure 18D). This aligns with the prevailing model in which oligodendrocyte-like (OC-like) and astrocyte-like (AC-like) lineages arise from stem-like precursors.
[0119] Having evaluated CytoTRACE 2 on baseline cases, we next used single-cellpotency predictions to define potency-associated markers for each cancer type. Aftervalidating the specificity of marker genes using pseudo-bulk samples (Figures 19F to19H), we queried their expression in bulk tumor transcriptomes. We then leveraged available clinical annotations to link potency markers to survival time.
[0120] In AML and oligodendroglioma, potency markers were strongly associated withpatient outcomes, with elevated signatures of the least and most mature cells presaging shorter and longer survival times, respectively. Moreover, in multivariable survivalmodels, CytoTRACE 2 markers of less mature cells far surpassed published signaturesof stemness in predicting shorter survival time, whether in AML (‘multipotent-like’) or oligodendroglioma (‘oligopotent-like’, Figure 18E). This suggested that CytoTRACE 2 predictions might have broad clinical relevance in human neoplasms.
[0121] To systematically test this, we collected 16.2k bulk tumor expression profilesfrom matching cancer types with overall survival data from The Cancer Genome Atlas (TCGA) and PRECOG (Figure 18C, right). Again, we observed highly significant survivalassociations across malignancies, with elevated multipotent- and oligopotent-likesignatures forecasting the poorest outcomes (Figure 18F). Importantly, such signatures were also considerably more prognostic than those learned from normal developmental systems, highlighting the value of defining potency-associated cellular states de novo with CytoTRACE 2.
[0122] We next assessed cancer-specific potency profiles using multivariableanalysis. When evaluated across tumor types, the majority of potency profiles were independently prognostic, both with respect to each other and in relation to proliferation genes, pluripotency signatures, stage, age, sex, tumor purity, tumor grade, mRNAsi (amachine learning-based pluripotency score), and TmS, a recently introduced measure oftumor RNA content. Other programs typically elevated in cancer, such as epithelial- mesenchymal transition (EMT), MYC-related processes, and recently described pan-cancer expression modules, were also insufficient to fully explain our findings, pointing toward distinct biology (Figure 19I). Developmental determinants of response to cancer therapy
[0123] Given increasing evidence that stem cell and plasticity programs contribute todrug resistance, we reasoned that potency predictions in cancer might associate withtherapy response. Toward this end, we explored several preliminary scenarios involving liquid and solid malignancies, multiple therapy types, and in vitro and in vivo settings. For example, in AML, a previously described 7-gene signature linked to cellular immaturity and differential response to 105 drugs (LinClass-7) was most enriched in multipotent-like AML cells, consistent with expectation (Figure 19J). In oligodendroglioma, CytoTRACE 2 confirmed a global reduction of stemness in malignant cells following treatment with mutant IDH inhibition, outperforming author-supplied gene signatures (Figure 19K).
[0124] In scRNA-seq profiles of melanoma, including new data generated in this work,signatures of immature malignant cells identified by CytoTRACE 2 were significantly higher in tumors with both previous and future resistance to immune checkpoint inhibitors (ICIs) (Figure 18G). These findings, which were not a surrogate for intra-tumoral HLAlevels (Figure 19L), were recapitulated in a pan-cancer analysis of bulk RNA-seq datacovering six cancer types, including 597 pre- and post-treatment human tumors (Figure18H) and 135 mouse tumor models (Figure 19M), treated with either single orcombination ICI therapy.
[0125] Collectively, these data further validate CytoTRACE 2, demonstrate its abilityto identify candidate differentiation states from poorly understood developmental systems, and provide a platform for the identification of novel diagnostics and therapeutic targets. Experimental Methods Human patient samples
[0126] All patient samples included in this study were collected with informed consentfor research use and were approved by the Washington University School of Medicine Institutional Review Board in accordance with the Declaration of Helsinki (2013). ThescRNA-seq cohort used in this analysis consisted of tumor tissue specimens resected from five patients with metastatic melanoma. Clinical classification of durable clinical benefit (DCB) versus no durable benefit (NDB) reflects the patient’s disease response at six months after immunotherapy and surgery and was determined by a board-certified radiation oncologist. Mice
[0127] C57BL / 6 mice were purchased from Jackson Laboratories and housed in theStanford Animal Facility. For all analyses shown in Figures 16D, 16E, and 17A-17E, 8- to12-week-old mice were used, with equal numbers of males and females. All animal procedures were conducted according to a protocol approved by the Stanford University APLAC committee (10868). Mice were maintained in-house under aseptic sterile conditions and supplied with autoclaved food and water. Flow Cytometry
[0128] For the analyses presented in Figures 16D, 17A, and 17B, mousehematopoietic stem cells (HSCs) and multipotent progenitors (MPPs) (cKit+ Lin– Sca-1+,termed ‘KLS’), common myeloid progenitors (CMPs) (cKit+ Lin– Sca1lo / – CD34med / hiCD16 / 32lo / –), and common lymphoid progenitors (CLPs) (cKitlo Lin– Sca1lo CD135+CD127+) were isolated as described previously1 (Figure S6A). In brief, hips, femurs, tibia, and humeri were harvested from C57BL / 6 mice. Bones were cleaned, cut, and flushed with a syringe filled with ice-cold FACS buffer (2% fetal bovine serum [FBS] in Hanks’ Balanced Salt Solution [HBSS] buffer). Cells in FACS buffer were filtered through a 40 ated in ammonium-chloride-potassium (ACK) lysis buffer for 5 minutes on ice. Cells were then spun down and resuspended in 400 μl FACS buffer per mouse. Lineage depletion beads (Miltenyi Biotec 130-110-470) were added to the cells (50 μl per mouse) and incubated for 10 min at 4°C. After incubation, the cells were loaded onto an LS magnetic separation column (Miltenyi Biotec 130-042-401), which was subsequently washed with 3 × 3 mL of FACS buffer. Before and after washing, pass-through cells were collected, spun down, and resuspended in FACS buffer. For the isolation of KLS and CMP cells, the following antibodies were used: anti-mouse lineagecocktail-A700 (BioLegend 133313, 5 μl per mouse), anti-CD117 (c-Kit)-BV395 (Thermo Fisher Scientific 363-1171-80, 1:100), anti-Sca1-BV605 (BioLegend 108133, 1:100), anti- CD34-eFluor 450 (Thermo Fisher Scientific 48-0341-80, 1:40), and anti-CD16 / 32-BV711 (BD Biosciences 740659, 1:100). Following the addition of the anti-CD34 antibody, cells were incubated on ice for 45 min before adding the remaining antibodies. The cells were then incubated with the remaining antibodies for an additional 20 min on ice, followed by washing and FACS analysis. For the isolation of CLP cells, the following antibodies were used: anti-mouse lineage cocktail-A700 (BioLegend 133313, 5 μl per mouse), anti-CD117 (c-Kit)-BV395 (Thermo Fisher Scientific 363-1171-80, 1:100), anti-Sca1-BV605 (BioLegend 108133, 1:100), anti-CD135-BV421 (BioLegend 135313, 1:100), and anti-CD127 (IL- -BV711 (BioLegend 135035, 1:100). The cells were incubated with theantibodies for 20 min followed by washing and FACS analysis. Flow cytometry was performed with a 100 μM nozzle on a BD FACSAria II using FACSDiva software.
[0129] Blood samples were collected from the same mice for the isolation of CD8a+ Tcells (CD3+ CD8a+) and B (CD19+) cells. Peripheral blood mononuclear cell (PBMC) isolation was performed using a SepMate™-15 tube (STEMCELL technologies 85415) according to the manufacturer's instructions. Enriched PBMCs were resuspended in FACS buffer and incubated with either T cell antibodies (anti-CD3-BV711, BioLegend 100241; anti-CD8a-BV605, BioLegend 100743) or B cell antibodies (anti-CD19-BV605, BioLegend 115539) on ice for 20 min. The cells were then washed with FACS buffer and analyzed on a BD FACSAria II using FACSDiva software. RNA isolation and real-time PCR
[0130] For the analyses presented in Figures 16D and 17B, 20,000 sorted cells fromeach bone marrow and blood population noted in “Flow cytometry” were lysed in RNA lysis buffer (RLT) and subjected to RNA extraction using the RNeasy Plus Micro Kit (Qiagen 74034). RNA was then reverse transcribed into cDNA with SuperScript III FirstStrand Synthesis kit (Thermo Fisher Scientific 11752-050) according to themanufacturer’s instructions. Real-time quantitative PCR was conducted on the QuantStudio 7 PRO Real-Time PCR System utilizing Power SYBR Green PCR MasterMix (Thermo Fisher Scientific 4368706). Actb was used as an internal control.In situ hybridization and immunofluorescence
[0131] Intestinal tissues analyzed in Figures 16E and 17C-17E were collected fromC57BL / 6 mice, cleaned with cold phosphate-buffered saline (PBS), and fixed in 10% erature compound (OCT) frozen sections were prepared for the RNAscope HiPlex12 Reagents Kit v2 assay (Advanced Cell Diagnostics 324409), which was performed according to the manufacturer's instructions with the following probes: Mm-Lgr5-T1 (Advanced Cell Diagnostics 312171-T1), Mm-Mki67-T2 (Advanced Cell Diagnostics 416771-T2), Mm- Fads1-T3 (Advanced Cell Diagnostics 801641-T3), Mm-Fads2-T4 (Advanced Cell Diagnostics 568621-T4), Mm-Fgfbp1-T5 (Advanced Cell Diagnostics 508831-T5), and Mm-Scd2-T7 (Advanced Cell Diagnostics 486111-T7). Protease Plus (Advanced Cell Diagnostics 322331) was used for tissue pre-treatment. Following the last round of in situ hybridization imaging, fluorophores were cleaved using fresh 10% cleaving solution v2. The intestinal tissues were then subjected to immunofluorescence staining. In brief, tissues were washed with PBS, permeabilized with 0.1% Triton X-100 in PBS, and then blocked with 5% bovine serum albumin (BSA) in PBS for 30 minutes at room temperature. The tissues were then incubated with anti-E-Cadherin-Alexa Fluor 488 antibody (BD Biosciences 560061, 1:50) diluted in staining buffer (5% BSA in PBS with 0.1% Triton X- 100) for 1 hour at room temperature, followed by washing and imaging. All fluorescence images were acquired on a Zeiss LSM 980 confocal microscope. To quantify colocalization, cells along the crypts-to-villi axis were first categorized into different cell zones, as described in the caption of Figure 17E. Mean fluorescence intensities were then determined using ImageJ (v1.53t). Tumor tissue dissociation
[0132] Of five primary human melanomas profiled in this work, four viablycryopreserved specimens (sample IDs 2321-07, 2204-07, 2130-07, 2109-07) were each thawed in a 37°C water bath and rinsed with cold sterile 1x PBS. Each sample was then manually minced into small fragments using a surgical scalpel and enzymatically dissociated in a digestion buffer containing 1 g collagenase (Sigma, C5138-1G), 0.1 gDNase I, Type IV (Sigma, D5025), 10 mL HEPES (10 mM; Gibco, 15630080), and 2500 U / L hyaluronidase (Sigma, H6254) in 1 L RPMI (Sigma, R8758). The digestion was carried out at 37°C for 10 min with continuous agitation. The supernatant was collected into a fresh tube and quenched with excess 10% FBS. The remaining tissue pellet was digested again by adding fresh digestion buffer and incubating with agitation at 37°C for tip to homogenize cell clumps and then quenched with excess 10% FBS. Digested cellsuspensions were combined, filtered through a 100 μm filter, pelleted at 500 × g for 5 minat 4°C, washed with cold sterile 1x PBS, and then resuspended in 1x PBS with 0.04% BSA for scRNA-seq.
[0133] For the remaining sample (ID DC2202A), a 4 mm punch biopsy specimen wastransported to the lab in DMEM on ice and immediately processed using the Tumor Dissociation Kit (Miltenyi Biotec, 130-095-929) according to the manufacturer’s recommended protocol. Briefly, the tumor sample was minced using sharp scissors to ~2 mm fragments, incubated in DMEM containing the recommended amounts of Enzymes H, R, and A in a gentleMACS C tube, mechanically dissociated using a gentleMACS dissociator using the h_tumor_01 program, and then incubated for 30 min at 37°C with a MACSmix rotator. The tumor suspension was further dissociated using the ‘h_tumor_02’ program and incubated for 30 min at 37°C using a MACSmix rotator, followed by dissociation using the ‘h_tumor_03’ program. Cells were resuspended with DMEM,filtered through a 70 μm strainer, centrifuged at 500 × g for 7 min, and resuspended in 1xPBS with 0.04% BSA before submitting for single-cell library preparation. Single-cell RNA-seq library preparation and sequencing
[0134] Single-cell libraries were prepared using the 10x Genomics Chromium SingleCell 5’ v2 Reagent Kit according to the manufacturer’s instructions. Briefly, single-cellsuspensions from tumor tissue samples (see “Tumor tissue dissociation”) wereresuspended in PBS with 0.04% BSA to a final concentration of 1,000 viable cells / L foroptimal single-cell capture. Cell viability was assessed using a hemocytometer and trypanblue staining. cDNA libraries were generated following the 10x Genomics v2 user guide, with PCR cycle numbers adjusted based on cDNA input concentrations. cDNA qualityand concentration were evaluated using an Agilent 2100 Bioanalyzer to confirm theintegrity of the libraries. cDNA libraries were then sequenced on an Illumina NovaSeq S4 platform at the Genome Technology Access Center at Washington University in St. Louis, with a target median sequencing depth of 50,000 reads per cell. The CytoTRACE 2 framework: Overview
[0135] Existing RNA-based surrogates of cellular differentiation status have notablelimitations for imputing absolute differentiation states and potency categories from scRNA-seq data. For example, CytoTRACE 1 employs gene counts as an unbiased strategy for identifying immature cells. Despite the utility of this approach, gene countsare subject to dataset-specific biases, making them suboptimal for absolute potencyassessment. Measures based on transcriptional entropy and RNA velocity also suffer from dataset-specific biases, a non-specific relationship to absolute differentiation status, or the requirement for continuous developmental processes within a narrowly defined time window.
[0136] Supervised machine learning (ML) models offer a potentially robust alternativeto the abovementioned strategies when adequate training data are available. However, ML methods also face key challenges when applied to scRNA-seq data, including sparsity, high dimensionality, and significant data heterogeneity encompassing both biological and technical variation. While deep learning is a promising subtype of ML, oftenachieving remarkable performance gains over other ML methods – especially in thepresence of significant complexity, noise, and uncertainty – most existing architectureslack inherent interpretability, limiting their broad applicability.
[0137] To address these challenges, we designed a novel deep learning frameworkthat can handle the complexities of single-cell potency assessment while achieving direct biological interpretability. Unlike most unsupervised methods that decompose single-cell expression data into a combination of previously known and simultaneously learned newgene programs, our approach, termed a Gene Set Binary Network (GSBN) fortransciptome assessment, is anchored to known phenotypic states but not known gene sets. As such, GSBNs have the flexibility to discover and recover new gene programs for known phenotypic states, such as potency categories, from scRNA-seq data. As part oftheir design, GSBNs are highly robust and fully interpretable, meaning they can be directly interrogated to extract meaningful markers for each phenotypic class of interest across datasets, platforms, and tissues. The CytoTRACE 2 framework: Technical description
[0138] CytoTRACE 2 consists of five high-level components, schematically depictedin Figs. 6C and 7B and described in detail below.Preprocessing: Ortholog mapping and rank-based normalization.Gene set binary networks: Identification of interpretable potency-associated genesets for each potency category. Enrichment assessment: Evaluation of gene set activation levels in single cells.Integration of scores: Integration of gene set activation levels, both within and acrossgene set binary networks. Postprocessing: Leveraging transcriptional covariance and uncertainty in modelpredictions to smooth single-cell potency scores and produce the final output. Core model architecture
[0139] Among these five components, Gene set binary networks, Enrichmentassessment, and Integration of scores constitute the CytoTRACE 2 core model, a neuralnetwork architecture consisting of a shared input layer; a set of gene set binary network (GSBN) modules, where denotes the number of potency categories; and a shared output layer (Fig. 7B). Within the core model, each GSBN module is trained to discriminate a single potency category and contains (i) a binary neural network (BNN) component, which encodes potency-associated gene sets, and (ii) downstream functionsto calculate and integrate gene set enrichment scores (Figs. 6C and 7B). Importantly,because weights in BNNs are constrained to binary rather than continuous values, BNNs also allow for more efficient computation and provide an implicit form of model regularization9.Preprocessing
[0140] Let input scRNA-seq dataset be an × gene expression matrix overgenes and cells. The following pre-processing steps prepare the input dataset for training or prediction.
[0141] First, gene symbols in are mapped and filtered using dictionary , a collectionof gene symbols that harmonizes all HGNC (human) and MGI (mouse) identifiers supported by CytoTRACE 2 (“Dictionary of input genes” below). Following this step, theresulting expression matrix, denoted , consists of = 14,271 genes and cells. Aspart of this process, any genes in not present in through mapping are set to zero. In the second step, is converted into dual representations: for the first, it is normalizedto counts or transcripts per million (CPM / TPM) and log2-adjusted, yielding an × matrix; for the second, it is mapped to rank space, yielding an × matrix , with the genesof each single-cell transcriptome assigned relative integer rank such that rank 1 corresponds to the gene with highest expression. While the log2 CPM / TPM representation maintains detailed transcriptomic information, the alternative encoding provided by rank space helps circumvent batch effects, mitigate the influence of extreme values and outliers, and reduce the risk of model overfitting. In tandem, these two representationsprovide an inherent regularization to model inputs. and are subsequently passed tothe CytoTRACE 2 core model where they jointly constitute the model input layer.
[0142] Inputs and are passed to each of gene set binary network (GSBN)modules within the CytoTRACE 2 core model. These modules begin by thresholding(Figure S1B) to learnable maximum rank , yielding × matrix :, = min , , .This rank trimming (see also “Model initialization and updates”) enables calculation of the rank-based enrichment score, described in “Enrichment assessment” below. Input remains the same.
[0143] Next, within each GSBN module, gene sets are learned in binary ×matrix , where is prespecified and all entries , {0,1}. constitutes thegene set selectionof the CytoTRACE 2 core model; it has a continuous equivalentused for model initialization and backpropagation (see also “Core model training”). Ateach forward iteration for model training, undergoes binarization:= binarize( , 0)where binarize denotes the following utility function:1, >binarize( , ) ,, = 0, , .Enrichment assessment
[0144] To quantify the enrichment of each gene set in the module (each column of), CytoTRACE 2 leverages two complementary measures: rank-based enrichmentscore (Score ) and expression-based enrichment score (Score ). Score aggregatesoverall expression activity of a given gene set in rank space whereas Score comparesthe average expression of genes in versus background levels. By integrating both scores, each providing a different axis of information, CytoTRACE 2 can learn more complex expression patterns while also achieving additional regularization through enrichment score competition. The two scores are defined as follows.
[0145] Score calculates the commonly used non-parametric UCell score11 for each geneset, or column of . For each cell 1 and module gene set 1 ,S S + 1 2Score , = 1 + , ,where S denotesof genes per gene setassigned nonzero weight in the binary weighting matrix: =
[0146] Score implementson Seurat’s AddModuleScore(AMS), computing the average expression of genes within a gene set subtracted by the aggregated expression of control, or background, feature sets12. To select background features, AMS groups genes into bins according to their average expression within a dataset. Then, for each gene, a “background” set of genes from the same average expression bin is sampled, ensuring that each gene is compared to other genes with similar average expression. Here, for computational efficiency and to avoidintroducing a dependency on dataset composition, we use our entire curated training cohort (see “Gold standard scRNA-seq datasets”) as the “dataset” in which to rank genes by average expression. We then compute a constant set of background genes to use foreach gene. We encode the mapping of genes to their background genes in the binary×matrix , where each row represents a gene as used in a gene set, and the thentry of row is 1 if gene is used as background for gene , and 0 otherwise.
[0147] In detail, we construct as follows. First, we compute the average log2CPM / TPM expression per gene across all cells from the training cohort. We then rank theresults and uniformly partition genes into = 24 bins of size according to rank,following the Seurat default12. Next, for each gene (each row of ), we randomly select without replacement a set of background genes following a Gaussian distribution withmean = and variance=where = 100, settingthat row to one. This approachprovides an additional regularizing effect compared to constant selection of a uniform number of background genes per gene. Note that left-multiplying a gene set matrix by maps the genes in the gene sets (columns) of to their corresponding background genes.
[0148] Then, given , for each cell 1 and module gene set 1 ,= , ,,,where the first termselected gene set genesin each cell of input gene expression matrix , and the second term calculates theaggregated average expression of background genes within the same cells.
[0149] The two resulting enrichment score matrices are subsequently concatenatedinto a single × 2 matrix := Score , Score ,
[0150] To transferspaces, CytoTRACE 2standardizes each score across cells, yielding × 2 matrix .Integration of scores
[0151] To convert the gene set enrichment scores to a single score per cell per GSBNmodule, the normalized scores are passed through a feedforward layer, termedthe enrichment layer in the CytoTRACE 2 core model, containing the associated length2 gene set enrichment score weight vector V and yielding length potency categoryscore vector . As part of this process, dropout is applied to reduce overfitting during model training, with a predetermined fraction of the normalized scores set at random to zero. From the weights in each V, concatenated across potency categories into matrix , the directionality and importance of each gene set can be interpreted (see “Interpretability” below).
[0152] The model then integrates across the potency category scores produced byeach GSBN module, concatenating the potency category score vectors into ×potency score matrix . This procedure represents the shared output layer of theCytoTRACE 2 core model.
[0153] To convert the logit entries of to likelihoods, the model applies a softmaxactivation function, yielding × matrix representing the likelihood of each cellbelonging to each of the six potency categories. The model then predicts cellular potency by assigning the potency category with highest likelihood for each cell, yielding lengthvector := argmax{ } ,
[0154] The vectorof the CytoTRACE 2 coremodel. However, the model also computes an absolute developmental potential from this set of likelihoods, termed the raw potency score RPS. For this aspect, we introduce length ordered vector to be multiplied by the potency category likelihood matrix:RPS == [0.0, 0.2, 0.4, 0.6, 0.8, 1.0]where RPS is the lengthscore vector. As the potency categories areordered based on their absolute developmental potential, the resulting raw potency score will be closer to one for higher potency categories, such as totipotent, and closer to zerofor lower potency categories, such as differentiated. As RPS directly incorporates modeluncertainty, it is passed to “Postprocessing” below in order to define a more granular developmental ordering. Postprocessing
[0155] Since the fully trained CytoTRACE 2 model predicts potency for each cellindividually, CytoTRACE 2 further processes the output (raw potency score RPS andpredicted potency categories ) to incorporate the neighborhood structure oftranscriptionally similar cells. We reasoned that doing so could further improve performance given our prior experience combining gene counts with transcriptional covariance in CytoTRACE 1. To this end, we devised and validated a three-step procedure using the training cohort, as described below. Importantly, this procedureimproves correlations with relative developmental orderings (see “Metrics” below) overRPS or alone without sacrificing the potency classification performance achieved by(Fig.7C).
[0156] In the first step, CytoTRACE 2 applies Markov diffusion to smooth (RPS) usingthe same implementation as CytoTRACE 12. In brief, the log2-adjusted CPM / TPM gene expression input L is used to create a Markov matrix from the transcriptional similarity between cells over the top 1,000 genes with highest dispersion2. This similarity matrix isthen used to smooth (RPS) with diffusion parameter as previously described2,yielding smoothed potency score (SPS) . Using the same sampling procedure describedin our previous work2, the running time of this step can be significantly reduced without loss of performance (Figure 7D). In this study, sampling was restricted to datasets with >10k cells (Tables 7A and 11A).
[0157] To reconcile SPS with predicted potency categories , in the second stepCytoTRACE 2 performs a binning procedure to maintain while preserving relative potency ordering within each category. To do so, CytoTRACE 2 first separates cells bytheir predicted potency category and assigns each cell 1 a rank ( , ) relativeto all cells sharing predicted potency category . For this transformation, within eachpotency category 1 , the cell with lowest potency score receives rank 1 while thecell with highest potency score receives maximum rank ( ). Cells are then arrangeduniformly by rank per potency category within equal length partitions of the unit interval,yielding binned smooth potency score SPS . Thus, the binned smooth potency score fordifferentiated cells extends from 0 to 1 / 6, unipotent from 1 / 6 to 2 / 6, and so on, with relative ordering within each bin matching that of the original smoothed potency score.
[0158] In the third step, to further smooth SPS while minimizing the impact on , andallowing for the preservation of rare cell states (Figure 10F), CytoTRACE 2 applies a variation of k-nearest neighbor (KNN) smoothing to datasets with >100 cells. Here, we introduce an efficient heuristic approach for adaptive neighborhood smoothing guided by two key assumptions: (i) cells with more similar gene expression profiles are more likely to share a potency phenotype; (ii) prediction errors for cells with the same ground truth potency exhibit a random distribution around a central mean. To balance these two considerations and identify an appropriate neighborhood size, we select adaptively for each cell according to the following process. First, given log2-adjusted CPM / TPM gene expression profiles for the selected cell, we standardize expression per cell to zero mean and unit variance, then perform dimension reduction of standardized gene expressionprofiles over all cells to the top 30 principal components (PCs). Using the top 30 PCs, wethen compute pairwise Euclidean distances for all cells, rescaling the resulting distances to unit maximum per cell of interest. Next, we define the neighborhood around each center cell through an iterative procedure, allowing a maximum neighborhood size of 30 cells. We start with the nearest cell to , denoted , and calculate the average potency scoreprediction for and , mapping the result to one of six broad potency categories, yielding. We repeat this calculation for the next two nearest cells to ( and ), yielding ,and compare and . If identical, we assume that we have sufficiently captured the= 3 (for the three non-self neighbors) and exiting the process. Ifnot identical, we repeat the procedure increasing the group size by one – in other words,comparing the nearest two cells to (yielding three total cells) with the next nearest threecells (i.e., , , and ). We repeat this process until the resulting potency categories arethe same between two groups, in which case we select to encompass all cells considered between the two groups, or until we exhaust our candidate nearest neighbor cells (i.e., reach a group size of 15). If concordance between nearest and next nearestgroups is not found, we keep our initial selection of = 3.
[0159] Once k is determined, we update our prediction for w according to the distance-weighted mean of neighborhood potencies to obtain the final potency score prediction: () SPS (1 )cytotrace2 =( ) (1 )where ( ) denotes the of center cell ,including the itself, andof cell to cell . Categoricalpotency predictions are updated based on the defined intervals above, yielding .
[0160] We found empirically that combining these three approaches yielded superiorperformance on the training cohort (Fig.7C). The CytoTRACE 2 framework: Technical description Loss function
[0161] For model training, we defined a loss function combining cross-entropy losswith an additional term penalizing gene set size based on the binary weighting matrixoriginating from each GSBN module, 1 . More precisely, we define the lossfunction as the sum of gene set size penalty loss and a prediction loss per cell := , , + ( , )In detail, given potency potency categories for cell (see “Gold standard,prediction loss as(, ) = ( , )where denotes the loss( , ) denotes the cross-entropy loss for cell . Loss weights for all cells are contained in the length weightingvector , which has unit sum and is constructed hierarchically to assign equal weight (i)to all broad potency categories, (ii) to all phenotypes within each broad potency category, and (iii) to all datasets contributing to each phenotype.
[0162] We defined gene set size penalty loss as1 12where | |matrix, denotesthe gene set size penalty weight, and serves as a scaling factor to make invariant tothe number of gene sets included in , with factor 12 selected to anchor the gene setsize penalty weight to the center of the range of hidden sizes tested (see “Hyperparameteroptimization”). This loss component serves to minimize the number of genes in each geneset while regularizing the training of the model. Model regularization
[0163] To promote model generalizability, we introduced two explicit regularizationaspects. We included a dropout layer to avoid model overfitting to specific enrichment scores (“Integration of scores”). A dropout layer randomly drops (sets to zero) units in ahidden layer of a neural network. This layer was applied to the normalized scoresduring training only. Additionally, a penalty term was added to the loss function toconstrain the number of genes in each gene set of (“Loss function”). Model initialization and updates
[0164] Model weights were initialized according to PyTorch v2.0.0 default except forthe binary weighting matrices, which were initialized at random with values sampled fromthe Gaussian distribution with mean 0.1 and standard deviation 0.055 to produce asparse initial binarization with approximately 500 genes selected per gene set.
[0165] Model training was performed with mini-batch learning using a batch size of1024. To balance batches and ensure equal representation for the model learning process, each batch was constructed via uniform sampling across datasets and phenotypes as implemented by torch.utils.data.WeightedRandomSampler in PyTorch.
[0166] Following initialization, forward propagation proceeded for each iteration asdescribed in “Core module architecture”, with parameters updated according to their definition. For numeric stability, the cutoff rank (“Gene set binary networks”) for trimminginput rank space expression matrix was not learned directly but rather computed as afunction of learnable parameter , which was initialized uniformly at random from0 1 per module and suitably scaled. As gene set enrichment score calculation(“Enrichment assessment”) requires a gene set pool larger than the gene set itself for comparison, was computed from in such a way as to ensure that the ranks of at least 10 more genes beyond the maximum gene set size of the module were preservedfollowing trimming to . Thus, at each iteration, the updated was scaled andconstrained as follows:= 10 + max S + 1000 max(0, )
[0167] Model at each iteration against ground truth, withthe loss function andand used to backpropagate updates tonetwork weights using PyTorch’s NAdam optimizer with custom learning rate lr = 0.001(see “Hyperparameter optimization” below) and otherwise default parameters. Given the role of inertia in successfully training binary neural networks, we employed cross-epoch gradient accumulation to dampen binary weight flipping and achieve a stabilizing effect. This approach additionally facilitates broader hyperparameter space exploration while validation-based early stopping (see “Model evaluation and stopping”) ensures that the most performant model encountered during training is retained. Backpropagation for the binary neural network component of each GSBN module was implemented with Straight- Through Estimator and hardtanh activation function. Model evaluation and stopping
[0168] We evaluated model validation performance over training and validation setsvia category-weighted accuracy, defined as the mean F1 score of predicted versus ground truthacross evaluable potency categories as implemented with. To do this, we first calculated the F1 score for each phenotype (standardized as in Table S1C) and dataset pair using metrics.f1_score precision_recall_fscore_support from sklearn v1.0.2. ModelsWe then averaged the resulting scores across datasets per phenotype, across phenotypes within each broad potency category, and across broad potency categories, yielding the final weighted accuracy. For the standard CytoTRACE 2 model, each validation set consisted of a single dataset; however, for the leave-clade-out model (see “Generalizability to unseen cell-type clades”), validation sets included all cells covering a clade, regardless of dataset. All models were trained for 40100 epochs with the best model weights by the highest score on the validation set after a minimum of 15 initial training epochs preserved and returned for the final model. Hyperparameter optimization
[0169] To evaluate the hyperparameter space of CytoTRACE 2, we performedBayesiana hyperparameter optimizationsweep over the training cohort using SigOpt(v8.8.2wandb (v0.16.4). We optimizedexplored the learning rate lr over{0.01,0.005,0.001,0.0005}{0.01,0.005,0.001,0.0005,0.0001}, number M of gene sets per broad potency category over {1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,32}{1,2,4,8,12,16,24,32,48}, gene set size {0,0.25,0.5}, and enrichment considering whether to use model expressionAMS enrichment score, model gene count, UCell enrichment score, or the concatenation combination of both as described in “Enrichment assessment” above. For every iteration of leave-one-dataset-out nested cross-validation, we trained models across 100500 different combinations of these hyperparameters sampled based on the Bayesian optimization algorithm of SigOpt.random hyperparameter search. To minimize overfitting to training data, we used a nested cross validation framework. While one dataset was held out from training and evaluated as a validation set, another dataset was also held out from training but used to determine the early stopping point as described in “Model evaluation and stopping”. We scored each hyperparameter combination according to the averageby weighted accuracy across the respectiveover model validation sets, where weighted accuracy for each validation dataset was defined as the average of the category-weighted F1 scores per standardized phenotype and evaluated using metrics.precision_recall_fscore_support from sklearn v1.0.2.S1C; see “Model evaluation and stopping”).
[0170] We observed that variation in hyperparameter values had minimal impact onperformance, underscoring overall model robustness (Fig. 7B). Nonetheless, and consistent with expectation, combining the two gene set enrichment scores achieved superior performance compared to either enrichment score alone (Fig. 7B, center). Moreover, when both were combined, the model was more robust to changes in other hyperparameters, such as learning rate, number of gene sets or gene set size penalty. These data supported the use of both enrichment scores rather than either enrichment score individually.
[0171] Final hyperparameter selection was done by a manual curation processidentifying values yielding consistently – albeit modestly – higher weighted accuracy. Inselecting the number of gene sets per potency category, we found that model performance increased with before plateauing (Fig. 7F). Final hyperparameter selection was done by a manual curation process identifying values yielding consistently– albeit modestly – higher weighted accuracy. In selecting the number of gene sets perpotency category, we found that model performance increased with before plateauing (Figure 7F, right); as such, we selected slightly larger than the number correspondingto the elbow of this curve. The final hyperparameters used were = 24 gene sets perpotency; = 0.5 dropout probability; = 0.01 gene set size penalty weight; and lr =0.001 learning rate.
[0172] Next, we evaluated the enrichment metrics. Among all models, we limited to 84models with hyperparameter values in ranges of plateau ( 2 gene sets per potency;= 0.5 dropout probability, 0.01, lr 0.001). AMS enrichment and both AMS andUCell enrichment achieved superior performance compared to UCell enrichment alone (Figure 7G). Given the potential to enhance generalizability, we therefore selected the combination of AMS and UCell enrichment metrics for the final model. Model ensembling
[0173] Models were trained via leave-one-dataset-out cross-validation for each of thetraining datasets, with final CytoTRACE 2 predictions in non-training data obtained as theresult of integrating predictions across the 19 resulting models followed by an additionalpostprocessing step. As described in “Integration of scores" above, each model yieldsa × potency category likelihood matrix . Models were integrated by entry-wiseaveraging of potency category likelihood matrices to yield a single potency categorylikelihood matrix from which potency category predictions and raw potencyscores weredescribed above, before passing them to “Postprocessing”.The CytoTRACE 2 framework: Dictionary of input genes
[0174] To create dictionary (“Preprocessing” above), all human gene symbols weremapped to their closest mouse orthologs, as determined by gene sequence similarity,using the GRCh38.p14 and GRCm39 annotation files available from Ensembl version 109, respectively. In cases where a single mouse gene was identified as the best hit for multiple human genes, the human gene with maximum sequence similarity to was selected and the remaining human gene(s) excluded from further consideration. Unique human gene symbols without orthologs by the above process were also included for completeness. To define a common subset, only genes present in at least 80% of datasets from an initial development cohort, a subset of the final training cohort, were retained. Combining these steps, was assembled with 14,271 unique gene symbols, including 13,750 orthologous pairs and 521 genes without orthologs in Ensembl via the mapping step above. The CytoTRACE 2 framework: Interpretability
[0175] The GSBN architecture of CytoTRACE 2 enables direct interrogation of thebinary weight matrices, consisting of gene sets associated with each potency category(Figs. 6C and 7B). By examining the orientation of the output layer weights for each geneset, we found that gene sets with positive weights (polarity) were highly enriched in a given potency category, whereas those with negative weights (polarity) were preferentially depleted (Fig. 14D). Additionally, we reasoned that genes repeatedly selected for a given potency category were more likely to be important for effective classification. As such, we designed a metric to quantify feature importance, assigning importance scores to genes according to the frequency at which they were selected in positively versus negatively weighted gene sets. Here, we incorporate gene selection frequency across all 17 training models computed by LOOCV over the training cohort datasets.
[0176] More formally, we define × feature importance score matrix containingthe feature importance score of each gene 1 for each potency category 1based on the gene set compositions and enrichment weights across models. Two enrichment weights correspond with each gene set, one per enrichment score type (see“Enrichment assessment”). Given gene set enrichment weight matrix of model , wecalculate the polarity Polarity( , , ) of gene set defined within model for potencycategory module as the sign of the average of these two weights. Then, relying onmodel binary weighting matrices to encode gene set composition, we construct feature importance score matrix entry-wise as ,= , [i, j] Polarity , ,where , [ , ] denotes the [ , ]th entry of the binary weighting matrix from module ofmodel .Gold standard scRNA-seq datasets: Overview
[0177] Classically defined potency levels, as depicted in Fig. 6A, are not directlyannotated in publicly available scRNA-seq datasets. Therefore, to train, validate, and benchmark CytoTRACE 2, we downloaded and curated 31 human and mouse scRNA- seq datasets from peer-reviewed studies with experimentally confirmed developmental states and assignable potency levels. As part of this selection process, we applied the following inclusion and exclusion criteria to enhance experimental rigor: Only functionally validated developmental states supported by lineage tracing or transplantation assays were considered for analysis. Datasets with transient cell changes, such as from metabolic activation or suppression, cell cycle transitions, or environmental perturbations were excluded, as these do not represent durable developmental processes. Datasets with irreconcilable technical batches resulting in major imbalances in the number of cells per phenotype were excluded. Single-nucleus RNA-sequencing datasets were excluded, as they do not capture cytoplasmic RNA and include immature transcripts.
[0178] Among datasets satisfying these conditions, author-supplied cell typeannotations were mapped to one of six standardized potency categories – totipotent,pluripotent, multipotent, oligopotent, unipotent, and differentiated – or NE (not evaluable)using established definitions (‘Potency annotations’ below). These potency categories were further subdivided into 24 granular categories, ranging from ‘1’ (least differentiated) to ‘24’ (most differentiated). Cellular phenotypes were hierarchically grouped into these categories based on potency, developmental timing and sequence, and self-renewal capacity.
[0179] Where possible, we also examined single-cell developmental states in adataset-specific manner and without regard to potency categories. Such ‘relative’ orderings ranged from 1 (least differentiated) to (most differentiated) in a given dataset, and exceeded the number of resolvable potency categories in some datasets, permitting a more granular assessment.Gold standard scRNA-seq datasets: Training and test datasets
[0180] Using the abovementioned criteria, we assembled a 33-dataset potency atlas(Figure 7A), from which we selected a training cohort consisting of seven human and 12 mouse scRNA-seq datasets from 13 studies. We ensured that all six broad potency categories were represented in both species along with a diverse array of biological (e.g., tissue types) and technical characteristics (e.g., sequencing platforms). As part of this effort, and to align with precedent in the field, we incorporated all human and mousescRNA-seq datasets (n = 13) with annotatable potency categories analyzed by Gulati etal.2To broadly cover tissue types, we also included cell phenotypes from the Tabula Muris scRNA-seq atlas16for which potency categories could be determined (15 tissue types, 43 phenotypes). The resulting training cohort encompasses 312,523 cells, 18 tissue types, 93 phenotypes, and six scRNA-seq platforms (Figure 6B).
[0181] The remaining datasets served as a held-out test cohort, which mirrors thetraining cohort with respect to species representation in each broad potency category. Consisting of three human and 11 mouse scRNA-seq datasets from 14 studies, the test cohort spans 93,535 cells, 73 phenotypes, nine tissue types, and seven scRNA-seq platforms, including two tissue types and 21 phenotypes that were absent from training(Figures 8A and 9C).
[0182] To augment these data, we annotated potency categories in 459,320 evaluablecells from Tabula Sapiens, a multi-tissue scRNA-seq atlas from postmortem human donor biopsies. However, given the confounding influence of postmortem intervals on human tissue mRNA levels, we hypothesized that Tabula Sapiens might exhibit reduced data quality. To test this, we calculated the ratio of mitochondrial (MTR) reads to total reads within each single-cell transcriptome as a proxy for overall data quality. Indeed, we calculated a mean MTR across all Tabula Sapiens tissue types, stratified by platform, of7.4% (median of medians) – nearly 90% higher than expected for human cell typesprofiled by scRNA-seq data (median of medians = 3.9%) and 78% higher than other human datasets in the training and test cohorts, both of which include embryonic tissues with high metabolic activity (median of medians = 4.2%). Accordingly, we omitted Tabula Sapiens from the primary test cohort and evaluated it as a secondary benchmark in Fig. 10A.
[0183] Collectively, these 31 gold standard datasets with newly annotated potencylevels represent a unique community resource for systematic characterization of absolute developmental states and their molecular programs in humans and mice. Depending on platform, all scRNA-seq expression matrices were normalized to transcripts per million (TPM) or counts per million (CPM) as appropriate. Full details of each dataset, including dataset name, accession number, PMID, species, platform, tissue type, number of cells, number of phenotypes, and number of potency levels.Gold standard scRNA-seq datasets: Potency annotations
[0184] We used the following annotation scheme to assign single-cell transcriptomesfrom gold standard datasets to broad potency categories. Where possible, we also annotated developmental orderings within each broad potency category, resulting in alarger repertoire of 24 granular potency levels spanning the full range of cellular ontogeny(n = 2 in ‘Totipotent’, n = 6 in ‘Pluripotent’, n = 7 in ‘Multipotent’, n = 4 in ‘Oligopotent’, n= 3 in ‘Unipotent’, n = 2 in ‘Differentiated’; Figure 7A). However, such finer-grainedorderings have sparse representation across species and scRNA-seq datasets; therefore, we restricted model training to the six broader categories. Granular potency levels are also indicated in the descriptions below as numeric values between “1” (highest potency) and “24” (lowest potency).
[0185] Totipotent. Totipotency was defined as the potential of a cell to give rise to allthe cell types in a body, including extraembryonic cells, with maternal support. In additionto the totipotent zygote (1), we included the 2-cell stage of mouse and human embryos inwhich initial zygotic division has formed a 2-cell embryo (2). This assignment wassupported by experiments in mammalian models showing development of physiologically normal organisms after re-implantation of zygote and blastomeres from the 2-cell stage.
[0186] Pluripotent. Pluripotency was defined as the potential of a cell to give rise to allthe cell types in a body except extraembryonic cells. Although there are some reports that suggest blastomeres from the 4-cell stage are totipotent, given mixed reports and uncertainty in humans, we conservatively assigned the 4-cell stage as the most primitivecell state in the pluripotent category (3). Chronologically, the 8-cell stage (4) and 16-cellstage (5)– referred to as a morula – come next. The embryo then progresses to the 32-cell stage (6) when it undergoes compaction, in which a central cavity forms and cellsseparate to give rise to the trophectoderm and inner cell mass, the latter of which is the origin of embryonic stem cells. As the pluripotent cells divide and specialize, theytransition from the early blastocyst through the mid-(7) and late-(8) blastocyst stages,which mark the final stages of pluripotency in our analysis.
[0187] Multipotent. Multipotency was defined as the potential of an immature cell togive rise to multiple cell types (at least four) across different lineages. Multipotent cells exist both during early embryogenesis for specification of germ layers (mesoderm, endoderm, ectoderm) and during the postnatal period as tissue-resident stem cell populations. Chronologically, the anterior primitive streak is the earliest source of multipotent cells (“9”), which sequentially gives rise to developmentally restricted cells in each germ layer (e.g., paraxial mesoderm (“10”), which then segments into somitomeres (“11”) and somites (“12”)). Development occurs in parallel across germ layers, so the granular order of multilineage precursors from different lineages was aligned to relative developmental time using author-provided labels for embryonic age or somite stage. For example, pre-cranial neural crest cells (pre-CNCCs; “12”) arise during the mouse 4- somite stage, so their developmental timeline matched early somites (“12”). These cells then give rise to CNCCs (“13”) and early multipotent precursors (“14”), such as the delaminating CNCC, dermomyotome, and sclerotome. Given their tissue-specific developmental potency, tissue-resident multipotent stem and progenitor cells, such as hematopoietic stem and progenitors, pancreatic progenitors, skeletal stem cells, radial glial cells, and others, were annotated as the least primitive among multipotent cells (“15”).
[0188] Oligopotent. Oligopotency was defined as the potential of an immature cell togive rise to more than one but fewer than four cell types across different lineages and / orcell types immediately downstream of known multipotent stem and progenitor cells. In our potency atlas, the most primitive cells in this category are unpurified hematopoietic progenitors, such as CD34+cells, which are enriched for restricted progenitors but also contain rare multipotent hematopoietic stem cells (“16”). Sequentially, cells were granularly ordered as oligopotent progenitor populations without rare multipotent cells (“17”), restricted bipotent stem and progenitor cells (“18”), and then basal cells (“19”), which are a heterogeneous mixture of predominantly bipotent progenitors, along with unipotent progenitors, differentiated cells, and rare multipotent cells capable of regeneration in response to injury.
[0189] Unipotent. Unipotency was defined as the potential of an immature cell topredominantly give rise to one mature downstream cell type. In our potency atlas, the most primitive cells in this category are unipotent stem-like cells, such as limbal stem cells, satellite stem cells, and nascent type II pneumocytes (“20”). Sequentially, cells weregranularly ordered as unipotent progenitors without self-renewal capacity (“21”), followedby immature cells transitioning toward differentiation (“22”).
[0190] Differentiated. Differentiated cells were defined as mature cells within a singlelineage, including non-terminally differentiated cells (“23”), which have limited potential to further differentiate or change phenotypic state in response to a stimulus (e.g., naïve B cells and naïve T cells upon antigen recognition; fibroblasts and other stromal cells in response to tissue injury), and terminally differentiated cells (“24”), which essentially have zero potential to make other cells, including of their same type.Gold standard scRNA-seq datasets: Additional annotation considerations
[0191] For cells with identical phenotypes but different author-supplied labels, weunified the annotations. For example, ‘HSC-MPPs’ from ‘HSC development (Smart-seq2)’ and ‘Hematopoietic stem cell progenitor (HSCP)’ from ‘HSPCs (C1)’ were annotated as ‘Hematopoietic stem and early progenitor’. To balance the representation of cells from distinct lineages within a given broad potency category, we also re-annotated related cell subsets sharing a common parental phenotype. For example, ‘CD4 helper T cells’ from ‘Peripheral blood (10x)’ and ‘CD8 Memory T cells’ from ‘BM-MNC (CITE-seq)’ were labeled as ‘T cell’. This was crucial when training CytoTRACE 2 since the probability ofsampling individual cells was weighted based on phenotype. In this way, each major phenotype contributed equally during model training regardless of the number of evaluable cells, mitigating the chance of overweighting and overfitting (see “Training and hyperparameter tuning” above). Performance assessment: Metrics
[0192] Two key metrics, each illustrated in Figures 6E and 7E, were used to quantifyreconstruction of known developmental orderings: absolute order and relative order. Absolute order quantifies cross-dataset performance, whereby predicted orderings from all cells with annotated potency levels are analyzed together, regardless of dataset, tissue type, or platform. Relative order quantifies performance within a given dataset and tissue type, akin to conventional pseudotime, and ranges from 1 (least differentiated) to (most differentiated) in each dataset. For both metrics, we applied weighted Kendall correlation( ) (wdm package v0.2.4 in R) to assess concordance between known and predicteddevelopmental orderings. Similar to our previous work2, ground truth phenotypes corresponding to less mature cells were coded with lower ranks (starting at 1); therefore, higher predictions of developmental order were ranked such that higher values received lower ranks, and vice versa.
[0193] For categorical predictions (CytoTRACE 2 and potency classificationbenchmarking outputs only), we evaluated potency classification performance as well. Binary correctness of predicted versus ground truth broad potency categories was assessed via mean multi-class F1 score, implemented with function f1_score fromsklearn.metrics with average = None (Figures 7C; 9D second from right; 10B to 10E leftbottom; 11A left; and 11B x-axis). To account for the magnitude of deviations from groundtruth potency, we also considered mean absolute error (MAE), assigning each broad potency class an integer label corresponding to the class ordering, with labels ranging from 1 (differentiated) to 6 (totipotent), and computing the absolute value of the differencebetween predicted and ground truth categories (Figures 9D far right; 10B to 10E rightbottom; 11A right; and 11B y-axis). For both metrics, scores were computed per groundtruth potency category then aggregated by mean across potencies.Performance assessment: Generalizability to unseen cell-type clades
[0194] To test the generalizability of CytoTRACE 2 to unseen developmental systems,we trained a version with a leave-clade-out framework (Figure 9E), grouping phenotypes into 19 developmental clades. Of note, to ensure representation of some totipotent and pluripotent phenotypes for all training sets, we partitioned embryonic phenotypes into two clades by alternating granular potency level annotation. For each clade, we trained an ensemble of two models over the remaining 18 clades, selecting at random 17 clades for training and one clade as a held-out validation set to be used for early stopping (see “Model evaluation and stopping”) for each model. We then applied the resulting ensembleto the unseen test clade, assessing performance across all held-out clades in Figure 9F.Performance assessment: Analysis of datasets with genome-wide chromatinaccessibility
[0195] To assess the robustness of the model to variation in the composition of thetraining cohort, we repeated the CytoTRACE 2 training process as described in “The CytoTRACE 2 framework” across a series of three randomized splits covering all 33 gold standard datasets (Table S2D). We partitioned the datasets at random into three folds, each containing 11 datasets. To ensure minimum adequate representation within each category, we confirmed that each fold contained at least one phenotype per broad potency category. Tabula Muris, which was divided into two sub-datasets according to platform for the original CytoTRACE 2 training cohort due to its size and diversity, was again divided, with one of its sub-datasets assigned to another fold at random. For each split, two folds were combined to form the training cohort and the remaining one left as atest set for evaluation (2:1 training-test split). Performance across the test sets of thesethree randomized splits, along with the original CytoTRACE 2 test set, was assessed by absolute order, relative order, mean multi-class F1 score, and MAE (see “Metrics”), showing strong consistency across folds (Figure 9D).Performance assessment: Analysis of datasets with genome-wide chromatinaccessibility
[0196] To assess single-cell chromatin accessibility as an alternative readout of cellpotency separate from our annotation scheme, we applied CytoTRACE 2 to three publicly available multiome datasets (Figures 9G and 9H).
[0197] The first dataset, profiling human hematopoietic cells with ReDeeM, wasobtained as a preprocessed and annotated Seurat object containing both scRNA and scATAC sequencing data. Raw scRNA-seq counts from the multiome assay were analyzed with CytoTRACE 2.
[0198] The second dataset, consisting of mouse intestinal epithelial cells profiled bythe 10x Single Cell Multiome platform, contains preprocessed paired single-nucleus RNA(snRNA) expression and single-nucleus ATAC (snATAC) data, and is accompanied byan unpaired scRNA-seq dataset profiling the same biological replicate. These data are accessible from the Gene Expression Omnibus (GEO) under accession numbers GSM6208867, GSM6208868, and GSM6208863, respectively. As cell type annotations were provided for the scRNA-seq but not multiome assay, we performed label transfer using the Seurat v5 default pipeline. CytoTRACE 2 was then applied to the raw snRNA- seq count matrix from the multiome assay.
[0199] The third dataset, consisting of preprocessed unpaired snATAC-seq andscRNA-seq data of human fetal pancreatic islets, was obtained from GEO accession numbers GSE202498 and GSE202497, respectively. For this dataset, the raw count matrix from the scRNA-seq experiment was used for CytoTRACE 2 predictions. Given that the snATAC-seq and scRNA-seq experiments were conducted separately and the single cells were not jointly profiled, we addressed this by averaging the CytoTRACE 2 predictions and chromatin accessibility data into phenotype-specific bins (using author- supplied labels), then comparing results across shared phenotypes.
[0200] For Figure 9G, we assessed the relationship between CytoTRACE 2 scoresand chromatin accessibility, calculating the coefficient of determination between mean CytoTRACE 2 scores and the mean number of peaks per phenotype. For the analyses in Figure 9H, we evaluated the concordance of CytoTRACE 2 scores and the number of ATAC-seq peaks with ground truth developmental potential at the single-cell level via absolute and relative order (see “Metrics”).Performance assessment: Analysis of mouse hematopoiesis clonal barcoding data
[0201] For the analyses presented in Figures 8C and 8D, we utilized preprocessedclonally barcoded scRNA-seq data of mouse hematopoiesis profiled by the “lineage and RNA recovery” (LARRY) heritable clonal barcoding system. In this setting, barcoded hematopoietic stem and multipotent progenitor cells were allowed to differentiate in vitro for two days, sampled for scRNA-seq, then transplanted into irradiated mice and sampled after nine and 16 days of hematopoiesis for additional scRNA-seq. We restricted analyses to cells with clonal barcodes present in at least two timepoints and at least two cells per barcode / timepoint pair. For any given upstream timepoint, to determine the number of unique downstream phenotypes arising from one or more sister cells sharing the same barcode, we used barcode-matching cells from the latest possible timepoint. One caveat of this approach is that barcodes of identical sequence can be dispersed among distinct oligopotent progenitor phenotypes, leading to inflated counts of daughter cell diversity. Nevertheless, median CytoTRACE 2 predictions were closely aligned with potency definitions (Figure 8D; “Potency annotations” above), indicating that such inflation is acceptably low. Moreover, in cases where an upstream cell was not author-annotated as ‘progenitor’ or ‘undifferentiated’, counts were restricted to downstream cells sharing the same clonal barcode and phenotype label (as such cells were invariably lineage-restricted and either unipotent or differentiated; Figure 8C). Robustness of CytoTRACE 2: Annotation error
[0202] To evaluate the robustness of CytoTRACE 2 to potential noise within potencyannotations, we trained models across two scenarios of training cohort annotation error, then evaluated model performance over the test cohort (see “Training and test datasets”). To simulate annotation error, we formulated label noise as a transition matrix, encoding the probability of perturbation from one potency to another (Figure 10A). Transition matrix perturbation probabilities were designed to follow a Gaussian distribution based on the rank distance between the original potency and perturbed potency. In detail, the probability that the potency label of cell transitions from true potency to perturbed potency1 ( )| = exp , , {1, 2, 3, 4, 5, 6}2 2where potencies potency categories.Standard20%, 50%, and 80% perturbation levels. Rows were normalized to unit sum for a net probability of one. For the first annotation error scenario, we considered cell-level annotation error and perturbedthe potency annotations of individual cells independently (Figure 10B). For the second,we considered phenotype-level annotation errors and simultaneously perturbed the potency annotations of entire standardized phenotypes (Figure 10C).Robustness of CytoTRACE 2: Variation in gene counts and UMI counts
[0203] To determine the influence of variable gene counts and UMI counts onCytoTRACE 2, we performed two experiments in which scRNA-seq expression data from all 14 datasets in the test cohort were perturbed by downsampling gene counts (Figure 10D) and all seven droplet-based datasets in the test cohort were perturbed by downsampling UMIs (Figure 10E). We assessed the robustness of the model to different gene counts by downsampling the expression data of each cell to the same number of genes: 2000, 1000, 750, 500, 250, 100. We selected the top genes by highest expression and set the expression of the remaining genes to zero. For any expression level ties atthe threshold, we selected the genes to include to reach the target gene count at random.The downsampling process for UMIs consisted of randomly sampling the expression data of each cell based on the transcriptome probability distribution, defined as the fractionalexpression of each gene after scaling the sum of UMIs in each cell to one. Then, usingthe raw count matrices, we downsampled the expression data of each cell to the same number of UMIs: 5000, 3000, 2000, 1000, 500, or 100 UMIs. Cells with UMIs lower than a given threshold were unaltered. We repeated each process for five replicates, then assessed performance for standard metrics as described above (see “Metrics”) relative to the CytoTRACE 2 predictions without perturbation.Robustness of CytoTRACE 2: Titration of cell type rarity
[0204] Given the inclusion of neighborhood-based smoothing in modelpostprocessing, we performed a titration experiment applying CytoTRACE 2 to test datasets with selected phenotypes downsampled to increasingly rare abundance. For eleven phenotypes spanning a range of potencies, we downsampled cells of the selected phenotype to predefined abundances of 50, 20, 10, 8, 5, 2, and 1 cell(s), leaving the remaining cells in the dataset unchanged. We repeated this titration process five times for each phenotype, observing robust predictions down to five cells per phenotype (Figure 10F). As such, we recommend that the final postprocessing step (adaptive k-nearest neighbor smoothing) be omitted when exceedingly rare cell states (comprised of <5 cells each) are of interest. Benchmarking: Cell type prediction for potency classificaiton
[0205] To evaluate CytoTRACE 2 against supervised machine-learning approachescommonly employed in cell type prediction tasks (Figures 11A and 11B), we selectedthree dedicated single-cell annotation methods with superior performance in a benchmarking study (scPred [SVM-based], SingleCellNet [random forest-based], and scmap) and five general-purpose classifiers (below), each trained to predict six broad potency labels based on single-cell expression profiles.
[0206] All tools were trained and tested over a series of four folds, including the originalCytoTRACE 2 training / test split (Figure 7A) along with three randomized splits (see “Randomization of training and test sets”), collectively encompassing all 33 gold standard datasets as described above, with classification performance per test cohort assessed bymean multiclass F1 score and MAE (Figure 11A and 11B; see “Metrics”). For all methods,expression data were first mapped into the uniform feature space used by CytoTRACE 2 (see “Preprocessing” and “Dictionary of input genes”). Unless otherwise specified, and for all general-purpose classifiers, expression data were then CPM / TPM normalized and log2-transformed, and subsequently standardized per cell to zero mean and unit variance. Other normalization schemes generally yielded worse performance and were thus omitted from further consideration (i.e., log2-adjusted CPM / TPM data, either used alone or with gene-level standardization). No explicit dataset integration or batch correction was performed. For general-purpose classifiers, versions were trained with and withoutsample weighting (computed as for CytoTRACE 2; see “Loss function”) for class imbalance mitigation, with the best performing version across all folds selected for each. All parameters were set to default values unless otherwise specified.
[0207] CytoTRACE 2. We applied CytoTRACE 2 with model ensembling andpostprocessing as described in the “The CytoTRACE 2 framework” to predict cell potency categories. Datasets containing more than 100k cells were processed in batches of 100,000 cells, and diffusion was applied in batches of 10,000 cells for datasets exceeding 10k cells.
[0208] scPred. A dedicated cell type classification method, scPred first performs adimension reduction, identifying principal components exhibiting significant variationacross classes, then, as the default option, applies a support vector machine approach for classification. Following the recommended pipeline for scPred (v1.9.2), we first normalized and scaled expression data using the NormalizeData() and ScaleData() functions in Seurat (v5.1.0), respectively. We then used scPred’s getFeatureSpace() function to identify class-informative principal components, trainModel() to train the default SVM with radial kernel model for each potency category (one-versus-rest), and scPredict() for classification. A relaxed probability threshold of 0 was used to avoid “unassigned” labels.
[0209] SingleCellNet. SingleCellNet performs cell type classification using a randomforest multi-class classification approach. Here, we trained the method over unnormalizedexpression data via the scn_train function of pySingleCellNet (v0.1.1) with nTopGenes=200, nTopGenePairs=200, nRand=100, nTrees=1000, stratify=False, and propOther=0.4. The scn_classify() function with nrand=0 was used for classification.
[0210] scmap. scmap uses a clustering approach to project cells onto a referencedataset for cell type classification. Following the recommended pipeline for scmap (v1.26.0), we log2-transformed expression data, then used selectFeatures() to select informative genes and indexCell() to create a scmap-cell index for the training dataset. For classification, we used scmapCell() to project the index onto the test dataset and scmapCell2Cluster() to obtain label assignments. A relaxed probability threshold of 0 was set to assign labels to as many cells as possible regardless of assignment confidence.
[0211] Logistic regression. We trained a logistic regression model to perform cellpotency classification using the SGDClassifier from scikit-learn (v1.4.2) with loss=‘log_loss’, default L2 regularization, and sample weights provided for class balancing. This function internally employs a one-versus-rest (OVR) strategy, training a separate binary classifier for each potency category and selecting the potency category with highest confidence at evaluation.
[0212] XGBoost. We trained and applied the XGBClassifier function from the XGBoostlibrary (v2.1.1) with default parameters and without sample weights. Like logistic regression, this method uses the OVR approach.
[0213] Linear SVM. We implemented a linear support vector machine (SVM) modelusing Scikit-learn's SGDClassifier with loss=‘hinge’ for linear support vector classification with OVR. Sample weights were provided during training.
[0214] Radial SVM. We implemented an additional SVM version using SVC fromscikit-learn (v1.4.2) with the default radial basis function (RBF) kernel and gamma=‘auto’. The default decision function, which employs an inference of OVR from one-versus-one (OVO) fits internally, was used. Sample weights were not provided during training.
[0215] Multinomial logistic regression. Using LogisticRegression from scikit-learn(v1.4.2) with multi_class=‘multinomial’, we fit a single logistic regression model for all potency categories simultaneously using cross-entropy loss and the ‘sag’ solver. A maximum number of iterations (max_iter=500) and tolerance (tol=1e-3) were set to ensure convergence. Sample weights were not provided during training.Benchmarking: Developmental potential inference and annotated gene sets
[0216] To rigorously assess performance on our compendium of 31 curated scRNA-seq datasets, we compared CytoTRACE 2 with eight published methods for predicting developmental potential from scRNA-seq data as well as nearly 19k previously annotated gene sets. Unless otherwise stated, all evaluated methods and gene sets were applied to scRNA-seq datasets individually, without batch correction or integration across datasets, with expression data normalized per author recommendations and with default parameters. All expression data were subset to the cells with known potency. Each tissue and platform pair of Tabula Sapiens and Tabula Muris datasets were run separately.
[0217] Several methods rely on human gene symbols, as noted below. For all suchinstances, we mapped mouse dataset gene symbols to their closest human orthologs, as determined by gene sequence similarity, using the GRCm39 and GRCh38.p14 annotationfiles available from Ensembl, respectively. In cases where a single human gene wasidentified as the best hit for multiple mouse genes, the mouse gene with maximum sequence similarity to was selected.
[0218] As several methods have slower running times, to promote an equitablecomparison while achieving computational feasibility, larger datasets were first downsampled. Tabula Muris dataset was downsampled to 30 cells per phenotype, separated by tissue and platform pair, and the ‘Immune cell atlas (10x)’, ‘Human breast 1 (10x)’41, ‘Human breast 2 (10x)’42, and Tabula Sapiens17datasets were downsampled to 100 cells per phenotype. Cell types in Tabula Sapiens17with less than 5 cells were removed after the prediction of each method to overcome the reduced data quality of Tabula Sapiens (“Training and test datasets”).
[0219] CytoTRACE 2. We applied CytoTRACE 2 with model ensembling andpostprocessing as described in the “The CytoTRACE 2 framework” to predict cell potency categories and scores. Datasets containing more than 100k cells were processed in batches of 100,000 cells, and diffusion was applied in batches of 10,000 cells for datasetsexceeding 10k cells. To evaluate the 19 scRNA-seq datasets included in the CytoTRACE2 training cohort, we trained a separate model for each over the remaining 16 datasets. All other datasets were evaluated with the primary version of CytoTRACE 2 trained over all training datasets.
[0220] CytoTRACE 1. CytoTRACE 1, the predecessor of CytoTRACE 2, introducedtranscriptional diversity quantified through gene counts as a correlate of developmental potential and exploited this concept to predict relative cellular potency from scRNA-seq. CytoTRACE 1 (v0.3.3) was applied with default parameters.
[0221] SCENT (SR). SCENT estimates relative cellular potency from scRNA-seq anda reference protein-protein interaction (PPI) network using single-cell signaling entropy (SR), a measure of the diversity of molecular pathway activity in a cell. SCENT (v1.0.3) was executed with the “net13Jun12” human PPI network provided with the package and otherwise default parameters. For mouse datasets, genes were first mapped to humanorthologs as described above. All gene symbols were converted to Entrez ID usingorg.Hs.eg.db (v3.15.0) in R. Gene expression matrices were normalized perdocumentation recommendation.
[0222] SCENT (CCAT). CCAT, implemented within the SCENT package, wasdeveloped as a highly efficient alternative to the original SCENT method, SCENT (SR).CCAT was applied with the same package, PPI network, and preprocessing stepsdescribed above (“SCENT (SR)”) with expression datasets prepared per documentation recommendation.
[0223] FitDevo. Similar to SCENT (CCAT), FitDevo infers cellular potency from thecorrelation between gene expression and a measure of gene weights. FitDevo (v1.2.0) was applied following tutorial instructions with binary gene weight matrix downloaded from the same source.
[0224] SLICE. SLICE relies on transcriptomic entropy for cellular potency predictionand lineage reconstruction, estimating entropy over functional groups of genes computed from Gene Ontology annotations. SLICE (v0.99.0) was applied according to demo details from the method’s GitHub page.
[0225] StemID. StemID infers cellular differentiation trajectories from scRNA-seq datawith a clustering-based algorithm analyzing links between clusters. StemID, implemented in RaceID (v0.1.4), was run according to documentation vignette instructions. For eachdataset, an SCseq object was initialized from each input gene expression matrix usingfilterData() with mintotal = 10. Ltree() and compentropy() were then applied consecutively to obtain the StemID score for cell potency.
[0226] scTour. scTour implements a deep learning architecture combining avariational autoencoder with a neural ordinary differential equation to reconstruct thedevelopmental trajectory of an input scRNA-seq dataset, oriented according to gene counts. scTour (v1.0.0) was trained and applied to each dataset individually per “Model training” documentation vignette instructions. When the raw count matrix was available for the dataset, the negative binomial conditioned likelihood loss function was used. Otherwise, the CPM / TPM expression matrix was log2-transformed, and the mean squared error loss function was used instead. Cell potency scores were obtained fromthe developmental pseudotime predictions extracted from the model training output with get_time().
[0227] mRNAsi. mRNAsi utilizes a one-class logistic regression framework toconstruct a cellular stemness index applicable to cell potency estimation from bulk andscRNA-seq data. All input gene expression matrices were CPM / TPM normalized and log2-transformed.
[0228] Gene sets. The predictive capacity of 18,706 annotated gene sets (17,810gene sets from MSigDB and 896 gene sets of transcription factor binding sites from ENCODE / ChEA) was assessed via gene set enrichment analysis. For each gene set,the AddModuleScore() function with default parameters from Seurat (v4.3.0) was appliedto each expression matrix normalized via Seurat’s NormalizeData() function.Benchmarking: Comparison to scVelo
[0229] As scVelo relies on splicing kinetics, necessitating the processing of rawsequencing data, we limited our analyses to nine gold standard datasets from the test cohort that were generated by platforms with built-in support by velocyto and for which raw sequencing data are publicly available. Raw FASTQ files for seven of these datasets, namely ‘BM-MNC (CITE-seq)’, ‘Retinal neurons (10x)’, ‘Pancreas (10x)’, ‘Peripheral glia (Smart-seq2)’, ‘Skeletal stem cell (C1)’, ‘HSCs and MPPs (inDrop)’, were obtained from the Sequence Read Archive (SRA) from NCBI, with study IDs SRP188993, SRP168426, SRP200419, SRP109011, SRP239468, and SRP094420, respectively. For ‘Peripheral glia (Smart-seq2)’, we analyzed sample IDs prefixed with ‘E12.5’. Notably, raw FASTQ files were only available for 227 of 473 cells in the ‘Skeletal stem cell (C1)’ dataset. For the remaining two datasets, ‘Mouse neurogenesis (10x)’, and ‘Mouse mature neural cell types (10x)’57, data were obtained as BAM files from SRA study ID SRP476153.
[0230] FASTQ files were downloaded using sra-tools v3.1.1 and processed withcutadapt v4.9 for adapter trimming of Smart-seq2 / C1 reads. For preprocessing of inDrop samples, dropest v0.8.6 was used. Reads were mapped and sorted BAM files were generated with STAR (v2.7.11b) and Cell Ranger (v8.0.1) using GRCm39 and GRCh38 reference genomes for mouse and human datasets, respectively. Loom files containingspliced, unspliced, and spanning reads were then generated from the BAM files along with corresponding GTFs using the velocyto.py v0.17.17 Python command line tool.
[0231] Following quantification of spliced / unspliced counts, the scVelo v0.3.1 Pythonvelocity estimation workflow was run. For all datasets, both a generalized dynamicalmodel and a differential kinetics adjusted model with grouping by the CytoTRACE 2standardized phenotypes were employed. With the exception of random_state inscvelo.pp.neighbors(), which was set to 0 to ensure reproducible results, all otherparameters were set to those in the respective vignettes, including min_shared_counts inscvelo.pp.filter_and_normalize(), which was set to 20 for dynamical models and 30 for differential kinetics models. Following velocity estimation, cell-internal latent time was inferred using scvelo.tl.latent_time(). The resulting outputs were then evaluated via absolute and relative order (see “Performance assessment” above), and CytoTRACE 2outputs were assessed over the same cells for comparison (Figures 11C and 11D).Application to induced pluripotent stem cell generation and differentiation
[0232] For the analyses presented in Figures 12B and 12C, we applied CytoTRACE2 to a time-resolved scRNA-seq dataset covering the reprogramming of mouse embryonic fibroblasts (MEFs) to induced pluripotent stem cells (iPSCs) under different protocols(Figure 12B and 12C). Author-supplied cell type labels are visualized in Figure 12B. Cellsannotated as MEF express with a standardized expression level (z-score) exceeding 1.96. We analyzed reprogramming efficiency across two protocols in which cells were either maintained in serum or transferred to a serum-free N2B272i medium on day eight. CytoTRACE 2 potency scores (Figure 12C) support a substantial increase in reprogramming efficiency with the latter protocol.
[0233] For the analysis in Figure 11E, we applied CytoTRACE 2 to a single-cell RNA-seq dataset covering in vitro differentiation of iPSCs into multipotent endoderm. We processed the data using Seurat (v4.3.0) with NormalizeData(), FindVariableFeatures(), ScaleData(), and RunPCA() with default parameters. Cell type assignment was performed based on collection day and PCA-defined pseudotime value partitions. Analysis of mouse embryogenesis
[0234] For the analyses presented in Figure 13A-13C, we downloaded and curatedsix publicly available scRNA-seq datasets spanning each embryonic day during mouse prenatal development. One dataset, which covers pre-implantation through earlyimplantation (E0.5–E4.5) (Deng et al.62), was obtained from the 19-dataset training cohortand evaluated using a CytoTRACE 2 model trained on the remaining 18 datasets to avoidoverfitting (see “Methods and annotated gene sets”). Four datasets coveringembryogenesis periods from implantation to organogenesis were previously assembled. Finally, a recent single-nucleus RNA-seq dataset covering organogenesis through birth (E8.75-P0) and generated by sci-RNA-seq3 was downloaded.
[0235] Since we compared CytoTRACE 2 against multiple methods with highlyvariable time complexity (“Methods and annotated gene sets”), all cells were randomly downsampled to 30 cells per author-supplied phenotype per time point, resulting in a combined dataset of 183,771 cells. This allowed us to balance considerations of performance versus computational efficiency. We ran each method on each dataset individually as described in “Methods and annotated gene sets” above. No dataset integration or batch normalization procedures were applied. For Organogenesis (E8.5) and Organogenesis (E8.5-P0), which were sequenced using sci-RNA-seq3, we used count data after running SCTransform of Seurat (v4.3.0) with default parameters. Due to the large size of the dataset, Organogenesis (E8.75-P0) was run with ten randomlydivided batches for SCENT (SR) and SLICE. Primordial germ cells were excluded owingto the wide range of potency levels.
[0236] For the analyses in Figures 13D to 13F, we leveraged a data-driven lineagetree of mouse embryogenesis encoded as a directed acyclic graph. Although the tree wasconstructed using a heuristic approach based on transcriptional covariance acrossembryonic time, it reflects many known parent-daughter relationships. It thus serves as aproxy for developmental potential. We defined ground truth as the distance from the root(zygote) to each daughter node (Figure 13D left). Using matching phenotype labelsbetween the tree and the data presented in Figure 13A, CytoTRACE 2 potency scores were averaged by phenotype, balanced first by time points within a given embryonic day(if any) and then by embryonic day. If the same phenotype was present in more than onedataset, we weighted equally by dataset. For each direct path in the tree (from root toleaf), the resulting scores were then converted to rank space (Figure 13D center). Toreconcile cases where a given node participates in multiple paths, we used the average rank for . CytoTRACE 1 predictions were processed in the same manner (Figure 13Dright). The resulting ranks were correlated with ground truth distances (distance from theroot) in Figure 13E. In Figure 13F, to evaluate CytoTRACE 2 in a continuous manner, we associated the mean potency score (calculated as above) for each node with the number of unique paths to all descendent leaves. Given variable sampling rates of downstream lineages in the organogenesis datasets (above), we omitted phenotypes from extraembryonic tissues in these analyses.Analysis of molecular programs: Visualization and specificity of molecularprograms learned by CytoTRACE 2
[0237] For the analyses presented in Figures 14B and 15A, we ran CytoTRACE 2 oneach of the training and test datasets, then extracted positive potency score matrix from each of the 19 models per dataset. is derived from the final layer of each GSBNmodule; obtained similarly to the potency score matrix in “Integration of scores” above,but by multiplying by only positive weights from the enrichment layer; and has a dimensionality of input cells by 114 when combined across the 19 models (=19 models in the training cohort × 6 potency categories per model). We concatenated matricesfrom the 19 models across all 33 datasets ( = 124,231 cells; downsampled as describedin “Methods and annotated gene sets” above) to produce , standardized across all cells in the training and test sets separately. We fit PCA from scikit-learn (v1.1.1) to the training set component of the resulting matrix, retaining the first three principal components, then applied the resulting projection to the training and test set components individually. Next, we repeated this process fitting UMAP from umap-learn (v0.4.6) to the PCA projection of the training component, then applying the resulting UMAP projection to the PCA projection of the training and test components individually. To adjust for differences in cell density that confound visualization, we averaged CytoTRACE 2 potency scores within each window of 0.5 UMAP units squared across the two components of UMAP space. The same procedure was applied to visualize the ground truth potency of each cell (Figure 14B, top).
[0238] We assessed the specificity of potency-associated gene sets in Figure 14C(center and bottom) using the abovementioned enrichment index (matrix in “Integration of Scores”). In brief, for a given potency category , we theenrichment indices for author-supplied phenotypes within each dataset, the median value of the resulting quantities in each ground truth potency category. Next, we calculated the pairwise difference,between the median enrichment index of and the median enrichment index of potency category and repeated this for all . Wethen calculated two test statistics: min( ) and mean( ). To simulate a null distribution,we permuted the phenotype-level enrichment indices, recomputed the median enrichment index for each ground truth potency category, and calculated both statistics. We repeated this process 10,000 times. To determine an empirical p-value for potency category , we tallied the proportion of times both statistics were as high (or higher) than the test statistics from the original data. We did this for each potency category in the training and test cohorts.
[0239] For the analysis presented in Figure 14D, we examined the expression of thetop 500 potency-associated markers learned by CytoTRACE 2 (matrix F in“Interpretability”) in gold standard training and test sets. We first filtered and mapped genesymbols in every dataset to CytoTRACE 2 input features (n = 14,271), then CPM / TPMnormalized as appropriate and log2-adjusted the data. Keeping training and test data separate, the expression matrices from each dataset were mean-aggregated into pseudo- bulk expression profiles by phenotype. We then further averaged shared phenotypes across datasets profiled by the same general platform (droplet / UMI or plate-seq / non- UMI), and finally, by species identifier (human or mouse). This resulted in a 14,271 × 237 matrix, with 14,271 genes (rows) and 237 phenotype, species, and platform combinations across training and test sets (columns). Using this matrix, we calculated the meanexpression of the top 500 positive / negative genes per potency category (matrix F in“Interpretability”), then unit variance normalized the resulting expression signatures across pseudo-bulk samples separately for training and test sets (Figure 14D).Analysis of molecular programs: Functional annotation analysis
[0240] To interpret potency-associated genes learned by CytoTRACE 2, we appliedfgsea (v1.25.1) to each rank-ordered gene list in with minSize=15 and otherwise defaultparameters. is an × matrix consisting of model importance scores for allevaluable genes (=14,271) in each of = 6 potency categories learned on the trainingcohort (“Interpretability” above;). Mouse and human MSigDb signatures from MH / H: hallmark gene sets, M2 / C2: curated gene sets including CGP and CP:WIKIPATHWAYS, CP:REACTOME, and CP:KEGG_MEDICUS; and M5 / C5: ontology gene sets. We ran fgsea on mouse and human gene sets separately, and human gene sets were limited to those with no counterpart in mouse gene sets. When running human gene sets, genes in were first mapped to human orthologs by dictionary (“Dictionary of input genes”). Weselected a subset of representative molecular signatures for display in Figure 15C, highlighting both canonical and poorly understood potency-related biology.
[0241] For Figures 15F and 15G, we evaluated conservation across tissues for the top100 positive markers per broad potency category, as determined by CytoTRACE 2 (see “Interpretability”). For this purpose, we analyzed pseudo-bulk-expression profiles of each phenotype / dataset pair in our 33-dataset potency atlas using single-sample GSEA (ssGSEA) from the GSVA package in R (v1.46.0) to mitigate technical variation. Once ssGSEA scores were obtained for all 14,271 genes in the CytoTRACE 2 feature space, we then averaged them, per gene, into the same 237 phenotype, species, and platform combinations described in “Top potency-associated genes learned by CytoTRACE 2”. Keeping training and test cohorts separate, we further averaged enrichment scores by developmental system (here, denoted ‘tissue’). For each gene and broad potency category , conservation was measured as the number of tissues for which gene showed highest mean enrichment in potency category as compared to other potency categories. We compared mean enrichment within each tissue for the available potency categories of tissue . For any potency categories lacking representation for tissue among goldstandard datasets, we inferred mean enrichment for such potency categories byaveraging over all evaluable tissues. To assess statistical significance, we repeated the above process on 100 randomly selected background genes across 100 iterations.Validation of CytoTRACE 2 by large-scale functional genomics
[0242] To assess the biological relevance of CytoTRACE 2 model features, weanalyzed large-scale in vivo CRISPR screening data of mouse hematopoiesis (Figure14E). These data encompass ~7k genes along with CasTLE –log10 p-values,representing the effect of knock out (KO) on HSC differentiation. Although these effectswere separately measured for lymphoid (n = 6,783 genes) and myeloid (n = 6,732 genes)lineages, directed –log10 p-values were well correlated between them (r = 0.78).Therefore, to create a single ordered list of KO effects, we combined directed –log10p- values for genes with effect score data in both lineages, keeping the most significant directed –log10 p-value for each gene (positive or negative). Contributions from each lineage were nearly perfectly balanced, with higher positive scores and higher negative scores implying that KO of a gene promotes or inhibits HSC differentiation, respectively.We then intersected the resulting vector with those within the CytoTRACE 2 model space,resulting in n = 5,757 genes. We applied fgsea (v1.25.1) to the rank-ordered list to jointlyevaluate the enrichment of the top 100 positive and negative multipotency markers from the CytoTRACE 2 feature matrix (Figure 14F).
[0243] To assess robustness across the number of top multipotency markers selected,we repeated the above process checking 50, 100, 200, and 500 markers (Figure 15D). We compared the median –log10 Q-value of gene set enrichment over the four gene set sizes against (i) the same process repeated for CytoTRACE 2 markers for all other potency categories, (ii) mouse HSC / MPP markers, and (iii) general cross-tissue multipotency markers. For (ii) and (iii), we leveraged ssGSEA scores from “Tissue conservation of top potency-associated genes” above, retaining only those scores that span all phenotype, species, and platform combinations from the training cohort that overlap the same 5,757 genes. For (ii), we restricted our analysis to data originating from the mouse immune system. We then mean-aggregated the ssGSEA scores by potency, subtracted the mean scores of non-multipotency categories from the scores of multipotent cells, and retained the top / bottom 500 genes. We confirmed significant enrichment of positive markers for hematopoietic stem cells using EnrichR. For (iii), we leveraged ssGSEA scores from all combinations of phenotype, species, and platform in the training set. For each gene, we then calculated pairwise AUC scores between multipotency andother potency categories and averaged the resulting scores. We used the top / bottom 500 genes by mean AUC for further analysis. Analysis of multipotency-associated programs
[0244] All WikiPathways gene sets from canonical pathways (CP) in M2 / C2 withpositive normalized enrichment scores in multipotency (see “Functional annotationanalysis” above) are presented in Figure 16A. For Figure 16C, we repeated the ssGSEAprocedure described in “Conservation of top potency-associated genes” using a gene setcomprised of the UFA factors, Fads1, Fads2, and Scd2. Mean-aggregated ssGSEAscores across 237 phenotype, species, and platform combinations in training and test sets are displayed in Figure 16C. For the statistical analysis, we applied the sameprocedure described for Figure 14C (“Visualization and specificity of potency programslearned by CytoTRACE 2”) but replaced the enrichment index with ssGSEA enrichment scores. AUCs of UFA genes (main text) were calculated for training and test sets separately using the ssGSEA scores described above, but after averaging the scores by tissue to address imbalances (as detailed in “Conservation of top potency-associated genes”). As described in “Validation of CytoTRACE 2 by large-scale functional genomics”, AUCs were first determined in a pairwise manner across all non-multipotency categories and then averaged.Candidate programs in human neoplasms: KP-Tracer analysis
[0245] For the analyses presented in Figures 18A and 18B, we assessed theassociation between CytoTRACE 2 potency predictions and clonal expansion in a Kras;Trp53(KP)-driven lung adenocarcinoma mouse model where tumor evolution was tracked with joint profiling of evolving lineage barcodes and single-cell transcriptomes. Three genotypes were studied: normal type (NT) (i.e., the standard KP model), Apc-perturbed (KPA), and Lkb1-perturbed (KPL). We obtained preprocessed expression datafor NT mice. Preprocessed expression data for KPA and KPL mice, along with inferred CNV data for NT mice, were obtained directly from the authors.
[0246] We partitioned cells for each tumor into two groups by clonal diversity using aprobabilistic coalescent model. We then calculated a clonal diversity score for each groupas the number of unique clones divided by the total number of cells per group. For each tumor, we compared the resulting scores for the two groups of cells, labeling the group with higher score as high clonal diversity and the group with lower score as low clonal diversity. To enrich for tumors in which descendant and precursor populations might both be present, we filtered for specimens in which the dominant cancer cell state, as defined by author-provided annotations, differed between high and low clonal diversity groups. This yielded 22 tumors (11 NT, 9 KPA, and 2 KPL) selected for further analysis. In support of this strategy, among selected tumors, CNV counts were significantly lower in high clonal diversity groups, suggesting that such groups arose earlier (Figure 19A, left). The same was not true in the tumors that were excluded (Figure 19A, right). Thus, within the evaluable sample set, cells with high clonal diversity appear to preferentially represent precursor populations, reinforcing the notion that selected tumors contain both precursor and descendant populations. We then applied CytoTRACE 2 to the scRNA-seq data of all selected tumors and assessed predicted potency by median CytoTRACE 2 score per clonal diversity group in each tumor (Figure 18B). We also compared the median CytoTRACE 2 score per clonal diversity group with the median number of inferred CNVs in the NT group (for which such data were available) (Figure 19B). Candidate programs in human neoplasms: Oligodendroglioma analysis
[0247] For Figure 7D, we applied CytoTRACE 2 to scRNA-seq profiles of sixoligodendrogliomas, with coordinates for the associated oligodendroglioma 2D lineage hierarchy embedding. We then assigned malignant oligodendroglioma cells to four transcriptional states and visualized the association of CytoTRACE 2 potency predictions with the author-supplied stemness score. For the latter, we separated cells according to the stemness score by partitioning them into successive intervals of 0.25 units. We thendisplayed CytoTRACE 2 potency scores as a function of each interval (Figure 18D, right).
[0248] For the analyses in Figure 19K, we assessed how CytoTRACE 2 captures theIDH inhibitor-induced differentiation of oligodendroglioma described by Spitzer et al. as compared to gene author-defined gene sets. We supplemented tumor data from Tirosh et al. with four additional oligodendroglioma samples. To focus on malignant cells, we selected the same CD45–cells analyzed by the authors for samples with flow-sorting dataavailable, removing normal oligodendrocytes (MGH170 and MGH229). For the other twosamples (BWH445 pre- and post-treatment), we applied Seurat (v4.3.0) as describedbelow (“Correlates of ICI response in melanoma scRNA-seq data”) and excluded clusters expressing macrophage or oligodendrocyte markers. We then applied AddModuleScore() from Seurat to calculate single-cell enrichment scores for author-provided gene sets and determined mean CytoTRACE 2 potency scores and AddModuleScore values in each tumor.Candidate programs in human neoplasms: sCRNA-seq tumor atlases
[0249] For the analyses presented in Figures 18C to 18F and 19C to 19J, wedownloaded 43 preprocessed scRNA-seq tumor atlases spanning 17 cancer types, 615 tumors, and 430,332 malignant cells from the Curated Cancer Cell Atlas website on June 28, 2023. We selected these datasets because they encompass diverse developmental origins, including solid and hematological malignancies, and have matching cancer typesin TCGA and PRECOG, making it possible to link potency-associated molecularprograms with clinical metadata at scale. For quality control, we excluded tumor samples with less than 10 malignant cells, cancer types with less than five available tumor samples, redundant samples based on duplicate identifiers, and cell line samples. We also excluded datasets with less than 10,000 genes in the CytoTRACE 2 input dictionary (“Dictionary of input genes” above).
[0250] In all cases, we leveraged author-supplied cell type annotations, includingclassifications of malignant and non-malignant cells from 3CA. We also standardized non- malignant cell phenotypes based on author-provided cell types and further classified them into three broader categories based on shared lineage: immune and stromal cells in thetumor microenvironment (TME) and non-TME non-malignant cells. All non-malignantcells in each dataset were uniformly downsampled without replacement to 100 cells per standardized phenotype per tumor type for marker discovery and copy number inference. Depending on platform, all datasets were standardized to log2 CPM or TPM space as appropriate.Candidate programs in human neoplasms: Single-cell potency prediction andanalysis
[0251] After curating scRNA-seq tumor atlases by the above protocol, we ranCytoTRACE 2 with default parameters (“Methods and annotated gene sets”) on all annotated malignant cells from each tumor sample.Candidate programs in human neoplasms: Copy number analysis
[0252] For Fig. 19D, we inferred copy number profiles of individual tumor samplesusing inferCNV (v1.12.0). As input, non-malignant cells were selected per tumor sample as described in “Identification of potency-associated markers”. Tumors with <100 malignant cells, <10 cells in any potency category, <2 predicted potency categories, or <100 non-malignant cells were excluded. For efficiency, malignant and non-malignant cells were further downsampled to a maximum of 1,000 cells each. Following thisprocess, we then ran inferCNV on each of the remaining 373 tumor samples with cutoff= 0.1 (droplet datasets) or cutoff = 1 (plate-seq datasets) and otherwise defaultparameters.
[0253] To assess the relationship between predicted potency and inferred copynumber patterns, we calculated a co-association index for each evaluable tumor sampleas follows. First, we applied PCA to the Euclidean distance matrix of all single cell inferCNV profiles for a given tumor sample. Next, we applied nn2 from the RANN package (v2.6.1) in R with = 2 to the first two PCs. For cells in each potency category , we tallied the number of nearest neighbors matching . To determine a null distribution, we then randomized potency category labels across all single cells and repeated the above process 10,000 times, allowing us to calculate an empirical p-value for the number ofnearest neighbors of with matching potency category labels expected by randomchance. We adjusted potency-specific p-values per tumor using the Benjamini-Hochberg method for multiple hypothesis testing correction, selected the minimum resulting value per tumor sample, then applied a Benjamini-Hochberg correction again across all tumorsamples. The resulting Q values are termed co-association indices.Candidate programs in human neoplasms: Identification of biomarkers
[0254] We employed the following meta-analytical procedure to identify potency-associated marker genes, prioritizing genes with consistent representation across tumor samples and datasets within a given cancer type. We began by combining, for each dataset, single-cell transcriptomes of non-malignant cells (downsampled as described above) with single-cell transcriptomes of malignant cells from each tumor sample. Following sample exclusions, only one dataset (‘Glioblastoma’) consisted of scRNA-seq data profiled on more than one platform; in this case, we separately combined normal cells from each platform to avoid technical biases. We then applied FindAllMarkers() in Seurat (v4.3.0) to each tumor sample separately, with only.pos = F, min.pct = -Inf, logfc.threshold = -Inf, min.cells.feature = -Inf, min.cells.group = 10, return.thresh = Inf, and otherwise default parameters. We stored the average log2 fold change (LFC) and adjusted p-value for cells of each potency category predicted by CytoTRACE 2 along with cells of non-malignant phenotypes standardized as described above (“scRNA-seq tumor atlases”). For non-malignant phenotypes, LFC values were further averaged within three broader phenotypic categories to balance major lineage compartments.
[0255] Next, we averaged LFC values of each phenotype (including potencycategories) for tumor samples of a given cancer type . This was done across datasets,filtering for common genes across datasets within cancer type , yielding an by matrix, containing LFC values with genes across phenotypes. To increase the specificityof top-scoring markers, we calculated the minimum delta between , and , over allphenotypes . This yielded an by matrix containingvalues for each gene and phenotype in cancer type . To select the top markers, we then created a ranked gene list for each phenotype in cancer type by sorting,indecreasing order. This was done for all genes with , > 0.1 and , > 0.5. The topgenes in this list were selected as potency-associated markerstype . If less than genes satisfied these criteria, we relaxed the,requirement and padded the list with the remaining genes sorted by decreasinguntil we obtained a maximum of genes.
[0256] For Figure 19E, we quantified the enrichment of AML cell-type-specific genesignatures (‘LSPC-Primed-Top100’, ‘LSPC-Quiescent’, ‘GMP-like-Top100’, and ‘Mono-like-Top100’;), each expected to be enriched in multipotent, multipotent, oligopotent, and unipotent / differentiated cells, respectively. We then calculated the LFC matrix of AML as described above, restricted it to CytoTRACE 2 potency labels, and normalized each gene in the matrix to mean zero and unit variance. Finally, using the resulting values, we calculated the average expression of all genes in each gene signature. Candidate programs in human neoplasms: Specificity of potency-associated markers
[0257] For Figure 19F, we confirmed the specificity of identified marker genes in thefollowing manner. Given a phenotype in a cancer type , we averaged LFC , acrossall cancer-type-specific potency markers . We then measured the average LFC for malignant cells with matching potency, malignant cells with non-matching potency (averaged across non-matching potency categories in ), and the following categories when present: adjacent normal cells, immune cells, and non-immune stromal cells. The resulting values were unit variance normalized for each evaluable potency category and cancer type pair.
[0258] For the analysis presented in Figure 19H, we tested the deconvolutionperformance of interrogating potency-associated marker genes in bulk tumors using a leave-one-dataset-out cross-validation approach (Figure 19G). Specifically, for each cancer type with at least two scRNA-seq datasets (11 of 17 cancer types), we serially held one dataset out. We then reconstituted tumor samples in the held-out dataset by averaging the expression profiles of all cells from each tumor (including non-malignant cells) in non-log linear CPM / TPM space, depending on platform. Using the non-held-out dataset(s), we calculated a new set of potency-associated marker genes as described in “Identification of potency-associated markers”. We then calculated the geometric mean of CPM / TPM expression levels of newly identified markers in held-out pseudo-bulksamples. To evaluate performance, we determined – for each potency-associated genesignature – the Pearson correlation between marker gene expression levels and therelative abundances of the same potency category (as imputed by CytoTRACE 2) in held- out pseudo-bulk samples. We then repeated this process for every fold.Candidate programs in human neoplasms: Bulk tumor expression datasets
[0259] We assembled preprocessed bulk RNA-seq and microarray expression profilesfrom TCGA and PRECOG for the same 17 cancer types analyzed from 3CA. Read countsfrom TCGA were obtained via gdc-client v1.5.0 in June 2020. Each count matrix wasnormalized to TPM and log2 adjusted. Ensembl gene IDs were converted to HGNC gene symbols using the MAGeCKFlute::TransGeneID R function (v1.99.2). In total, 52 duplicate samples were removed Expression data in log2space were analyzed unchanged. Datasets with a minimum expression value of at least 0 and a maximum value >30 were log2 adjusted, then quantile normalized using normalize.quantiles() from the preprocessCore package (v1.62.0) in R. Astrocytoma samples were obtained from TCGA-LGG and three astrocytoma bulk datasets from PRECOG. Oligodendroglioma samples were obtained from TCGA-LGG.Candidate Programs in human neoplasms: Survival analysis
[0260] To assess survival associations of predicted potency markers in TCGA andPRECOG, we first calculated the mean log2expression of each gene set. Next, we applied Cox proportional hazards regression using coxph() from the survival package (v3.5-6) in R to calculate z-scores relating gene set expression levels to overall survival. We did this, per dataset, for tumor samples of the same cancer type. As many cancer types in PRECOG contain multiple datasets, we integrated z-scores within a cancer type using Liptak’s method, with weights set to the square root of sample sizes. To integrate z-scores across TCGA and PRECOG for a given cancer type, we again applied Liptak’s method with weights set to the square root of sample sizes, and to combine the resulting meta-z-scores across evaluable cancer types for a given potency category, we applied Stouffer’s method as previously described. Any potency categories not identified in a given cancer type were omitted from the meta-z-score calculation. We expressed all meta-z-scores as –log10p-values with directionality for clarity.
[0261] Survival associations were calculated in various contexts using the meta-zintegration scheme detailed above. In Figure 18E, we benchmarked CytoTRACE 2 potency-associated markers against previously described cancer gene sets in oligodendroglioma and AML. We first calculated the average expression of CytoTRACE2 potency-associated markers per potency category and published gene sets. We thenassessed survival associations by multivariable Cox regression with all CytoTRACE 2 potency-associated markers and published gene sets included as covariates. For thepan-cancer analysis in Figure 18F, we calculated survival associations via univariableanalysis, we applied multivariable Cox regression with all evaluable potency categories included as covariates. To assess the impact of proliferation, we calculated the averageexpression of proliferation-related markers in each tumor. We then included proliferationas an additional covariate in bivariable Cox models incorporating each potency category. We repeated this process for pluripotency-associated markers as well. Several analyses were restricted to TCGA owing to the availability of additional covariates. For examplewe analyzed the influence of stage, age, and sex on potency-related survival associations using multivariable analysis, including pathologic stage and age as continuous variables and sex as a binary variable. We also evaluated potency-associated survival associations in TCGA after controlling for tumor grade, tumor purity (ABSOLUTE and ESTIMATE), mRNAsi, and TmS.Candidate programs in human neoplasms: Analysis of malignant cell transcripthallmarks
[0262] Briefly, using datasets processed as described in “scRNA-seq tumor atlases”,malignant cells of each tumor sample were scored for expression of each module program, for which at least 50% of genes were present in the expression matrix, using sigScores() from the scalop R package (v1.1.0). Module programs with scores >1 were considered expressed (Figure 19I). Candidate programs in human neoplasms: Correlates of predicted potency and drug resistance signature in AML
[0263] For the analysis in Figure S7J, we analyzed the associations of predictedpotency and expression of LinClass-7, a 7-gene signature associated with immaturity and drug resistance in AML. We calculated the average LFC across the genes in LinClass-7using the same pipeline as Figure S7E, described in “Identification of potency-associatedmarkers”.Candidate programs in human neoplasms: Correlates of ICI response in melanomascRNA-seq data
[0264] For the analysis presented in Fig. 18G, we investigated whether candidatepotency markers are associated with ICI response. To this end, we identified scRNA-seq profiles (SMART-seq2) from eight post-treatment ICI-resistant melanomas with at least 10 malignant cells each included in 3CA. To augment these data, we generated scRNA-seq data (10x Chromium) from five additional melanoma samples from patients treatedwith ICIs, including three patients in whom durable clinical benefit (DCB) was achieved (“Human patient samples” above). Using Seurat (v4.3.0) to preprocess these newly generated samples, we first removed cells with <200 detectable genes and mitochondrial gene ratios >0.25. We then ran NormalizeData(), FindNeighbors() with dims = 1:10, FindClusters() with default parameters, and RunUMAP() with dims = 1:10. We removedclusters with >50% of cells predicted as doublets using DoubletFinder (v2.0.3). We thenannotated clusters based on expression of canonical marker genes (positive markersunless indicated otherwise): MITF, S100B, MLANA, PMEL, or TYR = melanoma cells;EPCAM = epithelial cells; VWF or PECAM1 = endothelial cells; COL1A1 or COL3A1 =fibroblasts; ACTA2 or TAGLN = smooth muscle cells; CD79A or CD79B = B cells; IGKCor MZB1 = plasma cells; CD3D and CD4 or IL7R = CD4 T cells; CD3D and CD8A orCD8B = CD8 T cells; low CD3D and high GNLY = NK cells; CD68 or CD163 =macrophages; CD1C or CD209 = dendritic cells. Malignant and non-malignantclassifications were corroborated by inferCNV, applied as above (“Copy number analysis”).
[0265] Using melanoma-associated potency markers defined from 3CA datasets, forwhich markers of multipotent-like, oligopotent-like, and differentiated-like cells were defined, we first harmonized across platforms (Smart-seq2 and 10x Chromium) byretaining only those genes present in all 13 evaluable expression matrices. We thencalculated the mean log2expression, denoted LE, of each potency category in malignant cells from each sample. To combat technical variation across samples, we furthernormalized LE values using a potency category with “opposite” predicted developmentalpotential. More specifically, within each sample, we subtracted the LE of differentiated-like cells from the LE values of multipotent-like and oligopotent-like cells. Separately, weaveraged the LE values of multipotent-like and oligopotent-like cells, then subtracted theresulting quantities from differentiated-like cells.Candidate programs in human neoplasms: Correlates of ICI response in bulk tumorexpression datasets
[0266] For Fig. 18H, preprocessed bulk tumor RNA-seq data and correspondingclinical information from 597 patients with cancer treated with anti-PD-1 monotherapy oranti-PD-1 / anti-CTLA-4 combination therapy ICIs were obtained from public repositories,including TIGER and TIDE. For Fig. 14I, preprocessed bulk RNA-seq data from mousetumor models treated with ICIs, accompanied by treatment and response annotations,were downloaded from TISMO. All datasets were either obtained or processed as log2TPM prior to analysis, and only cancer types present in our 3CA analyses wereconsidered. Datasets, treatment pairs with fewer than 5 samples were excluded
[0267] To quantify the enrichment of candidate potency markers from each evaluablecancer type, we first standardized expression for each gene to zero mean, unit variancein each dataset, subsetting by ICI time point and treatment to mitigate batch effects. We then repeated the same procedure for scoring potency-associated markers as described above in “Correlates of ICI response in melanoma scRNA-seq data” to ensure consistency. Expression enrichments were averaged by dataset, ICI response status, andICI treatment in Figs. 118H and 19M.Quantification and statistical analysis
[0268] Relationships between two ordered variables were assessed by correlationtests or linear regression. Two-group comparisons were assessed using unpaired or paired tests, as appropriate. Differences across three or more groups were analyzedusing a Kruskal-Wallis test. Results with P < 0.05 were considered significant. Dataanalyses were performed with Python (v3.9.0) and R (v4.2.0+).Table 1. Biomarkers for Cellular Potency Phenotypes Top 500 genes positively associated with each potency phenotype as identified by feature importance weight, including all genes with feature importance weight tied with those of rank 500. Genes negatively associated with each potency phenotype are identified by the method as well (data not shown). FeaturePotency category Gene symbolimportance weight Totipotent Depdc7 73Totipotent Psip1 68Totipotent 67Totipotent 67Totipotent 67Totipotent 67Totipotent 66Totipotent 66Totipotent 66Totipotent 65Totipotent 65Totipotent 65Totipotent 64Totipotent 63Totipotent 63Totipotent 63Totipotent 62Totipotent 61Totipotent 61Totipotent 59Totipotent 59Totipotent 59Totipotent 59Totipotent 59Totipotent 58Totipotent 58Totipotent 57Totipotent 57Totipotent 57Totipotent 57Totipotent 56Totipotent 56Totipotent 56Totipotent 56Totipotent Celf1 55Totipotent Nanos1 55Totipotent Rps29 55Totipotent Tiparp 55Totipotent Usp16 55Totipotent Aspm 54Totipotent Cir1 54Totipotent Nipbl 54Totipotent Fth1 53Totipotent Zfp420 53Totipotent Gatad2b 52Totipotent Tuba4a 52Totipotent Usp34 52Totipotent Yap1 52Totipotent Zc3h13 52Totipotent Brwd1 51Totipotent Cdk12 51Totipotent Jak2 51Totipotent Kctd10 51Totipotent Usp7 51Totipotent Chd7 50Totipotent Gda 50Totipotent Macrod2 50Totipotent Map2k1 50Totipotent Mt1 50Totipotent Snapc1 50Totipotent Tsg101 50Totipotent Arih1 49Totipotent Dnajc3 49Totipotent Epc1 49Totipotent Jmy 49Totipotent Mis18a 49Totipotent Ralbp1 49Totipotent Dynll2 48Totipotent Eif5b 48Totipotent Epn2 48Totipotent Fa2h 48Totipotent Nup214 48Totipotent Srsf5 48Totipotent Usp1 48Totipotent Ccdc59 47Totipotent Gins2 47Totipotent Gtf2b 47Totipotent Jarid2 47Totipotent Kif3a 47Totipotent Rala 47Totipotent Rpl37 47Totipotent Atrx 46Totipotent Eps8 46Totipotent Esco2 46Totipotent Ggnbp2 46Totipotent Gtf2a1 46Totipotent Igsf11 46Totipotent Mbtd1 46Totipotent Rundc3b 46Totipotent Serf2 46Totipotent Spry4 46Totipotent Wtap 46Totipotent Atp6v1d 45Totipotent Calm1 45Totipotent Ipcef1 45Totipotent Mfsd2a 45Totipotent Srek1ip1 45Totipotent Tpd52 45Totipotent Twistnb 45Totipotent Atp6v1b2 44Totipotent Bnc1 44Totipotent Chd1 44Totipotent Dclk2 44Totipotent Eif2s2 44Totipotent Map9 44Totipotent Mpdz 44Totipotent Mpp5 44Totipotent Otud4 44Totipotent Parp12 44Totipotent Rps27 44Totipotent Runx1t1 44Totipotent Siah2 44Totipotent Zfand6 44Totipotent Ccdc15 43Totipotent Cnbp 43Totipotent Hhex 43Totipotent Iqgap2 43Totipotent Mxd1 43Totipotent Prkar1a 43Totipotent Usp2 43Totipotent Zfp281 43Totipotent Arpc2 42Totipotent Atp2b1 42Totipotent Pabpc4 42Totipotent Rnf141 42Totipotent Socs7 42Totipotent Spin1 42Totipotent Zcrb1 42Totipotent 42Totipotent 41Totipotent 41Totipotent 41Totipotent 41Totipotent 41Totipotent 41Totipotent 41Totipotent 41Totipotent 41Totipotent 41Totipotent 41Totipotent 40Totipotent 40Totipotent 40Totipotent 40Totipotent 40Totipotent 40Totipotent 40Totipotent 40Totipotent 40Totipotent 40Totipotent 40Totipotent 39Totipotent 39Totipotent 39Totipotent 39Totipotent 39Totipotent 39Totipotent 39Totipotent 39Totipotent 38Totipotent 38Totipotent 38Totipotent 38Totipotent 38Totipotent 38Totipotent 38Totipotent 38Totipotent 37Totipotent 37Totipotent 37Totipotent 37Totipotent 37Totipotent Crk 37Totipotent Dpf3 37Totipotent Fam131a 37Totipotent Hap1 37Totipotent Ktn1 37Totipotent Nefh 37Totipotent Plekha3 37Totipotent Pom121 37Totipotent Suds3 37Totipotent Tnrc6b 37Totipotent Ubc 37Totipotent Akap2 36Totipotent Arhgap44 36Totipotent Bag5 36Totipotent Daam1 36Totipotent Foxo3 36Totipotent Gng12 36Totipotent Grip1 36Totipotent Gtsf1 36Totipotent Hddc3 36Totipotent Krr1 36Totipotent Rab3d 36Totipotent Tmcc2 36Totipotent Tmem87a 36Totipotent Trim13 36Totipotent Uck2 36Totipotent Ywhab 36Totipotent Arid1a 35Totipotent Cgn 35Totipotent Cnn3 35Totipotent Ezh1 35Totipotent Klhl8 35Totipotent Lpin2 35Totipotent Ppp1r3d 35Totipotent Slc30a1 35Totipotent Tshz1 35Totipotent Zcchc2 35Totipotent Abl2 34Totipotent Ccdc181 34Totipotent Ccdc92 34Totipotent Cdc40 34Totipotent Cdc42se2 34Totipotent Esco1 34Totipotent Gfod1 34Totipotent Itm2b 34Totipotent Mat2b 34Totipotent Mepce 34Totipotent Rbpms2 34Totipotent 34Totipotent 34Totipotent 34Totipotent 34Totipotent 34Totipotent 34Totipotent 34Totipotent 34Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 33Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent 32Totipotent Sdf4 32Totipotent Slc39a10 32Totipotent 32Totipotent 32Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 31Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent Hmgcll1 30Totipotent Med10 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 30Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent 29Totipotent Acox1 28Totipotent Atg14 28Totipotent Crybg3 28Totipotent Cwf19l2 28Totipotent Ddx4 28Totipotent Dnajb6 28Totipotent Dyrk1a 28Totipotent Eif1 28Totipotent Eif2ak4 28Totipotent Elovl4 28Totipotent Fam184a 28Totipotent Foxj3 28Totipotent Golga4 28Totipotent Hpd 28Totipotent Ier2 28Totipotent Ipo8 28Totipotent Lars2 28Totipotent Lmo1 28Totipotent Lrrcc1 28Totipotent Lyar 28Totipotent Mkrn1 28Totipotent Nudt4 28Totipotent Otud3 28Totipotent Pik3c2b 28Totipotent Plk3 28Totipotent Ppp1r12a 28Totipotent Rab6a 28Totipotent Rab7 28Totipotent Reep5 28Totipotent Slc6a9 28Totipotent Spag1 28Totipotent Sprtn 28Totipotent Susd3 28Totipotent Tm7sf3 28Totipotent Tmem97 28Totipotent Tpr 28Totipotent Zcchc10 28Totipotent Arap2 27Totipotent Arfgef1 27Totipotent Baiap2 27Totipotent Bmpr1b 27Totipotent Cenpe 27Totipotent Cnot6l 27Totipotent Cpsf2 27Totipotent Elavl3 27Totipotent Etv5 27Totipotent Fam122a 27Totipotent Fbrs 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 27Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent 26Totipotent Hook1 26Totipotent Igf2bp3 26Totipotent Kpna3 26Totipotent Papd7 26Totipotent Plag1 26Totipotent Rab20 26Totipotent Rph3al 26Totipotent Sept7 26Totipotent Sgms2 26Totipotent Slc25a48 26Totipotent Slc35c1 26Totipotent Sorbs2 26Totipotent Stx6 26Totipotent Sub1 26Totipotent Tiam2 26Totipotent Ubb 26Totipotent Zar1 26Totipotent Zfyve9 26Totipotent Zswim4 26Totipotent Actr3 25Totipotent Atp6v1f 25Totipotent Bclaf1 25Totipotent Caml 25Totipotent Ccdc173 25Totipotent Cks1b 25Totipotent Cpe 25Totipotent Dhrs1 25Totipotent Dusp1 25Totipotent Exosc9 25Totipotent Gdap2 25Totipotent Gm527 25Totipotent Igsf3 25Totipotent Kif5b 25Totipotent Lrrn4 25Totipotent Map1lc3b 25Totipotent Mcu 25Totipotent Mest 25Totipotent Ninj1 25Totipotent Nr6a1 25Totipotent Nucks1 25Totipotent Nufip1 25Totipotent Pcm1 25Totipotent Pls1 25Totipotent Prelid2 25Totipotent Ptpn23 25Totipotent Rab2a 25Totipotent Reep2 25Totipotent Rfk 25Totipotent Rnf111 25Totipotent Slc9a3r1 25Totipotent Smarca2 25Totipotent Smu1 25Totipotent Ssx2ip 25Totipotent Stk3 25Totipotent Stk35 25Totipotent Usp54 25Totipotent Utp23 25Totipotent Zfp462 25Totipotent Zp3 25Pluripotent Slc6a8 219Pluripotent Pou5f1 218Pluripotent Bcat1 211Pluripotent Dppa5a 205Pluripotent Aebp2 202Pluripotent Dppa3 201Pluripotent Calcoco2 200Pluripotent Sgk1 195Pluripotent Bhmt 190Pluripotent Mta3 190Pluripotent Phc1 190Pluripotent Slc7a8 187Pluripotent Btrc 186Pluripotent Slc2a3 185Pluripotent Terf1 179Pluripotent Cdc20 179Pluripotent Fbxo15 178Pluripotent Tfcp2l1 177Pluripotent Ddah1 176Pluripotent Fam151a 175Pluripotent Dppa4 174Pluripotent Snx27 174Pluripotent Brd3 174Pluripotent Gldc 173Pluripotent Nanog 173Pluripotent Pla2g16 173Pluripotent Fbp2 172Pluripotent Psmc4 172Pluripotent Fabp3 169Pluripotent Cdc6 165Pluripotent Tex19.2 165Pluripotent G3bp2 162Pluripotent Hhla1 162Pluripotent Gata3 161Pluripotent L1td1 159Pluripotent Trim28 158Pluripotent Uba52 158Pluripotent Pip4k2a 157Pluripotent Bnip3l 157Pluripotent Igf2bp1 157Pluripotent Alppl2 156Pluripotent Atg2a 156Pluripotent S100pbp 156Pluripotent Rps9 155Pluripotent Ttc9c 152Pluripotent Aass 151Pluripotent B3galt1 151Pluripotent Gata6 151Pluripotent F11r 150Pluripotent Mrpl43 149Pluripotent Ppp2r1a 149Pluripotent Sox15 148Pluripotent Mylpf 147Pluripotent Thg1l 147Pluripotent Pttg1 146Pluripotent Rnf168 146Pluripotent Akap12 145Pluripotent B3gnt7 144Pluripotent Cyp2s1 144Pluripotent Cd84 143Pluripotent Igf2r 143Pluripotent Epas1 142Pluripotent Akap1 142Pluripotent Ash2l 141Pluripotent Mad2l2 141Pluripotent Jarid2 140Pluripotent Ccna1 140Pluripotent Abcf1 139Pluripotent Taldo1 139Pluripotent Echdc2 137Pluripotent Rbm47 137Pluripotent Cul2 137Pluripotent Tead4 137Pluripotent Ipo4 137Pluripotent Pkd1l3 137Pluripotent Pmm1 136Pluripotent Arih2 136Pluripotent Rpl18 135Pluripotent Qars 135Pluripotent Acer2 133Pluripotent Uimc1 133Pluripotent Mybl2 132Pluripotent Ubxn1 131Pluripotent Ptgfrn 130Pluripotent Lamp1 129Pluripotent Usp48 129Pluripotent Stk32b 129Pluripotent Snrpd3 128Pluripotent Eif3l 128Pluripotent Rpusd4 127Pluripotent Hyou1 126Pluripotent Pbdc1 126Pluripotent Ubxn11 126Pluripotent Grn 125Pluripotent Stip1 125Pluripotent Elp3 124Pluripotent Wdr45 124Pluripotent Gpm6b 122Pluripotent Riok3 122Pluripotent Olfml1 121Pluripotent Ptges 120Pluripotent Ybx1 119Pluripotent Sass6 119Pluripotent Pdgfa 119Pluripotent Cited1 118Pluripotent Atg13 116Pluripotent Coil 116Pluripotent Papola 116Pluripotent Klf5 116Pluripotent Sfn 116Pluripotent Clpx 116Pluripotent Nup133 116Pluripotent Rabggtb 116Pluripotent Snx5 116Pluripotent Lars2 115Pluripotent Svopl 115Pluripotent Rpl28 115Pluripotent Dab2 114Pluripotent Ak4 114Pluripotent Amt 114Pluripotent Csnk1e 114Pluripotent Fer 114Pluripotent Wdr47 113Pluripotent Atxn3 112Pluripotent Cxadr 112Pluripotent Slc25a4 112Pluripotent Tbrg1 112Pluripotent Hspa1b 112Pluripotent Slc25a3 112Pluripotent Atp6v0d1 111Pluripotent Mkrn1 110Pluripotent Flnb 110Pluripotent Rpl39l 110Pluripotent Tbc1d23 110Pluripotent Sema4b 110Pluripotent Mtpap 110Pluripotent Eef1b2 110Pluripotent Sord 109Pluripotent Fmr1nb 109Pluripotent Tmem62 109Pluripotent Cd59b 109Pluripotent Slc34a2 109Pluripotent Nr3c2 109Pluripotent Rpl9 109Pluripotent Rbpms 108Pluripotent Ube2v1 108Pluripotent Slc46a1 107Pluripotent Grpel1 107Pluripotent Lrpap1 106Pluripotent Plk1 106Pluripotent Cd9 106Pluripotent Rpl14 106Pluripotent Rbpj 105Pluripotent Ptprz1 105Pluripotent Sgta 105Pluripotent Cnot1 105Pluripotent Ulk1 104Pluripotent Tfap2c 104Pluripotent Rpl30 104Pluripotent Phactr1 103Pluripotent Ctsd 103Pluripotent Tbrg4 103Pluripotent Mcm10 103Pluripotent Zfp106 103Pluripotent Hist1h1a 103Pluripotent Drg2 103Pluripotent Mvp 103Pluripotent Elmo1 103Pluripotent Aars 102Pluripotent Znhit3 102Pluripotent Gcnt3 102Pluripotent Atp6v1c1 101Pluripotent Khdc1b 101Pluripotent 101Pluripotent 100Pluripotent 100Pluripotent 100Pluripotent 100Pluripotent 100Pluripotent 100Pluripotent 99Pluripotent 99Pluripotent 99Pluripotent 99Pluripotent 99Pluripotent 99Pluripotent 99Pluripotent 98Pluripotent 98Pluripotent 98Pluripotent 98Pluripotent 98Pluripotent 98Pluripotent 98Pluripotent 98Pluripotent 98Pluripotent 97Pluripotent 97Pluripotent 97Pluripotent 97Pluripotent 97Pluripotent 97Pluripotent 96Pluripotent 96Pluripotent 96Pluripotent 96Pluripotent 96Pluripotent 95Pluripotent 95Pluripotent 95Pluripotent 95Pluripotent 95Pluripotent 95Pluripotent 95Pluripotent 94Pluripotent 94Pluripotent 94Pluripotent Sugt1 93Pluripotent Alpl 93Pluripotent Gpd2 93Pluripotent Renbp 93Pluripotent Psmc1 93Pluripotent Spp1 93Pluripotent Polr1a 93Pluripotent Smyd5 93Pluripotent Rrp12 93Pluripotent Gpx4 93Pluripotent Pcbp1 92Pluripotent Kdm4a 92Pluripotent Kat7 92Pluripotent Ccdc30 92Pluripotent Macf1 92Pluripotent Ttyh2 92Pluripotent Sephs1 92Pluripotent Ddx25 92Pluripotent Arhgap5 92Pluripotent Tet1 91Pluripotent Nfatc2ip 91Pluripotent Rnd1 91Pluripotent Tmem41b 90Pluripotent Isyna1 90Pluripotent Gchfr 90Pluripotent Smn1 90Pluripotent Pdgfrl 90Pluripotent Pomp 90Pluripotent Ap2m1 90Pluripotent Serpinf1 90Pluripotent Tpd52 89Pluripotent Pard3 89Pluripotent Fn3krp 89Pluripotent Cacna2d1 89Pluripotent Rpap1 89Pluripotent Rabgap1l 89Pluripotent Ppat 89Pluripotent Diablo 89Pluripotent Mfge8 89Pluripotent Psap 88Pluripotent Egln1 88Pluripotent Hsf2bp 88Pluripotent Kalrn 88Pluripotent Abcf2 88Pluripotent Nefm 88Pluripotent Cbs 87Pluripotent Aqp3 87Pluripotent Rnf7 87Pluripotent Dync1h1 87Pluripotent Usp10 87Pluripotent Fgfr2 87Pluripotent Lin7c 87Pluripotent Exo1 87Pluripotent Gls2 87Pluripotent Haus2 87Pluripotent Clu 86Pluripotent Eif2b1 86Pluripotent Sh3bgrl3 86Pluripotent Pmaip1 86Pluripotent Slc44a1 86Pluripotent Mrpl20 86Pluripotent Mfn1 85Pluripotent Ctsa 85Pluripotent Gm10184 85Pluripotent Pfdn6 85Pluripotent Inpp5f 85Pluripotent Zbtb44 85Pluripotent Fam50a 85Pluripotent Trip12 84Pluripotent Ino80d 84Pluripotent Gdf3 84Pluripotent Pqbp1 84Pluripotent Slc39a4 84Pluripotent Tcf7l2 84Pluripotent Psmc5 84Pluripotent Rbbp6 84Pluripotent Wnk3 83Pluripotent Cldn7 83Pluripotent Neo1 83Pluripotent Zfp296 83Pluripotent Gnptab 83Pluripotent Ngdn 83Pluripotent Asns 83Pluripotent Pld3 82Pluripotent Cdipt 82Pluripotent Dnmt3b 82Pluripotent Smg7 82Pluripotent Eif3b 82Pluripotent Lpcat4 82Pluripotent Hist1h1e 82Pluripotent Dock10 82Pluripotent Hip1 81Pluripotent Mrps18b 81Pluripotent Bcas3 81Pluripotent Gpr176 81Pluripotent Slc5a6 81Pluripotent Lrrc8a 81Pluripotent Slc28a3 81Pluripotent Ppp1cc 81Pluripotent Prkd2 81Pluripotent Tardbp 81Pluripotent Rbm3 81Pluripotent Pdcl3 80Pluripotent Triml2 80Pluripotent Cln8 80Pluripotent Snx2 80Pluripotent Arpc4 80Pluripotent Vdac2 80Pluripotent Dcp2 79Pluripotent Ccdc88c 79Pluripotent Klf17 79Pluripotent Ankrd12 79Pluripotent Cds2 79Pluripotent Ankzf1 79Pluripotent Samd8 79Pluripotent Sfrp2 79Pluripotent Cndp2 79Pluripotent Esrp2 79Pluripotent Lap3 79Pluripotent Rps10 79Pluripotent Rpl19 79Pluripotent Rnasek 78Pluripotent Aldoc 78Pluripotent Lamtor2 78Pluripotent Sun1 78Pluripotent Apoc1 78Pluripotent Cdc14b 78Pluripotent Cluh 78Pluripotent Supt5 78Pluripotent Prpf31 78Pluripotent Ckb 77Pluripotent Rbpms2 77Pluripotent Nol9 77Pluripotent Usf1 77Pluripotent Rnf125 77Pluripotent Kdr 77Pluripotent Tax1bp3 77Pluripotent Mrto4 77Pluripotent Fn1 76Pluripotent Cecr2 76Pluripotent Hk2 76Pluripotent Gnpda1 76Pluripotent Myo5b 76Pluripotent Rpl10l 76Pluripotent Cops2 76Pluripotent Bysl 76Pluripotent Kif22 76Pluripotent Cops5 76Pluripotent Hat1 76Pluripotent Dnajb6 75Pluripotent Nudt4 75Pluripotent Adh1 75Pluripotent Arhgap1 75Pluripotent Mmp16 75Pluripotent Adcy5 75Pluripotent Cdk7 75Pluripotent Smad1 75Pluripotent Haus8 75Pluripotent Phb2 75Pluripotent Esrrb 74Pluripotent Psmd4 74Pluripotent Agtrap 74Pluripotent Sfr1 74Pluripotent Med14 74Pluripotent Rrp15 74Pluripotent Kif20a 74Pluripotent Rnf167 74Pluripotent Apex1 74Pluripotent Rab3ip 73Pluripotent Lgmn 73Pluripotent Stard10 73Pluripotent Npl 73Pluripotent Med12l 73Pluripotent Litaf 73Pluripotent Nrbf2 73Pluripotent Dnpep 73Pluripotent Atl2 73Pluripotent Rrm2 73Pluripotent Enpp1 73Pluripotent Gabarapl1 73Pluripotent Eng 73Pluripotent Tuba1a 73Pluripotent Ddx3y 73Pluripotent Igbp1 73Pluripotent Rps17 73Pluripotent Tma7 73Pluripotent Zkscan14 72Pluripotent Elf3 72Pluripotent Yif1b 72Pluripotent Dhrs3 72Pluripotent Ino80c 72Pluripotent Nup205 72Pluripotent Lrp2 72Pluripotent Tars 72Pluripotent Habp2 72Pluripotent Klhl21 72Pluripotent Cox7a2l 72Pluripotent Folr1 71Pluripotent Metap1 71Pluripotent Slc5a11 71Pluripotent Tle6 71Pluripotent Slc35e4 71Pluripotent Clpp 71Pluripotent Thy1 71Pluripotent Zfp691 71Pluripotent Lonp1 71Pluripotent Pagr1a 71Pluripotent Csf1r 71Pluripotent Rpl4 71Pluripotent Srsf2 71Pluripotent Polr3g 70Pluripotent Nr5a2 70Pluripotent Degs1 70Pluripotent Ap1m2 70Pluripotent Dcaf13 70Pluripotent Mapk6 70Pluripotent Trim25 70Pluripotent Sap18 70Pluripotent Ppp1r10 70Pluripotent Cars 70Pluripotent Timeless 70Pluripotent Agpat2 70Pluripotent Nup93 70Pluripotent Psmd13 70Pluripotent Ctrb1 70Pluripotent Mat2a 70Pluripotent Rps15a 70Pluripotent Slc7a2 69Pluripotent Cystm1 69Pluripotent Gpm6a 69Pluripotent Otud5 69Pluripotent Arl8b 69Pluripotent Bik 69Pluripotent Plaa 69Pluripotent Slc39a1 69Pluripotent Abcd4 69Pluripotent Dnajc11 69Pluripotent Rnf25 69Pluripotent Zmiz1 69Pluripotent Cdyl 69Pluripotent Snrpn 69Pluripotent Fkbp6 69Pluripotent Hspa4 69Pluripotent Dync1li2 69Pluripotent Dctd 69Pluripotent Mthfd2 69Pluripotent Sorbs2 68Pluripotent Usp44 68Pluripotent Tubb4b 68Pluripotent Tbpl1 68Pluripotent Enah 68Pluripotent Ddb1 68Pluripotent Fgf23 68Pluripotent Pmm2 68Pluripotent Hsf2 68Pluripotent Lrrc59 68Pluripotent Mybbp1a 68Pluripotent Psmc2 68Pluripotent Srp68 68Pluripotent Rps15 68Pluripotent Mtrr 67Pluripotent Nrxn1 67Pluripotent Sec23b 67Pluripotent Map7 67Pluripotent Atic 67Pluripotent Tecpr2 67Pluripotent Tyro3 67Pluripotent Tmem209 67Pluripotent Wfdc2 67Pluripotent Ccdc93 67Pluripotent H2afz 67Pluripotent Cebpz 66Pluripotent Pcsk1n 66Pluripotent Myo1b 66Pluripotent Neu1 66Pluripotent Ppp1r3b 66Pluripotent Tpd52l2 66Pluripotent Fgfr1 66Pluripotent Nup160 66Pluripotent Jade1 66Pluripotent Ybx3 66Pluripotent Sec24b 66Pluripotent Fos 66Multipotent Mdk 276Multipotent Prtg 270Multipotent Smoc2 268Multipotent Myoc 267Multipotent Spink2 259Multipotent Angptl1 258Multipotent Socs3 256Multipotent Ifitm3 253Multipotent Igf2bp1 250Multipotent Id3 250Multipotent Lin28a 247Multipotent Col1a2 242Multipotent Dkk1 241Multipotent Hmga2 238Multipotent Fads1 238Multipotent Maged2 237Multipotent Eef1g 234Multipotent Crabp2 234Multipotent Ogt 232Multipotent Hopx 230Multipotent Aplnr 230Multipotent Cps1 229Multipotent Cnn3 228Multipotent Fhl1 227Multipotent C1qtnf4 227Multipotent Has2 227Multipotent Cd34 225Multipotent Nr6a1 224Multipotent Mn1 223Multipotent Loxl2 222Multipotent Atf3 221Multipotent Pcolce 219Multipotent Slc12a2 219Multipotent Nktr 219Multipotent Mest 218Multipotent Pdgfa 216Multipotent Prom1 212Multipotent Srsf6 212Multipotent Lrrn4cl 212Multipotent Ifitm2 212Multipotent Pabpc1 211Multipotent Nrip1 209Multipotent Fkbp1a 207Multipotent Irf1 207Multipotent Entpd2 206Multipotent Tspan6 206Multipotent Msi2 205Multipotent Snrpn 204Multipotent Oat 204Multipotent Igfbp6 202Multipotent Itm2a 201Multipotent Mixl1 201Multipotent Lgals4 199Multipotent Vdr 199Multipotent Dsp 198Multipotent Kpnb1 197Multipotent Usp22 197Multipotent Tubb5 197Multipotent Socs2 196Multipotent Clec3b 196Multipotent Mmp2 196Multipotent Ttn 196Multipotent Tuba1a 195Multipotent Id1 195Multipotent Hspe1 195Multipotent Adnp 194Multipotent Col3a1 194Multipotent Rpl23 193Multipotent Fstl1 191Multipotent Zfp286 191Multipotent Fst 191Multipotent Nts 189Multipotent Cd81 187Multipotent Rps19 185Multipotent Ascl2 185Multipotent Tspan3 185Multipotent Serinc5 183Multipotent Gpc3 182Multipotent Cilp 181Multipotent Trp53 181Multipotent Timp2 181Multipotent Cxcl14 180Multipotent Hes1 180Multipotent Tpi1 180Multipotent Slit2 180Multipotent Mier3 179Multipotent Aldh1b1 179Multipotent Myc 178Multipotent Fbn2 178Multipotent Slc2a3 177Multipotent Arglu1 177Multipotent Tcf4 177Multipotent Krt19 177Multipotent Cntnap2 176Multipotent Lgals2 176Multipotent Hlf 175Multipotent Prelp 175Multipotent Lypd8 175Multipotent Itm2c 174Multipotent Bgn 174Multipotent Rgmb 174Multipotent Pcbp2 173Multipotent Spata13 173Multipotent Rps5 172Multipotent Msmo1 172Multipotent Lum 170Multipotent Uaca 170Multipotent Dcn 170Multipotent Car2 170Multipotent Hsp90ab1 169Multipotent Ran 169Multipotent Rpl41 169Multipotent Cdca7 169Multipotent Eif4b 168Multipotent Flt3 168Multipotent Rbmx 168Multipotent Tmprss4 168Multipotent Sertad1 168Multipotent Npm1 167Multipotent Ggh 167Multipotent Tmpo 167Multipotent Cdc23 166Multipotent Akap12 164Multipotent Prdx1 164Multipotent Rap2b 164Multipotent Cdk4 164Multipotent Sparc 164Multipotent Anxa1 164Multipotent Baalc 163Multipotent Sypl 163Multipotent Rps27l 162Multipotent Marcksl1 162Multipotent Calcrl 161Multipotent 161Multipotent 161Multipotent 161Multipotent 160Multipotent 160Multipotent 160Multipotent 159Multipotent 159Multipotent 158Multipotent 157Multipotent 157Multipotent 157Multipotent 156Multipotent 156Multipotent 156Multipotent 156Multipotent 156Multipotent 156Multipotent 154Multipotent 154Multipotent 154Multipotent 154Multipotent 154Multipotent 153Multipotent 153Multipotent 152Multipotent 151Multipotent 151Multipotent 151Multipotent 150Multipotent 150Multipotent 149Multipotent 148Multipotent 148Multipotent 148Multipotent 148Multipotent 148Multipotent 147Multipotent 147Multipotent 147Multipotent 147Multipotent 146Multipotent 146Multipotent 145Multipotent Rhoc 145Multipotent Ddit3 145Multipotent Ctnnb1 145Multipotent Poglut1 145Multipotent Glrx 145Multipotent Fbln2 144Multipotent Itih5 144Multipotent Casp4 144Multipotent Tnfaip6 144Multipotent Cmtm3 144Multipotent Lrwd1 144Multipotent Rbp1 144Multipotent Nid2 143Multipotent Hmgcs1 143Multipotent Fus 143Multipotent Cbx5 142Multipotent Fyttd1 142Multipotent Klf2 141Multipotent Rab34 141Multipotent Aplp2 141Multipotent Islr 141Multipotent Cyp1b1 141Multipotent Ids 140Multipotent Pgm3 140Multipotent Lrig3 140Multipotent Cttn 139Multipotent Rest 139Multipotent Igf1 139Multipotent Pycrl 139Multipotent Flrt2 139Multipotent Serpinf1 138Multipotent Sdc2 138Multipotent Kcnq1 138Multipotent Peg10 138Multipotent Pim1 138Multipotent Acat2 137Multipotent Hells 137Multipotent Dhcr24 137Multipotent Podn 137Multipotent Mgst3 137Multipotent N4bp2 137Multipotent Mcm6 136Multipotent Map3k8 136Multipotent Ankrd28 136Multipotent Pkp4 136Multipotent Ttc3 136Multipotent Fbln1 135Multipotent Srsf3 135Multipotent Aldh7a1 135Multipotent Ecm1 135Multipotent Snx10 135Multipotent Serpinh1 135Multipotent Calml4 134Multipotent Nono 134Multipotent Pea15a 134Multipotent F3 134Multipotent Anxa4 134Multipotent Tuba1b 133Multipotent Hnrnpa1 133Multipotent Efna5 133Multipotent Col5a2 133Multipotent Nr2c2 133Multipotent Ccnf 133Multipotent Ivns1abp 133Multipotent Xrcc5 132Multipotent Smarca1 132Multipotent Arid5a 131Multipotent Cacul1 130Multipotent Mfsd11 130Multipotent Casc4 130Multipotent Lcor 130Multipotent Txn1 130Multipotent Nhsl1 130Multipotent Gatad1 130Multipotent Dhrs4 130Multipotent Tm4sf20 130Multipotent Dclre1c 129Multipotent Rps3a1 129Multipotent Eya3 129Multipotent Fgl2 129Multipotent Armcx1 129Multipotent Dpt 129Multipotent Adamts1 129Multipotent Cdh11 129Multipotent Elp3 128Multipotent Glul 128Multipotent Apod 128Multipotent Mycn 127Multipotent Morf4l1 127Multipotent Serping1 127Multipotent Igfbp7 127Multipotent Noxa1 127Multipotent Meis1 127Multipotent Idi1 127Multipotent Shisa5 127Multipotent Pnisr 126Multipotent Fkbp9 126Multipotent Pcna 126Multipotent Cbr1 125Multipotent Ppp2r5c 125Multipotent Lgr5 125Multipotent Wnt3 125Multipotent Ldha 125Multipotent Echdc2 124Multipotent Relb 124Multipotent Akap13 124Multipotent Ppp1r1b 124Multipotent Gpx3 124Multipotent Ormdl1 124Multipotent Npdc1 123Multipotent St13 123Multipotent Bcl3 123Multipotent Ptgs2 123Multipotent Ablim1 123Multipotent Fbn1 123Multipotent Akna 122Multipotent Nodal 122Multipotent Mgst1 122Multipotent Nid1 121Multipotent Mov10 121Multipotent Areg 121Multipotent Arl3 121Multipotent Cyp51 121Multipotent Ywhae 121Multipotent Col5a1 121Multipotent Sp5 121Multipotent Crip1 120Multipotent Ranbp1 120Multipotent Drg1 119Multipotent Boc 119Multipotent Thrap3 119Multipotent Ctsk 119Multipotent Slc35a1 119Multipotent Tppp3 119Multipotent Msrb3 119Multipotent Sod1 119Multipotent Prmt5 118Multipotent Rpl26 118Multipotent Tfpi 118Multipotent Zc3h14 118Multipotent Dtwd1 118Multipotent Nav1 118Multipotent Enah 117Multipotent Eif5a 117Multipotent Paics 117Multipotent Olfml3 117Multipotent Ddx27 117Multipotent Col18a1 117Multipotent Uap1 116Multipotent Dnmt3a 116Multipotent Scfd1 116Multipotent Ptma 116Multipotent Prdx4 116Multipotent Enc1 116Multipotent Ndufa7 116Multipotent Stra8 116Multipotent Nme7 115Multipotent Xpo1 115Multipotent Rnf186 115Multipotent Sar1a 114Multipotent Tubb2a 114Multipotent Atp5c1 114Multipotent Rftn2 114Multipotent Enpp2 114Multipotent Plagl2 114Multipotent Tubb6 114Multipotent Wnt5a 114Multipotent Hnrnph1 113Multipotent Thbs4 113Multipotent Pls3 113Multipotent Usp21 113Multipotent Phf20 113Multipotent Pum1 112Multipotent Impact 112Multipotent Hspd1 112Multipotent Xrcc6 112Multipotent Calu 112Multipotent Fam110b 112Multipotent Wdr6 112Multipotent Tnrc6b 112Multipotent Pkm 112Multipotent Rps21 111Multipotent Nap1l1 111Multipotent Capza1 111Multipotent Prdx3 111Multipotent Stra6 111Multipotent Npr3 110Multipotent Bcdin3d 110Multipotent Cetn2 110Multipotent Msi1 110Multipotent Sptbn1 110Multipotent Rps23 110Multipotent Serpine1 110Multipotent Stmn1 110Multipotent Bin1 109Multipotent Thumpd1 109Multipotent Lysmd2 109Multipotent Sox9 109Multipotent Zbtb38 109Multipotent Bola2 108Multipotent Fam120b 108Multipotent Cd248 108Multipotent Pigr 108Multipotent Slc27a2 108Multipotent Glipr2 108Multipotent Supt16 108Multipotent Cd59b 107Multipotent Sfpq 107Multipotent Rnf6 107Multipotent Jun 107Multipotent Rnf43 107Multipotent H2afy2 107Multipotent Mat2a 106Multipotent Sall1 106Multipotent Mlh3 106Multipotent Plekhg2 106Multipotent Ubap2l 106Multipotent Cdx1 106Multipotent Ccl25 106Multipotent Dcbld2 106Multipotent Msh6 105Multipotent Hmgcs2 105Multipotent Gem 105Multipotent Rars2 105Multipotent Fam45a 105Multipotent Atxn10 105Multipotent Mfap2 105Multipotent Tia1 105Multipotent Lgals9 105Multipotent Pi16 104Multipotent Mier1 104Multipotent Foxo3 104Multipotent Igfbp4 104Multipotent Ube2v1 103Multipotent Igbp1 103Multipotent Kat6b 103Multipotent Fnip1 103Multipotent Snhg12 103Multipotent Rtp4 103Multipotent Adh5 103Multipotent Timm21 103Multipotent Llgl1 103Multipotent Rnf213 102Multipotent Ung 102Multipotent Atf4 102Multipotent Psmg4 102Multipotent Aspn 102Multipotent Cct2 102Multipotent C87436 102Multipotent Gng10 102Multipotent Htra3 102Multipotent Srsf2 101Multipotent Soat1 101Multipotent Gga2 101Multipotent Taf7 101Multipotent Pcmtd2 101Multipotent Inpp4a 101Multipotent Prss57 101Multipotent Axin2 101Multipotent Svil 101Multipotent Nfix 101Multipotent Sptan1 100Multipotent Rps3 100Multipotent Rfc3 100Multipotent Nedd8 100Multipotent Cpne8 100Multipotent Ces1d 100Multipotent Iars 100Multipotent Nrn1 100Multipotent Lpp 100Multipotent Ankrd10 100Multipotent Tbl1x 100Multipotent Hadh 100Multipotent Mpp5 100Multipotent Dctd 99Multipotent Lpcat3 99Multipotent Stag1 99Multipotent Anapc13 99Multipotent 99Multipotent 99Multipotent 99Multipotent 99Multipotent 99Multipotent 99Multipotent 99Multipotent 98Multipotent 98Multipotent 98Multipotent 98Multipotent 98Multipotent 98Multipotent 98Multipotent 98Multipotent 98Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 97Multipotent 96Multipotent 96Multipotent 96Multipotent 96Multipotent 96Multipotent 96Multipotent 96Multipotent 95Multipotent 95Multipotent 95Multipotent 95Multipotent 95Multipotent 95Multipotent Tgm2 95Multipotent Ndn 95Multipotent 95Multipotent 95Oligpotent 231Oligpotent 231Oligpotent 231Oligpotent 230Oligpotent 230Oligpotent 230Oligpotent 230Oligpotent 230Oligpotent 230Oligpotent 230Oligpotent 230Oligpotent 230Oligpotent 229Oligpotent 228Oligpotent 226Oligpotent 226Oligpotent 226Oligpotent 225Oligpotent 220Oligpotent 220Oligpotent 220Oligpotent 220Oligpotent 216Oligpotent 216Oligpotent 216Oligpotent 216Oligpotent 214Oligpotent 213Oligpotent 213Oligpotent 212Oligpotent 212Oligpotent 212Oligpotent 212Oligpotent 212Oligpotent 210Oligpotent 210Oligpotent 210Oligpotent 209Oligpotent 207Oligpotent 207Oligpotent 205Oligpotent 205Oligpotent Slpi 205Oligpotent Commd6 203Oligpotent Fndc3b 202Oligpotent Krt15 202Oligpotent Gpx1 200Oligpotent Rpl27 200Oligpotent Rpn1 200Oligpotent Plac8 199Oligpotent Thbs1 198Oligpotent Mylk 198Oligpotent Sftpa1 198Oligpotent Bhlhe41 197Oligpotent Tagln 196Oligpotent Cst7 196Oligpotent Anxa5 195Oligpotent Eef1d 194Oligpotent Dmkn 194Oligpotent Fhl2 193Oligpotent Rpl34 190Oligpotent Polr2l 190Oligpotent Rpl9 190Oligpotent Gm9844 189Oligpotent Rpl30 188Oligpotent Mef2c 188Oligpotent Slc38a2 188Oligpotent Afap1l1 186Oligpotent Lsm7 185Oligpotent Nfib 185Oligpotent Ogfrl1 185Oligpotent Mbd3 183Oligpotent Fabp5 183Oligpotent Hs2st1 182Oligpotent Itgb4 182Oligpotent Tes 181Oligpotent Mndal 180Oligpotent Sdc1 180Oligpotent Myb 180Oligpotent Clca2 179Oligpotent Lmo2 178Oligpotent Hdac2 178Oligpotent Clec11a 176Oligpotent Fbxo32 176Oligpotent Rpl35 176Oligpotent Rab31 174Oligpotent Rps13 174Oligpotent Jag1 173Oligpotent Eid1 173Oligpotent S100a8 173Oligpotent Tspo 173Oligpotent 171Oligpotent 171Oligpotent 170Oligpotent 169Oligpotent 168Oligpotent 166Oligpotent 165Oligpotent 165Oligpotent 165Oligpotent 165Oligpotent 164Oligpotent 163Oligpotent 163Oligpotent 163Oligpotent 163Oligpotent 162Oligpotent 161Oligpotent 161Oligpotent 160Oligpotent 160Oligpotent 158Oligpotent 158Oligpotent 158Oligpotent 158Oligpotent 157Oligpotent 156Oligpotent 156Oligpotent 156Oligpotent 156Oligpotent 156Oligpotent 156Oligpotent 155Oligpotent 155Oligpotent 155Oligpotent 155Oligpotent 154Oligpotent 154Oligpotent 154Oligpotent 153Oligpotent 152Oligpotent 152Oligpotent 152Oligpotent 151Oligpotent Hbegf 151Oligpotent Atxn1 151Oligpotent 151Oligpotent 151Oligpotent 150Oligpotent 150Oligpotent 150Oligpotent 148Oligpotent 148Oligpotent 147Oligpotent 147Oligpotent 147Oligpotent 147Oligpotent 146Oligpotent 146Oligpotent 146Oligpotent 146Oligpotent 146Oligpotent 146Oligpotent 146Oligpotent 145Oligpotent 145Oligpotent 145Oligpotent 144Oligpotent 144Oligpotent 144Oligpotent 143Oligpotent 143Oligpotent 143Oligpotent 142Oligpotent 142Oligpotent 141Oligpotent 141Oligpotent 141Oligpotent 140Oligpotent 140Oligpotent 139Oligpotent 139Oligpotent 139Oligpotent 138Oligpotent 138Oligpotent 138Oligpotent 138Oligpotent 137Oligpotent 137Oligpotent 137Oligpotent Rpl37 137Oligpotent Rrs1 137Oligpotent 136Oligpotent 136Oligpotent 136Oligpotent 136Oligpotent 135Oligpotent 135Oligpotent 134Oligpotent 134Oligpotent 134Oligpotent 133Oligpotent 132Oligpotent 132Oligpotent 131Oligpotent 131Oligpotent 131Oligpotent 131Oligpotent 131Oligpotent 130Oligpotent 130Oligpotent 130Oligpotent 130Oligpotent 130Oligpotent 129Oligpotent 129Oligpotent 129Oligpotent 129Oligpotent 129Oligpotent 129Oligpotent 129Oligpotent 128Oligpotent 128Oligpotent 128Oligpotent 127Oligpotent 127Oligpotent 126Oligpotent 126Oligpotent 126Oligpotent 126Oligpotent 126Oligpotent 125Oligpotent 124Oligpotent 124Oligpotent 123Oligpotent 123Oligpotent Rpl3 123Oligpotent Trp63 123Oligpotent Tmem107 122Oligpotent Krt6a 122Oligpotent Adprh 122Oligpotent Grsf1 122Oligpotent Mtmr14 122Oligpotent P4hb 122Oligpotent Fam162a 121Oligpotent Gca 121Oligpotent Canx 121Oligpotent Cd93 121Oligpotent Wdr83os 121Oligpotent Tacstd2 121Oligpotent Ccnd2 120Oligpotent Gypc 120Oligpotent Fbl 120Oligpotent Flcn 120Oligpotent Atp1b3 119Oligpotent Dusp2 119Oligpotent Gnptg 119Oligpotent Sin3b 119Oligpotent Irf6 119Oligpotent Spi1 118Oligpotent Zc3h8 118Oligpotent Msln 118Oligpotent Rpl19 118Oligpotent Arhgdib 118Oligpotent Ext2 117Oligpotent Rpl6 117Oligpotent Tmco6 117Oligpotent Fxyd3 117Oligpotent Naca 117Oligpotent Trip11 117Oligpotent Baz1a 117Oligpotent Ssr1 116Oligpotent Llph 116Oligpotent Zbtb11 115Oligpotent Mgat1 115Oligpotent Cryab 115Oligpotent Rps24 115Oligpotent H2afy 114Oligpotent Rgs18 114Oligpotent Sel1l 114Oligpotent Eny2 114Oligpotent Col4a3 114Oligpotent Cyr61 114Oligpotent Rps3a1 113Oligpotent Stmn1 113Oligpotent 113Oligpotent 113Oligpotent 113Oligpotent 113Oligpotent 113Oligpotent 113Oligpotent 112Oligpotent 112Oligpotent 112Oligpotent 112Oligpotent 112Oligpotent 112Oligpotent 112Oligpotent 111Oligpotent 111Oligpotent 111Oligpotent 111Oligpotent 111Oligpotent 111Oligpotent 111Oligpotent 110Oligpotent 110Oligpotent 110Oligpotent 110Oligpotent 110Oligpotent 110Oligpotent 109Oligpotent 109Oligpotent 109Oligpotent 109Oligpotent 109Oligpotent 109Oligpotent 109Oligpotent 109Oligpotent 109Oligpotent 109Oligpotent 109Oligpotent 108Oligpotent 108Oligpotent 108Oligpotent 108Oligpotent 108Oligpotent 108Oligpotent Pnrc1 108Oligpotent Cnrip1 108Oligpotent 107Oligpotent 107Oligpotent 107Oligpotent 107Oligpotent 107Oligpotent 106Oligpotent 106Oligpotent 106Oligpotent 106Oligpotent 106Oligpotent 106Oligpotent 106Oligpotent 106Oligpotent 105Oligpotent 105Oligpotent 105Oligpotent 105Oligpotent 105Oligpotent 105Oligpotent 105Oligpotent 105Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 104Oligpotent 103Oligpotent 103Oligpotent 103Oligpotent 103Oligpotent 103Oligpotent 102Oligpotent 102Oligpotent Nek7 102Oligpotent Lmo1 102Oligpotent Tmem213 102Oligpotent Fosl1 102Oligpotent Eef1b2 101Oligpotent Pias1 101Oligpotent Slc35f5 101Oligpotent Uxt 101Oligpotent Eif4a1 101Oligpotent Mefv 101Oligpotent Rictor 101Oligpotent Eif5 101Oligpotent Ccdc85b 101Oligpotent Ppp2cb 101Oligpotent F3 100Oligpotent Smarcc1 100Oligpotent Tgif1 100Oligpotent Hook2 100Oligpotent Impdh2 100Oligpotent N4bp2l2 100Oligpotent Upf1 100Oligpotent Gabarapl1 100Oligpotent Ctnnb1 99Oligpotent Pja1 99Oligpotent Atp6v1h 99Oligpotent Arc 99Oligpotent Pqlc1 99Oligpotent Trib1 99Oligpotent Egfl6 99Oligpotent Clmn 99Oligpotent Rab2a 99Oligpotent Fst 98Oligpotent Hnrnpa1 98Oligpotent Ppp1cb 98Oligpotent Spg11 98Oligpotent Hsp90b1 98Oligpotent Snhg8 98Oligpotent Pigc 98Oligpotent Gm1673 98Oligpotent Xxylt1 98Oligpotent Sphk1 98Oligpotent Srp14 98Oligpotent Man1a 97Oligpotent Btbd2 97Oligpotent Akap9 97Oligpotent Tpst1 97Oligpotent Tfap2c 97Oligpotent Cdh3 97Oligpotent Usp25 97Oligpotent Ripk3 97Oligpotent Naa38 97Oligpotent S100a9 97Oligpotent Mzt2 96Oligpotent Odf2l 96Oligpotent Runx1 96Oligpotent Rpl11 96Oligpotent Sod2 96Oligpotent Fam117b 96Oligpotent Zfp655 96Oligpotent Plxnc1 96Oligpotent Rpl10a 95Oligpotent Card10 95Oligpotent Alg14 95Oligpotent Rbm34 95Oligpotent Gpd2 95Oligpotent Mllt1 95Oligpotent Rassf1 95Oligpotent Rasal3 95Oligpotent Rpl24 95Oligpotent Rps12 95Oligpotent Sept7 95Oligpotent Eif3l 94Oligpotent Rabggta 94Oligpotent Rpl12 94Oligpotent Nedd4 94Oligpotent Ftsj3 94Oligpotent Zfp318 94Oligpotent Tcerg1 94Oligpotent Gmppb 94Oligpotent Nucb2 93Oligpotent Minos1 93Oligpotent Heatr5b 93Oligpotent Haus7 93Oligpotent Kcnk12 93Oligpotent Pak1ip1 93Oligpotent Rbm27 93Oligpotent Cebpz 93Oligpotent Etfb 93Oligpotent Lamp2 92Oligpotent Tmem248 92Oligpotent Spns3 92Oligpotent Foxred2 92Oligpotent Ggps1 92Oligpotent Cd38 92Oligpotent Chchd5 92Oligpotent Lrpprc 91Oligpotent Qpct 91Oligpotent Glt8d1 91Oligpotent Tspan32 91Oligpotent Mrto4 91Oligpotent Klhl21 91Oligpotent Iars2 91Oligpotent Ptk7 91Oligpotent Fbxo28 91Oligpotent Xkr8 91Oligpotent Dhrs11 91Oligpotent Ager 91Oligpotent Stxbp2 91Oligpotent Nudt14 90Oligpotent Lgmn 90Oligpotent Snrnp48 90Oligpotent Ccdc90b 90Oligpotent Dhx16 90Oligpotent Rnf44 90Oligpotent Rpl32 90Oligpotent Zfp1 90Oligpotent Ltb 90Oligpotent Med1 89Oligpotent Gypa 89Oligpotent Inip 89Oligpotent Capg 89Oligpotent Fkbp2 89Oligpotent Pa2g4 89Oligpotent Parg 89Oligpotent Mid1 88Oligpotent Dst 88Oligpotent Prkdc 88Oligpotent Slc7a11 88Oligpotent Ccdc126 88Oligpotent Plcd3 88Oligpotent Lyl1 88Oligpotent Cln6 88Oligpotent Pdia4 88Oligpotent Tmem39a 88Oligpotent Bcar1 88Oligpotent Rab27a 88Oligpotent Dcaf13 88Unipotent Xist 202Unipotent Car4 202Unipotent St3gal4 202Unipotent Trp53i11 202Unipotent Ncam1 201Unipotent Lgals3 201Unipotent Lat 201Unipotent Dll3 200Unipotent Btg2 200Unipotent Tfrc 199Unipotent Nrp1 199Unipotent S100a16 198Unipotent Elavl4 198Unipotent Mall 198Unipotent Ina 197Unipotent Vpreb3 197Unipotent Rbp7 194Unipotent Cldn13 194Unipotent Cd3g 193Unipotent Ccnd3 192Unipotent Cadm1 191Unipotent Chodl 191Unipotent Tubb3 191Unipotent Parvb 191Unipotent Fdps 191Unipotent Il7r 191Unipotent F13a1 190Unipotent Nefl 190Unipotent Pdzk1ip1 190Unipotent Aldh1a3 188Unipotent Cmpk1 188Unipotent Arpp21 187Unipotent Scd2 187Unipotent Rrm2 186Unipotent Fn3k 184Unipotent Lmna 183Unipotent Arl5c 182Unipotent Des 182Unipotent Anxa2 181Unipotent Serpinh1 181Unipotent Lcn2 181Unipotent Tshz1 181Unipotent Sfta2 180Unipotent Insm1 179Unipotent Fam178b 179Unipotent Tmem170b 179Unipotent Nmu 178Unipotent Fahd1 177Unipotent Crlf1 176Unipotent Hoxb8 175Unipotent Nefm 175Unipotent Cx3cr1 175Unipotent Vpreb2 175Unipotent Clcn3 175Unipotent Ucp2 174Unipotent Tmsb4x 174Unipotent Hist1h2ac 173Unipotent Cxcl1 173Unipotent Eif4e3 173Unipotent Ndrg2 173Unipotent Saa1 172Unipotent Jund 172Unipotent Cpa3 172Unipotent Bhlhe41 171Unipotent Tubb1 171Unipotent Mmp7 171Unipotent Slc39a8 170Unipotent Mfng 169Unipotent Wfdc2 168Unipotent Car1 168Unipotent Krt19 168Unipotent Rarres1 166Unipotent Bach2 166Unipotent Cd200 166Unipotent Carhsp1 166Unipotent Atg4d 165Unipotent Mcur1 165Unipotent Hist1h1b 165Unipotent Cxcr6 164Unipotent Kcnh2 164Unipotent H1f0 164Unipotent Ywhaz 163Unipotent Tcf19 163Unipotent Slc43a3 162Unipotent Gdpd1 162Unipotent Fkbp8 162Unipotent Actb 162Unipotent Tagln3 161Unipotent Ltf 160Unipotent Ppbp 160Unipotent Atp6v1b2 159Unipotent Smdt1 159Unipotent Prss8 158Unipotent Stom 158Unipotent Synj1 158Unipotent Ccr2 156Unipotent Fscn1 156Unipotent Abcb10 156Unipotent Blvrb 155Unipotent Rapsn 155Unipotent Igf2bp3 155Unipotent Stxbp1 154Unipotent Sox11 154Unipotent Tifab 154Unipotent Cldn3 154Unipotent Eif2ak1 153Unipotent Syngr1 153Unipotent Cd82 152Unipotent Hist1h1c 152Unipotent Dmbt1 151Unipotent Me2 151Unipotent Minpp1 150Unipotent Cdk2ap2 150Unipotent Gsdmc2 149Unipotent Ctsz 149Unipotent Arf1 149Unipotent Mid1ip1 148Unipotent Thy1 148Unipotent Cd151 147Unipotent Myh9 147Unipotent Arl6ip5 146Unipotent Tfdp1 146Unipotent Skap1 146Unipotent Myl4 146Unipotent Gata1 145Unipotent Tln1 145Unipotent Elavl3 145Unipotent Ndufc1 145Unipotent Cd3e 145Unipotent Ddah2 144Unipotent S100g 144Unipotent Trim58 144Unipotent Cbfa2t2 144Unipotent Ankrd9 144Unipotent Ermap 144Unipotent Cd79b 144Unipotent Crip1 144Unipotent Mki67 143Unipotent Nhlh1 143Unipotent Cebpg 143Unipotent Stat1 143Unipotent Hist1h1d 142Unipotent Gna11 142Unipotent Aqp5 142Unipotent Cd47 142Unipotent Ank1 141Unipotent Rnase12 141Unipotent Alas1 140Unipotent Ccl9 140Unipotent Actr10 140Unipotent Ndufaf3 140Unipotent Mzb1 140Unipotent Rhag 139Unipotent E2f2 139Unipotent Enpep 139Unipotent Cd3d 139Unipotent Arl4d 138Unipotent Akirin2 138Unipotent Ethe1 138Unipotent Fgfr4 138Unipotent Grap2 138Unipotent Syvn1 138Unipotent Dusp11 138Unipotent Fgfbp1 138Unipotent Impdh1 137Unipotent Myh14 137Unipotent Wsb2 137Unipotent Fasn 137Unipotent Slc25a10 136Unipotent R3hdm4 136Unipotent Cst6 136Unipotent Lck 136Unipotent Ube2f 135Unipotent Rpl22 135Unipotent Zeb2 134Unipotent Rap2b 134Unipotent Slc25a51 134Unipotent Ccne1 134Unipotent Msn 133Unipotent Asb6 133Unipotent Rag1 133Unipotent Tmem2 133Unipotent Meg3 132Unipotent Btbd17 132Unipotent Alg2 132Unipotent Mob3c 132Unipotent Ythdc2 132Unipotent S100a8 131Unipotent Serpinb6a 131Unipotent Klf1 131Unipotent Acly 131Unipotent Ddx23 131Unipotent Hist1h3h 131Unipotent Fbxo42 131Unipotent Neurod1 130Unipotent Fam53b 130Unipotent Gnas 130Unipotent Phf21b 130Unipotent Gp9 130Unipotent Irf4 130Unipotent H3f3b 130Unipotent Dnajb4 129Unipotent Kdm7a 129Unipotent Ptbp3 129Unipotent Hist1h2ae 129Unipotent Mapkapk3 129Unipotent Uck2 128Unipotent Arhgap27 128Unipotent Bicd2 128Unipotent Cpox 127Unipotent Camp 127Unipotent Chmp2a 127Unipotent Trps1 127Unipotent Sftpd 126Unipotent Tmem50a 126Unipotent Pdcd10 126Unipotent Drp2 126Unipotent Rgs10 126Unipotent Rarg 126Unipotent Cdkn1a 126Unipotent Cxcl17 126Unipotent Idh1 126Unipotent Hist1h1e 125Unipotent Sla2 125Unipotent Igsf21 125Unipotent Anp32e 125Unipotent Nhsl2 125Unipotent Sfpq 125Unipotent Ube2m 124Unipotent Pcbp1 124Unipotent Pigq 124Unipotent Taf15 124Unipotent Wdr18 123Unipotent Reg3g 123Unipotent Tmem259 123Unipotent Hsd11b1 123Unipotent Mgam 123Unipotent Ywhae 123Unipotent Tspan17 123Unipotent Cox7a1 122Unipotent Slc18a3 122Unipotent Rufy1 122Unipotent Igf2 122Unipotent Nppc 122Unipotent Map2k3 122Unipotent Arg2 122Unipotent Ublcp1 122Unipotent Matn4 121Unipotent Hspa13 121Unipotent Clec1b 121Unipotent Smim3 121Unipotent Pklr 121Unipotent Cnst 121Unipotent Nrros 120Unipotent Brpf1 119Unipotent Sh3glb1 119Unipotent Mfsd1 119Unipotent Lrmp 118Unipotent Uqcrh 118Unipotent Dok3 118Unipotent Map4k2 118Unipotent E2f8 118Unipotent Btg1 118Unipotent Cnnm4 118Unipotent Phf20l1 118Unipotent Ptprs 117Unipotent Maoa 117Unipotent Flnc 117Unipotent Rcor2 117Unipotent Mum1 117Unipotent Fbxo5 117Unipotent Ppp1r14a 116Unipotent Slc35a4 116Unipotent Gpcpd1 116Unipotent Cep44 116Unipotent Fam110a 116Unipotent Coasy 116Unipotent Pkp3 115Unipotent Rb1cc1 115Unipotent Wasf2 115Unipotent Gla 115Unipotent Scarb2 115Unipotent Acp5 114Unipotent Reg3a 114Unipotent Rps17 114Unipotent Ppp2r5a 114Unipotent Mfn2 114Unipotent Kcnn4 113Unipotent Cux1 113Unipotent Cntn2 113Unipotent Rnf5 113Unipotent Hdac7 113Unipotent Smarcd1 113Unipotent Treml1 113Unipotent Pxmp4 113Unipotent Insig1 113Unipotent Peli1 113Unipotent Sephs1 113Unipotent Aagab 113Unipotent Sdc4 112Unipotent Nol7 112Unipotent Glod5 111Unipotent Clec4a1 111Unipotent Musk 111Unipotent C3 111Unipotent Sash3 111Unipotent Slc12a6 111Unipotent Ccl2 111Unipotent Rfesd 111Unipotent Apoe 110Unipotent St6galnac5 110Unipotent Dbn1 110Unipotent Taf6 110Unipotent Tma7 110Unipotent Ssbp2 110Unipotent Pvr 110Unipotent Pcbp2 110Unipotent Zap70 110Unipotent Itsn2 110Unipotent Zfp639 110Unipotent Prelid1 110Unipotent Cald1 109Unipotent Actn4 109Unipotent Bbc3 109Unipotent Ndufc2 109Unipotent Lmo4 109Unipotent Relt 109Unipotent Npc2 109Unipotent Hoxb3 109Unipotent Arhgef18 109Unipotent Spry2 109Unipotent Nucb1 109Unipotent Anpep 109Unipotent Fth1 109Unipotent Ccdc85b 108Unipotent Tprgl 108Unipotent Mafk 108Unipotent Sowaha 108Unipotent Nfil3 108Unipotent Fbxo3 108Unipotent Gngt2 108Unipotent Isg20 108Unipotent Akap2 107Unipotent Mnx1 107Unipotent Eif1b 107Unipotent Man2b2 107Unipotent Gclc 107Unipotent Pirb 107Unipotent Gtf3c1 107Unipotent Sit1 107Unipotent Ahdc1 107Unipotent Krit1 107Unipotent Arhgef1 107Unipotent Xk 107Unipotent Ermard 107Unipotent Ikbkg 107Unipotent Cryab 106Unipotent St18 106Unipotent Sema6a 106Unipotent Hoxb4 106Unipotent Trp53inp2 106Unipotent Pcnp 106Unipotent Dctn6 106Unipotent Ddx54 106Unipotent Eif4g3 106Unipotent Caprin1 106Unipotent Tuba8 105Unipotent Slc43a2 105Unipotent Lrrfip2 105Unipotent Dpysl3 105Unipotent Hist2h2ac 105Unipotent Irf2 105Unipotent Tnfaip8l2 105Unipotent Tbpl1 105Unipotent Birc5 105Unipotent Ank3 105Unipotent Pccb 105Unipotent Rhoj 105Unipotent Fam213b 105Unipotent Scn1b 105Unipotent Tsr1 105Unipotent Ly9 105Unipotent Aimp2 105Unipotent Rab1b 105Unipotent Vprbp 105Unipotent Gmpr 104Unipotent Dhx40 104Unipotent Il17rc 104Unipotent Bcl7a 104Unipotent Arl5b 104Unipotent Afmid 104Unipotent S100a4 104Unipotent Rbm3 103Unipotent Tm6sf2 103Unipotent Pcyt2 103Unipotent Aes 103Unipotent Serpinb3a 103Unipotent B4galnt1 103Unipotent Marcks 103Unipotent Hemgn 103Unipotent Prkcb 103Unipotent Serping1 103Unipotent Cebpb 103Unipotent Ano6 102Unipotent Ndufa4 102Unipotent Fpgs 102Unipotent Smim1 102Unipotent Ehbp1l1 102Unipotent Mpc2 102Unipotent Zbtb18 102Unipotent Btf3 102Unipotent Cndp2 102Unipotent Scmh1 102Unipotent Fam89a 102Unipotent Lpcat3 102Unipotent Naa50 102Unipotent Rpgrip1 102Unipotent Atp13a3 102Unipotent Pitrm1 102Unipotent Chkb 102Unipotent Clptm1 101Unipotent Arl5a 101Unipotent Itm2a 101Unipotent Gars 101Unipotent Enox2 101Unipotent Klhl23 101Unipotent Azgp1 101Unipotent S100a9 100Unipotent Slc43a1 100Unipotent Onecut2 100Unipotent Prkar1a 100Unipotent Ldlrad3 100Unipotent Pou2af1 100Unipotent Pf4 100Unipotent Hist1h2bj 100Unipotent Znhit2 100Unipotent Ldlrad4 100Unipotent Rheb 100Unipotent Pcif1 100Unipotent Sept7 99Unipotent Mob1b 99Unipotent Map2k2 99Unipotent Chst3 99Unipotent Tpcn1 99Unipotent Cenpq 99Unipotent Rfc4 99Unipotent Ets1 99Unipotent Lsm14a 99Unipotent Ybx3 98Unipotent Idh2 98Unipotent Cluh 98Unipotent Rrp7a 98Unipotent Acrbp 98Unipotent Nhlh2 98Unipotent Maged1 98Unipotent Cmtr1 98Unipotent Atp6v1c1 98Unipotent Acadm 98Unipotent Arhgap19 98Unipotent Cdkn1b 98Unipotent Vdac3 98Unipotent Nup160 98Unipotent Mib1 97Unipotent Ckmt1 97Unipotent Itgb3 97Unipotent Chd9 97Unipotent Sh3bgrl2 97Unipotent Fermt3 96Unipotent Dpysl2 96Unipotent Cox15 96Unipotent Pank2 96Unipotent Tgfb1i1 96Unipotent Rhoa 96Unipotent Cyp11a1 96Unipotent Tradd 96Unipotent Bysl 96Unipotent Cxcr4 96Unipotent Atp2a3 95Unipotent Dgkd 95Unipotent Ywhah 95Unipotent Mpeg1 95Unipotent Hpse 95Unipotent Ifngr1 95Unipotent Arl4c 95Unipotent Hbq1b 95Unipotent Manf 95Unipotent Ptp4a2 95Unipotent Dph3 95Unipotent Ankmy1 95Unipotent Tut1 95Unipotent Rwdd4a 95Unipotent Bfsp2 95Unipotent Fn1 95Unipotent Eif1a 94Unipotent Wwtr1 94Unipotent Slc34a2 94Unipotent Nipal1 94Unipotent Prim2 94Unipotent Fxyd6 94Unipotent Eif6 94Unipotent Gbe1 94Unipotent Cd5 94Unipotent Traf4 94Unipotent Kcnj8 94Unipotent Crmp1 94Unipotent Prkab2 94Unipotent Actr1a 94Unipotent Mvb12a 94Unipotent Rer1 94Unipotent Lrwd1 94Unipotent Rcl1 94Unipotent Prpf39 94Differentiated Gm9844 217Differentiated Clps 212Differentiated Rps12 211Differentiated Klk1 211Differentiated Tspan13 211Differentiated Spink4 211Differentiated Gimap7 211Differentiated Agr2 211Differentiated Tyrobp 211Differentiated Cd74 210Differentiated Guk1 209Differentiated Rpl17 209Differentiated Areg 209Differentiated Cd24a 208Differentiated Stard10 208Differentiated Tm4sf4 208Differentiated Reg4 207Differentiated Junb 207Differentiated Fcgbp 207Differentiated Mapt 207Differentiated Tff3 207Differentiated Hepacam2 207Differentiated Sell 207Differentiated Rnase1 206Differentiated Gstk1 206Differentiated Lbh 205Differentiated Ndrg2 203Differentiated Tymp 203Differentiated Gpr183 203Differentiated Ccr7 203Differentiated Cd83 203Differentiated Ltb 203Differentiated Cpe 203Differentiated Aplp1 202Differentiated Guca2b 202Differentiated Gabarap 201Differentiated Timp3 201Differentiated Qsox1 201Differentiated Sct 200Differentiated Bhlhe40 200Differentiated Ltc4s 200Differentiated Fxyd3 200Differentiated Fth1 199Differentiated Alox5ap 199Differentiated Fau 198Differentiated Nbeal1 198Differentiated Sel1l3 198Differentiated Fcer2a 198Differentiated Gatm 197Differentiated Gnao1 196Differentiated Chgb 195Differentiated Rpl12 195Differentiated Cst3 195Differentiated Pcsk1n 195Differentiated Apoe 194Differentiated Krt7 194Differentiated Scg3 194Differentiated Sh3bgrl3 194Differentiated Myadm 193Differentiated Ly6d 193Differentiated Ctrb1 192Differentiated Try5 192Differentiated Rhov 192Differentiated Tnfaip3 191Differentiated Coro1a 190Differentiated Tm4sf1 190Differentiated Krt18 190Differentiated Sparcl1 190Differentiated Malat1 189Differentiated S1pr1 189Differentiated Tspan7 189Differentiated Nbr1 188Differentiated Gimap4 187Differentiated Cyp4b1 186Differentiated Ids 186Differentiated Stk17b 186Differentiated Rasd1 185Differentiated B2m 185Differentiated Isg15 185Differentiated Ggt1 185Differentiated Gpm6a 185Differentiated Atoh1 184Differentiated Sdhaf1 184Differentiated Anxa13 183Differentiated C1qa 183Differentiated Basp1 183Differentiated Hpd 182Differentiated Tulp4 181Differentiated Spdef 181Differentiated 180Differentiated 180Differentiated 180Differentiated 180Differentiated 179Differentiated 179Differentiated 178Differentiated 177Differentiated 176Differentiated 176Differentiated 175Differentiated 175Differentiated 175Differentiated 175Differentiated 175Differentiated 174Differentiated 174Differentiated 174Differentiated 173Differentiated 173Differentiated 173Differentiated 172Differentiated 172Differentiated 172Differentiated 172Differentiated 172Differentiated 171Differentiated 171Differentiated 171Differentiated 170Differentiated 170Differentiated 169Differentiated 169Differentiated 169Differentiated 168Differentiated 168Differentiated 168Differentiated 168Differentiated 167Differentiated 167Differentiated 166Differentiated 166Differentiated 166Differentiated 165Differentiated Thra 165Differentiated Krt32 165Differentiated 165Differentiated 165Differentiated 165Differentiated 165Differentiated 164Differentiated 164Differentiated 164Differentiated 163Differentiated 163Differentiated 163Differentiated 162Differentiated 162Differentiated 162Differentiated 161Differentiated 161Differentiated 161Differentiated 160Differentiated 160Differentiated 160Differentiated 160Differentiated 160Differentiated 160Differentiated 160Differentiated 160Differentiated 159Differentiated 159Differentiated 158Differentiated 158Differentiated 157Differentiated 157Differentiated 157Differentiated 157Differentiated 156Differentiated 156Differentiated 156Differentiated 156Differentiated 156Differentiated 156Differentiated 156Differentiated 156Differentiated 156Differentiated 155Differentiated 155Differentiated 154Differentiated Yaf2 154Differentiated Scg5 154Differentiated Kif5c 154Differentiated C1qb 154Differentiated Dazap1 154Differentiated Fyn 154Differentiated Mocs2 154Differentiated Cfh 153Differentiated Ier5 153Differentiated Pdk2 153Differentiated Dnajc10 153Differentiated Rnf11 152Differentiated Plaur 152Differentiated Psap 151Differentiated Tada3 151Differentiated Slc40a1 150Differentiated Fbxl15 150Differentiated Lair1 150Differentiated Smpd3 149Differentiated Il1r1 149Differentiated Scamp5 149Differentiated Gramd3 149Differentiated Litaf 149Differentiated Camk2b 148Differentiated Ikbkap 148Differentiated Pim3 147Differentiated Txndc15 147Differentiated Tcta 147Differentiated Slc2a5 147Differentiated Pdgfrb 147Differentiated Celf4 147Differentiated Fbln1 147Differentiated Rnpepl1 146Differentiated Il13ra1 146Differentiated Metrnl 146Differentiated Kit 146Differentiated Bcl11a 146Differentiated Ide 145Differentiated Ptprn2 145Differentiated Gcnt3 145Differentiated Reep6 145Differentiated Sept4 144Differentiated Hist3h2a 144Differentiated Gpm6b 144Differentiated Vamp2 144Differentiated Pklr 143Differentiated Eif1 143Differentiated Ech1 143Differentiated Cel 142Differentiated Cox4i1 142Differentiated Tspan1 142Differentiated Ppp1r15b 142Differentiated Coq2 142Differentiated Apob 142Differentiated Senp6 142Differentiated Uba52 141Differentiated Camk2n1 141Differentiated Cdk5r1 141Differentiated Trappc9 141Differentiated Ezr 141Differentiated Ncf1 141Differentiated Rcn2 141Differentiated Pmm2 140Differentiated Nrxn2 140Differentiated Hpgds 139Differentiated Vopp1 139Differentiated Rps6ka4 139Differentiated Sorl1 139Differentiated Trem1 139Differentiated Tmc5 139Differentiated Cdc42ep4 138Differentiated Cd19 138Differentiated Pde4dip 138Differentiated Dll4 138Differentiated Ednrb 138Differentiated Adam8 138Differentiated Srgap3 138Differentiated Fcnb 138Differentiated Gprasp1 137Differentiated Pbxip1 137Differentiated Serpini1 137Differentiated Tanc2 137Differentiated Pcsk2 137Differentiated Scn3a 137Differentiated Etv1 137Differentiated Atp6v0e2 137Differentiated Bloc1s2 136Differentiated Spint1 136Differentiated Ppm1m 136Differentiated Lime1 136Differentiated Pcp4l1 136Differentiated Hmgcs2 136Differentiated Gcg 135Differentiated Tpd52 135Differentiated Prmt2 135Differentiated Ash1l 135Differentiated Ern2 135Differentiated Serinc1 135Differentiated Rbp4 135Differentiated Aldoc 135Differentiated Yif1b 135Differentiated Agap1 135Differentiated Itm2b 134Differentiated Utrn 134Differentiated Srsf4 134Differentiated Gm11127 134Differentiated Serpina1c 134Differentiated S100a1 134Differentiated Acp2 134Differentiated Strn4 134Differentiated Cd164 134Differentiated Ugdh 134Differentiated Mxd1 134Differentiated Lmna 133Differentiated Dhps 133Differentiated Atg3 133Differentiated Irf8 133Differentiated Krt4 133Differentiated Cxcr5 133Differentiated Enc1 133Differentiated Pln 132Differentiated Mdm2 132Differentiated Tmem176a 132Differentiated Fgd2 132Differentiated Ghr 132Differentiated Josd2 132Differentiated Tox3 132Differentiated Chordc1 131Differentiated Ilvbl 131Differentiated Smpx 131Differentiated Tnfaip8 131Differentiated Cyba 131Differentiated Apoc3 131Differentiated Gpr160 131Differentiated Peg3 131Differentiated Tgm2 131Differentiated Lpgat1 131Differentiated Mmp7 130Differentiated Dcx 130Differentiated Ddx41 130Differentiated Uba2 130Differentiated Gpsm3 130Differentiated Ablim1 130Differentiated Stat4 130Differentiated Cdc42se2 129Differentiated Ankrd12 129Differentiated Nppa 129Differentiated Tc2n 129Differentiated Samd8 129Differentiated Ms4a1 129Differentiated Smad7 128Differentiated Rpl41 128Differentiated Zfp750 128Differentiated Iapp 128Differentiated Osbpl2 128Differentiated Hif1a 128Differentiated Foxa3 128Differentiated Them4 127Differentiated Pck1 127Differentiated Dapl1 127Differentiated Sorbs1 127Differentiated Hsd17b6 127Differentiated Zmynd11 127Differentiated Pde2a 126Differentiated Acox1 126Differentiated Atp1a2 126Differentiated Enpp5 126Differentiated Rnf144a 126Differentiated Nudt4 126Differentiated Hoxb2 126Differentiated Tsc22d3 126Differentiated Tmeff1 125Differentiated Cish 125Differentiated Itpr1 125Differentiated Upp1 125Differentiated Rassf6 125Differentiated Pak3 125Differentiated Itgax 125Differentiated Adh1 125Differentiated Ets1 124Differentiated Tst 124Differentiated Nr1d1 124Differentiated Crim1 124Differentiated Ttc22 124Differentiated Gcc2 124Differentiated Ace 124Differentiated Hmgn3 124Differentiated Krt8 123Differentiated Ank2 123Differentiated Xpnpep2 123Differentiated Grina 123Differentiated Myo1b 123Differentiated Cyp2f2 122Differentiated Cfp 122Differentiated L1cam 122Differentiated Unc5b 122Differentiated Apoa1 122Differentiated Cd55 122Differentiated Pnmal2 121Differentiated Fars2 121Differentiated Lyar 121Differentiated Il4ra 121Differentiated Abr 121Differentiated Ergic1 121Differentiated Cd7 120Differentiated Kif1b 120Differentiated Sacm1l 120Differentiated Cyp2c55 120Differentiated Gucy1a3 120Differentiated Dio1 120Differentiated Txndc5 120Differentiated Maged1 119Differentiated Sbk1 119Differentiated Pfn1 119Differentiated Dcps 119Differentiated F8a 119Differentiated Cx3cl1 119Differentiated Lman1 119Differentiated Pik3r5 119Differentiated Ctso 119Differentiated Btg1 118Differentiated Enho 118Differentiated Nostrin 118Differentiated Mt3 118Differentiated Furin 118Differentiated Ceacam1 118Differentiated Ccr5 118Differentiated March8 118Differentiated Kcnk1 118Differentiated Mettl1 118Differentiated Cxcl12 118Differentiated Crip1 117Differentiated Azgp1 117Differentiated Klc1 117Differentiated Mlph 117Differentiated Apbb2 117Differentiated Tom1l2 117Differentiated Epm2aip1 117Differentiated Stc2 117Differentiated Ypel3 116Differentiated Smap2 116Differentiated Nme3 116Differentiated Prkaa2 116Differentiated Flnb 116Differentiated Notch2 116Differentiated Ttc3 116Differentiated Spns2 116Differentiated Pla2g7 116Differentiated Pcbd1 116Differentiated Osgin1 116Differentiated Fpr1 116Differentiated Tmem230 115Differentiated Stard7 115Differentiated Rab4b 115Differentiated Yipf5 115Differentiated Smim6 115Differentiated Csf2ra 115Differentiated Stat5b 115Differentiated Jak1 115Differentiated Fhl2 115Differentiated Cd3e 114Differentiated Scarb2 114Differentiated Vps37b 114Differentiated Stap2 114Differentiated Avil 114Differentiated Tpsb2 114Differentiated Lrrc32 114Differentiated Slc25a25 114Differentiated Pdia5 114Differentiated Aldh3a2 114Differentiated Cog7 113Differentiated Spock2 113Differentiated Arpp19 113Differentiated Pclo 113Differentiated Il1rn 113Differentiated Rdh12 113Differentiated Gpc6 113Differentiated Mgll 112Differentiated Pkdcc 112Differentiated Gns 112Differentiated Klrb1 112Differentiated Ddx24 112Differentiated Dmpk 112Differentiated Slc9a3 112Differentiated Slc22a5 112Differentiated Nrxn3 112Differentiated Tfap2b 112Differentiated Nek6 112Differentiated Atp6ap2 111Differentiated Dlc1 111Differentiated Rab15 111Differentiated Dnase1 111Differentiated Tfpi 111Differentiated Elavl3 110Differentiated Rbm42 110Differentiated Cfd 110Differentiated Eml4 110Differentiated Dusp23 110Differentiated Ptms 109Differentiated Rab3ip 109Differentiated Tcn2 109Differentiated Clrn3 109Differentiated Zfp428 109Differentiated Tmem220 109Differentiated Slc2a2 109Differentiated Cgref1 109Differentiated Ncr1 109Differentiated Zfp36l1 108Differentiated Sf3b5 108Differentiated Rhobtb3 108Differentiated Rpl23a 108Differentiated Pla1a 108Differentiated Mep1b 108Differentiated Apc2 108Differentiated Mgat3 108Differentiated Ndufa12 108Differentiated Tubb2a 108Differentiated Syce2 108Differentiated Mef2c 108Differentiated Isg20 107Differentiated Gstm5 107Differentiated Entpd3 107Differentiated Tgoln1 107Differentiated Acaa1b 107Differentiated Prcc 107Differentiated Bcl7b 107Differentiated Igflr1 107Differentiated Pfdn5 107Differentiated Arl1 107Differentiated Golgb1 107Differentiated Nbea 107Differentiated Emp3 107RRM1 NPR3IGF2BP2 CAPRIN1DTL TCEAL4C5orf42 LRCH3TCF4 CENPJMAD2L1 MCM2STT3B PAN3MCM5 LRBAUTP20 TEX30C1orf186 PRSS57SSBP1 RPLP0P2GADD45GIP1 SNHG8RSL24D1 CTR9NLK BDH2CCDC90B RPL22L1RPL22L1 MRPL30HNRNPC CLNS1AOARD1 WDR43GAPT ZCCHC2VPS13B TNS3NCOA3 MYH9RRM2B MMAALSM10 USP4DUSP7 HIP1HDAC7 NFKB1FBXO8 GSECATAD2B SLC8B1PTPRA NASPARC NAPON2 NAFGR MARCOMTSS1 FYB1POU2F2 CCR1CPPED1 ASAH1DAZAP2 ANXA5SLA SMIM25ARRB2 HCKJAML JAMLHCK CYBAORC6 HMGB2KIF22 AURKBDDX39A NDC80ECT2 NCAPG2HIST1H1E CENPKPRIM1 ASF1BGAPDH HELLSPTMA HIST1H1BTCF19 FAM111BANGPTL2 SHDSOX8 GPR17RPS19 TRAPPC2BCAS2 NAHIST1H2BC NARPS12 NAFOS NAACER3 NAIFIT2 NARND3 BAALCTSPAN7 FADS2RTN1 HRSP12PKIA BLCAPEZR BTBD17ELP4 MORN2LMO3 KLHL9C6orf1 UBQLN2SEMA6D ZC2HC1ATXNDC16 ITGB8TTYH2 DICER1-AS1DUT KIAA0101RMI2 MTHFD1PXMP2 TMEM106CE2F1 IDH2CHEK1 GGCTPCCA PSMC3SDF2L1 PHGDHCARHSP1 COPS3SPATA33 HACD3YEATS4 CFAP20RPL18 RPS25DHX9 ABCE1COA3 TPT1ZNF439 C5orf56LINC00114 PLIN3HSD17B11 FHITSTK24 SMARCA4PGM2L1 POU2AF1ADRBK1 BZRAP1-AS1PREX1 SLC38A1BSDC1 CPNE1GRAP UNKZFP36 MS4A1EVL CD69S100A4 KLF6EPC1 CYB561A3TIMP1 TIMP1IFITM2 KLF10TBC1D9 HELZ2ARRDC1 RASAL3KRT10 DDX24IER2 SDHAF1DCN EHFNDRG2 TUFT1FAM107B VDAC2CYP1B1 PGM3MRPS9 UAP1CHN1 TRAP1PSMC5 CRYZL1GART MRPL45SCRG1 ENPP5HSF2 COA3ETV6 ZHX1PDS5B CSTF3PTK7 C21orf58HNRNPM ZNF326DEK ARL6IP1NUCKS1 GGHPCNA RHEBGGCT KIF2CPSMA4 PKMYT1NDUFS8 SUMO2LRR1 NANDUFB11 NAMORF4L2 NAZFP36 HEBP2CTSD ESR1GPRC5A MAL2NFKBIA KIAA1324EFNA1 PCBD1CXCL10 AZGP1BLVRB PLEKHF2YIPF2 ATP6V1G1BIK ARFIP2CASP2 TARBP1ATRX EREGEIF4B SOX4GATAD1 GPIARID1A FMR1TTC19 POF1BIDH1 GANTCERG1 POMPRSL1D1 CCT3CCT5 PTPN11ILF2 ARL4ACCNB1 EIF3CLCDT1 UTP18ZCCHC17 NAP1L4PHF6 KPNB1UBE2C DBICENPW UBE2CCDKN3 PSMA7TFF1 CLDN3FXYD3 KRT18PLAUR CLDN4TINAGL1 ITGB4KRT8 FERMT1RHOC LLGL2MYH14 STK24S100A4 EFNB1ADM ADIPOR1MYADM DIO3OSPPP1R14D HOXB8ST14 CLDN23PLK1 PRC1TMPO TPX2C19orf48 DTYMKPHGDH SMC2KIF20A SNRPGBUB1 NASPHNRNPA2B1 TRIB2VRK1 RFC3HNRNPA1 GMNNDCX SNTG1GFAP NOVA1NRXN1 NFIASLAIN1 CELF5TF NRXN2SEPT4 ST7LHPP PDZRN3HOPX SV2ASNTG1 CACNG4NKAIN3 THRBKLHDC8A DAB1DCC NOL4RNF168 GPIRRM1 RPL7AHSPA8 FKBP4HNRNPM MYCCBX5 DTYMKUSP1 TIMM13EPCAM UTP6SDHA WDR45BC6orf62 BLMHMRPS33 IFI27GPR87 KLK8MRPL33 SOX15ANXA8 MRPL36FSCN1 PDLIM1GTF3C6 PSMB7RPLP1 HMOX2HRAS MT1EMYLK NOP10MMP9 C1RDDIT4 FFAR3SGMS1 SCARB2TRIM22 ATF4PERP CEBPAIFIT3 SEZ6L2TNFRSF21 NATNF NAARL6IP5 NASPRR1B S100PZG16B SPRR1BLCN2 CEACAM5KLK12 TMEM159MUC20 TJP2MXD1 LCE3DATP12A RAB11ATNFRSF11B TRIM16TF TM4SF1BCAT1 UPK1BKRT75 COPB1LURAP1L SCGB3A1KLF13 ZC2HC1AESD CENPUSPC24 TRIM59STAMBP SERBP1IPO4 KARSGNB1 CBX3MRPL44 PRKAG2-AS1ZFYVE19 TPX2MTPN E2F1HMGB1 WBP4TIPIN SAP30SERPINE1 EIF3MPHLDA3 NARARRES2 NAGDF15 NAHLA-A NDUFA4L2IGKC FABP7B2M CNDP2HP HCFC1R1SLC6A13 CA9HLA-DRB5 CCND1PTGR1 PTTG1IPPCK1 PHPT1MSRB1 PCBD1LSM2 APOA2HNRNPH1 AKR1C2HMGB2 GPX3TAF12 RFC4MDM4 SUCLG1GCLC NUDT5NXT1 RHOBTB1ITM2A CENPUCACYBP TRMT11VSNL1 NANFIA NAMRPS14 NAPOLR2H SLC16A3TAGLN2 PLCG2ZNF581 CIARTTGFB1I1 KDSRRHOA FOXRED1DAZAP2 TUBB2BMET NANONO NAMIF NAUGT2B7 EIF2S2POLR2I ST20CMSS1 PDAP1TMEM33 NAMORN2 NATOMM22 NAKCTD9 NALAMTOR2 NASNRPE NALGALS3 HPCCL3 CD24SAA4 ADH1BSCP2 C5APOB C1RPTGDS PXMP2RNF181 CYP2A6LAPTM5 BPHLSERPINC1 HPGDMCM2 HES6PRKDC MCM2RRM1 TUBA1BLRPPRC LAPTM4BFLOT2 RBBP4CRYZL1 TCP1KLHL23 FKBP4AHCY SMC3TULP3 MAD2L1APEX2 SMC2POLR2B NUDCDDX5 HNRNPA3SULT1E1 RPLP1HTRA1 RPL10AECI2 COX6B1SH3BGRL HSP90B1RPLP1 DDX19BALDH3A2 VSNL1FBLN1 RRP15UQCRC2 TTF2LOXL4 RPL15GNG5 NAIDI1 NARPL26L1 NAPSMB9 NAERH NAHSD17B6 NAXBP1 CD24HLA-A P4HBOCIAD2 JUPABCC3 PDXDC1ANKRD12 S100A11SLC9A3R2 LINC00342PCSK1N TMED4CEBPB TMEM45BARSD PGM2L1ZWINT YWHAERCN1 HMGA1ERH PSMG1PTGES3 NDUFV2POLD2 AURKAIP1HMGB2 H2AFVMRPL51 MZT2BP4HB COPS5GSTO1 MTHFD1NENF CENPVTMEM59 PAR-SNSMUG1 KCNJ13ACP6 NAPTPN12 NAEIF5 NAESRG NAZNF124 NAPEX13 NATRAF1 MPZSESN1 AHNAKHMHA1 NDUFS8SPHK1 P4HA2LNPEP SNAPINARHGAP25 LINC00473RRAS GBP4TMEM123 THAP11STRADA TSPAN17C1QC ATP6V1FIL4I1 SDC4IFI30 LRPAP1LOC646214 HDDC2HERC2P4 ROPN1HLA-C LAMTOR5ZNF225 FAM84BBTBD6 C11orf49TRPM1 FARP2TUBA1B STMN1STMN1 RRM2RRM2 TUBA1BNUDT1 TMPOHSPD1 PRR11MYBL2 SOD1DDX39A SPC25HNRNPAB TXNHNRNPC VBP1PSMD4 ADAPNN ATAD5TUBB4B SNRPAAHNAK NLGN4XS100A11 PMEPA1YPEL3 IGLL5TIMP1 NATGOLN2 NAGLUL NAUBE2C UBE2CTUBA1B PBKPBK HMGN2DUT CENPKPKMYT1 TUBBH2AFY CENPHARL6IP6 TMEM54ANP32E BUB1BWDR76 DERASGOL2 NUDCD2CACYBP XRCC2LOC100506844 SSRP1CD320 LOC100506548THOC7 ZNF593C12orf45 RPS29MORC2-AS1 SCYL3SNHG16 WDR45F8 HADHMED28 NAPWP2 NAZNF736 NAZNF32 C17orf76-AS1GLRB NACADLL3 SMARCE1SCAF8 NATAGLN3 NALINC00673 NABTRC NAFAM174A NARIMKLB NAJUNB MT1XNFKBIA C1orf61TF GPRC5BPON2 CA12MLC1 B3GNT1PIGS SLC22A17LIPA AKR1C3LCAT NAV1CLDN11 PSAT1ADHFE1 THY1BLCAP CRIP2RPH3A SEZ6LOGN CCNB2CDKN3 DUTATAD2 ID1TWISTNB NDC80EVA1B CENPMCDCA4 MIS18AWASF2 ACAT2PMP22 KIF23CENPN NDUFB11RPL22L1 S100A1HSP90B1 DPCDRPL36A RPL23RPS6 NACCDC59 NARPL24 NARNF113A KRT19SERPINE1 PRELID1AURKAIP1 METTL23CYB5A NAGALK1 NADXO NAC21orf59 NAHBP1 NAGOLGA2 NATPSB2 POLR2BCXCR4 CD72SAT1 EAF2HES1 ZNF316TSC22D3 OSBPL2BANK1 CEP44CD72 FAM43AISG20 EMSYBLOC1S2 MSI2TOP2A CBX5SMC4 CENPKNUCKS1 CDCA3RACGAP1 BARD1IGF1 RAD54LTMEM97 DCLRE1CSMCHD1 NAGPX7 NACHTF18 NAMT2A NAMT1E NAPRG4 NAISG15 WFDC2GRN ABHD11VTCN1 RBP1C4orf48 CRB3TNFRSF14 PORNUPR1 AMNLTBP3 NR2F6BCAS1 PCDH7AMN HTATIP2CENPW SPDL1CDCA3 EIF4EBP1GGH MKI67GMNN SETEIF5A WDR34PSMA1 TUBB4BSLC35F2 SLC35F2TXNDC9 NUTF2RAD51C PPM1GSNAI1 BTF3DYNLL1 POLR2KPPP1R14B LUC7L3ERCC1 NATRMT112 NAEWSR1 NASLU7 NANDUFA2 NAWDR66 NAS100P FXYD3SPINK1 MUC1TFF3 ELF3ST14 ANKRD36CIFI6 SPINT1MUC5AC LAMB3PI3 PHLDA2DHRS9 CLINT1CD164 MACC1TMEM106B TMEM45BMAFF TMEM63ARDH10 PROM1FOS MSH6AEBP1 IL17RBMRPS6 PARP1DNAJC9 TMEM97YWHAH LIG1OCLN CPSF7LMO4 FEN1HUWE1 LSM2SUPT16H SNRPD3SMC4 MAD2L1MAD2L1 TUBA1CHIST1H4C ASPMHNRNPK BUB3CBX5 CCNFGPSM2 CALM3NPM1 PSMA4SLIRP THOC7COLCA1 RPL23SDHD NASNRPG NASTK3 NADHCR24 PPP3CASAT1 SLC39A6KRT8 MARCKSL1BAIAP2 APLP2SLC39A6 HMG20BSLC9A3R2 SMCO4CXCL1 NPYNR4A1 SEMA3CSERP1 CERS4
Claims
WHAT IS CLAIMED IS:
1. A computational method to be performed by a computational system to determineclass composition of a collection of biological entities, comprising:receiving omic data results of one or more entities, each omic data resultcomprising omic features; constructing a matrix of omic features x entities derived from the omic data results;and determining a class status of each entity by inputting the matrix of omic features xentities as features within a trained computational framework that comprises a set of oneor more binary networks encoding trained binarization of omic features within one or more omic feature sets to predict the class status of each entity, wherein the class composition comprises two or more class status that characterize class composition.
2. The computational method of claim 1, wherein each binary network learns whichomic features predict the class status based on binarization of the matrix of omic features x entities.
3. The computational method of claim 2, wherein each binary network further learnsa prediction power of each omic feature for predicting the class status.
4. The computational method of claim 2 or 3, wherein the computational frameworkcomprises two or more computational modules, each computational module comprises abinary network for binarization of the matrix of omic features x entities to assess a classstatus, wherein each computational module is trained to predict a likelihood of the class status.
5. The computational method of claim 2 or 3, wherein the computational frameworkconsists of a single module to assess two or more class status with a binary network for binarization of the matrix of omic features x entities, wherein the module is trained to predict a likelihood of each class status of the class composition.
6. The computational method of claim 4 or 5, wherein each module yields a set ofone or more class status scores for each entity.
7. The computational method of claim 6, wherein the class status score comprisesan omic enrichment score that characterizes enrichment of omic features within the entity.
8. The method of claim 7 further comprising:for each entity, ranking and trimming the omic features based quantity within theentity’s omic data results.
9. The computational method of claim 8, wherein the omic enrichment score is basedon the rank of the omic features and the learned prediction power of the omic features.
10. The method of claim 7 further comprising:for each entity, normalizing the omic features to a data point associated with the entity’s omic data results.
11. The computational method of claim 10, wherein the omic enrichment score isbased on an aggregation of omic features.
12. The computational method of any one of claims 6-11 further comprising:integrating the omic enrichment scores of each entity to yield a module class score.
13. The computational method of claim 12 further comprising:determining a continuous class score, wherein the continuous class score is determined by aggregating the module class scores of each computational module of the computational framework.
14. The computational method of any one of claims 6-11 further comprising:determining a continuous class score, wherein the continuous class score is determined by a regression model that integrates the omic enrichment scores of each entity.
15. The computational method of claim 13 or 14 further comprising:classifying each entity into a class status, wherein each entity is classified basedon the value of the entity’s enrichment score within the continuous class score.
16. The computational method of any one of claims 6 further comprising:classifying each entity into a class status, wherein each entity is classified basedon the value of the entity’s class status score.
17. The computational method of any one of claims 1-16 further comprising:applying a smoothing process to one or more of: values of the omic features prior to binarization; values of resulting class status scores; or values of resulting enrichment scores; wherein the smoothing process smoothens the values for each entity based on the values of similar entities.
18. The method of claim 17, wherein the smoothing process uses a Markov diffusionprocess or a nearest neighbor’s process to determine similarity among the entities and to smoothen the values.
19. The computational method of any one of claims 1-18, wherein the omic data isgene expression data from RNA-seq, methylation pattern data from methyl-seq, histone patter data from ChIP-seq, open chromatin data from ATAC-seq, or protein expression data from proteomic analysis.
20. The computational method of any one of claims 1-19, wherein each entity isdefined as one of the following: a cell or a sample.
21. A computational method to be performed by a computational system to determinepotency composition of a collection of cells, comprising:receiving transcriptomic data results of a collection of cells, each transcriptomicdata result comprising gene expression features; constructing a matrix of genes x cells derived from the transcriptomic data results;and determining a potency of each cell by inputting the matrix of genes x cells asfeatures within a trained computational framework that comprises a set of one or more binary networks encoding trained binarization of omic features within one or more omicfeature sets to predict the potency status of each cell, wherein the potency compositioncomprises two or more potency status that characterize potency composition.
22. The computational method of claim 21, wherein each binary network learns whichgenes predict the potency status based on binarization of the matrix of genes x cells.
23. The computational method of claim 22, wherein each binary network further learnsa prediction power of each gene for predicting the potency status.
24. The computational method of claim 22 or 23, wherein the computational frameworkcomprises two or more computational modules, each computational module comprises abinary network for binarization of the matrix of omic features x entities to assess a potencystatus, wherein each computational module is trained to predict a likelihood of the potency status.
25. The computational method of claim 22 or 23, wherein the computational frameworkconsists of a single module to assess two or more potency status with a binary network for binarization of the matrix of omic features x entities, wherein the module is trained to predict a likelihood of each potency status of the potency composition.
26. The computational method of claim 24 or 25, wherein each module yields a set ofone or more potency status scores for each cell.
27. The computational method of claim 26, wherein the potency status scorecomprises gene enrichment score that characterizes enrichment of gene features withinthe cell.
28. The method of claim 27 further comprising:for each cell, ranking and trimming the gene features based on quantity within theentity’s transcriptomic data results.
29. The computational method of claim 28, wherein the gene enrichment score isbased on the rank of the gene features and the learned prediction power of the gene features.
30. The method of claim 27 further comprising:for each entity, normalizing the gene features to total gene count or total transcriptcount of the entity’s transcriptomic data results.
31. The computational method of claim 30, wherein the gene enrichment score isbased on an aggregation of gene features.
32. The computational method of any one of claims 26-31 further comprising:integrating the gene enrichment scores of each cell to yield a module potency score.
33. The computational method of claim 32 further comprising:determining a continuous potency score, wherein the continuous potency score is determined by aggregating the module potency scores of each computational module of the computational framework.
34. The computational method of any one of claims 26-31 further comprising:determining a continuous potency score, wherein the continuous potency score is determined by a regression model that integrates the gene enrichment scores of each cell.
35. The computational method of claim 33 or 34 further comprising:classifying each cell into a potency status, wherein each cell is classified based onthe value of the cell’s enrichment score within the continuous potency score.
36. The computational method of any one of claims 26 further comprising:classifying each cell into a potency status, wherein each cell is classified based onthe value of the cell’s potency status score.
37. The computational method of any one of claims 21-36 further comprising:applying a smoothing process to one or more of: values of the gene features prior to binarization; values of resulting potency status scores; or values of resulting enrichment scores; wherein the smoothing process smoothens the values for each cell based on the values of similar cells.
38. The method of claim 37, wherein the smoothing process uses a Markov diffusionprocess or a nearest neighbor’s process to determine similarity among the cells and to smoothen the values.
39. The computational method of any one of claims 21-38, wherein the potencycomposition comprises the following potency classes: totipotent, pluripotent, multipotent,oligopotent, unipotent, and differentiated40. The computational method of any one of claims 21-39, wherein the collection ofcells comprises cancer cells.
41. The computational method of any one of claims 21-40, wherein the collection ofcells comprises developmental cells.
42. The computational method of any one of claims 21-41, wherein the collection ofcells comprises quiescent stem cells, mitotic stem cells, and senescent cells.
43. A method to treat cancer based on potency phenotype, comprising:obtaining omic data of cancer cells of sample collected from an individual havingcancer; classifying each cell with a potency status for each cancer cell; andadministering a treatment to the individual with a treatment appropriate based ona collective potency status of the cancer cells.
44. The method of claim 43, wherein administering a treatment to the individualcomprises forgoing administration of immune checkpoint inhibitors when the number ofcells classified with a potency status of multipotent, oligopotent, or unipotent is over athreshold.
45. The method of claim 43, wherein administering a treatment to the individualcomprises administration of immune checkpoint inhibitors when the number of cellsclassified with a potency status status multipotent, oligopotent, or unipotent is below athreshold.