Methods for classifying the differentiation state of cells and related compositions of differentiated cells - Patents.com

JP2025512442A5Pending Publication Date: 2026-04-17ASPEN NEUROSCIENCE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ASPEN NEUROSCIENCE INC
Filing Date
2023-04-14
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively classify and identify differentiated states in cell populations in vitro, especially in cell replacement therapy, cells with specific differentiated states.

Method used

A computing device is provided that includes a processor and memory for calculating similarity scores by reference data sets and test data sets to classify the differentiation state of cells. The reference data set contains information on the differences in gene expression between different differentiation states, and the test data set contains the gene expression levels of the cells to be classified.

Benefits of technology

Accurate classification of the differentiation states of cell populations in vitro is achieved, and methods for selecting and implanting cells in specific differentiated states are provided, suitable for the treatment of neurodegenerative diseases such as Parkinson's disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are methods for classifying the differentiation state of an in vitro population of cells, such as an in vitro population of neural cells, as well as methods for selecting and / or transplanting an in vitro population of cells having a desired differentiation state. Also provided herein are computing devices for performing the provided methods and related compositions, products, and kits, including for use in methods of treating a subject having a disease or condition, such as a neurodegenerative disease, such as Parkinson's disease.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Application No. 63 / 331,783, entitled "METHODS OF CLASSIFYING THE DIFFERENTIATION STATE OF CELLS AND RELATED COMPOSITIONS OF DIFFERENTIATED CELLS," filed April 15, 2022, and U.S. Provisional Application No. 63 / 353,525, entitled "METHODS OF CLASSIFYING THE DIFFERENTIATION STATE OF CELLS AND RELATED COMPOSITIONS OF DIFFERENTIATED CELLS," filed June 17, 2022, the contents of each of which are incorporated by reference in their entirety herein for all purposes.

[0002] FIELD OF THEINVENTION The present disclosure relates to methods for classifying the differentiation state of an in vitro population of cells, such as an in vitro population of neural cells, and methods for selecting and / or transplanting an in vitro population of cells having a desired differentiation state. Also provided herein are computing devices for performing the provided methods and related compositions, products and kits, including for use in methods of treating a subject having a disease or condition, such as a neurodegenerative disease, e.g., Parkinson's disease. [Background technology]

[0003] Various methods for differentiating pluripotent stem cells into lineage-specific cell populations and the resulting cell compositions are intended to be utilized in cell replacement therapy for patients with diseases resulting in loss of function of defined cell populations. In some aspects, it is desirable to administer cells with a specific differentiation state. Improved methods of sorting and identifying such cells are needed. Summary of the Invention

[0004] In some embodiments, provided herein is a computing device for classifying a differentiation state of an in vitro population of cells, the computing device comprising a memory including: a first reference dataset comprising a representation of gene expression levels for one or more genes that are differentially expressed between cells in a first differentiation state and cells in a second differentiation state; and a second reference dataset comprising a representation of gene expression levels for one or more genes that are differentially expressed between cells in the second differentiation state and cells in a third differentiation state.

[0005] In some of any of the provided embodiments, the computing device further comprises a processor that executes instructions stored in the memory to perform a method, the method including: (a) receiving as input a test dataset including expression levels for genes expressed in one or more test cells included in the in vitro population of cells, the expression levels in the test dataset including expression levels for (i) one or more of the genes whose expression level representations are included in the first reference dataset, and (ii) one or more of the genes whose expression level representations are included in the second reference dataset; (b) calculating, using the test dataset and the first reference dataset, a first similarity score indicative of whether the differentiation state of the test cell is more similar to the first differentiation state or more similar to the second differentiation state; (c) calculating, using the test dataset and the second reference dataset, a second similarity score indicative of whether the differentiation state of the test cell is more similar to the second differentiation state or more similar to a third differentiation state; and (d) classifying the differentiation state of the one or more test cells based on one or both of the first similarity score and the second similarity score.

[0006] In some embodiments, the classification is based on one of the first similarity score and the second similarity score. In some embodiments, the classification is based on the lower of the first similarity score and the second similarity score. In some embodiments, the classification is based on the higher of the first similarity score and the second similarity score. In some embodiments, the classification is based on the first similarity score. In some embodiments, the classification is based on the second similarity score.

[0007] In some embodiments, the classification is based on both the first similarity score and the second similarity score.

[0008] In some of any of the provided embodiments, the computing device further comprises a processor that executes instructions stored in the memory to perform a method, the method including: (a) receiving as input a test dataset including expression levels for genes expressed in one or more test cells included in the in vitro population of cells, the expression levels in the test dataset including expression levels for (i) one or more of the genes whose expression level representations are included in the first reference dataset, and (ii) one or more of the genes whose expression level representations are included in the second reference dataset; (b) calculating, using the test dataset and the first reference dataset, a first similarity score indicative of whether the differentiation state of the test cell is more similar to the first differentiation state or more similar to the second differentiation state; (c) calculating, using the test dataset and the second reference dataset, a second similarity score indicative of whether the differentiation state of the test cell is more similar to the second differentiation state or more similar to a third differentiation state; and (d) classifying the differentiation state of the one or more test cells based on the first similarity score and the second similarity score.

[0009] In some of any of the provided embodiments, the memory further includes a control dataset including a representation of gene expression levels for one or more genes expressed in the cell at one or more control differentiation states, which may be the same or different from one of the first, second, or third differentiation states. In some of any of the provided embodiments, the test dataset includes gene expression levels for one or more of the genes whose expression level representations are included in the control dataset, and the instructions include calculating a degree of correlation between the representation of gene expression levels for the one or more genes in the control dataset and the gene expression levels for the one or more genes in the test dataset to calculate a correlation score, and classifying the differentiation state of the one or more test cells based on the correlation score and one or both of the first similarity score and the second similarity score.

[0010] In some embodiments, the classification is based on the correlation score and one of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the lower of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the higher of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the first similarity score. In some embodiments, the classification is based on the correlation score and the second similarity score.

[0011] In some embodiments, the classification is based on the correlation score and both the first similarity score and the second similarity score.

[0012] In some of any of the provided embodiments, the memory further includes a control dataset including a representation of gene expression levels for one or more genes expressed in the cell at one or more control differentiation states, where the control differentiation states can be the same or different from one of the first, second, or third differentiation states. In some of any of the provided embodiments, the test dataset includes gene expression levels for one or more of the genes whose expression level representations are included in the control dataset, and the instructions include calculating a degree of correlation between the representation of gene expression levels for the one or more genes in the control dataset and the gene expression levels for the one or more genes in the test dataset to calculate a correlation score, and classifying the differentiation state of the one or more test cells is based on the first similarity score, the second similarity score, and the correlation score.

[0013] In some of any of the provided embodiments, a correlation score is calculated prior to calculating the first similarity score and the second similarity score, and the method terminates if the correlation score of the test cell does not meet a predetermined cutoff value.

[0014] In some of any of the provided embodiments, the control dataset includes gene expression levels that are normalized by counts per million mapped reads (CPM) and filtered to include only gene expression levels that exceed a threshold CPM value. In some of any of the provided embodiments, the control dataset includes a centroid of the gene expression levels of one or more genes in the control dataset. In some of any of the provided embodiments, the correlation score is calculated by normalizing the gene expression levels of one or more genes in the test dataset and calculating the correlation between the gene expression levels of one or more genes in the test dataset and the centroid. In some of any of the provided embodiments, the control dataset includes a coefficient of variation (CV) value of the gene expression levels of one or more genes in the control dataset, and the correlation to the centroid is weighted by the inverse of the CV value.

[0015] In some of the embodiments provided, the in vitro population of cells is derived from a culture of cells differentiated from pluripotent cells subjected to appropriate differentiation conditions. In some of the embodiments provided, the first differentiation state is earlier in the stem cell differentiation pathway than the second differentiation state. In some of the embodiments provided, the second differentiation state is earlier in the stem cell differentiation pathway than the third differentiation state. In some of the embodiments provided, the first differentiation state is in a cell differentiation pathway that parallels the cell differentiation pathway of the second differentiation state.

[0016] In some of the embodiments provided, the population of cells is selected from the group consisting of stem cell derived cardiomyocytes, stem cell derived skeletal muscle cells, stem cell derived renal tubule cells, stem cell derived red blood cells, stem cell derived smooth muscle cells, stem cell derived lung cells, stem cell derived thyroid cells, stem cell derived pancreatic cells, stem cell derived epidermal cells, stem cell derived pigment cells, and stem cell derived neural cells. In some of the embodiments provided, the population of cells is stem cell derived neural cells. In some of the embodiments provided, the second differentiation state is a determined dopaminergic neural cell differentiation state. In some of the embodiments provided, the second differentiation state is a cell with suitability for engraftment.

[0017] In some of any of the provided embodiments, the second differentiation state is a hematopoietic progenitor cell differentiation state.

[0018] In some of any of the provided embodiments, the first reference dataset comprises a representation of gene expression levels for one or more genes selected from Table E1. In some of any of the provided embodiments, the second reference dataset comprises a representation of gene expression levels for one or more genes selected from Table E2. In some of any of the provided embodiments, the first reference dataset comprises a representation of gene expression levels for at least 20 genes selected from Table E1. In some of any of the provided embodiments, the second reference dataset comprises a representation of gene expression levels for at least 20 genes selected from Table E2. In some of any of the provided embodiments, the first reference dataset comprises a representation of gene expression levels for at least 50 genes selected from Table E1. In some of any of the provided embodiments, the second reference dataset comprises a representation of gene expression levels for at least 50 genes selected from Table E2.

[0019] In some of any of the provided embodiments, at least one of the first, second, and third differentiation states is characterized using an in vitro assay. In some of any of the provided embodiments, at least one of the first, second, and third differentiation states is characterized using an in vivo assay. In some of any of the provided embodiments, the in vivo assay includes determining whether the reference cells can survive, engraft, and / or innervate tissue when administered to an animal or human subject. In some of any of the provided embodiments, the in vivo assay includes determining whether the reference cells improve or reverse symptoms of a neurodegenerative disease when transplanted into an animal or human subject.

[0020] In some of any of the provided embodiments, the animal subject comprises an animal model of Parkinson's disease. In some of any of the provided embodiments, the memory further comprises one or more additional reference datasets, each of the additional reference datasets comprising a representation of gene expression levels for one or more genes that are differentially expressed between cells in the second differentiation state and cells in the additional differentiation state, and the processor executes instructions for calculating, using the additional reference datasets, one or more additional similarity scores indicative of whether the differentiation state of the test cell is more similar to the second differentiation state or one of the one or more additional differentiation states, and classifying the differentiation state of the one or more test cells is based on the first similarity score, the second similarity score, and the one or more additional similarity scores.

[0021] In some of any of the provided embodiments, the representation of gene expression levels in the first reference data set and / or the second reference data set is obtained using machine learning. In some of any of the provided embodiments, the machine learning comprises principal component analysis. In some of the provided embodiments, the representation of gene expression levels in the first reference data set and / or the second reference data set comprises normalized gene expression levels. In some of the provided embodiments, if the first and second similarity scores indicate that the differentiation state of the one or more test cells is more similar to the second differentiation state, the differentiation state of the one or more test cells is classified as being the second differentiation state. In some of the provided embodiments, if the first similarity score indicates that the differentiation state of the one or more test cells is more similar to the second differentiation state, the differentiation state of the one or more test cells is classified as being the second differentiation state. In some of the provided embodiments, if the second similarity score indicates that the differentiation state of the one or more test cells is more similar to the second differentiation state, the differentiation state of the one or more test cells is classified as being the second differentiation state.

[0022] In some embodiments, a method for selecting a population of cells having a desired differentiation state includes: (a) calculating a first similarity score using a test dataset and a first reference dataset, wherein the first reference dataset comprises a representation of gene expression levels for one or more genes that are differentially expressed between cells in a first differentiation state and cells in a second differentiation state, and the test dataset comprises expression levels for genes expressed in one or more test cells comprised in an in vitro population of cells, the expression levels in the test dataset comprising an expression level representation for one or more of the genes comprised in the first reference dataset, and the first similarity score indicates that the differentiation state of the test cells is more similar to the first differentiation state or the second differentiation state. Also provided herein are methods comprising: (a) determining whether a differentiation state of the test cell is more similar to the second differentiation state or the third differentiation state; (b) calculating a second similarity score using the test dataset and a second reference dataset, wherein the second reference dataset comprises representations of gene expression levels for one or more genes that are differentially expressed between cells in the second differentiation state and cells in a third differentiation state, and the expression levels in the test dataset comprise expression levels for one or more of the genes whose expression level representations are included in the second reference dataset, the second similarity score indicating whether the differentiation state of the test cell is more similar to the second differentiation state or the third differentiation state; and (c) classifying the differentiation state of the one or more test cells based on one or both of the first similarity score and the second similarity score.

[0023] In some embodiments, the classification is based on one of the first similarity score and the second similarity score. In some embodiments, the classification is based on the lower of the first similarity score and the second similarity score. In some embodiments, the classification is based on the higher of the first similarity score and the second similarity score. In some embodiments, the classification is based on the first similarity score. In some embodiments, the classification is based on the second similarity score.

[0024] In some embodiments, the classification is based on both the first similarity score and the second similarity score.

[0025] In some embodiments, a method for selecting a population of cells having a desired differentiation state includes: (a) calculating a first similarity score using a test dataset and a first reference dataset, wherein the first reference dataset comprises a representation of gene expression levels for one or more genes that are differentially expressed between cells in a first differentiation state and cells in a second differentiation state, and the test dataset comprises expression levels for genes expressed in one or more test cells comprised in an in vitro population of cells, the expression levels in the test dataset comprising an expression level representation for one or more of the genes comprised in the first reference dataset, and the first similarity score indicates whether the differentiation state of the test cells is more similar to the first differentiation state or the second differentiation state. Also provided herein are methods comprising: (a) calculating a second similarity score using the test dataset and a second reference dataset, the second reference dataset comprising representations of gene expression levels for one or more genes that are differentially expressed between cells in the second differentiation state and cells in a third differentiation state, the expression levels in the test dataset comprising expression levels for one or more of the genes whose expression level representations are included in the second reference dataset, the second similarity score indicating whether the differentiation state of the test cells is more similar to the second differentiation state or the third differentiation state; and (b) calculating a second similarity score using the test dataset and a second reference dataset, the second reference dataset comprising representations of gene expression levels for one or more genes that are differentially expressed between cells in the second differentiation state and cells in a third differentiation state, the second similarity score indicating whether the differentiation state of the test cells is more similar to the second differentiation state or the third differentiation state; and (c) classifying the differentiation state of the one or more test cells based on the first similarity score and the second similarity score.

[0026] In some of any of the provided embodiments, the test dataset includes gene expression levels for one or more genes, the control dataset including a representation of gene expression levels for one or more genes expressed in cells in a control differentiation state, which may be the same as or different from one of the first, second or third differentiation states, and the method further includes calculating a degree of correlation between the representation of gene expression levels for the one or more genes in the control dataset and the gene expression levels for the one or more genes in the test dataset to calculate a correlation score, and classifying the differentiation state of the one or more test cells is based on the correlation score and one or both of the first similarity score and the second similarity score.

[0027] In some embodiments, the classification is based on the correlation score and one of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the lower of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the higher of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the first similarity score. In some embodiments, the classification is based on the correlation score and the second similarity score.

[0028] In some embodiments, the classification is based on the correlation score and both the first similarity score and the second similarity score.

[0029] In some of any of the provided embodiments, the test dataset includes gene expression levels for one or more genes, wherein the control dataset includes a representation of gene expression levels for one or more genes expressed in cells in a control differentiation state, which may be the same as or different from one of the first, second or third differentiation states, and the method further includes calculating a degree of correlation between the representation of gene expression levels for the one or more genes in the control dataset and the gene expression levels for the one or more genes in the test dataset to calculate a correlation score, and classifying the differentiation state of the one or more test cells is based on the first similarity score, the second similarity score and the correlation score.

[0030] In some of any of the provided embodiments, a correlation score is calculated prior to calculating the first similarity score and the second similarity score, and the method terminates if the correlation score of the test cell does not meet a predetermined cutoff value.

[0031] In some of any of the provided embodiments, the control dataset includes gene expression levels that are normalized by counts per million mapped reads (CPM) and filtered to include only gene expression levels that exceed a threshold CPM value. In some of any of the provided embodiments, the control dataset includes a centroid of the gene expression levels of one or more genes in the control dataset. In some of any of the provided embodiments, the correlation score is calculated by normalizing the gene expression levels of one or more genes in the test dataset and calculating the correlation between the gene expression levels of one or more genes in the test dataset and the centroid. In some of any of the provided embodiments, the control dataset includes a coefficient of variation (CV) value of the gene expression levels of one or more genes in the control dataset, and the correlation to the centroid is weighted by the inverse of the CV value.

[0032] In some of any of the provided embodiments, the first differentiation state is earlier in the stem cell differentiation pathway than the second differentiation state. In some of any of the provided embodiments, the second differentiation state is earlier in the stem cell differentiation pathway than the third differentiation state. In some of any of the provided embodiments, the first differentiation state is in a cell differentiation pathway that parallels the cell differentiation pathway of the second differentiation state.

[0033] In some of any of the provided embodiments, the population of cells is selected from the group consisting of stem cell derived cardiomyocytes, stem cell derived skeletal muscle cells, stem cell derived renal tubule cells, stem cell derived red blood cells, stem cell derived smooth muscle cells, stem cell derived lung cells, stem cell derived thyroid cells, stem cell derived pancreatic cells, stem cell derived epidermal cells, stem cell derived pigment cells, and stem cell derived neural cells. In some of any of the provided embodiments, the population of cells is stem cell derived neural cells.

[0034] In some of the embodiments provided, the second differentiation state is a differentiation state of a determined dopaminergic neuronal cell. In some of the embodiments provided, the second differentiation state is a differentiation state of a cell that is compatible with engraftment.

[0035] In some of any of the provided embodiments, the second differentiation state is a hematopoietic progenitor cell differentiation state.

[0036] In some of any of the provided embodiments, the first reference dataset comprises a representation of gene expression levels for one or more genes selected from Table E1. In some of any of the provided embodiments, the second reference dataset comprises a representation of gene expression levels for one or more genes selected from Table E2. In some of any of the provided embodiments, the first reference dataset comprises a representation of gene expression levels for at least 20 genes selected from Table E1. In some of any of the provided embodiments, the second reference dataset comprises a representation of gene expression levels for at least 20 genes selected from Table E2. In some of any of the provided embodiments, the first reference dataset comprises a representation of gene expression levels for at least 50 genes selected from Table E1. In some of any of the provided embodiments, the second reference dataset comprises a representation of gene expression levels for at least 50 genes selected from Table E2.

[0037] In some of any of the provided embodiments, at least one of the first, second, and third differentiation states is characterized using an in vivo assay. In some of any of the provided embodiments, the in vivo assay includes determining whether the reference cells can survive, engraft, and / or innervate tissue when administered to an animal or human subject.

[0038] In some of any of the provided embodiments, the in vivo assay comprises determining whether the reference cells improve or reverse symptoms of the neurodegenerative disease when transplanted into an animal or human subject. In some of any of the provided embodiments, the animal subject comprises an animal model of Parkinson's disease. In some of any of the provided embodiments, the method further comprises calculating one or more additional similarity scores using one or more additional reference datasets, each of the additional reference datasets comprising a representation of gene expression levels for one or more genes that are differentially expressed between cells in the second differentiation state and cells in the additional differentiation state, the one or more additional similarity scores indicating whether the differentiation state of the test cells is more similar to the second differentiation state or one of the one or more additional differentiation states, and classifying the differentiation state of the one or more test cells is based on the first similarity score, the second similarity score, and the one or more additional similarity scores.

[0039] In some of the embodiments provided, the representation of gene expression levels in the first reference data set and / or the second reference data set is obtained using machine learning. In some of the embodiments provided, the machine learning comprises principal component analysis. In some of the embodiments provided, the representation of gene expression levels in the first reference data set and / or the second reference data set comprises normalized gene expression levels.

[0040] In some of any of the provided embodiments, the method further comprises classifying the differentiation state of the one or more test cells as the second differentiation state if the first similarity score indicates that the differentiation state of the one or more test cells is more similar to the second differentiation state. In some of any of the provided embodiments, the method further comprises classifying the differentiation state of the one or more test cells as the second differentiation state if the second similarity score indicates that the differentiation state of the one or more test cells is more similar to the second differentiation state. In some of any of the provided embodiments, the method further comprises selecting an in vitro population of cells comprising one or more test cells classified as having the second differentiation state as having the desired differentiation state.

[0041] In some of any of the provided embodiments, the method further includes classifying the differentiation state of the one or more test cells as the second differentiation state if the first and second similarity scores indicate that the differentiation state of the one or more test cells is more similar to the second differentiation state. In some of any of the provided embodiments, the method further includes selecting an in vitro population of cells comprising the one or more test cells classified as having the second differentiation state as having the desired differentiation state.

[0042] In some embodiments, a method for selecting a population of cells having a desired differentiation state includes the steps of: (a) detecting a differential expression level of one or more of AC010247.2, ANKRD33B, APC2, AQP4, ASCL1, AURKB, BARHL2, CACNA1G, CAPN6, CBLN1, CCNB2, CDH1, CDH20, CHGA, COL1A1, COL1A2, COL22A1, COL4A1, CRABP1, DBX1, or COL1B1 for one or more test cells in an in vitro population of cells; , DCN, DCX, DDC, DOCK10, E2F4, EDNRB, ESRP1, EZH2, FABP7, FBLN1, FLRT3, FOXA2, FOXM1, GAP43, GFAP, GFRA1, GJA1, GLRA2, HE S1, HES2, HES5, ITGA5, JPH4, LDHA, LIN28A, LIX1, LMX1A, LUM, NCAM1, NES, NEUROG2, NGFR, NKX2-2, NMNAT2, NPTX1, NR4A2, NR4 A2(NURR1), NSG2, NFYA, OLFM3, OLIG1, OLIG2, OTX2, P4HA1, PBX1, PDGFRA, PIEZO2, PITX3, PLP1, PMEL, PMP2, POSTN, POU2F2, PPP2R2B, PRTG, PTTG1, REST, RET, RFX4, RFX4, SALL4, SIN3A, SLC16A3, SLC18A2, SLC1A, SLC1A2, SLC1A3, SLC4A4, SMAD4, SNA Also provided herein is a method comprising: (a) obtaining a test dataset comprising gene expression levels of one or more genes selected from P25, SOX10, SOX2, SOX9, STMN2, SUZ12, SV2B, SYN1, SYT1, SYT13, TH, TOP2A, TPH1, TPM2, and TXNIP; and (b) applying the gene expression levels as input to a process configured to predict whether a population of cells has a desired differentiation state.

[0043] In some of any of the embodiments provided, the in vitro population of cells comprises stem cell-derived neural cells. In some of any of the embodiments provided, the desired differentiation state is a determined differentiation state of dopaminergic neural cells. In some of any of the embodiments provided, the desired differentiation state is a differentiation state of cells that have a suitability for engraftment.

[0044] In some of any of the embodiments provided, the desired differentiation state is that of a hematopoietic progenitor cell.

[0045] In some embodiments, a method for selecting a population of cells predicted to exhibit neurite outgrowth following transplantation into a brain region comprises: (a) detecting a subset of AC010247.2, ANKRD33B, APC2, AQP4, ASCL1, AURKB, BARHL2, CACNA1G, CAPN6, CBLN1, CCNB2, CDH1, CDH20, CHGA, COL1A1, COL1A2, COL22A1, COL4A1, CRA1, COL22A2, COL22A3, COL22A4, COL22A5, COL22A6, COL22A7, COL22A8, COL22A9, COL22B1, COL22B2, COL22B3, COL22B4, COL22B5, COL22B6, COL22B7, COL22B8, COL22B9, COL22B1 ... BP1, DBX1, DCN, DCX, DDC, DOCK10, E2F4, EDNRB, ESRP1, EZH2, FABP7, FBLN1, FLRT3, FOXA2, FOXM1, GAP43, GFAP, GFRA1, GJA1, GLR A2, HES1, HES2, HES5, ITGA5, JPH4, LDHA, LIN28A, LIX1, LMX1A, LUM, NCAM1, NES, NEUROG2, NGFR, NKX2-2, NMNAT2, NPTX1, NR4A2, NR4A2(NURR1), NSG2, NFYA, OLFM3, OLIG1, OLIG2, OTX2, P4HA1, PBX1, PDGFRA, PIEZO2, PITX3, PLP1, PMEL, PMP2, POSTN, POU2F2 , PPP2R2B, PRTG, PTTG1, REST, RET, RFX4, RFX4, SALL4, SIN3A, SLC16A3, SLC18A2, SLC1A, SLC1A2, SLC1A3, SLC4A4, SMAD4, SNAP2 Also provided herein is a method, comprising: (a) obtaining a test dataset comprising gene expression levels of one or more genes selected from the group consisting of: 5, SOX10, SOX2, SOX9, STMN2, SUZ12, SV2B, SYN1, SYT1, SYT13, TH, TOP2A, TPH1, TPM2, and TXNIP; and (b) applying the gene expression levels as input to a process configured to predict whether a population of cells will exhibit neurite outgrowth following transplantation into a brain region.

[0046] In some of any of the provided embodiments, the one or more genes are at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or more of AC010247.2, ANKRD33B, APC2, AQP4, ASCL 1, AURKB, BARHL2, CACNA1G, CAPN6, CBLN1, CCNB2, CDH1, CDH20, CHGA, COL1A1, COL1A2, COL22A1, COL4A1, CRABP1, DBX1, DCN, DCX, DDC, DOCK10, E2F4, EDNRB, ESRP1, EZH2, FABP7, FBLN1, FLRT3, FOXA2, FOXM1, GAP43, GFAP, GFRA1, GJA1, GLRA2, HES1, HES2, HES5, ITGA5, JPH4, LDHA, LIN28A, LIX1, LMX1A, LUM, NCAM1, NES, NEUROG2, NGFR, NKX2-2, NMNAT2, NPTX1, NR 4A2, NR4A2(NURR1), NSG2, NFYA, OLFM3, OLIG1, OLIG2, OTX2, P4HA1, PBX1, PDGFRA, PIEZO2, PITX3, PLP1, PMEL, PMP2, PO These include STN, POU2F2, PPP2R2B, PRTG, PTTG1, REST, RET, RFX4, RFX4, SALL4, SIN3A, SLC16A3, SLC18A2, SLC1A, SLC1A2, SLC1A3, SLC4A4, SMAD4, SNAP25, SOX10, SOX2, SOX9, STMN2, SUZ12, SV2B, SYN1, SYT1, SYT13, TH, TOP2A, TPH1, TPM2, and TXNIP.

[0047] In some of any of the provided embodiments, the process includes a machine learning model. In some of any of the provided embodiments, the machine learning model is trained using gene expression levels of one or more genes. In some of any of the provided embodiments, one or more outputs of the machine learning model are used to predict whether a population of cells has a desired differentiation state. In some of any of the provided embodiments, one or more outputs of the machine learning model are used to predict whether a population of cells will exhibit neurite outgrowth after transplantation in the brain region. In some of any of the provided embodiments, the method further includes classifying a differentiation state of one or more test cells based on one or more outputs of the machine learning model. In some of any of the provided embodiments, the method further includes predicting whether a test cell will exhibit neurite outgrowth after transplantation in the brain region based on one or more outputs of the machine learning model. In some of any of the provided embodiments, the method further includes selecting an in vitro population of cells comprising one or more test cells classified as having a desired differentiation state. In some of any of the provided embodiments, the method further includes selecting an in vitro population of cells comprising one or more test cells predicted to exhibit neurite outgrowth after transplantation in the brain region.

[0048] Also provided herein in some embodiments is a method for transplanting a population of cells having a desired differentiation state into a subject, the method comprising: (a) selecting a population of cells having the desired differentiation state using any of the methods provided; and (b) transplanting the population of cells into the subject. In some of any of the embodiments provided, the cells having the desired differentiation state are dopaminergic cells, and the population of cells is transplanted into a brain region of the subject. In some of any of the embodiments provided, the cells having the desired differentiation state are derived from a culture of cells differentiated from pluripotent cells under conditions that cause the cells to neuronally differentiate.

[0049] In some of any of the embodiments provided, the cells having the desired differentiation state are hematopoietic progenitor cells, and the population of cells is transplanted into a brain region of the subject. In some of any of the embodiments provided, the cells having the desired differentiation state are derived from a culture of cells differentiated from pluripotent cells under conditions that cause the cells to neuronally differentiate.

[0050] Also provided herein, in some embodiments, is a pharmaceutical composition comprising a pharmaceutical carrier and a population of cells having a desired differentiation state, where the cells are selected using any of the methods provided.

[0051] In some of any of the embodiments provided, the cells having the desired differentiation state are neural cells suitable for treating a neurodegenerative disease when transplanted into the brain of a subject in need of such treatment. In some of any of the embodiments provided, the neural cells comprise determined dopaminergic cells. In some of any of the embodiments provided, the neural cells comprise engraftable neural cells.

[0052] In some of any of the embodiments provided, the neural cells comprise hematopoietic progenitor cells.

[0053] Also provided herein in some embodiments is a method for training a machine learning model to classify a differentiation state of an in vitro population of cells, the method comprising: (a) obtaining, for a reference population of a plurality of cells, gene expression levels for one or more genes that are differentially expressed between cells in a first differentiation state and cells in a second differentiation state, and applying the gene expression levels as inputs to train a first machine learning model to predict whether the in vitro population of cells comprises one or more test cells having a differentiation state more similar to the first differentiation state or the second differentiation state; and (b) obtaining, for the reference population of a plurality of cells, gene expression levels for one or more genes that are differentially expressed between cells in the second differentiation state and cells in a third differentiation state, and applying the gene expression levels as inputs to train a second machine learning model to predict whether the in vitro population of cells comprises one or more test cells having a differentiation state more similar to the second differentiation state or the third differentiation state.

[0054] Also provided herein in some embodiments is a method of training a machine learning model to classify a differentiation state of an in vitro population of cells, the method comprising: (a) selecting one or more genes that are differentially expressed between cells in a first differentiation state and cells in a second differentiation state, and applying as input the expression levels of the selected genes relative to a reference population of a plurality of cells to train a first machine learning model to predict whether the in vitro population of cells comprises one or more test cells having a differentiation state more similar to the first differentiation state or the second differentiation state; and (c) selecting one or more genes that are differentially expressed between cells in the second differentiation state and cells in a third differentiation state, and applying as input the expression levels of the selected genes relative to a reference population of a plurality of cells to train a second machine learning model to predict whether the in vitro population of cells comprises one or more test cells having a differentiation state more similar to the second differentiation state or the third differentiation state.

[0055] In some of any of the provided embodiments, the method further includes obtaining gene expression levels for one or more genes expressed in cells in a control differentiation state, which can be the same as or different from one of the first, second, or third differentiation states, and applying the gene expression levels as input to train a control machine learning model to predict whether the in vitro population of cells contains one or more test cells that are similar to cells in the control differentiation state.

[0056] Also provided herein are pharmaceutical compositions comprising a pharmaceutical carrier and a population of neural cells, wherein the cells are selected using any of the methods provided.

[0057] Also provided herein is an in vitro stem cell derived neural cell population comprising cells expressing one or more genes selected from the group consisting of CCNB2, AURKB, PTTG1, TOP2A, NEUROG2, HES1, REST, E2F4, FOXM1, SIN3A, NFYA, LIN28A, FLRT3, ITGA5, NES, SOX2, SOX9 and RFX4. In some embodiments, the in vitro stem cell derived neural cell population is one in which (1) at least one gene from the one or more genes is selected from the group consisting of CCNB2, AURKB, PTTG1, TOP2A, NEUROG2, HES1, REST, E2F4, FOXM1, SIN3A, NFYA, LIN28A, FLRT3 and ITGA5, and (2) at least one gene from the one or more genes is selected from the group consisting of NES, SOX2, SOX9 and RFX4. In some embodiments, at least one of the one or more genes is REST.

[0058] In some of any of the embodiments of the in vitro stem cell derived neural cell populations, at least 50% of the cells in the population express one or more genes.

[0059] In some of any of the embodiments of the in vitro stem cell derived neural cell populations, at least 60% of the cells in the population express one or more genes.

[0060] In some of any of the embodiments of the in vitro stem cell derived neural cell populations, at least 70% of the cells in the population express one or more genes.

[0061] In some of any of the embodiments of the in vitro stem cell derived neural cell populations, at least 80% of the cells in the population express one or more genes.

[0062] In some of any of the embodiments of the in vitro stem cell derived neural cell populations, at least 90% of the cells in the population express one or more genes.

[0063] In some of the embodiments of any of the in vitro stem cell derived neural cell populations, the cells in the population express EN1 and CORIN.In some embodiments, less than 20% of the total cells in the composition express TH.In some embodiments, less than 10% of the total cells in the composition express TH.

[0064] In some embodiments of any of the in vitro stem cell derived neural cell populations, expression is RNA expression. In some embodiments, RNA expression is measured by RNA sequencing.

[0065] In some of any of the embodiments of the in vitro stem cell derived neural cell populations, the populations have been differentiated in vitro from pluripotent stem cells (PSCs).

[0066] In some of the embodiments of any of the in vitro stem cell derived neural cell populations, the one or more genes are overexpressed in the cells of the population compared to iPSCs. In some embodiments, the one or more genes are overexpressed in the cells of the population compared to the cells of the precursor population differentiated from iPSCs. In some embodiments, the one or more genes are overexpressed in the cells of the population compared to the cells of the mature committed dopaminergic neural cell population differentiated from iPSCs. In some embodiments, the mature committed dopaminergic neural cell expresses LMX1A and / or NR4A2 (NURR1). In some embodiments, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the cells in the committed dopaminergic neural cell population express LMX1A and / or NR4A2. In some embodiments, overexpression is a positive log2 fold change of greater than or about 1.5-fold, greater than or about 2.0-fold, greater than or about 3.0-fold, greater than or about 4.0-fold, or greater than or about 5-fold.

[0067] In some of the embodiments of any of the in vitro stem cell derived neural cell populations, the one or more genes are genes that are reduced in expression in the cells of the population compared to iPSCs. In some embodiments, the one or more genes are genes that are reduced in expression in the cells of the population compared to cells of a precursor population differentiated from iPSCs. In some embodiments, the one or more genes are genes that are reduced in expression in the cells of the population compared to cells of a mature committed dopaminergic neural cell population differentiated from iPSCs. In some embodiments, the mature committed dopaminergic neural cells express LMX1A and / or NR4A2 (NURR1). In some embodiments, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the cells in the committed dopaminergic neural cell population express LMX1A and / or NR4A2. In some embodiments, decreased expression is a negative log2 fold change of more than or about 1.5-fold, more than or about 2.0-fold, more than or about 3.0-fold, more than or about 4.0-fold, or more than or about 5-fold.

[0068] In some of any of the embodiments of the in vitro stem cell derived neural cell populations, less than 30%, less than 20%, or less than 10% of the cells in the population express LMX1A and / or NR4A2.

[0069] In some of the embodiments of any of the in vitro stem cell derived neural cell populations, the cells in the population can engraft and innervate other cells in vivo. In some embodiments, the cells in the population can exhibit neurite outgrowth when administered to the brain of a subject. In some embodiments, the cells in the population can produce dopamine, and optionally do not produce or substantially do not produce norepinephrine.

[0070] In some of the embodiments of any of the in vitro stem cell derived neural cell populations, the population comprises at least 5 million total cells, at least 10 million total cells, at least 15 million total cells, at least 20 million total cells, at least 30 million total cells, at least 40 million total cells, at least 50 million total cells, at least 100 million total cells, at least 150 million total cells, or at least 200 million total cells.In some embodiments, the population is from 5 million or about 5 million total cells to 200 million or about 200 million total cells, from 5 million or about 5 million total cells to 150 million or about 150 million total cells, from 5 million or about 5 million total cells to 100 million or about 100 million total cells, from 5 million or about 5 million total cells to 50 million or about 50 million total cells, from 5 million or about 5 million total cells to 25 million or about 25 million total cells, from 50 million or about 5 million total cells to 50 million or about 50 million total cells, 10 million or about 5 million total cells to 10 million or about 10 million total cells, 10 million or about 10 million total cells to 200 million or about 200 million total cells, 10 million or about 10 million total cells to 150 million or about 150 million total cells, 10 million or about 10 million total cells to 100 million or about 100 million total cells, 10 million or about 10 million total cells to 50 million or about 50 million total cells, 10 million or about 10 million total cells to 50 million or about 50 million total cells, 10,000,000 total cells to 25 million or about 25 million total cells, 25 million or about 25 million total cells to 200 million or about 200 million total cells, 25 million or about 25 million total cells to 150 million or about 150 million total cells, 25 million or about 25 million total cells to 100 million or about 100 million total cells, 25 million or about 25 million total cells to 50 million or about 50 million total cells, 50 million or about 50 million total cells to 200 million or about 200 million total cells, 50 million or about 50 million total cells to 150 million or about 150 million total cells, 50 million or about 50 million total cells to 100 million or about 100 million total cells, 100 million or about 100 million total cells to 200 million or about 200 million total cells, 100 million or about 100 million total cells to 150 million or about 150 million total cells, or 150 million or about 150 million total cells to 200 million or about 200 million total cells.

[0071] In some of any of the embodiments of the in vitro stem cell derived neural cell populations, at least about 70%, 75%, 80%, 85%, 90%, or 95% of the total cells in the composition are viable.

[0072] Also provided herein is a pharmaceutical composition comprising a pharmaceutical carrier and an in vitro stem cell derived neural cell population provided herein.

[0073] In some embodiments of the pharmaceutical compositions provided, the composition comprises a cryoprotectant, hi some embodiments, the cryoprotectant is selected from the group consisting of glycerol, propylene glycol, and dimethyl sulfoxide (DMSO).

[0074] In any of the embodiments of the provided pharmaceutical compositions, the composition is for use in treating a neurodegenerative disease or condition in a subject, optionally the neurodegenerative disease or condition comprises loss of dopaminergic neurons. In some embodiments, the neurodegenerative disease or condition comprises loss of dopaminergic neurons in the substantia nigra, optionally the SNc. In some embodiments, the neurodegenerative disease or condition is Parkinson's disease. In some embodiments, the neurodegenerative disease or condition is parkinsonism.

[0075] In embodiments of any of the provided pharmaceutical compositions, the composition is for use in treating a neurodegenerative disease or condition in a subject, wherein the neurodegenerative disease or condition involves loss of microglial cells. In some embodiments, the neurodegenerative disease or condition is Parkinson's disease. In some embodiments, the neurodegenerative disease or condition is parkinsonism. In some embodiments, the neurodegenerative disease or condition is an age-related neurodegenerative disease. In some embodiments, the neurodegenerative disease or condition is Alzheimer's disease. In some embodiments, the neurodegenerative disease or condition is frontotemporal dementia.

[0076] Also provided herein are methods of treatment comprising implanting a therapeutically effective amount of any of the provided pharmaceutical compositions into a brain region of a subject in need of treatment. In some embodiments, the number of cells implanted into a subject is about 0.25×10 6 Cells ~ approx. 20×10 6 cells, approximately 0.25 x 10 6 Cells ~ approx. 15 x 10 6 cells, approximately 0.25 x 10 6 Cells ~ approx. 10×10 6 cells, approximately 0.25 x 10 6 Cells ~ approx. 5 x 10 6 cells, approximately 0.25 x 10 6 cells ~ approx. 1 x 10 6 cells, approximately 0.25 x 10 6 Cells ~ approx. 0.75×10 6 cells, approximately 0.25 x 10 6 Cells ~ approx. 0.5×10 6 cells, approximately 0.5 x 10 6 Cells ~ approx. 20×10 6 cells, approximately 0.5 x 10 6 Cells ~ approx. 15 x 10 6 cells, approximately 0.5 x 10 6 Cells ~ approx. 10×10 6 cells, approximately 0.5 x 10 6 Cells ~ approx. 5 x 10 6 cells, approximately 0.5 x 10 6 cells ~ approx. 1 x 10 6 cells, approximately 0.5 x 10 6 Cells ~ approx. 0.75×10 6 cells, approximately 0.75×10 6 Cells ~ approx. 20×10 6 cells, approximately 0.75×10 6 Cells ~ approx. 15 x 10 6 cells, approximately 0.75×10 6 Cells ~ approx. 10×10 6 cells, approximately 0.75×10 6 Cells ~ approx. 5 x 10 6 cells, approximately 0.75×10 6 cells ~ approx. 1 x 10 6 cells, approximately 1 x 10 6 Cells ~ approx. 20×10 6 cells, approximately 1 x 10 6 Cells ~ approx. 15 x 10 6 cells, approximately 1 x 106 Cells ~ approx. 10×10 6 cells, approximately 1 x 10 6 Cells ~ approx. 5 x 10 6 cells, approximately 5 x 10 6 Cells ~ approx. 20×10 6 cells, approximately 5 x 10 6 Cells ~ approx. 15 x 10 6 cells, approximately 5 x 10 6 Cells ~ approx. 10×10 6 cells, approximately 10 x 10 6 Cells ~ approx. 20×10 6 cells, approximately 10 x 10 6 Cells ~ approx. 15 x 10 6 cells, or approximately 15 x 10 6 Cells ~ approx. 20×10 6 It is a cell.

[0077] In any of the embodiments of the provided methods of treatment, the subject has a neurodegenerative disease or condition. In some embodiments, the neurodegenerative disease or condition comprises loss of dopaminergic neurons. In some embodiments, the subject has lost at least 50%, at least 60%, at least 70%, or at least 80% of dopaminergic neurons. In some embodiments, the subject has lost at least 50%, at least 60%, at least 70%, or at least 80% of dopaminergic neurons in the substantia nigra (SN), optionally in the SN pars compacta (SNc). In some embodiments, the neurodegenerative disease or condition is Parkinsonism. In some embodiments, the neurodegenerative disease or condition is Parkinson's disease.

[0078] In any of the embodiments of the provided methods of treatment, the subject has a neurodegenerative disease or condition. In some embodiments, the neurodegenerative disease or condition comprises loss of microglial cells. In some embodiments, the neurodegenerative disease or condition is Parkinsonism. In some embodiments, the neurodegenerative disease or condition is Parkinson's disease. In some embodiments, the neurodegenerative disease or condition is an age-related neurodegenerative disease. In some embodiments, the neurodegenerative disease or condition is Alzheimer's disease. In some embodiments, the neurodegenerative disease or condition is frontotemporal dementia.

[0079] In any of the embodiments of the provided methods, the transplantation into the brain region is the substantia nigra. In some embodiments, the transplantation is by stereotactic injection. In some embodiments, the cells of the pharmaceutical composition are autologous to the subject. [Brief description of the drawings]

[0080] [Figure 1A] 1 shows a decision tree of an exemplary method for using gene expression levels to identify a cell population in a desired differentiation state (e.g., a mid-differentiation state, such as a committed state). In this exemplary method, the gene expression levels of a test cell population are first evaluated to determine whether the expression levels are similar to the expression levels of a reference cell population used during method development. If the expression levels are not too dissimilar or novel, the expression levels are then evaluated to determine whether they are more consistent with the expression levels of a population of early state cells (e.g., progenitor cells) or the expression levels of a population of mid-state cells (e.g., committed cells). If the expression levels are more consistent with the expression levels of a population of mid-state cells, the expression levels are finally evaluated to determine whether they are more consistent with the expression levels of a population of late state cells (e.g., committed cells) or the expression levels of a population of mid-state cells. If the gene expression levels are more consistent with the gene expression levels of a population of mid-state cells, the test population is identified as such.

[0081] [Figure 1B]We show how the provided methods can be used to identify cells in an intermediate differentiation state, such as a determined state. As shown in FIG. 1B, the provided methods can be used for multiple target cell types and multiple in vitro differentiation protocols. Different differentiation protocols within the same target cell type can give different optimal intermediate timing. This intermediate stage of differentiation may be when the cell population is most suitable for transplantation, such as for treating a disease or condition. As shown herein, the method trained on gene expression levels of cells from a first differentiation protocol can also be used to identify cells in an intermediate differentiation state in a second differentiation protocol. The time in days (d) shown in FIG. 1B is only one example.

[0082] [Figure 2A] A flowchart of the training and use of an exemplary machine learning method for identifying a population of intermediate state cells (e.g., committed cells) using gene expression levels is shown. FIG. 2A shows a flowchart for determining a cutoff value of the novelty score indicating whether the gene expression levels of a test cell population are similar to the gene expression levels of a reference cell population used to train the method, and how this cutoff value can be applied to the test cell population. FIG. 2B shows a flowchart for training a first model that distinguishes between early state cells (e.g., progenitor cells) and intermediate state cells (e.g., committed cells; Model A); and for training a separate second model that distinguishes between late state cells (e.g., committed cells) and intermediate state cells (e.g., committed cells; Model B), and how both models can be applied to the test cell population. These procedures applied to the reference cell populations taken at different time points during the neural differentiation protocol are described in Example 1. [Figure 2B]A flowchart of the training and use of an exemplary machine learning method for identifying a population of intermediate state cells (e.g., committed cells) using gene expression levels is shown. FIG. 2A shows a flowchart for determining a cutoff value of the novelty score indicating whether the gene expression levels of a test cell population are similar to the gene expression levels of a reference cell population used to train the method, and how this cutoff value can be applied to the test cell population. FIG. 2B shows a flowchart for training a first model that distinguishes between early state cells (e.g., progenitor cells) and intermediate state cells (e.g., committed cells; Model A); and for training a separate second model that distinguishes between late state cells (e.g., committed cells) and intermediate state cells (e.g., committed cells; Model B), and how both models can be applied to the test cell population. These procedures applied to the reference cell populations taken at different time points during the neural differentiation protocol are described in Example 1.

[0083] [Figure 3A] Results of a machine learning method trained using reference cell populations harvested at different time points during the neural differentiation protocol are shown in Figures 3A-3F. Results for neuronal cell populations are shown in Figure 3G. Results for glial test cell populations are shown in Figure 3H. Results for test cell populations of various cell types are shown in Figure 3H. [Figure 3B] Results of a machine learning method trained using reference cell populations harvested at different time points during the neural differentiation protocol are shown in Figures 3A-3F. Results for neuronal cell populations are shown in Figure 3G. Results for glial test cell populations are shown in Figure 3H. Results for test cell populations of various cell types are shown in Figure 3H. [Figure 3C] Results of a machine learning method trained using reference cell populations harvested at different time points during the neural differentiation protocol are shown in Figures 3A-3F. Results for neuronal cell populations are shown in Figure 3G. Results for glial test cell populations are shown in Figure 3H. Results for test cell populations of various cell types are shown in Figure 3H. [Figure 3D]Results of a machine learning method trained using reference cell populations harvested at different time points during the neural differentiation protocol are shown in Figures 3A-3F. Results for neuronal cell populations are shown in Figure 3G. Results for glial test cell populations are shown in Figure 3H. Results for test cell populations of various cell types are shown in Figure 3H. [Figure 3E] Results of a machine learning method trained using reference cell populations harvested at different time points during the neural differentiation protocol are shown in Figures 3A-3F. Results for neuronal cell populations are shown in Figure 3G. Results for glial test cell populations are shown in Figure 3H. Results for test cell populations of various cell types are shown in Figure 3H. [Figure 3F] Results of a machine learning method trained using reference cell populations harvested at different time points during the neural differentiation protocol are shown in Figures 3A-3F. Results for neuronal cell populations are shown in Figure 3G. Results for glial test cell populations are shown in Figure 3H. Results for test cell populations of various cell types are shown in Figure 3H. [Figure 3G] Results of a machine learning method trained using reference cell populations harvested at different time points during the neural differentiation protocol are shown in Figures 3A-3F. Results for neuronal cell populations are shown in Figure 3G. Results for glial test cell populations are shown in Figure 3H. Results for test cell populations of various cell types are shown in Figure 3H. [Figure 3H] Results of a machine learning method trained using reference cell populations harvested at different time points during the neural differentiation protocol are shown in Figures 3A-3F. Results for neuronal cell populations are shown in Figure 3G. Results for glial test cell populations are shown in Figure 3H. Results for test cell populations of various cell types are shown in Figure 3H.

[0084] [Figure 4A] We show the results of the machine learning method trained using a reference cell population harvested at different time points during the microglial differentiation protocol. Figures 4A-4D show the results for the reference cell population. Figures 5A-5D show the validation results using a test cell population that was not used for model training. [Figure 4B]We show the results of the machine learning method trained using a reference cell population harvested at different time points during the microglial differentiation protocol. Figures 4A-4D show the results for the reference cell population. Figures 5A-5D show the validation results using a test cell population that was not used for model training. [Figure 4C] We show the results of the machine learning method trained using a reference cell population harvested at different time points during the microglial differentiation protocol. Figures 4A-4D show the results for the reference cell population. Figures 5A-5D show the validation results using a test cell population that was not used for model training. [Figure 4D] We show the results of the machine learning method trained using a reference cell population harvested at different time points during the microglial differentiation protocol. Figures 4A-4D show the results for the reference cell population. Figures 5A-5D show the validation results using a test cell population that was not used for model training. [Figure 5A] We show the results of the machine learning method trained using a reference cell population harvested at different time points during the microglial differentiation protocol. Figures 4A-4D show the results for the reference cell population. Figures 5A-5D show the validation results using a test cell population that was not used for model training. [Figure 5B] We show the results of the machine learning method trained using a reference cell population harvested at different time points during the microglial differentiation protocol. Figures 4A-4D show the results for the reference cell population. Figures 5A-5D show the validation results using a test cell population that was not used for model training. [Figure 5C] We show the results of the machine learning method trained using a reference cell population harvested at different time points during the microglial differentiation protocol. Figures 4A-4D show the results for the reference cell population. Figures 5A-5D show the validation results using a test cell population that was not used for model training. [Figure 5D]We show the results of the machine learning method trained using a reference cell population harvested at different time points during the microglial differentiation protocol. Figures 4A-4D show the results for the reference cell population. Figures 5A-5D show the validation results using a test cell population that was not used for model training. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0085] In some embodiments, methods are provided herein for classifying the differentiation state of a population of cells. In some embodiments, methods are also provided herein for selecting a population of cells having a desired differentiation state, e.g., a population of cells classified by any of the methods provided as having a desired differentiation state. In some embodiments, methods are also provided herein for transplanting a population of cells having a desired differentiation state, e.g., a population of cells classified or selected according to any of the methods provided.

[0086] In some embodiments, the methods provided involve classifying the differentiation state of a population of cells. In some embodiments, the classifying is based on characteristics of one or more test cells of the population of cells. In some embodiments, the classifying is based on gene expression levels of one or more test cells of the population of cells.

[0087] In some embodiments, methods are also provided herein for identifying a population of cells predicted to show neurite outgrowth after transplantation into a brain region. In some embodiments, methods are also provided herein for selecting a population of cells predicted to show neurite outgrowth after transplantation into a brain region, for example a population of cells identified as such by any of the methods provided. In some embodiments, methods are also provided herein for transplanting a population of cells predicted to show neurite outgrowth after transplantation into a brain region, for example a population of cells identified or selected as such according to any of the methods provided.

[0088] In some embodiments, the population is an in vitro population of cells.

[0089] In some embodiments, the method includes calculating a first similarity score and a second similarity score using the gene expression levels. In some embodiments, the classification is based on one or both of the first and second similarity scores. In some embodiments, the classification is based on one of the first similarity score and the second similarity score. In some embodiments, the classification is based on the lower of the first similarity score and the second similarity score. In some embodiments, the classification is based on the higher of the first similarity score and the second similarity score. In some embodiments, the classification is based on the first similarity score. In some embodiments, the classification is based on the second similarity score. In some embodiments, the first similarity score indicates whether the differentiation state of the population of cells is more similar to the first differentiation state or the second differentiation state. In some embodiments, the second similarity score indicates whether the differentiation state of the population of cells is more similar to the second differentiation state or the third differentiation state.

[0090] In some embodiments, the method includes calculating a first similarity score and a second similarity score using gene expression levels. In some embodiments, the classification is based on the first and second similarity scores. In some embodiments, the first similarity score indicates whether the differentiation state of the population of cells is more similar to the first differentiation state or the second differentiation state. In some embodiments, the second similarity score indicates whether the differentiation state of the population of cells is more similar to the second differentiation state or the third differentiation state.

[0091] The first, second and third differentiation states can be in the same or different stem cell differentiation pathways. In some embodiments, the first, second and third differentiation states are all in the same stem cell differentiation pathway. In some embodiments, the second differentiation state is an intermediate differentiation state for the first and third differentiation pathways. For example, in some embodiments, the first differentiation state is earlier in the stem cell differentiation pathway than the second differentiation state, and the second differentiation state is earlier in the stem cell differentiation pathway than the third differentiation state. In other embodiments, the second and third differentiation states are in different stem cell differentiation pathways, and the first differentiation state is a differentiation state of a cell that can differentiate into either the second or third differentiation state.

[0092] The provided methods allow for the determination of cell identity, e.g., cell differentiation state, when a single or a few characteristics or features, such as gene expression markers or functional properties, are not available (e.g., unknown) or cannot be practically used to determine cell identity, e.g., cell differentiation state. In some embodiments, a particular cell population differentiated from pluripotent stem cells, including determined dopaminergic cells, may be cells at a differentiation stage where the cells cannot be identified by one or a few characteristics or features. In some embodiments, differentiating cells may enter a differentiation state where no definitive biomarkers can be used to determine the cell identity, e.g., differentiation state. Pluripotent stem cells can be positively identified using definitive biomarkers, e.g., expression levels of certain genes, and differentiated cells can be positively identified based on functional markers, but individual markers for the identification of cells at various transient stages throughout differentiation are unknown. Without such markers, it has previously been difficult to characterize, define, and / or identify pre-differentiated cells with specific cell phenotypes. In some embodiments, the methods provided herein overcome the lack of a single or a few characteristics or features (e.g., biomarkers) by examining a group of genes and their expression levels. Such an approach does not rely on knowledge of individual marker genes, but instead uses a whole transcriptome approach in characterizing and identifying the differentiation state of cells.

[0093] Induced pluripotent stem cells (iPSCs) are believed to be useful as cell therapy at least for their ability to differentiate into specialized cell types. For example, iPSCs, like pluripotent stem cells, can be differentiated into specific cell types that can be used to replace diseased or damaged tissues. In some cases, therapeutic treatments can include administering (e.g., injection) differentiating cells that have not entered a terminally differentiated state to a subject. The inability to determine the identity of the differentiated cells throughout the differentiation process can result in uncertainty about the success of the process. For example, it may be necessary to run the differentiation process to completion to determine whether the differentiation process is successful. Thus, without the ability to determine whether the differentiating cells have progressed through transient stages as necessary, the differentiation process can be time-consuming and inefficient, which may prevent treatment of the subject, for example, if the differentiation process fails. In some embodiments, the methods provided improve the differentiation process by allowing for the determination of cell identity throughout the differentiation state, which can be used to determine whether the cells undergoing the differentiation process have differentiated appropriately and / or according to defined criteria. As an example, if it is determined that the cells are not properly differentiated, the process can be terminated and, optionally, restarted with a different iPSC clone from the subject.

[0094] In certain cell therapies using cells differentiated from pluripotent stem cells, it is advantageous to use cells in the middle stage of the differentiation process.In some embodiments, the method and device are useful for identifying cells in the middle stage that are most effective when used in cell therapy.As an example, neural cells obtained by differentiation from pluripotent stem cells may be more adaptable to engraftment in the brain of the subject undergoing treatment when they are in the middle stage between the early stage (e.g., that of progenitor cells) and the late stage (e.g., that of committed cells).

[0095] In some embodiments, also provided herein is a computing device, including one for performing any of the provided methods. In some embodiments, also provided herein is a composition, article of manufacture, and kit comprising a population of cells, including a population of cells classified by any of the provided methods as having a desired differentiation state. In some embodiments, also provided herein is a method of transplanting a population of cells having a desired differentiation state, e.g., classified according to any of the provided methods, into a subject.

[0096] All publications, including patent documents, scientific literature, and databases, referred to in this application are incorporated herein by reference in their entirety for all purposes to the same extent as if each individual publication was individually incorporated by reference. To the extent that a definition set forth herein is contrary to or otherwise inconsistent with a definition set forth in a patent, patent application, published application, or other publication incorporated herein by reference, the definition set forth herein shall take precedence over the definition incorporated herein by reference.

[0097] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0098] I. Definition Unless otherwise defined, all technical terms, notations, and other technical and scientific terms or terminology used herein are intended to have the same meaning as commonly understood by those of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms having a commonly understood meaning are defined herein for clarity and / or ready reference, and the inclusion of such definitions herein should not necessarily be construed as representing something substantially different to what is generally understood in the art.

[0099] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, "a" or "an" means "at least one" or "one or more." It is understood that the embodiments and variations described herein include "consisting of" and / or "consisting essentially of" embodiments and variations.

[0100] Throughout this disclosure, various aspects of the claimed subject matter are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity, and should not be construed as rigidly limiting the scope of the claimed subject matter. Thus, the description of a range should be construed as specifically disclosing all possible subranges as well as individual numerical values ​​within that range. For example, when a range of values ​​is provided, it is understood that each intervening value between the upper and lower limits of that range, and any other specified or intervening value in the specified range, is encompassed within the claimed subject matter. The upper and lower limits of these smaller ranges may be independently included in the smaller ranges, and are also encompassed by the claimed subject matter, subject to any specifically excluded limits within the specified range. When a specified range includes one or both of the limits, ranges excluding either or both of those inclusive limits are also encompassed by the claimed subject matter. This applies regardless of the breadth of the range.

[0101] As used herein, the term "about" refers to a normal error range for each value that is easily recognized. Reference herein to a value or parameter of "about" includes (and describes) an embodiment that is directed to the value or parameter itself. For example, a description that refers to "about X" includes a description of "X."

[0102] As used herein, the statement that a cell or population of cells "expresses" or "is positive" for a particular marker refers to the detectable presence of the particular marker on or within the cell. When referring to a surface marker, the term refers to the presence of surface expression detected by flow cytometry, for example, by staining with an antibody that specifically binds to the marker and detecting the antibody, and the staining is detectable by flow cytometry at a level substantially greater than that detected by performing the same procedure with an isotype-matched control under otherwise identical conditions, and / or at a level substantially similar to cells known to be positive for the marker, and / or at a level substantially higher than cells known to be negative for the marker. When referring to an intracellular marker, such as a transcription product or translation product, the term refers to the presence of a detectable transcription or translation product, for example, the product is detected at a level substantially greater than that detected by performing the same procedure with a control under otherwise identical conditions, and / or at a level substantially similar to cells known to be positive for the marker, and / or at a level substantially higher than cells known to be negative for the marker.

[0103] As used herein, the statement that a cell or population of cells "does not express" or "negative" for a particular marker refers to the substantial absence of detectable presence of the particular marker on or within the cell. When referring to a surface marker, the term refers to the absence of surface expression detected by flow cytometry, for example, by staining an antibody that specifically binds to the marker and detecting the antibody, where the staining is not detected by flow cytometry at a level substantially greater than that detected by performing the same procedure with an isotype-matched control under otherwise identical conditions, and / or at a level substantially lower than that of cells known to be positive for the marker, and / or at a level substantially similar to that of cells known to be negative for the marker. When referring to an intracellular marker, such as a transcription product or translation product, the term refers to the absence of detectable transcription or translation product, for example, where the product is not detected at a level substantially greater than that detected by performing the same procedure with a control under otherwise identical conditions, and / or at a level substantially lower than that of cells known to be positive for the marker, and / or at a level substantially similar to that of cells known to be negative for the marker.

[0104] The term "expression" or "expressed" as used herein with respect to a gene refers to the transcription and / or translation product of that gene. The expression level of a DNA molecule in a cell can be determined based on the amount of corresponding mRNA present in the cell or the amount of protein encoded by the DNA produced by the cell (Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88).

[0105] As used herein, the term "stem cell" refers to a cell characterized by the ability to self-renew by mitotic cell division and the ability to differentiate into any of several cell types. Among mammalian stem cells, embryonic and somatic stem cells can be distinguished. Embryonic stem cells are present in blastocysts and give rise to embryonic tissues, whereas somatic stem cells are present in adult tissues for the purpose of tissue regeneration and repair.

[0106] "Self-renewal" refers to the ability of a cell to divide and generate at least one daughter cell that has the self-renewal properties of the parent cell. The second daughter cell can commit to a specific differentiation pathway. For example, a self-renewing hematopoietic stem cell can divide and form one daughter stem cell and another daughter cell that is committed to differentiation in the myeloid or lymphatic pathway.

[0107] As used herein, the term "progenitor cell" refers to a cell that has the potential to differentiate into any of multiple cell types, but has lost the ability to self-renew compared to stem cells. For example, a progenitor cell upon cell division can produce two daughter cells that exhibit a more differentiated (e.g., restricted) phenotype.

[0108] As used herein, the term "non-self-renewing cell" refers to a cell that undergoes cell division to produce daughter cells, neither of which has the differentiation potential of the parent cell type, e.g., to produce differentiated daughter cells.

[0109] As used herein, the term "adult stem cell" refers to undifferentiated cells found in an individual after embryonic development. Adult stem cells proliferate by cell division to replenish dying cells and regenerate damaged tissues. Adult stem cells have the ability to divide and create other cells similar to themselves or create more differentiated cells. Although adult stem cells are associated with the expression of pluripotency markers such as Rex1, Nanog, Oct4, or Sox2, they do not have the ability to differentiate into all three germ layers of cell types of pluripotent stem cells.

[0110] As used herein, the term "pluripotent" or "pluripotency" refers to cells that have the capacity, under appropriate conditions, to give rise to progeny that can undergo differentiation into cell types that collectively exhibit characteristics associated with cell lineages from the three germ layers (endoderm, mesoderm, and ectoderm). Pluripotent stem cells can contribute to tissues of prenatal, postnatal, or adult organisms.

[0111] As used herein, the term "pluripotent stem cell characteristics" refers to cellular characteristics that distinguish pluripotent stem cells from other cells. The expression or non-expression of a particular combination of molecular markers is an example of a pluripotent stem cell characteristic. More specifically, human pluripotent stem cells may express at least some, and optionally all, of the markers from the following non-limiting list: SSEA-3, SSEA-4, TRA-1-60, TRA-1-81, TRA-2-49 / 6E, ALP, Sox2, E-cadherin, UTF-1, Oct4, Lin28, Rex1, and Nanog. Cell morphology associated with pluripotent stem cells is also a pluripotent stem cell characteristic.

[0112] As used herein, the terms "induced pluripotent stem cell", "iPS" and "iPSC" refer to pluripotent stem cells artificially obtained (e.g., through artificial manipulation) from non-pluripotent cells. "Non-pluripotent cells" may be cells that have less ability to self-renew and differentiate than pluripotent stem cells. The less potent cells may be adult stem cells, tissue-specific progenitor cells, or primary or secondary cells.

[0113] The term "specification" or "designated" as provided herein refers to a cell or tissue fate narrowed to a limited number of specific cell types. A specific cell can still change its specific fate until it reaches a committed state. A specific cell can differentiate autonomously (e.g., alone) when placed in a developmental pathway-neutral environment, such as a petri dish or test tube. At the specification stage, cell commitment can still be altered. When a specific cell is transplanted into a population of different specific cells, the fate of the transplant can be altered by interactions with its new neighboring cells.

[0114] As used herein, "committed state" refers to a cell that has only one cell type that it can differentiate into. For example, a committed dopaminergic cell may not yet be a dopaminergic neuron itself, and may or may not express definitive markers of a dopaminergic neuron, but cannot become another type of neuron. Committed cells can also differentiate autonomously when placed in an area of ​​the embryo that is unrelated to the cell. For example, an unrelated area of ​​a committed dopaminergic cell is any organ or tissue other than the brain. Committed cells can also differentiate autonomously when placed in a cluster of different designated cells in a petri dish.

[0115] As used herein, the terms "differentiated" or "committed" refer to one or more cells that have acquired a cell-type specific function.

[0116] "Neuronal progenitor cells" are cells that tend to differentiate into neural or glial cells and do not have the pluripotent potential of stem cells. Neuronal precursors are cells that are committed to neuronal or glial lineages and are characterized by expressing one or more marker genes specific to neuronal or glial lineages. The terms "neural" and "neuron" are used according to their common meaning in the art and can be used interchangeably throughout this specification.

[0117] As used herein, "dopaminergic cells" or "differentiated dopaminergic cells" refer to cells that can synthesize the neurotransmitter dopamine. In some embodiments, the dopaminergic cells are A9 dopaminergic cells. The term "A9 dopaminergic cells" refers to the most dense group of dopaminergic cells in the human brain, located in the pars compacta of the substantia nigra of the midbrain of healthy adult humans.

[0118] The term "committed dopaminergic cells" as used herein refers to cells that differentiate into dopaminergic neurons and cannot differentiate into non-dopaminergic cells. A "committed dopaminergic cell" is a cell that can differentiate into a dopaminergic neuron regardless of its environment. A committed dopaminergic cell may express Foxa2 or Nurrl. A committed dopaminergic cell may not express serotonin.

[0119] As used herein, the term "reprogramming" refers to the process of dedifferentiating a non-pluripotent cell into a cell that exhibits characteristics of a pluripotent stem cell.

[0120] As used herein, the term "cell culture" may refer to an in vitro population of cells present outside of an organism. Cell cultures may be established from primary cells isolated from a cell bank or an animal, or from secondary cells derived from one of these sources, and immortalized for long-term in vitro culture.

[0121] As used herein, the terms "culture," "cultivate," "grow," "grow," "maintain," "maintaining," "proliferation," "proliferating," and the like, when referring to cell culture itself or the process of culturing, may be used interchangeably to mean that cells are maintained outside the body (e.g., ex vivo) under conditions suitable for survival. Cultured cells may be allowed to survive, and as a result of the culturing, the cells may grow, differentiate, or divide.

[0122] As used herein, a composition refers to any mixture of two or more products, substances, or compounds, including cells, which may be a solution, suspension, liquid, powder, paste, aqueous, non-aqueous, or any combination thereof.

[0123] The term "pharmaceutical composition" refers to a composition suitable for pharmaceutical use, such as in a mammalian subject (e.g., a human). A pharmaceutical composition typically comprises an effective amount of an active agent (e.g., cells) and a carrier, excipient, or diluent. The carrier, excipient, or diluent is typically a pharma- ceutically acceptable carrier, excipient, or diluent, respectively.

[0124] "Pharmaceutically acceptable carrier" refers to an ingredient in a pharmaceutical formulation, other than an active ingredient, that is non-toxic to a subject. Pharmaceutically acceptable carriers include, but are not limited to, buffers, excipients, stabilizers, or preservatives.

[0125] The term "package insert" is used to refer to instructions typically included in the commercial packaging of a therapeutic product that contain information regarding the indications, uses, dosage, administration, concomitant therapy, contraindications, and / or warnings regarding the use of such therapeutic product.

[0126] As used herein, a "subject" is a mammal, such as a human or other animal, typically a human.

[0127] II. Methods for Classifying or Identifying Cells In some embodiments, provided herein are methods for classifying the differentiation state of an in vitro population of cells. In some embodiments, the methods provided are for identifying an in vitro population of cells having a desired differentiation state. In some embodiments, the methods provided are for selecting an in vitro population of cells having a desired differentiation state.

[0128] In some embodiments, methods are also provided herein for predicting whether an in vitro population of cells will exhibit neurite outgrowth after transplantation into a brain region. In some embodiments, the methods provided are for identifying an in vitro population of cells that exhibit neurite outgrowth after transplantation into a brain region. In some embodiments, the methods provided are for selecting an in vitro population of cells that exhibit neurite outgrowth after transplantation into a brain region.

[0129] In some embodiments, the provided method is a computer-implemented method. In some embodiments, the provided method is performed by a computing device. In some embodiments, the provided method is performed by any of the provided computing devices, such as any of those described in Section III.

[0130] In some embodiments, the methods provided provide information regarding, among other things, whether an in vitro population of cells (e.g., a population of neural cells) includes cells that are committed to differentiate into a particular functional cell type (e.g., including committed dopaminergic cells), or whether the in vitro population of cells includes cells from an earlier stage (e.g., pluripotent stem cells, neural progenitor cells), a later stage (e.g., committed dopaminergic cells), or other differentiated cell type. In some embodiments, the methods provided predict whether an in vitro population of cells will differentiate into a particular cell type (e.g., into dopaminergic cells). In some embodiments, cells identified by the methods provided are committed to differentiate into a particular functional cell type (e.g., into dopaminergic cells). Whether a cell is committed to differentiate into a particular functional cell type (e.g., whether the cell is a committed dopaminergic cell) can be further verified in vitro or in vivo by fully differentiating the cells. The methods provided also encompass identifying cells that are pluripotent stem cells, particular cells, differentiated neuronal types other than committed dopaminergic cells, or other differentiated cell types.

[0131] In some embodiments, the provided methods include receiving as input a test dataset comprising one or more test cell features. Exemplary test cells are described in Section II-C. In some embodiments, the provided methods include receiving as input a test dataset comprising expression levels for genes expressed in one or more test cells. The gene expression levels can be assessed using any of the methods described in Section II-D.

[0132] In some embodiments, the methods provided include calculating a first similarity score and a second similarity score. In some embodiments, the first similarity score indicates whether the differentiation state of the test cell is more similar to the first differentiation state or more similar to the second differentiation state. In some embodiments, the second similarity score indicates whether the differentiation state of the test cell is more similar to the second differentiation state or more similar to the third differentiation state. Exemplary methods for calculating the first and second similarity scores are described in Section II-A. Exemplary first, second, and third differentiation states are described in Section II-C.

[0133] In some embodiments, the differentiation state of the one or more test cells is classified based on one or both of the first and second similarity scores. In some embodiments, methods provided include classifying the differentiation state of the one or more test cells based on one or both of the first and second similarity scores.

[0134] In some embodiments, the differentiation state of the one or more test cells is classified based on the first and second similarity scores. In some embodiments, methods provided include classifying the differentiation state of the one or more test cells based on the first and second similarity scores.

[0135] In some embodiments, the differentiation state of the one or more test cells is classified as being the second differentiation state if one or both of the first and second similarity scores indicate that the differentiation state of the one or more test cells is more similar to the second differentiation state. In some embodiments, methods provided include classifying the differentiation state of the one or more test cells as being the second differentiation state if one or both of the first and second similarity scores indicate that the differentiation state of the one or more test cells is more similar to the second differentiation state.

[0136] In some embodiments, the classification is based on one of the first similarity score and the second similarity score. In some embodiments, the classification is based on the lower of the first similarity score and the second similarity score. In some embodiments, the classification is based on the higher of the first similarity score and the second similarity score. In some embodiments, the classification is based on the first similarity score. In some embodiments, the classification is based on the second similarity score.

[0137] In some embodiments, the differentiation state of the one or more test cells is classified as being the second differentiation state if the first and second similarity scores indicate that the differentiation state of the one or more test cells is more similar to the second differentiation state. In some embodiments, a method provided includes classifying the differentiation state of the one or more test cells as being the second differentiation state if the first and second similarity scores indicate that the differentiation state of the one or more test cells is more similar to the second differentiation state.

[0138] In some embodiments, the differentiation state of the one or more test cells is classified as being a second differentiation state and the in vitro population of cells is identified as having the desired differentiation state. In some embodiments, the differentiation state of the one or more test cells is classified as being a second differentiation state and the methods provided include identifying the in vitro population of cells as having the desired differentiation state.

[0139] In some embodiments, the methods provided include selecting an in vitro population of cells having a desired differentiation state for use in treating a disease or condition in a subject. In some embodiments, the in vitro population of cells having a desired differentiation state is selected for transplantation into a subject. In some embodiments, the methods provided include transplanting the in vitro population of cells having a desired differentiation state into a subject, for example, according to any of the methods described in Section VI.

[0140] In some embodiments, the provided method also includes calculating a correlation score using the one or more test cell features and the control data set. In some embodiments, classifying the differentiation state of the one or more test cells is based on the correlation score and one or both of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and one of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the lower of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the higher of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the first similarity score. In some embodiments, the classification is based on the correlation score and the second similarity score. Exemplary methods for calculating the correlation score are described in Section II-B.

[0141] In some embodiments, the provided method also includes calculating a correlation score using the one or more test cell features and the control dataset. In some embodiments, classifying the differentiation state of the one or more test cells is based on the first similarity score, the second similarity score, and the correlation score. Exemplary methods for calculating the correlation score are described in Section II-B.

[0142] In some embodiments, the provided method includes the use of a trained machine learning model. In some embodiments, the first and second similarity scores are determined using a first and a second machine learning model, respectively. In some embodiments, the first and the second similarity scores are determined based on one or more outputs of the first and the second machine learning model, respectively. Exemplary model types of the first and the second machine learning models are described in Section II-A-4. In some embodiments, each of the first and the second machine learning models is trained using characteristics, e.g., gene expression levels, of a plurality of reference cell populations. Exemplary reference cell populations are described in Section II-C.

[0143] Also provided herein, in some embodiments, are methods for training machine learning models that can be used to classify the differentiation state of in vitro populations of cells.

[0144] In some embodiments, the method includes training a first and a second machine learning model. In some embodiments, the first and the second machine learning models are trained using gene expression levels. Exemplary genes included and / or selected for model training are described in Section II-A-3. Exemplary model types for the first and the second machine learning models are described in Section II-A-4. The gene expression levels can be evaluated according to any of the methods described in Section II-D.

[0145] In some embodiments, the provided method includes obtaining gene expression levels for one or more genes that are differentially expressed between cells in a first differentiation state and cells in a second differentiation state for a plurality of reference cell populations. In some embodiments, the method includes selecting genes that are differentially expressed between cells in the first differentiation state and cells in the second differentiation state. In some embodiments, the method includes obtaining expression levels of one or more of the selected genes for the plurality of reference cell populations. In some embodiments, the gene expression levels of the plurality of reference cell populations are applied as inputs to train a first machine learning model. In some embodiments, one or more outputs of the trained first machine learning model can be used to classify the differentiation state of one or more test cells. In some embodiments, one or more outputs of the trained first machine learning model can be used to calculate a first similarity score that indicates whether the differentiation state of the test cell is more similar to the first differentiation state or the second differentiation state.

[0146] In some embodiments, the provided method includes obtaining gene expression levels for one or more genes that are differentially expressed between cells in a second differentiation state and cells in a third differentiation state for a plurality of reference cell populations. In some embodiments, the method includes selecting genes that are differentially expressed between cells in a second differentiation state and cells in a third differentiation state. In some embodiments, the method includes obtaining expression levels of one or more of the selected genes for a plurality of reference cell populations. In some embodiments, the gene expression levels of the plurality of reference cell populations are applied as inputs to train a second machine learning model. In some embodiments, one or more outputs of the trained second machine learning model can be used to classify the differentiation state of one or more test cells. In some embodiments, one or more outputs of the trained second machine learning model can be used to calculate a second similarity score indicating whether the differentiation state of the test cell is more similar to the second differentiation state or the third differentiation state. Exemplary reference cell populations and the first, second and third differentiation states are described in Section II-C.

[0147] In some embodiments, the method further comprises obtaining gene expression levels for one or more genes expressed in cells in a control differentiation state for a plurality of reference cell populations. The control differentiation state may be the same or different from one of the first, second, or third differentiation states. Exemplary control differentiation states are described in Section II-C. In some embodiments, the method further comprises applying the gene expression levels of the one or more genes as inputs to train a control machine learning model. In some embodiments, one or more outputs of the trained control machine learning model can be used to classify the differentiation state of one or more test cells. In some embodiments, one or more outputs of the trained control machine learning model can be used to determine whether the differentiation state of the test cells is similar to the control differentiation state.

[0148] A. Similarity Score In some embodiments, the provided method includes calculating a first similarity score and a second similarity score. In some embodiments, the differentiation state of the one or more test cells is classified based on one or both of the first and second similarity scores. In some embodiments, the differentiation state of the one or more test cells is classified as being the second differentiation state if one or both of the first and second similarity scores indicate that the differentiation state of the one or more test cells is similar to the second differentiation state. In some embodiments, the provided method includes classifying the differentiation state of the one or more test cells as being the second differentiation state if one or both of the first and second similarity scores indicate that the differentiation state of the one or more test cells is similar to the second differentiation state.

[0149] In some embodiments, the classification is based on one of the first similarity score and the second similarity score. In some embodiments, the classification is based on the lower of the first similarity score and the second similarity score. In some embodiments, the classification is based on the higher of the first similarity score and the second similarity score. In some embodiments, the classification is based on the first similarity score. In some embodiments, the classification is based on the second similarity score.

[0150] In some embodiments, the provided method includes calculating a first similarity score and a second similarity score. In some embodiments, the differentiation state of the one or more test cells is classified based on the first and second similarity scores. In some embodiments, the differentiation state of the one or more test cells is classified as being the second differentiation state if the first and second similarity scores indicate that the differentiation state of the one or more test cells is similar to the second differentiation state. In some embodiments, the provided method includes classifying the differentiation state of the one or more test cells as being the second differentiation state if the first and second similarity scores indicate that the differentiation state of the one or more test cells is similar to the second differentiation state.

[0151] In some embodiments, the first and second similarity scores are calculated using gene expression levels of the test dataset. In some embodiments, the gene expression levels of the test dataset are compared with gene expression levels contained in the first and second reference datasets. In some embodiments, the first and second similarity scores are calculated using representations of gene expression levels contained in the first and second reference datasets, respectively. In some embodiments, the representations are obtained by machine learning. In some embodiments, the first and second reference datasets include first and second machine learning models, respectively, and the first and second similarity scores are calculated by applying the gene expression levels of the test dataset as inputs to the first and second machine learning models, respectively, and are based on one or more outputs of the first and second machine learning models, respectively.

[0152] In some embodiments, one or both of the first and second similarity scores is a binary output (e.g., 0 or 1, or -1 or 1) indicating whether the differentiation state of the one or more test cells is the second differentiation state. In some embodiments, one or both of the first and second similarity scores is a non-binary output.

[0153] Depending on the type of non-binary output, the first and second similarity scores above or below a predefined threshold level may indicate that the differentiation state of one or more test cells is a second differentiation state. The predefined threshold level may be the same or different for the first and second similarity scores. Any suitable method for setting the predefined threshold level may be used. For example, in some embodiments, the predefined threshold level of the first similarity score is set based on a plurality of first similarity scores calculated using gene expression levels of a plurality of reference cell populations used to obtain a representation of gene expression levels of the first reference dataset, e.g., used to train a first machine learning model of the first reference dataset. In some embodiments, the predetermined threshold level is set to a value that separates the first similarity score of the reference cell population comprising cells in a first differentiation state from the first similarity score of the reference cell population comprising cells in a second differentiation state by an accuracy metric, e.g., precision, recall, accuracy, or F1 score, of at least 0.5, 0.6, 0.7, 0.8, 0.85, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99. Similarly, in some embodiments, the predetermined threshold level of the second similarity score is set based on a plurality of second similarity scores calculated using gene expression levels of a plurality of reference cell populations used to obtain a representation of gene expression levels of the second reference dataset, e.g., used to train a second machine learning model of the second reference dataset. In some embodiments, the predetermined threshold level is set to a value that separates the second similarity score of a reference cell population comprising cells in a second differentiation state from the second similarity score of a reference cell population comprising cells in a third differentiation state by an accuracy metric, e.g., precision, recall, accuracy, or F1 score, of at least 0.5, 0.6, 0.7, 0.8, 0.85, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99.

[0154] In some embodiments, one or both of the first and second similarity scores are probabilities of the differentiation state of the one or more test cells being the second differentiation state. In some embodiments, a probability exceeding a predetermined probability threshold level indicates that the differentiation state of the one or more test cells is the second differentiation state. The predetermined probability threshold level may be the same or different for the first and second similarity scores. In some embodiments, the predetermined probability threshold level is 0.5, about 0.5, greater than 0.5, or greater than about 0.5. In some embodiments, the predetermined probability threshold level is 0.55, about 0.55, greater than 0.55, or greater than about 0.55. In some embodiments, the predetermined probability threshold level is 0.6, about 0.6, greater than 0.6, or greater than about 0.6. In some embodiments, the predetermined probability threshold level is 0.65, about 0.65, greater than 0.65, or greater than about 0.65. In some embodiments, the predetermined probability threshold level is 0.7, about 0.7, greater than 0.7, or greater than about 0.7. In some embodiments, the predetermined probability threshold level is 0.75, about 0.75, greater than 0.75, or greater than about 0.75. In some embodiments, the predetermined probability threshold level is 0.8, about 0.8, greater than 0.8, or greater than about 0.8. In some embodiments, the predetermined probability threshold level is 0.85, about 0.85, greater than 0.85, or greater than about 0.85. In some embodiments, the predetermined probability threshold level is 0.9, about 0.9, greater than 0.9, or greater than about 0.9. In some embodiments, the predetermined probability threshold level is 0.91, about 0.91, greater than 0.91, or greater than about 0.91. In some embodiments, the predetermined probability threshold level is 0.92, about 0.92, greater than 0.92, or greater than about 0.92. In some embodiments, the predetermined probability threshold level is 0.93, about 0.93, greater than 0.93, or greater than about 0.93.In some embodiments, the predetermined probability threshold level is 0.94, about 0.94, greater than 0.94, or greater than about 0.94. In some embodiments, the predetermined probability threshold level is 0.95, about 0.95, greater than 0.95, or greater than about 0.95. In some embodiments, the predetermined probability threshold level is 0.96, about 0.96, greater than 0.96, or greater than about 0.96. In some embodiments, the predetermined probability threshold level is 0.97, about 0.97, greater than 0.97, or greater than about 0.97. In some embodiments, the predetermined probability threshold level is 0.98, about 0.98, greater than 0.98, or greater than about 0.98. In some embodiments, the predetermined probability threshold level is 0.99, about 0.99, greater than 0.99, or greater than about 0.99.

[0155] In some embodiments, one or both of the first and second similarity scores are compared to a predetermined threshold level. In some embodiments, one of the first and second similarity scores is compared to a predetermined threshold level. In some embodiments, the similarity score that is compared to the predetermined threshold level is based on which similarity score is closest to the predetermined threshold level. In some aspects, the similarity score that is compared to the predetermined threshold level is selected such that if the selected similarity score indicates that the differentiation state of the test cell is more similar to the second differentiation state, the other similarity scores are also expected to indicate that the differentiation state of the test cell is more similar to the second differentiation state.

[0156] 1. First Similarity Score In some aspects, the provided method includes calculating a first similarity score indicating whether the differentiation state of the test cell is more similar to the first differentiation state or the second differentiation state. In some embodiments, the first similarity score is calculated using a first reference data set comprising gene expression levels for one or more genes that are differentially expressed between cells in the first differentiation state and cells in the second differentiation state. In some embodiments, the first similarity score is calculated using a first reference data set comprising a representation of gene expression levels for one or more genes that are differentially expressed between cells in the first differentiation state and cells in the second differentiation state.

[0157] In some embodiments, the first reference data set comprises gene expression levels for one or more genes that are differentially expressed between cells in a first differentiation state and cells in a second differentiation state.In some embodiments, the gene expression levels are normalized gene expression levels.In some embodiments, the first similarity score is obtained by comparing the gene expression levels of the first reference data set with the gene expression levels of the test data set.

[0158] In some embodiments, the first reference data set comprises a representation of gene expression levels for one or more genes that are differentially expressed between cells in a first differentiation state and cells in a second differentiation state.In some embodiments, the representation of gene expression levels is obtained by machine learning.In some embodiments, the representation of gene expression levels is obtained by training a first machine learning model using the gene expression levels of one or more genes.

[0159] In some embodiments, the first similarity score is calculated using a first reference dataset comprising a first machine learning model trained using gene expression levels of one or more genes differentially expressed between cells in a first differentiation state and cells in a second differentiation state. In some embodiments, the first similarity score is calculated by providing gene expression levels of a test dataset as input to the first machine learning model or a process comprising the first machine learning model. For example, in some embodiments, the gene expression levels of the test dataset are normalized or transformed before being provided as input to the first machine learning model. In some embodiments, the first similarity score is an output of the first machine learning model. In some embodiments, the first similarity score is calculated using one or more outputs of the first machine learning model.

[0160] In some embodiments, the representation of the gene expression levels of the first reference dataset is obtained using the gene expression levels of a plurality of reference cell populations. In some embodiments, the first machine learning model is trained using the gene expression levels of a plurality of reference cell populations. In some embodiments, the plurality of reference cell populations includes at least one reference cell population comprising cells in a first differentiation state and at least one reference cell population comprising cells in a second differentiation state. In some embodiments, the plurality of reference cell populations includes a plurality of reference cell populations comprising cells in a first differentiation state and a plurality of reference cell populations comprising cells in a second differentiation state. In some embodiments, the reference cell populations, for example those comprising cells in a first or second differentiation state, are any of those described in Section II-C.

[0161] 2. Second Similarity Score In some aspects, the provided method includes calculating a second similarity score indicating whether the differentiation state of the test cell is more similar to the second differentiation state or the third differentiation state. In some embodiments, the second similarity score is calculated using a second reference data set comprising gene expression levels for one or more genes that are differentially expressed between cells in the second differentiation state and cells in the third differentiation state. In some embodiments, the second similarity score is calculated using a second reference data set comprising a representation of gene expression levels for one or more genes that are differentially expressed between cells in the second differentiation state and cells in the third differentiation state.

[0162] In some embodiments, the second reference data set comprises the gene expression levels of one or more genes that are differentially expressed between cells in the second differentiation state and cells in the third differentiation state.In some embodiments, the gene expression levels are normalized gene expression levels.In some embodiments, the second similarity score is obtained by comparing the gene expression levels of the second reference data set with the gene expression levels of the test data set.

[0163] In some embodiments, the second reference data set comprises a representation of gene expression levels for one or more genes that are differentially expressed between cells in the second differentiation state and cells in the third differentiation state.In some embodiments, the gene expression level representation is obtained by machine learning.In some embodiments, the gene expression level representation is obtained by training a second machine learning model using the gene expression levels of one or more genes.

[0164] In some embodiments, the second similarity score is calculated using a second reference dataset comprising a second machine learning model trained using gene expression levels of one or more genes differentially expressed between cells in the second differentiation state and cells in the third differentiation state. In some embodiments, the second similarity score is calculated by providing the gene expression levels of the test dataset as input to the second machine learning model or a process comprising the second machine learning model. For example, in some embodiments, the gene expression levels of the test dataset are normalized or transformed before being provided as input to the second machine learning model. In some embodiments, the second similarity score is an output of the second machine learning model. In some embodiments, the second similarity score is calculated using one or more outputs of the second machine learning model.

[0165] In some embodiments, the representation of the gene expression levels of the second reference dataset is obtained using the gene expression levels of the multiple reference cell populations. In some embodiments, the second machine learning model is trained using the gene expression levels of the multiple reference cell populations. In some embodiments, the multiple reference cell populations include at least one reference cell population comprising cells in a second differentiation state and at least one reference cell population comprising cells in a third differentiation state. In some embodiments, the multiple reference cell populations include a multiple reference cell population comprising cells in a second differentiation state and a multiple reference cell population comprising cells in a third differentiation state. In some embodiments, the reference cell population, for example comprising cells in a second or third differentiation state, is any of those described in Section II-C.

[0166] 3. Exemplary Genes The one or more genes of the first and / or second reference data set, for example, the one or more genes used to train the first and / or second machine learning model, including any of the provided methods including training the first and / or second machine learning model, can be selected based on any suitable criteria. The criteria can include that the one or more genes are expressed above a minimum threshold level in a relevant cell population, for example, in a reference cell population including cells of the first, second and / or third differentiation state, or in any combination of these reference cell populations. The criteria can also include that the one or more genes are differentially expressed between relevant cell populations (for example, between cells of the first and second differentiation state, or between cells of the second and third differentiation state), for example, differentially expressed by a threshold fold change level, have a certain statistical significance, or such that each of the one or more genes individually predicts the differentiation state.

[0167] In some embodiments, the one or more genes of the first reference dataset or the one or more genes selected to train the first machine learning model include genes whose expression levels increase from the first differentiation state to the second differentiation state. In some embodiments, the one or more genes of the first reference dataset or the one or more genes selected to train the first machine learning model include genes whose expression levels decrease from the first differentiation state to the second differentiation state. In some embodiments, the one or more genes of the first reference dataset or the one or more genes selected to train the first machine learning model include genes whose expression levels increase from the first differentiation state to the second differentiation state and genes whose expression levels decrease from the first differentiation state to the second differentiation state.

[0168] In some embodiments, the one or more genes of the second reference dataset or the one or more genes selected to train the second machine learning model include genes whose expression levels increase from the second differentiation state to the third differentiation state. In some embodiments, the one or more genes of the second reference dataset or the one or more genes selected to train the second machine learning model include genes whose expression levels decrease from the second differentiation state to the third differentiation state. In some embodiments, the one or more genes of the second reference dataset or the one or more genes selected to train the second machine learning model include genes whose expression levels increase from the second differentiation state to the third differentiation state and genes whose expression levels decrease from the second differentiation state to the third differentiation state.

[0169] In some embodiments, one or more genes of the first reference dataset or one or more genes selected to train the first machine learning model are the same as one or more genes of the second reference dataset or one or more genes selected to train the second machine learning model. In some embodiments, one or more genes of the first reference dataset or one or more genes selected to train the first machine learning model are different from one or more genes of the second reference dataset or one or more genes selected to train the second machine learning model. In some embodiments, some of the one or more genes of the first reference dataset or the genes selected to train the first machine learning model are included in the one or more genes of the second reference dataset or one or more genes selected to train the second machine learning model. In some embodiments, none of the one or more genes of the first reference dataset or the genes selected to train the first machine learning model are included in the one or more genes of the second reference dataset or one or more genes selected to train the second machine learning model.

[0170] In some embodiments, one or more genes of the first and / or second reference dataset or one or more genes selected for training the first and / or second machine learning model comprise a plurality of genes. In some embodiments, the plurality of genes comprises 2 genes, comprises about 2 genes, comprises more than 2 genes, or comprises more than about 2 genes. In some embodiments, the plurality of genes comprises 3 genes, comprises about 3 genes, comprises more than 3 genes, or comprises more than about 3 genes. In some embodiments, the plurality of genes comprises 4 genes, comprises about 4 genes, comprises more than 4 genes, or comprises more than about 4 genes. In some embodiments, the plurality of genes comprises 5 genes, comprises about 5 genes, comprises more than 5 genes, or comprises more than about 5 genes. In some embodiments, the plurality of genes comprises 6 genes, comprises about 6 genes, comprises more than 6 genes, or comprises more than about 6 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 7 genes, or comprises more than about 7 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 8 genes, or comprises more than 8 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 9 genes, or comprises more than 9 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 10 genes, or comprises more than 10 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 12 genes, or comprises more than 12 genes, or comprises more than 12 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 14 genes, or comprises more than 14 genes, or comprises more than 14 genes. In some embodiments, the plurality of genes comprises 16 genes, comprises about 16 genes, comprises more than 16 genes, or comprises more than about 16 genes.In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 18 genes, or comprises more than about 18 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 20 genes, or comprises more than 20 genes, or comprises more than about 20 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 25 genes, or comprises more than 25 genes, or comprises more than 25 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 30 genes, or comprises more than 30 genes, or comprises more than 30 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 35 genes, or comprises more than 35 genes, or comprises more than 35 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 40 genes, or comprises more than 40 genes, or comprises more than 40 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 45 genes, or comprises more than about 45 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 55 genes, or comprises more than 55 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 60 genes, or comprises more than 60 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 62 genes, or comprises more than 62 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 64 genes, or comprises more than 64 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 64 genes, or comprises more than 64 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 66 genes, or comprises more than 66 genes.In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 68 genes, or comprises more than about 68 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 70 genes, or comprises more than 70 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 80 genes, or comprises more than 80 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 90 genes, or comprises more than 90 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 100 genes, or comprises more than 100 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 110 genes, or comprises more than 110 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 120 genes, or comprises more than about 120 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 130 genes, or comprises more than 130 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 140 genes, or comprises more than 140 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 150 genes, or comprises more than 150 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 150 genes, or comprises more than 150 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 160 genes, or comprises more than 160 genes. In some embodiments the plurality of genes comprises 170 genes, comprises about 170 genes, comprises more than 170 genes, or comprises more than about 170 genes.In some embodiments, the plurality of genes comprises, comprises about, comprises more than 180 genes, or comprises more than about 180 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 190 genes, or comprises more than 190 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 200 genes, or comprises more than 200 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 250 genes, or comprises more than 250 genes, or comprises more than 250 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises more than 300 genes, or comprises more than 300 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 350 genes, or comprises more than about 350 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 400 genes, or comprises more than 400 genes, or comprises more than about 400 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 450 genes, or comprises more than 450 genes, or comprises more than about 450 genes. In some embodiments, the plurality of genes comprises, comprises about, comprises, or comprises more than 500 genes, or comprises more than 500 genes, or comprises more than about 500 genes.

[0171] In some embodiments, the one or more genes of the first and / or second reference datasets include genes with a minimum expression level in cells of the first, second and / or third differentiation state. In some embodiments, the one or more genes selected to train the first and / or second machine learning model are selected to have a minimum expression level in cells of the first, second and / or third differentiation state. In some embodiments, the one or more genes include genes with a read count, e.g., a count per million mapped reads (CPM) or log2CPM of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18 or 20.

[0172] In some embodiments, one or more genes of the first and / or second reference dataset are differentially expressed genes with a particular statistical significance. In some embodiments, one or more genes selected for training the first and / or second machine learning model are selected to be differentially expressed with a particular statistical significance. In some embodiments, one or more genes are differentially expressed genes with an associated p-value of less than 0.05. In some embodiments, one or more genes are differentially expressed genes with an associated p-value of less than 0.01. In some embodiments, one or more genes are differentially expressed genes with an associated p-value of less than 0.001. In some embodiments, one or more genes are differentially expressed genes with an associated p-value of less than 0.0001. In some embodiments, the p-value is an adjusted p-value. In some embodiments, the p-value is adjusted for multiple comparisons. Any suitable multiple comparison procedure can be used. In some embodiments, the p-value is a Bonferroni corrected p-value. In some embodiments, the p-value is a false discovery rate (FDR) adjusted p-value. In some embodiments, the p-values ​​are Holm-Bonferroni corrected p-values.

[0173] In some embodiments, the one or more genes of the first reference dataset, or the one or more genes selected to train the first machine learning model, are selected from the genes listed in Table E1. In some embodiments, the one or more genes include 10 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes include 20 or more genes selected from the genes selected from the genes listed in Table E1. In some embodiments, the one or more genes include 30 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes include 40 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes include 50 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes include 60 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes include 70 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes include 80 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes comprise 90 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes comprise 100 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes comprise 200 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes comprise 300 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes comprise 400 or more genes selected from the genes listed in Table E1. In some embodiments, the one or more genes comprise 500 or more genes selected from the genes listed in Table E1.

[0174] In some embodiments, the one or more genes of the second reference dataset, or the one or more genes selected to train the second machine learning model, are selected from the genes listed in Table E2. In some embodiments, the one or more genes include 10 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes include 20 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes include 30 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes include 40 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes include 50 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes include 60 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes include 70 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes include 80 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes comprise 90 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes comprise 100 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes comprise 200 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes comprise 300 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes comprise 400 or more genes selected from the genes listed in Table E2. In some embodiments, the one or more genes comprise 500 or more genes selected from the genes listed in Table E2.

[0175] In some embodiments, one or more genes of the first and / or second reference dataset are genes that are differentially expressed by at least a certain amount. In some embodiments, one or more genes selected to train the first and / or second machine learning model are selected to be differentially expressed by at least a particular amount. In some embodiments, one or more genes are genes that show at least a threshold-fold increase or decrease in gene expression levels. In some embodiments, one or more genes are genes that show at least a threshold-fold increase or decrease in gene expression levels and have a particular statistical significance, such as any of the associated p-values ​​described herein. In some embodiments, one or more genes are genes that show at least a 1-fold increase or decrease in gene expression levels. In some embodiments, one or more genes are genes that show at least a 2-fold increase or decrease in gene expression levels. In some embodiments, one or more genes are genes that show at least a 3-fold increase or decrease in gene expression levels. In some embodiments, one or more genes are genes that show at least a 4-fold increase or decrease in gene expression levels. In some embodiments, one or more genes are genes that show at least a 5-fold increase or decrease in gene expression levels. In some embodiments, the one or more genes are genes that exhibit at least a 6-fold increase or decrease in gene expression levels. In some embodiments, the one or more genes are genes that exhibit at least a 7-fold increase or decrease in gene expression levels. In some embodiments, the one or more genes are genes that exhibit at least an 8-fold increase or decrease in gene expression levels. In some embodiments, the one or more genes are genes that exhibit at least a 9-fold increase or decrease in gene expression levels. In some embodiments, the one or more genes are genes that exhibit at least a 10-fold increase or decrease in gene expression levels.

[0176] In some embodiments, the one or more genes of the first and / or second reference dataset are genes that individually predict cells having one or another differentiation state, e.g., a first or second differentiation state for the first reference dataset and a second or third differentiation state for the second reference dataset. In some embodiments, the one or more genes selected to train the first and / or second machine learning model are genes that individually predict cells having one or another differentiation state, e.g., a first or second differentiation state for the first machine learning model and a second or third differentiation state for the second machine learning model. The predictiveness of a gene can be evaluated using any suitable accuracy metric, e.g., precision, recall, accuracy, or F1 score. In some embodiments, the predictiveness of a gene is its accuracy in classifying the differentiation state of a cell based on a threshold expression level of the gene, where one or more cells having an expression level of the gene higher than the threshold are classified as having one differentiation state, and one or more cells having an expression level of the gene lower than the threshold are classified as having another differentiation state. In some embodiments, the accuracy is at least 80%. In some embodiments, the accuracy is at least 82%. In some embodiments, the accuracy is at least 84%. In some embodiments, the accuracy is at least 86%. In some embodiments, the accuracy is at least 88%. In some embodiments, the accuracy is at least 90%. In some embodiments, the accuracy is at least 92%. In some embodiments, the accuracy is at least 94%. In some embodiments, the accuracy is at least 96%. In some embodiments, the accuracy is at least 98%. In some embodiments, the accuracy is 100%.

[0177] 4. Example Machine Learning Model A variety of machine learning models are suitable for use in classifying the differentiation state of cells based on gene expression levels and are within the scope of the present disclosure.In some embodiments, the machine learning models of the first and second reference data sets are the same type of machine learning models, for example, both are logistic regression models.In some embodiments, the machine learning models of the first and second reference data sets are different types of machine learning models, for example, one logistic regression model and one support vector machine classifier.Similarly, the first and second machine learning models trained according to any of the methods provided can be the same or different types of machine learning models.

[0178] Any suitable method for training a machine learning model can be used, including those described in Hastie et al., The Elements of Statistical Learning (2016); and Abu-Mostafa et al., Learning from Data (2012). Exemplary machine learning models are also described in Hastie et al., The Elements of Statistical Learning (2016); and Abu-Mostafa et al., Learning from Data (2012).

[0179] Further exemplary machine learning models are provided in this section. The machine learning models of the first and second reference datasets or the first and second machine learning models trained according to any of the provided methods can be any of the exemplary machine learning models described herein.

[0180] In some embodiments, the machine learning model comprises a supervised machine learning model. In some embodiments, the machine learning model comprises an unsupervised machine learning model. In some embodiments, the machine learning model comprises a semi-supervised machine learning model. In some embodiments, the machine learning model comprises a clustering method.

[0181] In some embodiments the machine learning model comprises a regression model. In some embodiments the machine learning model comprises a classification model. In some embodiments the machine learning model comprises a binary classification model. In some embodiments the machine learning model comprises a multi-class classification model.

[0182] In some embodiments the machine learning model comprises a linear model, hi some embodiments the machine learning model comprises a non-linear model.

[0183] In some embodiments, the machine learning model comprises a logistic regression model. In some embodiments, the machine learning model comprises a linear regression model. In some embodiments, the machine learning model comprises a multiple linear regression model. In some embodiments, the machine learning model comprises a polynomial regression model. In some embodiments, the machine learning model comprises a quantile regression model. In some embodiments, the machine learning model comprises a principal components regression model. In some embodiments, the machine learning model comprises a partial least mean regression model. In some embodiments, the machine learning model comprises a support vector regression model. In some embodiments, the machine learning model comprises an ordinal regression model. In some embodiments, the machine learning model comprises a Poisson regression model. In some embodiments, the machine learning model comprises a negative binomial regression model. In some embodiments, the machine learning model comprises a quasi-Poisson regression model. In some embodiments, the machine learning model comprises a linear discriminant analysis (LDA) model. In some embodiments, the machine learning model comprises a naive Bayes classifier. In some embodiments, the machine learning model comprises a perceptron. In some embodiments, the machine learning model comprises a support vector machine (SVM). In some embodiments, the machine learning model comprises a quadratic classifier. In some embodiments, the machine learning model comprises a decision tree. In some embodiments, the machine learning model comprises a random forest. In some embodiments, the machine learning model comprises a neural network.

[0184] In some embodiments, the machine learning model comprises a connectivity-based clustering method. In some embodiments, the machine learning model comprises hierarchical clustering. In some embodiments, the machine learning model comprises a centroid-based clustering method. In some embodiments, the machine learning model comprises k-means clustering. In some embodiments, the machine learning model comprises a distribution-based clustering method. In some embodiments, the machine learning model comprises a Gaussian mixture model. In some embodiments, the machine learning model comprises a density-based clustering method. In some embodiments, the machine learning model comprises DBSCAN. In some embodiments, the machine learning model comprises OPTICS. In some embodiments, the machine learning model comprises a grid-based clustering method. In some embodiments, the machine learning model comprises STING. In some embodiments, the machine learning model comprises CLIQUE.

[0185] In some embodiments, the machine learning model comprises factor analysis. In some embodiments, the machine learning model comprises network component analysis. In some embodiments, the machine learning model comprises linear discriminant analysis. In some embodiments, the machine learning model comprises independent component analysis (ICA). In some embodiments, the machine learning model comprises principal component analysis (PCA). In some embodiments, the machine learning model comprises sparse PCA. In some embodiments, the machine learning model comprises robust PCA.

[0186] In some embodiments, the machine learning model includes non-negative matrix factorization (NMF). In some embodiments, the machine learning model includes traditional NMF. In some embodiments, the machine learning model includes discriminative NMF. In some embodiments, the machine learning model includes regularized NMF. In some embodiments, the machine learning model includes graph regularized NMF. In some embodiments, the machine learning model includes bootstrap sparse NMF.

[0187] In some embodiments, the machine learning model comprises Kernel PCA. In some embodiments, the machine learning model comprises Generalized Discriminant Analysis (GDA). In some embodiments, the machine learning model comprises an Autoencoder. In some embodiments, the machine learning model comprises T-Distributed Stochastic Neighbor Embedding (t-SNE). In some embodiments, the machine learning model comprises a Manifold Learning technique. In some embodiments, the machine learning model comprises Isomap. In some embodiments, the machine learning model comprises Locally Linear Embedding (LLE). In some embodiments, the machine learning model comprises Hessian LLE. In some embodiments, the machine learning model comprises Laplacian Eigenmap. In some embodiments, the machine learning model comprises Graph-Based Kernel PCA. In some embodiments, the machine learning model comprises Uniform Manifold Approximation and Projection (UMAP).

[0188] In some embodiments, the machine learning model comprises a penalized machine learning model. In some embodiments, the machine learning comprises a penalized version of any of the aforementioned models. A penalized machine learning model is a model in which coefficient estimates are regularized or constrained towards zero. In some embodiments, the machine learning model comprises a ridge regression model. In some embodiments, the machine learning model comprises a lasso regression model. In some embodiments, the machine learning model comprises an elastic net regression model.

[0189] In some embodiments the machine learning model comprises an ensemble model. In some embodiments the ensemble model comprises a boosting algorithm. In some embodiments the ensemble model comprises a bagging algorithm.

[0190] In some embodiments, the machine learning model comprises an ensemble model that includes any combination of multiple of any of the aforementioned models.

[0191] B. Correlation Score In some embodiments, the test dataset includes gene expression levels for one or more genes whose expression levels are included in the control dataset. In some embodiments, the test dataset includes gene expression levels for one or more genes having expression level representations included in the control dataset.

[0192] In some embodiments, the methods provided further include calculating a correlation score. In some embodiments, the correlation score indicates a similarity of the gene expression levels in the test dataset to the gene expression levels in the control dataset. In some embodiments, the correlation score indicates a similarity of the gene expression levels in the test dataset to the representation of the gene expression levels in the control dataset.

[0193] In some embodiments, calculating the correlation score comprises calculating the degree of correlation between the gene expression levels or representations thereof in the control dataset and the gene expression levels in the test dataset. Any suitable measure of the degree of correlation can be used, including Pearson correlation coefficient, Spearman rank correlation, and mutual information.

[0194] In some embodiments, a differentiation state of the one or more test cells is classified based on the correlation score and one or both of the first similarity score and the second similarity score. In some embodiments, the methods provided further include classifying a differentiation state of the one or more test cells based on the correlation score and one or both of the first similarity score and the second similarity score.

[0195] In some embodiments, the classification is based on the correlation score and one of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the lower of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the higher of the first similarity score and the second similarity score. In some embodiments, the classification is based on the correlation score and the first similarity score. In some embodiments, the classification is based on the correlation score and the second similarity score.

[0196] In some embodiments, if the correlation score indicates a dissimilarity between the gene expression levels in the test dataset and the gene expression levels or their representations in the control dataset, the differentiation state of the one or more test cells is not classified as a desired differentiation state. In some embodiments, if the degree of correlation does not exceed a predetermined cutoff value, the differentiation state of the one or more test cells is not classified as a desired differentiation state. In some embodiments, if the correlation score indicates that the correlation or explained variance between the gene expression levels or their representations in the control dataset and the gene expression levels in the test dataset is less than or about 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95, the differentiation state of the one or more test cells is not classified as a desired differentiation state.

[0197] In some embodiments, the differentiation state of the one or more test cells is classified based on the first similarity score, the second similarity score, and the correlation score. In some embodiments, the provided method further comprises classifying the differentiation state of the one or more test cells based on the first similarity score, the second similarity score, and the correlation score. In some embodiments, if the correlation score indicates a dissimilarity between the gene expression levels in the test dataset and the gene expression levels or representations thereof in the control dataset, the differentiation state of the one or more test cells is not classified as being a desired differentiation state. In some embodiments, if the correlation does not exceed a predetermined cutoff value, the differentiation state of the one or more test cells is not classified as being a desired differentiation state. In some embodiments, if the correlation score indicates that the correlation or explained variance between the gene expression levels or representations thereof of the control dataset and the gene expression levels of the test dataset is less than or about 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95, the differentiation state of the one or more test cells is not classified as being the desired differentiation state.

[0198] In some embodiments, the correlation score is calculated before, simultaneously, or after the calculation of the first and second similarity scores. In some embodiments, the correlation score is calculated before the calculation of the first and second similarity scores. In some embodiments, the provided method ends if the correlation score indicates dissimilarity between the gene expression levels in the test data set and the gene expression levels or their representations in the control data set. In some embodiments, the provided method ends if the correlation does not exceed a predetermined cutoff value.

[0199] In some embodiments, one or more genes of the control dataset include genes expressed in the cell at a control differentiation state. In some embodiments, the control differentiation state is any of the differentiation states described in Section II-C. In some embodiments, the control differentiation state is the same as one of the first, second and third differentiation states. In some embodiments, the control differentiation state is different from the first, second and third differentiation states.

[0200] In some embodiments, the one or more genes of the control dataset include genes expressed in cells at any of a plurality of control differentiation states. In some embodiments, the one or more genes of the control dataset include genes expressed in cells at each of a plurality of control differentiation states. In some embodiments, each of the plurality of control differentiation states is independently selected from any of the differentiation states described in Section II-C. In some embodiments, the plurality of control differentiation states includes a first, a second, and a third differentiation state.

[0201] In some embodiments, the gene expression level or its expression in the control dataset is based on the gene expression level of a plurality of reference cell populations.In some embodiments, the plurality of reference cell populations comprises the reference cell population whose gene expression level is used to train the first and second machine learning models, or comprises the reference cell population similar to that used to train the first and second machine learning models, for example, from the same cell type or the same stem cell differentiation pathway.Thus, in some aspects, the calculation of correlation score allows the comparison of the test dataset to the gene expression level of cells across the first, second and / or third differentiation state.

[0202] In some embodiments, the plurality of reference cell populations are different from, e.g., do not include, the reference cell populations whose gene expression levels are used to train the first and second machine learning models.For example, in some embodiments, the plurality of reference cell populations include cell types and / or differentiation states whose gene expression levels are different from the reference cell populations whose gene expression levels are used to train the first and second machine learning models.Thus, in some aspects, the calculation of correlation score allows the comparison of the test data set with the gene expression levels of cells other than the cells of the first, second and / or third differentiation states.

[0203] In some embodiments, one or more genes of the control dataset include genes that have at least a minimum expression level in cells of a control differentiation state. In some embodiments, one or more genes of the control dataset include genes that have at least a minimum expression level in cells of any of a plurality of control differentiation states. In some embodiments, one or more genes of the control dataset include genes that have at least a minimum expression level in cells of each of a plurality of control differentiation states. In some embodiments, one or more genes of the control dataset are expressed at at least a minimum expression level on average across a plurality of cell populations of a control differentiation state or a plurality of control differentiation states.

[0204] In some embodiments, one or more genes of the control dataset include genes with expression levels above a threshold. In some embodiments, one or more genes of the control dataset are filtered to include only genes whose expression levels exceed a threshold. In some embodiments, one or more genes of the control dataset include genes whose expression levels exceed a threshold on average across a plurality of cell populations of a control differentiation state or a plurality of control differentiation states. In some embodiments, the threshold is a threshold CPM value. In some embodiments, the threshold is a threshold log2 CPM value. In some embodiments, the threshold is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, or 20, or is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, or 20. In some embodiments, the threshold is 10 CPM or about 10 CPM. In some embodiments, the threshold is log2 CPM or is about 10 log2 CPM.

[0205] In some embodiments, the gene expression levels of the control dataset include a representation of the gene expression levels of one or more genes. In some embodiments, the gene expression levels of the control dataset include normalized gene expression levels. In some embodiments, the gene expression levels of the control dataset are normalized by CPM, e.g., CPM expression levels. In some embodiments, the gene expression levels of the control dataset are log-transformed. In some embodiments, the gene expression levels of the control dataset are log2-transformed.

[0206] In some embodiments, the gene expression level of the control dataset or a representation thereof comprises an average gene expression level of a plurality of reference cell populations. In some embodiments, a correlation measure is calculated between the gene expression level in the test dataset and the average gene expression level contained in the control dataset. In some embodiments, the average gene expression level comprises the centroid of the gene expression levels. In some embodiments, a correlation measure is calculated between the gene expression level in the test dataset and the centroid of the gene expression levels in the control dataset.

[0207] In some embodiments, the gene expression levels of the control dataset further comprise a measure of variance of the gene expression levels of the multiple reference cell populations. Any suitable measure of variance can be used, including standard deviation, range, interquartile range, mean absolute difference, median absolute deviation, mean absolute deviation, distance standard deviation, coefficient of variation (CV), interquartile coefficient of variance, relative mean difference, entropy, variance, and mean variance ratio. In some embodiments, the measure of variance is standard deviation. In some embodiments, the measure of variance is coefficient of variation (CV).

[0208] In some embodiments, the correlation measure is a weighted correlation value. In some embodiments, the correlation value is weighted by a measure of variance. In some embodiments, the correlation value is weighted by the inverse of a measure of variance. In some embodiments, the correlation measure is a 1 / CV weighted correlation value. In some embodiments, the correlation measure is a 1 / CV weighted correlation value calculated between the centroid values ​​of gene expression levels in the test dataset and gene expression levels in the control dataset.

[0209] C. Cell populations In some embodiments, the provided methods include classifying the differentiation state of a test cell. In some embodiments, the differentiation state of the test cell is classified based on a representation of gene expression levels, e.g., a machine learning model, based on gene expression levels from a plurality of reference cell populations. In some embodiments, the machine learning model used in the provided methods is trained using gene expression levels from a plurality of reference cell populations. In some embodiments, the machine learning model is trained to classify the differentiation state of the test cell using gene expression levels from a plurality of reference cell populations.

[0210] 1. Reference Cell Population In some aspects, the multiple reference cell populations include cells of known identity, e.g., cells of known cell type and / or differentiation state. For example, in some embodiments, the multiple reference cell populations used to train a first machine learning model include cells known to have a first or second differentiation state. Similarly, in some embodiments, the multiple reference cell populations used to train a second machine learning model include cells known to have a second or third differentiation state. In some embodiments, information about the known identity of the multiple reference cell populations is used in training the machine learning model or in establishing a criterion for determining whether the first and second similarity scores indicate whether the differentiation state of the test cell is more similar to the second differentiation state.

[0211] In some embodiments, the plurality of reference cell populations are derived from a culture of cells differentiated from pluripotent cells subjected to suitable differentiation conditions. The provided methods can be performed with reference cell populations produced according to any differentiation method. Exemplary differentiation methods are described in Section II-C.

[0212] In some embodiments, the plurality of reference cell populations comprises cells differentiated under conditions into dopaminergic neurons, hi some embodiments, the plurality of reference cell populations comprises cells differentiated according to any of the methods described in Section II-C.

[0213] In some embodiments, the pluripotent stem cells are induced pluripotent stem cells (iPSCs). In some embodiments, the iPSCs are generated from fibroblasts collected from a healthy human subject. In some embodiments, the iPSCs are generated from fibroblasts collected from a human subject with Parkinson's disease. Exemplary methods for iPSC generation are described in Section II-C.

[0214] In some embodiments, the cells of the reference cell population comprise pluripotent stem cells. In some embodiments, the pluripotent stem cells are induced pluripotent stem cells (iPSCs). In some embodiments, the iPSCs are generated from fibroblasts collected from a healthy human subject. In some embodiments, the iPSCs are generated from fibroblasts collected from a human subject with Parkinson's disease. In some embodiments, the iPSCs are generated from fibroblasts collected from a human subject susceptible to developing Parkinson's disease. Exemplary methods for iPSC generation are described in Section II-C.

[0215] In some embodiments, the cells of the reference cell population include cells differentiated under conditions that lead to neural cells, such as floor plate midbrain progenitor cells, committed dopaminergic cells, or dopaminergic neurons. In some embodiments, the cells of the reference cell population include cells differentiated according to any of the methods described in Section II-C. In some embodiments, the cells of the reference cell population include committed dopaminergic cells. In some embodiments, the cells of the reference cell population include dopaminergic neurons, such as committed dopaminergic neurons. In some embodiments, the cells of the reference cell population include iPSCs, such as cells derived from the iPSCs described above, cultured under conditions that promote differentiation into dopaminergic neurons.

[0216] In some embodiments, the cells of the reference cell population comprise dopaminergic neurons that express markers of midbrain dopaminergic neurons, such as expression of FOXA2 or tyrosine hydroxylase (TH). In some embodiments, the cells of the reference cell population comprise cells that express TH (TH+). In some embodiments, the cells of the reference cell population comprise cells that express FOXA2 (FOXA2+). In some embodiments, the cells of the reference cell population comprise cells that express TH and FOXA2 (TH+FOXA2+).

[0217] In some embodiments, the cells of the reference cell population include cells that are determined to be or can be dopaminergic neurons, i.e., determined to be dopaminergic cells, as confirmed based on one or more characteristics indicating that the cells may have functional activity of dopaminergic neurons but may not yet express or express at high levels markers of dopaminergic neurons. For example, the cells may exhibit lower levels of TH than dopaminergic neurons, but still exhibit one or more characteristics of determined dopaminergic cells indicating that the cells may have functional activity of dopaminergic neurons. In some embodiments, the one or more characteristics include the activity of surviving, engrafting, and / or innervating other cells when administered in vivo, for example, to an animal model. In some embodiments, the cells of the reference cell population include cells that can innervate host tissue after transplantation into an animal or human subject. In some embodiments, the cells of the reference cell population include cells that exhibit neurite outgrowth after transplantation into an animal or human subject. In some embodiments, the cells of the reference cell population include cells that survive after transplantation into an animal or human subject. In some embodiments, cells of the reference cell population include cells that engraft following transplantation into an animal or human subject.

[0218] In some embodiments, the cells of the reference cell population include cells that have a therapeutic effect for treating a neurodegenerative disease. In some embodiments, the transplanted cells improve or reverse the symptoms of the neurodegenerative disease. In some embodiments, the neurodegenerative disease is Parkinson's disease. In some embodiments, the cells improve Parkinson's symptoms when transplanted into a subject, e.g., the substantia nigra of a patient in need thereof.

[0219] In some embodiments, the cells of the reference cell population include cells screened for therapeutic efficacy for treating a neurodegenerative disease, for example, cells determined in an animal model of a neurodegenerative disease. In some embodiments, the neurodegenerative disease is Parkinson's disease. In some embodiments, the reference cells are screened using an animal model of Parkinson's disease. Any suitable animal model of Parkinson's disease can be used for screening. In some embodiments, the animal model is a lesion model in which 6-hydroxydopamine (6-OHDA) is unilaterally stereotaxically injected into the animal's substantia nigra. In some embodiments, the animal model is a lesion model in which 6-OHDA is unilaterally stereotaxically injected into the animal's medial forebrain bundle. In some embodiments, the cells are transplanted into the substantia nigra of the animal model. In some embodiments, a behavioral assay is performed to screen the therapeutic efficacy of transplantation into the animal model. In some embodiments, the behavioral assay includes monitoring amphetamine-induced rotation behavior. In some embodiments, the cells are determined to reduce, decrease or reverse Parkinson's disease model brain lesions in the model. In some embodiments, the cells may include cells that do not reduce, decrease, or reverse Parkinson's disease model brain pathology in this model. The reference cell population may include various cells that exhibit various or different therapeutic effects for treating a neurodegenerative disease, for example in an animal model.

[0220] 2. Test Cells In some aspects, the test cell is a cell of unknown identity, e.g., unknown cell type and / or differentiation state. In some embodiments, the test cell is known or suspected to be in a particular stem cell differentiation pathway, but the differentiation state within the pathway is unknown. In some aspects, the provided method allows for the cell type and / or differentiation state of the test cell to be determined based on the gene expression level of the test cell. Based on this determination, the in vitro population of cells containing the test cell can be classified as having a particular cell type and / or differentiation state.

[0221] In some embodiments, the in vitro population of cells containing the test cells is derived from a culture of cells differentiated from pluripotent cells subjected to suitable differentiation conditions. The provided methods can be performed with test cells generated according to any differentiation method. Exemplary differentiation methods and in vitro populations of cells are described in Section II-C.

[0222] In some embodiments, the cell is a neural cell derived from a stem cell. In some embodiments, the test cell comprises a cell differentiated under conditions into a dopaminergic neuron. In some embodiments, the test cell comprises a cell differentiated according to any of the methods described in Section II-C.

[0223] In some embodiments, the pluripotent stem cells are induced pluripotent stem cells (iPSCs). In some embodiments, the iPSCs are generated from fibroblasts collected from a healthy human subject. In some embodiments, the iPSCs are generated from fibroblasts collected from a human subject with Parkinson's disease. Exemplary methods for iPSC generation are described in Section II-C.

[0224] In some embodiments, the test cells are derived from an in vitro population of cells that are, or are suspected of being, on a different differentiation pathway than the reference cell population, hi some embodiments, the test cells are derived from an in vitro population of cells that have been, or are suspected of having been, generated using a different differentiation method than the differentiation method used to generate the reference cell population.

[0225] In some embodiments, the test cells are derived from an in vitro population of cells that are or are suspected to be on the same differentiation pathway as the reference cell population. In some embodiments, the test cells are derived from an in vitro population of cells that have been or are suspected to have been produced using the same differentiation method as the differentiation method used to produce the reference cell population. In some embodiments, the test cells are derived from an in vitro population of cells that are or are suspected to be on the same differentiation pathway as the reference cell population, but that have been or are suspected to have been produced using a different differentiation method than that used to produce the reference cell population.

[0226] 3. Exemplary Differentiation Pathways, Methods, and Conditions In some embodiments, the first differentiation state is earlier in the stem cell differentiation pathway than the second differentiation state. In some embodiments, the first differentiation state is later in the stem cell differentiation pathway than the second differentiation state. In some embodiments, the first and second differentiation states are from different stem cell differentiation pathways. In some embodiments, the first differentiation state is in a cell differentiation pathway that is parallel to the cell differentiation pathway of the second differentiation state. In some embodiments, the cell differentiation pathway is a divergent pathway, for example, such that the first and second differentiation states are different cell types.

[0227] In some embodiments, the second differentiation state is earlier in the stem cell differentiation pathway than the third differentiation state. In some embodiments, the second differentiation state is later in the stem cell differentiation pathway than the third differentiation state. In some embodiments, the second and third differentiation states are derived from different stem cell differentiation pathways. In some embodiments, the second differentiation state is in a cell differentiation pathway that is parallel to the cell differentiation pathway of the third differentiation state. In some embodiments, the cell differentiation pathway is a divergent pathway, for example, such that the second and third differentiation states are different cell types.

[0228] In some embodiments, the first, second and third differentiation states are all in the same stem cell differentiation pathway. In some embodiments, the first, second and third differentiation states are in different stem cell differentiation pathways. In some embodiments, the first and second differentiation states are in one stem cell differentiation pathway, and the first and third differentiation states are in another stem cell differentiation pathway, for example, a pathway in which the cells in the first differentiation state are precursor cells that can differentiate into cells in the second or third differentiation state. In some embodiments, the second and third differentiation states are of different cell types.

[0229] In some embodiments, the first, second, and third differentiation states are all within the same stem cell differentiation pathway. In some embodiments, the second differentiation state is an intermediate differentiation state between the first and third differentiation states.

[0230] Exemplary in vitro populations and differentiation states of cells are provided in this section. Although this section is organized based on a specific stem cell differentiation pathway, combinations of the first, second, third and control differentiation states across multiple cell types or pathways are also contemplated and disclosed herein. For example, in some embodiments, the first, second and third differentiation states are all the same cell type (e.g., neuron). In other embodiments, at least one of the first, second and third differentiation states can be a different cell type from the remaining differentiation states (e.g., the first differentiation state is a neuronal differentiation state, and the second and third differentiation states are cardiac differentiation states).

[0231] In some embodiments, the test cells are derived from an in vitro population of stem cell-derived cardiomyocytes (see, e.g., Le and Chong, Cell Death Discovery 2:16052 (2016)). In some embodiments, the stem cell-derived cardiomyocytes express Nkx2.5 and / or Isl-1. Exemplary methods for differentiating stem cell-derived cardiomyocytes in vitro are described in U.S. Pat. Nos. 9,234,176, 20170058263, Vahdat et al., Scientific Reports 9:16006 (2019), Laflamme et al. (2007) Nature Biotechnology 25:1015-24, and Wu et al. (2021) Biosci Rep 41(6):BSR20200833. In some embodiments, the first, second, third, and / or control differentiation state is a differentiation state of a cardiac progenitor cell, or a differentiation state of a committed or determined cardiomyocyte, endothelial cell, vascular smooth muscle cell, or cardiac fibroblast. In some embodiments, the first differentiation state is a differentiation state of a cardiac progenitor cell, the second differentiation state is a determined differentiation state of a cardiac myocyte, endothelial cell, vascular smooth muscle cell, or cardiac fibroblast, and the third differentiation state is a committed differentiation state corresponding to the second differentiation state. In some embodiments, the second differentiation state is a determined differentiation state of a cardiac myocyte. In some embodiments, the third differentiation state is a committed differentiation state of a cardiac myocyte. In some embodiments, the cells having a desired differentiation state, for example, identified as having a second differentiation state, can be used to treat degenerative diseases such as ischemic cardiomyopathy and conduction system diseases (such as sinus node dysfunction and atrioventricular block), or congenital heart diseases such as atrial or ventricular septal defects.

[0232] In some embodiments, the test cells are derived from an in vitro population derived from a culture of cells differentiated from pluripotent cells subjected to a differentiation protocol to induce differentiation of PSCs, e.g., iPSCs, into cardiomyocytes according to any of the methods described herein, e.g., as described in Laflamme et al. (2007) Nature Biotechnology 25:1015-24 or Wu et al. (2021) Biosci Rep 41(6):BSR20200833. In some embodiments, the cells in the second differentiation state are at any of days 14-21 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are at or before day 13, day 12, day 11, or day 10 of the differentiation protocol. In some embodiments, the cells in the third differentiation state are at or after day 22, day 30, day 40, day 50, day 60, or day 70 of the differentiation protocol. In some embodiments, the cells in the second differentiation state are anywhere between days 14-21 of the differentiation protocol, and the cells in the third differentiation state are anywhere between days 22 or later, 30 or later, 40 or later, 50 or later, 60 or later, or 70 or later of the differentiation protocol. In some embodiments, the cells in the first differentiation state are on or before day 13, 12 or earlier, 11 or earlier, or 10 or earlier of the differentiation protocol, and the cells in the second differentiation state are anywhere between days 14-21 of the differentiation protocol, and the cells in the third differentiation state are on or after day 22, 30 or later, 40 or later, 50 or later, 60 or later, or 70 or later of the differentiation protocol. In some embodiments, the cells in the third differentiation state are anywhere between days 70-126 of the differentiation protocol.

[0233] In some embodiments, the test cells are derived from an in vitro population of stem cell-derived skeletal muscle cells (see, e.g., Relaix et al., Nature Communications 12:692 (2021)). In some embodiments, the stem cell-derived skeletal muscle cells express PAX7 and / or PAX3. Exemplary methods for differentiating stem cell-derived skeletal muscle cells in vitro are described in WO 2001011011 and U.S. Pat. No. 9,789,136. In some embodiments, the first, second, third and / or control differentiation state is a differentiation state of a skeletal muscle progenitor cell, a committed skeletal muscle cell or a committed skeletal muscle cell. In some embodiments, the first differentiation state is a differentiation state of a skeletal muscle progenitor cell, the second differentiation state is a differentiation state of a committed skeletal muscle cell, and the third differentiation state is a differentiation state of a committed skeletal muscle cell. In some embodiments, cells identified as having a desired differentiation state, e.g., a second differentiation state, can be used to treat a muscle disorder, e.g., a myopathy, e.g., polymyositis, dermatomyositis, Duchenne muscular dystrophy; fibrositis; myasthenia gravis; rhabdomyolysis; amyotrophic lateral sclerosis; or sarcopenia.

[0234] In some embodiments, the test cells are derived from an in vitro population of stem cell-derived smooth muscle cells. An exemplary method for in vitro differentiation of stem cell-derived smooth muscle cells is described in US Pat. No. 7,531,355. In some embodiments, the first, second, third, and / or control differentiation state is a differentiation state of smooth muscle progenitor cells, committed smooth muscle cells, or determined smooth muscle cells. In some embodiments, the first differentiation state is a differentiation state of smooth muscle progenitor cells, the second differentiation state is a differentiation state of determined smooth muscle cells, and the third differentiation state is a differentiation state of committed smooth muscle cells. In some embodiments, cells identified as having a desired differentiation state, e.g., having a second differentiation state, can be used to reconstitute tissues containing smooth muscle progenitor cells (such as the urinary tract, epithelial pathways, or bladder) or to treat disorders affecting smooth muscle function, e.g., urinary incontinence, bladder disease, vascular disorders, bowel disorders, vesicoureteral reflux, or other disorders of smooth muscle function.

[0235] In some embodiments, the test cells are derived from an in vitro population of stem cell-derived endothelial cells. Exemplary methods for in vitro differentiation of stem cell-derived endothelial cells are described in U.S. Patent Nos. 10,041,036, 10,563,175, 10,828,337, 10,767,161, 9,938,499 and 10,947506. In some embodiments, the first, second, third and / or control differentiation state is a differentiation state of endothelial progenitor cells, a committed endothelial cell or a determined endothelial cell. In some embodiments, the first differentiation state is a differentiation state of endothelial progenitor cells, the second differentiation state is a differentiation state of determined endothelial cells and the third differentiation state is a differentiation state of committed endothelial cells.

[0236] In some embodiments, the test cells are derived from an in vitro population of stem cell-derived tubular cells (see, e.g., Chambers and Wingert, World J Stem Cells 2016;8(11):367-375). An exemplary method for differentiating stem cell-derived tubular cells in vitro is described in Ribeiro et al., Stem Cells Int. 2020:8894590. In some embodiments, the first, second, third, and / or control differentiation state is a differentiation state of a tubular progenitor cell or a migrated or committed podocyte, a proximal S1 cell, a proximal S2 cell, a proximal S3 cell, a proximal tubule cell, a DTL type 1 cell, a DTL type 2 cell, a DTL type 3 cell, an ascending thin limb cell, a MTAL limb cell, a CTAL cell, a macula densa cell, a distal convoluted tubule cell, a CNT cell, a PC(CCD) cell, a PC(OMCD) cell, a type A IC cell, a type B IC cell, or an IMCD cell. In some embodiments, the first differentiation state is a differentiation state of a tubular progenitor cell. The second differentiation state is of a committed podocyte, proximal S1 cell, proximal S2 cell, proximal S3 cell, proximal tubule cell, DTL type 1 cell, DTL type 2 cell, DTL type 3 cell, ascending thin limb cell, MTAL limb cell, CTAL cell, macula densa cell, distal tortuous tubule cell, CNT cell, PC(CCD) cell, PC(OMCD) cell, type A IC cell, type B IC cell, or IMCD cell, and the third differentiation state is a committed differentiation state corresponding to the second differentiation state. In some embodiments, the cells having a desired differentiation state, e.g., identified as having a second differentiation state, can be used to treat acute kidney injury, chronic kidney disease, refractory systemic lupus erythematosus or lupus nephritis, or for kidney transplantation (see, e.g., Wong, World J Stem Cells, 2021;13(7):914-933).

[0237] In some embodiments, the test cells are derived from an in vitro population of stem cell-derived erythroid cells. Exemplary methods for in vitro differentiation of stem cell-derived erythroid cells are described in US Pat. No. 1,027,211. In some embodiments, the first, second, third and / or control differentiation states are erythroid progenitor, committed erythroid or committed erythroid differentiation states. In some embodiments, the first differentiation state is an erythroid progenitor differentiation state, the second differentiation state is a committed erythroid differentiation state and the third differentiation state is a committed erythroid differentiation state. In some embodiments, the cells identified as having a desired differentiation state, e.g., having a second differentiation state, can be used to treat a disorder characterized by a deficiency of erythrocytes, e.g., to treat a subject having an autoimmune disorder, an immune deficiency, or any other disease or disorder that would benefit from a blood transfusion.

[0238] In some embodiments, the test cell is derived from an in vitro population of stem cell-derived lung cells (see, for example, Leeman et al., Curr Top Dev Biol 2014;107:207-233). In some embodiments, the stem cell-derived lung cells express Nkx2.1. Exemplary methods for in vitro differentiation of stem cell-derived lung cells are described in U.S. Pat. No. 11214769 and WO2015108893. In some embodiments, the first, second, third and / or control differentiation state is a differentiation state of a lung progenitor cell or a committed or committed airway epithelial cell, such as a goblet cell, a ciliated cell, a Clara cell, a neuroendocrine cell (neuroendocrine body), a basal cell, a metaphase cell (or parabasal cell), a serous cell, a brush cell, a cancer cell, a non-ciliated columnar cell, a dysplastic (e.g., squamous or cranial mucosa cell, bronchiolar metaplasia) cell; or an alveolar cell, such as a type 1 or type 2 pneumocyte or a cuboidal non-ciliated cell. In some embodiments, the first differentiation state is that of a lung progenitor cell, the second differentiation state is that of a committed airway epithelial cell, e.g., a committed goblet cell, a ciliated cell, a Clara cell, a neuroendocrine cell (neuroendocrine body), a basal cell, a metaphase cell (or parabasal cell), a serous cell, a brush cell, a cancer cell, a non-ciliated columnar cell, a dysplastic (e.g., squamous or cranial mucosa cell, bronchiolar metaplasia) cell; or a committed alveolar cell, e.g., a committed type 1 or type 2 pneumocyte or a cuboidal non-ciliated cell, and the third differentiation state is a committed differentiation state corresponding to the second differentiation state.In some embodiments, cells identified as having a desired differentiation state, e.g., a second differentiation state, can be used to treat a respiratory disorder, e.g., cystic fibrosis, respiratory distress syndrome, acute respiratory distress syndrome, pulmonary tuberculosis, cough, bronchial asthma, cough based on airway hyperresponsiveness (bronchitis, influenza syndrome, asthma, obstructive pulmonary disease, etc.), influenza syndrome, anti-cough, airway hyperresponsiveness, tuberculosis disease, asthma (airway inflammatory cell infiltration, airway hyperresponsiveness, bronchoconstriction, hypermucus secretion), chronic obstructive pulmonary disease, emphysema, pulmonary fibrosis, idiopathic pulmonary fibrosis, cough, reversible airway obstruction, adult respiratory disease syndrome, Fargeon's disease, farmer's lung, bronchopulmonary dysplasia, airway obstruction, emphysema, allergic bronchopulmonary aspergillosis, allergic bronchitis, bronchiectasis, occupational asthma, reactive airway disease syndrome, interstitial lung disease, or parasitic lung disease.

[0239] In some embodiments, the test cells are derived from an in vitro population of stem cell derived thyroid cells. In some embodiments, the stem cell derived thyroid cells express Pax-8 and / or NKX2-1. Exemplary methods for in vitro differentiation of stem cell derived thyroid cells are described in Fierabracci, Journal of Endocrinology 213(1):1-13 (2012). In some embodiments, the first, second, third and / or control differentiation state is a differentiation state of thyroid progenitor cells or a committed or committed follicular or C cell. In some embodiments, the first differentiation state is a differentiation state of thyroid progenitor cells. The second differentiation state is a committed follicular or C cell differentiation state, and the third differentiation state is a committed differentiation state corresponding to the second differentiation state. In some embodiments, the cells identified as having a desired differentiation state, e.g., having a second differentiation state, can be used to treat a thyroid disorder, e.g., goiter, adenoma, hypothyroidism, or an autoimmune disease.

[0240] In some embodiments, the test cells are derived from an in vitro population of stem cell derived pancreatic cells. In some embodiments, the stem cell derived pancreatic cells express Pdx1. In some embodiments, the stem cell derived pancreatic cells are endocrine cells. In some embodiments, the stem cell derived pancreatic cells are exocrine cells. Exemplary methods for differentiating stem cell derived pancreatic cells in vitro are described in U.S. Pat. No. 8,859,286, WO 2011011300, WO 2014105543, WO 2013095953, U.S. Pat. No. 9,157,062, and Balboa et al. (2022) Nature Biotechnology 40:1042-55. In some embodiments, the first, second, third, and / or control differentiation states are differentiation states of pancreatic progenitor cells or committed or determined exocrine cells (e.g., acinar or ductal cells) or endocrine cells (e.g., beta, alpha, delta, or PP cells). In some embodiments, the first differentiation state is a pancreatic progenitor cell differentiation state, the second differentiation state is a committed exocrine cell (e.g., acinar or ductal cell) or endocrine cell (e.g., beta, alpha, delta or PP cell), and the third differentiation state is a committed differentiation state corresponding to the second differentiation state. In some embodiments, the second differentiation state is a committed beta cell differentiation state. In some embodiments, the third differentiation state is a committed beta cell differentiation state. In some embodiments, cells identified as having a desired differentiation state, e.g., having the second differentiation state, can be used to treat a pancreatic disorder, a metabolic disorder, or a disease involving inappropriate production or use of insulin, such as type 1 diabetes.

[0241] In some embodiments, the test cells are derived from an in vitro population from a culture of cells differentiated from pluripotent cells that are subjected to a differentiation protocol to induce differentiation of PSCs, e.g., iPSCs, into beta cells by any of the methods described herein, e.g., as described in Balboa et al. (2022) Nature Biotechnology 40:1042-55. In some embodiments, the cells in the second differentiation state are at any of days 21-35 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are at or before day 20, at or before day 19, at or before day 18, or at or before day 17 of the differentiation protocol. In some embodiments, the cells in the third differentiation state are at or after day 36, at or after day 38, at or after day 40, at or after day 42, or at or after day 44 of the differentiation protocol. In some embodiments, the cells in the second differentiation state are anywhere between days 21-35 of the differentiation protocol, and the cells in the third differentiation state are anywhere between day 36 or later, day 38 or later, day 40 or later, day 42 or later, or day 44 or later of the differentiation protocol. In some embodiments, the cells in the first differentiation state are anywhere between day 20 or earlier, day 19 or earlier, day 18 or earlier, or day 17 or earlier of the differentiation protocol, and the cells in the second differentiation state are anywhere between day 21-35 of the differentiation protocol, and the cells in the third differentiation state are anywhere between day 36 or later, day 38 or later, day 40 or later, day 42 or later, or day 44 or later of the differentiation protocol. In some embodiments, the cells in the third differentiation state are anywhere between day 56-98 of the differentiation protocol.

[0242] In some embodiments, the test cells are derived from an in vitro population of stem cell-derived epidermal cells (see, e.g., Jackson et al., Stem Cell Research & Therapy 8:155 (2017)). Exemplary methods for in vitro differentiation of stem cell-derived epidermal cells are described in U.S. Pat. No. 9,404,122. In some embodiments, the first, second, third, and / or control differentiation states are epidermal progenitor cells or committed or determined differentiation states of caratinocytes, melanocytes, or Langerhans cells. In some embodiments, the first differentiation state is an epidermal progenitor cell differentiation state, the second differentiation state is a determined differentiation state of caratinocytes, melanocytes, or Langerhans cells, and the third differentiation state is a committed differentiation state corresponding to the second differentiation state. In some embodiments, cells identified as having a desired differentiation state, e.g., having a second differentiation state, can be used to treat skin damage or disorders, such as burns, chronic wounds, or stable vitiligo.

[0243] In some embodiments, the test cells are derived from an in vitro population of stem cell derived pigment cells. In some embodiments, the stem cell derived pigment cells are retinal pigment cells. In some embodiments, the stem cell derived pigment cells are melanocytes. Exemplary methods for in vitro differentiation of stem cell derived pigment cells are described in WO2005070011, WO2011149762, WO2014121077, WO2009051671 and WO2008129554. In some embodiments, the first, second, third and / or control differentiation state is a pigment precursor cell or a committed or determined retinal pigment cell or melanocyte differentiation state. In some embodiments, the first differentiation state is a pigment precursor cell differentiation state, the second differentiation state is a determined retinal pigment cell or melanocyte differentiation state, and the third differentiation state is a committed differentiation state corresponding to the second differentiation state. In some embodiments, cells identified as having a desired differentiation state, e.g., a second differentiation state, can be used to treat a degenerative disease, such as a retinal degenerative disease, e.g., macular degeneration.

[0244] A nerve cell In some embodiments, the test cell is derived from an in vitro population of stem cell-derived neural cells.Exemplary methods for in vitro differentiation of stem cell-derived neural cells are described in International Publication No. 2014176606, US Patent No. 8460931, US Patent No. 10273453, International Publication No. 2012095730, US Patent No. 9309495, US Patent No. 20190249140, US Patent No. 20180298326, International Publication No. 2009148170, International Publication No. 2021146349, International Publication No. 2021216623, International Publication No. 2021216622. WO 2013104752, WO 2010096496, WO 2013067362, WO 2016196661, WO 2015143342, and U.S. Patent Application Publication No. 20160348070.

[0245] In some embodiments, the method of differentiating neural cells from stem cells can be a method of differentiating pluripotent stem cells, such as iPSCs, into any neural cell type, using any available or known method for inducing differentiation of pluripotent stem cells, such as iPSCs.In some embodiments, the method induces differentiation of pluripotent stem cells into floor-plate midbrain progenitor cells, committed dopaminergic cells, and / or dopaminergic neurons.Any available or known method can be used for inducing differentiation of pluripotent stem cells into floor-plate midbrain progenitor cells, committed dopaminergic cells, and / or dopaminergic neurons.

[0246] In some embodiments, the method induces differentiation of pluripotent stem cells into glial cells, hi some embodiments, the glial cells are selected from the group consisting of microglial cells, astrocytes, oligodendrocytes and ependymal cells.

[0247] In some embodiments, the test cells are derived from an in vitro population of stem cell-derived microglial cells. In some embodiments, the method induces differentiation of pluripotent stem cells into microglial or microglia-like cells. Any available known method for inducing differentiation of pluripotent stem cells into microglial or microglia-like cells can be used. Exemplary methods of inducing differentiation of pluripotent stem cells into microglial or microglia-like cells can be found, for example, in McQuade et al. (2018) Molecular Neurodegeneration 13:67; Abud et al., Neuron (2017), Vol. 94:278-293; Douvaras et al., Stem Cell Reports (2017), Vol. 8:1516-1524; Muffat et al., Nature Medicine (2016), Vol. 22(11):1358-1367; and Pandya et al., Nature Neuroscience (2017), Vol. 20(5):753-759. In some embodiments, the first, second, third and / or control differentiation states are iPSC, hematopoietic progenitor or microglial cell differentiation states. In some embodiments, the first differentiation state is that of iPSCs, the second differentiation state is that of hematopoietic progenitor cells, and the third differentiation state is that of microglial cells. In some embodiments, cells identified as having a desired differentiation state, e.g., the second differentiation state, can be used to treat Parkinson's disease, parkinsonism, age-related neurodegenerative diseases, Alzheimer's disease, or frontotemporal dementia.

[0248] In some embodiments, the test cells are derived from an in vitro population derived from a culture of cells differentiated from pluripotent cells that are subjected to a differentiation protocol to induce differentiation of PSCs, e.g., iPSCs, into microglial cells, e.g., according to any of the methods described herein. In some embodiments, the cells in the second differentiation state are anywhere from day 28 to day 35 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are on or before day 27, on or before day 26, on or before day 25, or on or before day 24 of the differentiation protocol. In some embodiments, the cells in the third differentiation state are on or after day 36, on or after day 37, on or after day 38, on or after day 39, on or after day 40, on or after day 41, on or after day 42, or on or after day 43 of the differentiation protocol. In some embodiments, the cells in the second differentiation state are anywhere between days 28-35 of the differentiation protocol, and the cells in the third differentiation state are anywhere between days 36 or later, 37 or later, 38 or later, 39 or later, 40 or later, 41 or later, 42 or later, or 43 or later of the differentiation protocol. In some embodiments, the cells in the first differentiation state are on or before day 27, 26 or earlier, 25 or earlier, or 24 or earlier of the differentiation protocol, and the cells in the second differentiation state are anywhere between days 28-35 of the differentiation protocol, and the cells in the third differentiation state are on or after day 36, 37 or later, 38 or later, 39 or later, 40 or later, 41 or later, 42 or later, or 43 or later of the differentiation protocol. In some embodiments, the cells in the third differentiation state are anywhere between days 49-63 of the differentiation protocol.

[0249] In some embodiments, the method induces the differentiation of pluripotent stem cells into astrocytes. Any available known method for inducing the differentiation of pluripotent stem cells into astrocytes can be used. Exemplary methods for inducing the differentiation of pluripotent stem cells into astrocytes can be found, for example, in TCW et al., Stem Cell Reports (2017), Vol.9:600-614, including the methods described in the references cited therein, for example, in Table 1. Exemplary methods of inducing differentiation of pluripotent stem cells into astrocytes may, in some embodiments, include the use of commercially available kits and methods provided for the use of such kits, such as Astrocyte Medium, Catalog #1801 (ScienCell Research Laboratories, Carlsbad, CA); Astrocyte Medium, Catalog # A1261301 (ThermoFisher Scientific Inc, Waltham, MA); and AGM Astrocyte Growth Medium Bullet Kit, Catalog # CC-3186 (Lonza, Basel, Switzerland).

[0250] In some embodiments, the method induces the differentiation of pluripotent stem cells into oligodendrocytes. Any available known method for inducing the differentiation of pluripotent stem cells into oligodendrocytes can be used. Exemplary methods for inducing the differentiation of pluripotent stem cells into oligodendrocytes can be found, for example, in Ehrlich et al., PNAS (2017), Vol. 114 (11): E2243-E2252; Douvaras et al., Stem Cell Reports (2014), Vol. 3 (2): 250-259; Stacpoole et al., Stem Cell Reports (2013), Vol. 1 (5): 437-450; Wang et al., Cell Stem Cell (2013), Vol. 12 (2): 252-264; and Piao et al., Cell Stem Cell (2015), Vol. 16 (2): 198-210.

[0251] In some embodiments, the test cells are derived from an in vitro population of stem cell-derived GABAergic neurons. Exemplary methods for in vitro differentiation of stem cell-derived GABAergic neurons are described in Maroof et al. (2013) Cell Stem Cell 12(5):573-586. US Patent Application Publication No. 2020 / 0002679(A1), US Patent Application Publication No. 20110183912(A1), and US Patent Application Publication No. 20140248696(A1). In some embodiments, the first, second, third, and / or control differentiation state is a differentiation state of GABAergic neural progenitor cells, committed GABAergic neurons, or determined GABAergic neurons. In some embodiments, the first differentiation state is a GABAergic neuronal progenitor differentiation state, the second differentiation state is a committed GABAergic neuronal differentiation state, and the third differentiation state is a committed GABAergic neuronal differentiation state. In some embodiments, cells identified as having a desired differentiation state, e.g., having the second differentiation state, can be used to treat epilepsy.

[0252] In some embodiments, the test cells are derived from an in vitro population from a culture of cells differentiated from pluripotent cells subjected to a differentiation protocol to induce differentiation of PSCs (e.g., iPSCs) into inhibitory neurons (e.g., GABAergic neurons). Exemplary methods for differentiating inhibitory neurons in vitro are described in Kang et al. (2017) Sci Rep 7:12233 and Nicholas et al. (2013) Cell Stem Cell 12(5):573-86. In some embodiments, the cells in the second differentiation state are anywhere from week 5 to 10 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are at week 4 or earlier, week 3 or earlier, or week 2 or earlier of the differentiation protocol. In some embodiments, the cells in the third differentiation state are at week 12 or later, week 14 or later, week 16 or later, week 18 or later, or week 20 or later of the differentiation protocol. In some embodiments, the cells in the second differentiation state are anywhere between 5-10 weeks of the differentiation protocol and the cells in the third differentiation state are after 12 weeks, after 14 weeks, after 16 weeks, after 18 weeks, or after 20 weeks of the differentiation protocol. In some embodiments, the cells in the first differentiation state are on or before week 4, on or before week 3, or on or before week 2 of the differentiation protocol and the cells in the second differentiation state are on or before week 5-10 of the differentiation protocol and the cells in the third differentiation state are on or after week 12, on or after week 14, on or after week 16, on or after week 18, or after 20 weeks of the differentiation protocol. In some embodiments, the cells in the third differentiation state are on or after week 20-30 of the differentiation protocol.

[0253] In some embodiments, the method induces differentiation of iPSCs into floor plate midbrain progenitor cells, committed dopaminergic cells, and / or dopaminergic neurons. In some embodiments, the method involves (a) performing a first incubation comprising culturing pluripotent stem cells in a non-adherent culture vessel under conditions to generate cell spheroids, where starting at the beginning of the first incubation (day 0), the first incubation exposes the cells to (i) an inhibitor of TGF-β / Activin-Nodal signaling, (ii) at least one activator of Sonic Hedgehog (SHH) signaling, (iii) an inhibitor of bone morphogenetic protein (BMP) signaling, and (iv) an inhibitor of glycogen synthase kinase 3β (GSK3β) signaling, and (b) performing a second incubation comprising culturing the cells of the spheroids in a substrate-coated culture vessel under conditions to cause neuronal differentiation of the cells.

[0254] In some embodiments, the method involves exposing the pluripotent stem cells to (a) an inhibitor of bone morphogenetic protein (BMP) signaling; (b) an inhibitor of TGF-β / activin-nodal signaling; (c) at least one activator of Sonic Hedgehog (SHH) signaling. In some embodiments, the method further comprises exposing the pluripotent stem cells to at least one inhibitor of GSK3β signaling. In some embodiments, the exposure to the inhibitor of BMP signaling and the inhibitor of TGF-β / activin-nodal signaling occurs while the pluripotent stem cells are bound to a substrate. In some embodiments, the pluripotent stem cells are attached to a substrate while exposed to the inhibitor of BMP signaling, the inhibitor of TGF-β / activin-nodal signaling, and at least one activator of SHH signaling. In some embodiments, the pluripotent stem cells are bound to a substrate while exposed to the at least one inhibitor of GSK3β signaling. In some embodiments, the pluripotent stem cells are in a non-adherent culture vessel under conditions that produce cell spheroids while being exposed to at least one BMP signaling inhibitor, TGF-β / Activin-Nodal signaling inhibitor, and activator of SHH signaling. In some embodiments, the pluripotent stem cells are in a non-adherent culture vessel under conditions that produce cell spheroids while being exposed to at least one inhibitor of GSK3β signaling.

[0255] In some embodiments, the non-adherent culture vessel allows three-dimensional formation of cell aggregates. In some embodiments, the iPSCs are cultured in a non-adherent culture vessel, such as a multi-well plate, to generate cell aggregates (e.g., spheroids). In some embodiments, the iPSCs are cultured in a non-adherent culture vessel, such as a multi-well plate, to generate cell aggregates (e.g., spheroids) at about day 7 of the method. In some embodiments, the cell aggregates (e.g., spheroids) express at least one of PAX6 and OTX2 at or by about day 7 of the method.

[0256] In some embodiments, the first incubation is from about day 0 to about day 6. In some embodiments, the first incubation comprises culturing the pluripotent stem cells in culture medium ("medium"). In some embodiments, the first incubation comprises culturing the pluripotent stem cells in medium from about day 0 to about day 6. In some embodiments, the first incubation comprises culturing the pluripotent stem cells in medium to induce differentiation of the PSCs into floor plate midbrain progenitor cells.

[0257] In some embodiments, the medium is supplemented with a serum replacement that contains minimal non-human derived components (e.g., KnockOut™ serum replacement). In some embodiments, the serum replacement is provided to the medium at 5% (v / v) for at least a portion of the first incubation. In some embodiments, the serum replacement is provided to the medium at 5% (v / v) on days 0 and 1. In some embodiments, the serum replacement is provided to the medium at 2% (v / v) for at least a portion of the first incubation. In some embodiments, the serum replacement is provided to the medium at 2% (v / v) on days 2-6. In some embodiments, the serum replacement is provided to the medium at 5% (v / v) on days 0 and 1, and at 2% (v / v) on days 2-6.

[0258] In some embodiments, the medium is further supplemented with a small molecule, such as any of those described above, in some embodiments, the small molecule is selected from the group consisting of a Rho-associated protein kinase (ROCK) inhibitor, an inhibitor of TGF-β / activin-nodal signaling, at least one activator of Sonic Hedgehog (SHH) signaling, an inhibitor of bone morphogenetic protein (BMP) signaling, an inhibitor of glycogen synthase kinase 3β (GSK3β), and combinations thereof.

[0259] In some embodiments, the medium is supplemented with a Rho-associated protein kinase (ROCK) inhibitor for one or more days when the cells are passaged. In some embodiments, the medium is supplemented with a ROCK inhibitor every day the cells are passaged. In some embodiments, the medium is supplemented with a ROCK inhibitor on day 0.

[0260] In some embodiments, the ROCK inhibitor is selected from the group consisting of fasudil, ripasudil, netarsudil, RKI-1447, Y-27632, GSK429286A, Y-30141, and combinations thereof. In some embodiments, the ROCK inhibitor is a small molecule. In some embodiments, the ROCK inhibitor selectively inhibits p160ROCK. In some embodiments, the ROCK inhibitor is Y-27632, which has the following formula: [ka]

[0261] In some embodiments, the medium is supplemented with an inhibitor of TGF-β / activin-nodal signaling. In some embodiments, the medium is supplemented with an inhibitor of TGF-β / activin-nodal signaling until about day 7 (e.g., day 6 or day 7). In some embodiments, the medium is supplemented with an inhibitor of TGF-β / activin-nodal signaling from about day 0 to day 6, inclusive.

[0262] In some embodiments, the inhibitor of TGF-β / activin-nodal signaling is a small molecule. In some embodiments, the inhibitor of TGF-β / activin-nodal signaling can reduce or block transforming growth factor beta (TGFβ) / activin-nodal signaling. In some embodiments, the inhibitor of TGF-β / activin-nodal signaling inhibits ALK4, ALK5, ALK7, or a combination thereof. In some embodiments, the inhibitor of TGF-β / activin-nodal signaling inhibits ALK4, ALK5, and ALK7. In some embodiments, the inhibitor of TGF-β / activin-nodal signaling does not inhibit ALK2, ALK3, ALK6, or a combination thereof. In some embodiments, the inhibitor does not inhibit ALK2, ALK3, or ALK6. In some embodiments, the inhibitor of TGF-β / activin-nodal signaling is SB431542 (e.g., CAS 301836-41-9, molecular formula: C22H18N4O3, and name: 4-[4-(1,3-benzodioxol-5-yl)-5-(2-pyridinyl)-1H-imidazol-2-yl]-benzamide), which has the following formula: [ka]

[0263] In some embodiments, the inhibitor of TGF-β / activin-nodal signaling is a small molecule. In some embodiments, the inhibitor of TGF-β / activin-nodal signaling can reduce or block transforming growth factor beta (TGFβ) / activin-nodal signaling. In some embodiments, the inhibitor of TGF-β / activin-nodal signaling inhibits ALK4, ALK5, ALK7, or a combination thereof. In some embodiments, the inhibitor of TGF-β / activin-nodal signaling inhibits ALK4, ALK5, and ALK7. In some embodiments, the inhibitor of TGF-β / activin-nodal signaling does not inhibit ALK2, ALK3, ALK6, or a combination thereof. In some embodiments, the inhibitor does not inhibit ALK2, ALK3, or ALK6. In some embodiments, the inhibitor of TGF-β / activin-nodal signaling is SB431542 (e.g., CAS 301836-41-9, molecular formula: C22H18N4O3, and name: 4-[4-(1,3-benzodioxol-5-yl)-5-(2-pyridinyl)-1H-imidazol-2-yl]-benzamide), which has the following formula: [ka]

[0264] In some embodiments, the at least one activator of SHH signaling is an activator of the hedgehog receptor smoothened. In some embodiments, the at least one activator of SHH signaling is a small molecule. In some embodiments, the at least one activator of SHH signaling is palmorfamine (e.g., CAS483367-10-8), which has the following formula: [ka]

[0265] In some embodiments, the cells are exposed to purmorphamine at a concentration of about 10 μM. In some embodiments, the cells are exposed to purmorphamine at a concentration of about 10 μM until day 7 (e.g., day 6 or day 7). In some embodiments, the cells are exposed to purmorphamine at a concentration of about 10 μM from about day 0 to about day 6, inclusive.

[0266] In some embodiments, at least one activator of SHH signaling is SHH protein and palmorphamin. In some embodiments, the cells are exposed to a concentration of SHH protein and palmorphamin until about day 7 (e.g., day 6 or day 7). In some embodiments, the cells are exposed to SHH protein and palmorphamin from about day 0 to about day 6, inclusive. In some embodiments, the cells are exposed to a concentration of 100 ng / mL SHH protein and 10 μM palmorphamin until about day 7 (e.g., day 6 or day 7). In some embodiments, the cells are exposed to 100 ng / mL SHH protein and 10 μM palmorphamin from about day 0 to about day 6, inclusive.

[0267] In some embodiments, the medium is supplemented with an inhibitor of BMP signaling. In some embodiments, the medium is supplemented with an inhibitor of BMP signaling until about day 7 (e.g., day 6 or day 7). In some embodiments, the medium is supplemented with an inhibitor of BMP signaling from about day 0 to day 6, inclusive.

[0268] In some embodiments, the inhibitor of BMP signaling is a small molecule. In some embodiments, the inhibitor of BMP signaling is selected from LDN193189 or K02288. In some embodiments, the inhibitor of BMP signaling can inhibit "Small Mothers Against Decapentaplegic" SMAD signaling. In some embodiments, the inhibitor of BMP signaling inhibits ALK1, ALK2, ALK3, ALK6, or a combination thereof. In some embodiments, the inhibitor of BMP signaling inhibits ALK1, ALK2, ALK3, and ALK6. In some embodiments, the inhibitor of BMP signaling inhibits BMP2, BMP4, BMP6, BMP7, and activin cytokine signals, and subsequently inhibits SMAD phosphorylation of Smad1, Smad5, and Smad8. In some embodiments, the inhibitor of BMP signaling is LDN193189. In some embodiments, the inhibitor of BMP signaling is LDN193189, which has the following formula (e.g., IUPAC name 4-(6-(4-(piperazin-1-yl)phenyl)pyrazolo[1,5-a]pyrimidin-3-yl)quinoline, which has the chemical formula C25H22N6): [ka]

[0269] In some embodiments, cells are exposed to LDN193189 at a concentration of about 0.1 μM. In some embodiments, cells are exposed to LDN193189 at a concentration of about 0.1 μM up to about day 7 (e.g., day 6 or day 7). In some embodiments, cells are exposed to LDN193189 at a concentration of about 0.1 μM from about day 0 to about day 6, inclusive.

[0270] In some embodiments, the medium is supplemented with an inhibitor of GSK3β signaling. In some embodiments, the medium is supplemented with an inhibitor of GSK3β signaling until about day 7 (e.g., day 6 or day 7). In some embodiments, the medium is supplemented with an inhibitor of GSK3β signaling from about day 0 to day 6, inclusive.

[0271] In some embodiments, the inhibitor of GSK3β signaling is selected from the group consisting of lithium ion, valproic acid, iodotubercidin, naproxen, famotidine, curcumin, olanzapine, CHIR99012, and combinations thereof. In some embodiments, the inhibitor of GSK3β signaling is a small molecule. In some embodiments, the inhibitor of GSK3β signaling inhibits glycogen synthase kinase 3β enzyme. In some embodiments, the inhibitor of GSK3β signaling inhibits GSK3α. In some embodiments, the inhibitor of GSK3β signaling modulates TGF-β and MAPK signaling. In some embodiments, the inhibitor of GSK3β signaling is an agonist of Wingless / Integrated (Wnt) signaling. In some embodiments, the inhibitor of GSK3β signaling has an IC50=6.7 nM for human GSK3β. In some embodiments, the inhibitor of GSK3β signaling is CHIR99021 (e.g., “3-[3-(2-carboxyethyl)-4-methylpyrrole-2-methylidenyl]-2-indolinone”, or IUPAC name 6-(2-(4-(2,4-dichlorophenyl)-5-(4-methyl-1H-imidazol-2-yl)pyrimidin-2-ylamino)ethylamino)nicotinonitrile), which has the following formula: [ka]

[0272] In some embodiments, cells are exposed to CHIR99021 at a concentration of about 2.0 μM. In some embodiments, cells are exposed to CHIR99021 at a concentration of about 2.0 μM until about day 7 (e.g., day 6 or 7). In some embodiments, cells are exposed to CHIR99021 at a concentration of about 2.0 μM from about day 0 to about day 6, inclusive.

[0273] In some embodiments, at least about 50% of the medium is replaced daily from about day 2 to about day 6. In some embodiments, about 50% of the medium is replaced daily, every other day, or every third day from about day 2 to about day 6. In some embodiments, about 50% of the medium is replaced daily from about day 2 to about day 6. In some embodiments, at least about 75% of the medium is replaced on day 1. In some embodiments, about 100% of the medium is replaced on day 1. In some embodiments, the replacement medium contains about 2-fold concentrated small molecules compared to the concentration of the small molecules in the medium on day 0.

[0274] In some embodiments, the first incubation comprises culturing the pluripotent stem cells in a "basic induction medium." In some embodiments, the first incubation comprises culturing the pluripotent stem cells in a basal induction medium from about day 0 to about day 6. In some embodiments, the first incubation comprises culturing the pluripotent stem cells in a basal induction medium to induce differentiation of the PSCs into floor plate midbrain progenitor cells.

[0275] In some embodiments, the basal induction medium is formulated to contain Neurobasal™ medium and DMEM / F12 medium in a 1:1 ratio, supplemented with N-2 and B27 supplements, non-essential amino acids (NEAA), GlutaMAX™, L-glutamine, β-mercaptoethanol, and insulin. In some embodiments, the basal induction medium is further supplemented with any of the small molecules as described above.

[0276] In some embodiments, cell aggregates (e.g., spheroids) generated after a first incubation of pluripotent stem cells in a non-adherent culture vessel are displaced or dissociated prior to a second incubation of the cells on the substrate (adherent culture).

[0277] In some embodiments, the first incubation is performed to generate cell aggregates (e.g., spheroids) expressing at least one of PAX6 and OTX2. In some embodiments, the first incubation generates cell aggregates (e.g., spheroids) expressing PAX6 and OTX2. In some embodiments, the first incubation generates cell aggregates (e.g., spheroids) at or by about day 7 of the method. In some embodiments, the first incubation generates cell aggregates (e.g., spheroids) expressing at least one of PAX6 and OTX2 at or by about day 7 of the method. In some embodiments, the first incubation generates cell aggregates (e.g., spheroids) expressing PAX6 and OTX2 at or by about day 7 of the method.

[0278] In some embodiments, the cell aggregates (e.g., spheroids) generated by the first incubation are dissociated prior to a second incubation of the cells on the substrate. In some embodiments, the cell aggregates (e.g., spheroids) generated by the first incubation are dissociated to generate a cell suspension. In some embodiments, the cell suspension generated by dissociation is a single cell suspension. In some embodiments, the dissociation occurs when the spheroid cells express at least one of PAX6 and OTX2. In some embodiments, the dissociation occurs when the spheroid cells express PAX6 and OTX2. In some embodiments, the dissociation occurs at about day 7. In some embodiments, the cell aggregates (e.g., spheroids) are dissociated by enzymatic dissociation. In some embodiments, the enzyme is selected from the group consisting of actase, dispase, collagenase, and combinations thereof. In some embodiments, the enzyme comprises actase. In some embodiments, the enzyme is actase. In some embodiments, the enzyme is dispase. In some embodiments, the enzyme is collagenase.

[0279] In some embodiments, the cell aggregates or a cell suspension generated therefrom are transferred to a substrate-coated culture vessel for a second incubation. In some embodiments, the cell aggregates (e.g., spheroids) or a cell suspension generated therefrom are transferred to a substrate-coated culture vessel after dissociation of the cell aggregates (e.g., spheroids). In some embodiments, the transfer is performed immediately after dissociation. In some embodiments, the transfer is performed on about day 7.

[0280] In some embodiments, the cell aggregates (e.g., spheroids) are not dissociated before the second incubation. In some embodiments, the entire cell aggregates (e.g., spheroids) are transferred to a substrate-coated culture vessel for the second incubation. In some embodiments, the transfer is performed when the spheroid cells express at least one of PAX6 and OTX2. In some embodiments, the transfer is performed when the spheroid cells express PAX6 and OTX2. In some embodiments, the transfer is performed at about day 7.

[0281] In some embodiments, the second incubation involves culturing the cells of the spheroids in a culture vessel coated with a substrate comprising laminin, collagen, entactin, heparin sulfate proteoglycan, or a combination thereof, and exposing the cells to (i) an inhibitor of BMP signaling and (ii) an inhibitor of GSK3β signaling starting on day 7, and exposing the cells to (i) brain-derived neurotrophic factor (BDNF), (ii) ascorbic acid, (iii) glial cell line-derived neurotrophic factor (GDNF), (iv) dibutyryl cyclic AMP (dbcAMP), (v) transforming growth factor beta-3 (TGFβ3), and (vi) an inhibitor of Notch signaling starting on day 11. In some embodiments, the method further comprises harvesting the differentiated cells.

[0282] In some embodiments, the substrate-coated culture vessel is a culture vessel having a surface to which cells can adhere. In some embodiments, the substrate-coated culture vessel is a culture vessel having a surface to which a substantial number of cells adhere. In some embodiments, the substrate is a basement membrane protein. In some embodiments, the substrate is laminin, collagen, entactin, heparin sulfate proteoglycan, or a combination thereof. In some embodiments, the substrate is laminin. In some embodiments, the substrate is collagen. In some embodiments, the substrate is entactin. In some embodiments, the substrate is heparin sulfate proteoglycan. In some embodiments, the substrate is a recombinant protein. In some embodiments, the substrate is recombinant laminin. In some embodiments, the substrate-coated culture vessel is exposed to poly-L-ornithine. In some embodiments, the substrate-coated culture vessel is exposed to poly-L-ornithine prior to use for cell culture.

[0283] In some embodiments, the substrate-coated culture vessel allows monolayer cell culture. In some embodiments, cells derived from the cell aggregates (e.g., spheroids) generated by the first incubation are cultured in monolayer culture on the substrate-coated plate. In some embodiments, cells derived from the cell aggregates (e.g., spheroids) generated by the first incubation are cultured to generate a monolayer culture of cells positive for one or more of LMX1A, FOXA2, EN1, CORIN, and combinations thereof. In some embodiments, cells derived from the cell aggregates (e.g., spheroids) generated by the first incubation are cultured to generate a monolayer culture of cells in which at least some of the cells are positive for EN1 and CORIN. In some embodiments, cells derived from the cell aggregates (e.g., spheroids) generated by the first incubation are cultured to generate a monolayer culture of cells in which at least some of the cells are TH+. In some embodiments, at least some of the cells become TH+ by or at about day 25. In some embodiments, cells derived from the cell aggregates (e.g., spheroids) generated by the first incubation are cultured to generate a monolayer culture of cells in which at least some of the cells are TH+FOXA2+. In some embodiments, at least some of the cells become TH+FOXA2+ by or at about day 25.

[0284] In the method, the second incubation involves culturing the cells of the spheroids in a substrate-coated culture vessel under conditions that induce neuronal differentiation of the cells. In some embodiments, the cells of the spheroids are plated on the substrate-coated culture vessel on about day 7.

[0285] In some embodiments, the second incubation is from about day 7 until the cells are harvested. In some embodiments, the cells are harvested after about day 16. In some embodiments, the cells are harvested from about day 16 to about day 30. In some embodiments, the cells are harvested from about day 18 to about day 25. In some embodiments, the cells are harvested on about day 18. In some embodiments, the cells are harvested on about day 25. In some embodiments, the second incubation is from about day 7 to about day 18. In some embodiments, the second incubation is from about day 7 to about day 25.

[0286] In some embodiments, the second incubation involves culturing cells derived from the cell aggregates (eg, spheroids) in culture medium ("medium").

[0287] In some embodiments, the second incubation involves culturing the cells in medium from about day 7 until harvest or recovery. In some embodiments, the cells are cultured in medium to produce determined dopaminergic cells or dopaminergic neurons.

[0288] In some embodiments, the medium is supplemented with a serum replacement that contains minimal non-human derived components (e.g., KnockOut™ serum replacement). In some embodiments, the medium is supplemented with serum replacement from about day 7 to about day 10. In some embodiments, the medium is supplemented with about 2% (v / v) serum replacement. In some embodiments, the medium is supplemented with about 2% (v / v) serum replacement from about day 7 to about day 10.

[0289] In some embodiments, the medium is further supplemented with a small molecule, hi some embodiments, the small molecule is selected from the group consisting of a Rho-associated protein kinase (ROCK) inhibitor, an inhibitor of bone morphogenetic protein (BMP) signaling, an inhibitor of glycogen synthase kinase 3β (GSK3β), and combinations thereof.

[0290] In some embodiments, the medium is supplemented with a Rho-associated protein kinase (ROCK) inhibitor for one or more days when the cells are passaged. In some embodiments, the medium is supplemented with a ROCK inhibitor every day the cells are passaged. In some embodiments, the medium is supplemented with a ROCK inhibitor on day 7, day 16, day 20, or a combination thereof. In some embodiments, the medium is supplemented with a ROCK inhibitor on day 7. In some embodiments, the medium is supplemented with a ROCK inhibitor on day 16. In some embodiments, the medium is supplemented with a ROCK inhibitor on day 20. In some embodiments, the medium is supplemented with a ROCK inhibitor on days 7 and 16. In some embodiments, the medium is supplemented with a ROCK inhibitor on days 16 and 20. In some embodiments, the medium is supplemented with a ROCK inhibitor on days 7, 16, and 20.

[0291] In some embodiments, the ROCK inhibitor is fasudil, ripasudil, netarsudil, RKI-1447, Y-27632, GSK429286A, Y-30141, or a combination thereof. In some embodiments, the ROCK inhibitor is a small molecule. In some embodiments, the ROCK inhibitor selectively inhibits p160ROCK. In some embodiments, the ROCK inhibitor is Y-27632, which has the following formula: [ka]

[0292] In some embodiments, the cells are exposed to Y-27632 at a concentration of about 10 μM. In some embodiments, the cells are exposed to Y-27632 at a concentration of about 10 μM on day 7, day 16, day 20, or a combination thereof. In some embodiments, the cells are exposed to Y-27632 at a concentration of about 10 μM on day 7. In some embodiments, the cells are exposed to Y-27632 at a concentration of about 10 μM on day 16. In some embodiments, the cells are exposed to Y-27632 at a concentration of about 10 μM on day 20. In some embodiments, the cells are exposed to Y-27632 at a concentration of about 10 μM on days 7 and 16. In some embodiments, the cells are exposed to Y-27632 at a concentration of about 10 μM on days 16 and 20. In some embodiments, the cells are exposed to Y-27632 at a concentration of about 10 μM on days 7, 16, and 20.

[0293] In some embodiments, the medium is supplemented with an inhibitor of BMP signaling. In some embodiments, the medium is supplemented with an inhibitor of BMP signaling from about day 7 to about day 11 (e.g., day 10 or day 11). In some embodiments, the medium is supplemented with an inhibitor of BMP signaling from about day 7 to day 10 (inclusive).

[0294] In some embodiments, the inhibitor of BMP signaling is a small molecule. In some embodiments, the inhibitor of BMP signaling is LDN193189 or K02288. In some embodiments, the inhibitor of BMP signaling can inhibit "Small Mothers Against Decapentaplegic" SMAD signaling. In some embodiments, the inhibitor of BMP signaling inhibits ALK1, ALK2, ALK3, ALK6, or a combination thereof. In some embodiments, the inhibitor of BMP signaling inhibits ALK1, ALK2, ALK3, and ALK6. In some embodiments, the inhibitor of BMP signaling inhibits BMP2, BMP4, BMP6, BMP7, and activin cytokine signals, and subsequently inhibits SMAD phosphorylation of Smad1, Smad5, and Smad8. In some embodiments, the inhibitor of BMP signaling is LDN193189. In some embodiments, the inhibitor of BMP signaling is LDN193189, which has the following formula (e.g., IUPAC name 4-(6-(4-(piperazin-1-yl)phenyl)pyrazolo[1,5-a]pyrimidin-3-yl)quinoline, which has the chemical formula C25H22N6): [ka]

[0295] In some embodiments, the cells are exposed to LDN193189 at a concentration of about 0.1 μM. In some embodiments, the cells are exposed to LDN193189 at a concentration of about 0.1 μM from about day 7 to about day 11 (e.g., day 10 or day 11). In some embodiments, the cells are exposed to LDN193189 at a concentration of about 0.1 μM from about day 7 to about day 10, inclusive.

[0296] In some embodiments, the medium is supplemented with an inhibitor of GSK3β signaling. In some embodiments, the medium is supplemented with an inhibitor of GSK3β signaling from about day 7 to about day 13 (e.g., day 12 or day 13). In some embodiments, the medium is supplemented with an inhibitor of GSK3β signaling from about day 7 to day 12 (inclusive).

[0297] In some embodiments, the inhibitor of GSK3β signaling is selected from lithium ion, valproic acid, iodotubercidin, naproxen, famotidine, curcumin, olanzapine, CHIR99012, or a combination thereof. In some embodiments, the inhibitor of GSK3β signaling is a small molecule. In some embodiments, the inhibitor of GSK3β signaling inhibits glycogen synthase kinase 3β enzyme. In some embodiments, the inhibitor of GSK3β signaling inhibits GSK3α. In some embodiments, the inhibitor of GSK3β signaling modulates TGF-β and MAPK signaling. In some embodiments, the inhibitor of GSK3β signaling is an agonist of Wingless / Integrated (Wnt) signaling. In some embodiments, the inhibitor of GSK3β signaling has an IC50=6.7 nM for human GSK3β. In some embodiments, the inhibitor of GSK3β signaling is CHIR99021 (e.g., “3-[3-(2-carboxyethyl)-4-methylpyrrole-2-methylidenyl]-2-indolinone”, or IUPAC name 6-(2-(4-(2,4-dichlorophenyl)-5-(4-methyl-1H-imidazol-2-yl)pyrimidin-2-ylamino)ethylamino)nicotinonitrile), which has the following formula: [ka]

[0298] In some embodiments, the cells are exposed to CHIR99021 at a concentration of about 2.0 μM. In some embodiments, the cells are exposed to CHIR99021 at a concentration of about 2.0 μM from about day 7 to about day 13 (e.g., day 12 or day 13). In some embodiments, the cells are exposed to CHIR99021 at a concentration of about 2.0 μM from about day 7 to about day 12, inclusive.

[0299] In some embodiments, the medium is supplemented with Brain-Derived Neurotrophic Factor (BDNF). In some embodiments, the medium is supplemented with BDNF beginning at about day 11. In some embodiments, the medium is supplemented with BDNF from about day 11 until harvesting or recovery. In some embodiments, the medium is supplemented with BDNF from about day 11-18. In some embodiments, the medium is supplemented with BDNF from about day 11-25.

[0300] In some embodiments, the medium is supplemented with about 20 ng / mL of BDNF beginning at about day 11. In some embodiments, the medium is supplemented with 20 ng / mL of BDNF from about day 11 until harvest or recovery. In some embodiments, the medium is supplemented with about 20 ng / mL of BDNF from about day 11-18. In some embodiments, the medium is supplemented with about 20 ng / mL of BDNF from about day 11-25.

[0301] In some embodiments, the medium is supplemented with glial cell line derived neurotrophic factor (GDNF). In some embodiments, the medium is supplemented with GDNF beginning at about day 11. In some embodiments, the medium is supplemented with GDNF from about day 11 until harvest or recovery. In some embodiments, the medium is supplemented with GDNF from about day 11-18. In some embodiments, the medium is supplemented with GDNF from about day 11-25.

[0302] In some embodiments, the medium is supplemented with about 20 ng / mL of GDNF beginning at about day 11. In some embodiments, the medium is supplemented with 20 ng / mL of GDNF from about day 11 until harvest or recovery. In some embodiments, the medium is supplemented with about 20 ng / mL of GDNF from about day 11-18. In some embodiments, the medium is supplemented with about 20 ng / mL of GDNF from about day 11-25.

[0303] In some embodiments, the medium is supplemented with ascorbic acid. In some embodiments, the medium is supplemented with ascorbic acid beginning at about day 11. In some embodiments, the medium is supplemented with ascorbic acid from about day 11 until harvest or recovery. In some embodiments, the medium is supplemented with ascorbic acid from about day 11-18. In some embodiments, the medium is supplemented with ascorbic acid from about day 11-25.

[0304] In some embodiments, the medium is supplemented with about 0.2 mM ascorbic acid beginning at about day 11. In some embodiments, the medium is supplemented with about 0.2 mM ascorbic acid from about day 11 until harvest or recovery. In some embodiments, the medium is supplemented with about 0.2 mM ascorbic acid from about day 11-18. In some embodiments, the medium is supplemented with about 0.2 mM ascorbic acid from about day 11-25.

[0305] In some embodiments, the medium is supplemented with dibutyryl cyclic AMP (dbcAMP). In some embodiments, the medium is supplemented with dbcAMP beginning at about day 11. In some embodiments, the medium is supplemented with dbcAMP from about day 11 until harvest or recovery. In some embodiments, the medium is supplemented with dbcAMP from about day 11-18. In some embodiments, the medium is supplemented with dbcAMP from about day 11-25.

[0306] In some embodiments, the medium is supplemented with about 0.5 mM dbcAMP beginning at about day 11. In some embodiments, the medium is supplemented with about 0.5 mM dbcAMP from about day 11 until harvest or recovery. In some embodiments, the medium is supplemented with about 0.5 mM dbcAMP from about day 11-18. In some embodiments, the medium is supplemented with about 0.5 mM dbcAMP from about day 11-25.

[0307] In some embodiments, the medium is supplemented with transforming growth factor beta 3 (TGFβ3). In some embodiments, the medium is supplemented with TGFβ3 beginning at about day 11. In some embodiments, the medium is supplemented with TGFβ3 from about day 11 until harvest or recovery. In some embodiments, the medium is supplemented with TGFβ3 from about day 11-18. In some embodiments, the medium is supplemented with TGFβ3 from about day 11-25.

[0308] In some embodiments, the medium is supplemented with about 1 ng / mL TGFβ3 beginning at about day 11. In some embodiments, the medium is supplemented with 1 ng / mL TGFβ3 from about day 11 until harvest or recovery. In some embodiments, the medium is supplemented with about 1 ng / mL TGFβ3 from about day 11-18. In some embodiments, the medium is supplemented with about 1 ng / mL TGFβ3 from about day 11-25.

[0309] In some embodiments, the medium is supplemented with an inhibitor of Notch signaling. In some embodiments, the medium is supplemented with an inhibitor of Notch signaling beginning at about day 11. In some embodiments, the medium is supplemented with an inhibitor of Notch signaling from about day 11 until harvest or recovery. In some embodiments, the medium is supplemented with an inhibitor of Notch signaling from about day 11-18. In some embodiments, the medium is supplemented with an inhibitor of Notch signaling from about day 11-25.

[0310] In some embodiments, the inhibitor of Notch signaling is selected from Cowanin, PF-03084014, L685458, LY3039478, DAPT, or a combination thereof. In some embodiments, the inhibitor of Notch signaling inhibits gamma secretase. In some embodiments, the inhibitor of Notch signaling is a small molecule. In some embodiments, the inhibitor of Notch signaling is DAPT, which has the formula: [ka]

[0311] In some embodiments, the medium is supplemented with about 10 μM DAPT beginning at about day 11. In some embodiments, the medium is supplemented with 10 μM DAPT from about day 11 until harvest or recovery. In some embodiments, the medium is supplemented with about 10 μM DAPT from about day 11-18. In some embodiments, the medium is supplemented with about 10 μM DAPT from about day 11-25.

[0312] In some embodiments, beginning at about day 11, the medium is supplemented with about 20 ng / mL BDNF, about 20 ng / mL GDNF, about 0.2 mM ascorbic acid, about 0.5 mM dbcAMP, about 1 ng / mL TGFβ3, and about 10 μM DAPT. In some embodiments, from about day 11 until harvest or recovery, the medium is supplemented with about 20 ng / mL BDNF, about 20 ng / mL GDNF, about 0.2 mM ascorbic acid, about 0.5 mM dbcAMP, about 1 ng / mL TGFβ3, and about 10 μM DAPT. In some embodiments, from about day 11 to day 18, the medium is supplemented with about 20 ng / mL BDNF, about 20 ng / mL GDNF, about 0.2 mM ascorbic acid, about 0.5 mM dbcAMP, about 1 ng / mL TGFβ3, and about 10 μM DAPT. In some embodiments, from about day 11 to about day 25, the medium is supplemented with about 20 ng / mL BDNF, about 20 ng / mL GDNF, about 0.2 mM ascorbic acid, about 0.5 mM dbcAMP, about 1 ng / mL TGFβ3, and about 10 μM DAPT.

[0313] In some embodiments, serum replacement is provided to the medium from about day 7 to about day 10. In some embodiments, serum replacement is provided to the medium at 2% (v / v) from day 7 to day 10.

[0314] In some embodiments, from about day 7 to about day 16, at least about 50% of the medium is replaced daily. In some embodiments, from about day 7 to about day 16, about 50% of the medium is replaced daily, every other day, or every second day. In some embodiments, from about day 7 to about day 16, about 50% of the medium is replaced daily. In some embodiments, starting at about day 17, at least about 50% of the medium is replaced every other day. In some embodiments, starting at about day 17, at least about 50% of the medium is replaced every other day. In some embodiments, starting at about day 17, about 50% of the medium is replaced every other day. In some embodiments, starting at about day 17, about 50% of the medium is replaced every other day. In some embodiments, starting at about day 17, about 50% of the medium is replaced every other day. In some embodiments, the replacement medium contains small molecules concentrated about 2-fold compared to the concentration of small molecules in the medium on day 0.

[0315] In some embodiments, the second incubation involves culturing cells derived from the cell aggregates (e.g., spheroids) in a "basic induction medium." In some embodiments, the second incubation involves culturing cells derived from the cell aggregates (e.g., spheroids) in a "maturation medium." In some embodiments, the second incubation involves culturing cells derived from the cell aggregates (e.g., spheroids) in a basal induction medium and then in a maturation medium.

[0316] In some embodiments, the second incubation involves culturing the cells in basal induction medium from about day 7 to about day 10. In some embodiments, the second incubation involves culturing the cells in maturation medium beginning at about day 11. In some embodiments, the second incubation involves culturing the cells in basal induction medium from about day 7 to about day 10, and then in maturation medium beginning at about day 11. In some embodiments, the cells are cultured in maturation medium to produce committed dopaminergic cells or dopaminergic neurons.

[0317] In some embodiments, the basal induction medium is formulated to contain Neurobasal™ medium and DMEM / F12 medium in a 1:1 ratio, supplemented with N-2 and B27 supplements, non-essential amino acids (NEAA), GlutaMAX™, L-glutamine, β-mercaptoethanol, and insulin.

[0318] In some embodiments, maturation medium is formulated to contain Neurobasal™ medium supplemented with N-2 and B27 supplements, non-essential amino acids (NEAA), and GlutaMAX™.

[0319] In some embodiments, the cells are cultured in basal induction medium from about day 7 to about day 11 (e.g., day 10 or day 11). In some embodiments, the cells are cultured in basal induction medium from about day 7 to about day 10, inclusive. In some embodiments, the cells are cultured in maturation medium beginning on about day 11. In some embodiments, the cells are cultured in basal induction medium from about day 7 to about day 10, and then the cells are cultured in maturation medium beginning on about day 11. In some embodiments, the cells are cultured in maturation medium from about day 11 until the cells are harvested or collected. In some embodiments, the cells are harvested between day 16 and day 27. In some embodiments, the cells are harvested between day 18 and day 25. In some embodiments, the cells are harvested on day 18. In some embodiments, the cells are harvested on day 25.

[0320] In some embodiments, the test cells are derived from an in vitro population derived from a culture of cells differentiated from pluripotent cells that are subjected to a differentiation protocol to induce differentiation of PSCs, e.g., iPSCs, into dopaminergic neurons, e.g., according to any of the methods described herein.

[0321] In some embodiments, the cells in the second differentiation state are anywhere between days 15-21 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are on or before day 14, on or before day 13, on or before day 12, or on or before day 11 of the differentiation protocol. In some embodiments, the cells in the third differentiation state are on or after day 22, on or after day 23, on or before day 24, or on or after day 25 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are on or before day 14, on or before day 13, on or before day 12, or on or before day 11 of the differentiation protocol, the cells in the second differentiation state are anywhere between days 15-21 of the differentiation protocol, and the cells in the third differentiation state are on or after day 22, on or after day 23, on or after day 24, or on or after day 25 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are on or before day 10-14 of the differentiation protocol. In some embodiments, cells in the third differentiation state are at or after day 25 of the differentiation protocol.

[0322] In some embodiments, the cells in the second differentiation state are at any of days 16-18 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are at or before day 15, day 14, day 13, day 12, or day 11 of the differentiation protocol. In some embodiments, the cells in the third differentiation state are at or after day 19, day 20, day 11, day 22, day 23, day 24, or day 25 of the differentiation protocol. In some embodiments, cells in a first differentiation state are on or before day 15, on or before day 14, on or before day 13, on or before day 12, or on or before day 11 of the differentiation protocol, cells in a second differentiation state are on any of days 16-18 of the differentiation protocol, and cells in a third differentiation state are on or after day 19, on or after day 20, on or after day 21, on or after day 22, on or after day 23, on or after day 24, on or after day 25 of the differentiation protocol. In some embodiments, cells in a first differentiation state are on any of days 11-13 of the differentiation protocol. In some embodiments, cells in a third differentiation state are on or after day 25 of the differentiation protocol.

[0323] In some embodiments, the cells in the second differentiation state are on any of days 17-19 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are on or before day 16, 15, 14, 13, 12, or 11 of the differentiation protocol. In some embodiments, the cells in the third differentiation state are on or after day 20, 11, 22, 23, 24, or 25 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are on or before day 16, on or before day 15, on or before day 14, on or before day 13, on or before day 12, or on or before day 11 of the differentiation protocol, the cells in the second differentiation state are on any of days 17-19 of the differentiation protocol, and the cells in the third differentiation state are on or after day 20, on or after day 21, on or after day 22, on or after day 23, on or after day 24, on or after day 25 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are on any of days 12-14 of the differentiation protocol. In some embodiments, the cells in the third differentiation state are on or after day 25 of the differentiation protocol.

[0324] In some embodiments, the cells in the second differentiation state are at any of days 15-17 of the differentiation protocol. In some embodiments, the cells in the first differentiation state are at or before day 14, day 13, day 12, or day 11 of the differentiation protocol. In some embodiments, the cells in the third differentiation state are at or after day 18, day 19, day 20, day 11, day 22, day 23, day 24, or day 25 of the differentiation protocol. In some embodiments, cells in the first differentiation state are on or before day 14, on or before day 13, on or before day 12, or on or before day 11 of the differentiation protocol, cells in the second differentiation state are on any of days 15-17 of the differentiation protocol, and cells in the third differentiation state are on or after day 18, on or after day 19, on or after day 20, on or after day 21, on or after day 22, on or after day 23, on or after day 24, on or after day 25 of the differentiation protocol. In some embodiments, cells in the first differentiation state are on any of days 10-12 of the differentiation protocol. In some embodiments, cells in the third differentiation state are on or after day 30 of the differentiation protocol.

[0325] 4. Pluripotent stem cells In some embodiments, the test cell and / or the reference cell population are produced from pluripotent stem cells. Various sources of pluripotent stem cells can be used, including embryonic stem (ES) cells and induced pluripotent stem cells (iPSCs). In some embodiments, the pluripotent stem cells are iPSCs. iPSCs can be made by a process known as reprogramming, in which non-pluripotent cells are effectively "dedifferentiated" to an embryonic stem cell-like state by engineering them to express genes such as OCT4, SOX2 and KLF4 (Takahashi and Yamanaka Cell (2006) 126:663-76). In some embodiments, the pluripotent stem cells are iPSCs artificially derived from a non-pluripotent cell of a subject. In some embodiments, the non-pluripotent cell is a fibroblast. In some embodiments, the subject is a human. In some embodiments, the subject is a human with Parkinson's disease.

[0326] In some embodiments, pluripotency refers to cells that have the ability, under appropriate conditions, to give rise to progeny that can undergo differentiation into cell types that collectively exhibit characteristics associated with cell lineages from the three germ layers (endoderm, mesoderm, and ectoderm). Pluripotent stem cells can contribute to tissues of prenatal, postnatal, or adult organisms. Standard art-accepted tests, such as the ability to form teratomas in 8-12 week old SCID mice, can be used to establish the pluripotency of a cell population. However, identification of various pluripotent stem cell characteristics can also be used to identify pluripotent cells. In some embodiments, pluripotent stem cells can be distinguished from other cells by specific characteristics, including the expression or non-expression of certain combinations of molecular markers. More specifically, human pluripotent stem cells may express at least some, and optionally all, of the markers from the following non-limiting list: SSEA-3, SSEA-4, TRA-1-60, TRA-1-81, TRA-2-49 / 6E, ALP, Sox2, E-cadherin, UTF-1, Oct4, Lin28, Rex1, and Nanog. In some embodiments, a characteristic of a pluripotent stem cell is a cell morphology associated with a pluripotent stem cell.

[0327] Methods for generating iPSCs are known. For example, mouse iPSCs were reported in 2006 (Takahashi and Yamanaka), and human iPSCs were reported in late 2007 (Takahashi et al. and Yu et al.). Mouse iPSCs exhibit key characteristics of pluripotent stem cells, including expression of stem cell markers, formation of tumors containing cells from all three germ layers, and the ability to contribute to many different tissues when injected into mouse embryos at the very early stage of development. Human iPSCs also express stem cell markers and can generate cells characteristic of all three germ layers.

[0328] In some embodiments, non-pluripotent cells (e.g., fibroblasts) derived from a patient with Parkinson's disease (PD) are reprogrammed to become iPSCs before differentiating into neural cells. In some embodiments, fibroblasts can be reprogrammed to iPSCs by transforming the fibroblasts with genes (OCT4, SOX2, NANOG, LIN28, and KLF4) cloned into plasmids (see, e.g., Yu, et al., Science DOI:10.1126 / science.1172482). In some embodiments, non-pluripotent fibroblasts derived from a patient with PD are reprogrammed to differentiate into committed dopaminergic cells and / or dopaminergic neurons, for example, by reprogramming the cells using non-integrating Sendai virus (e.g., using the CTS™ CytoTune™-iPS 2.1 Sendai Reprogramming Kit). In some embodiments, the resulting differentiated cells are then administered to the patient from whom they were derived in an autologous stem cell transplant. In some embodiments, the PSCs (e.g., iPSCs) are allogeneic to the subject being treated, i.e., the PSCs are derived from an individual different from the subject to which the differentiated cells are administered. In some embodiments, non-pluripotent cells (e.g., fibroblasts) derived from another individual (e.g., an individual without a neurodegenerative disorder such as Parkinson's disease) are reprogrammed to become iPSCs prior to differentiation into committed dopaminergic cells and / or dopaminergic neurons. In some embodiments, reprogramming is achieved, at least in part, by using a non-integrating Sendai virus to reprogram the cells (e.g., using the CTS™ CytoTune™-iPS 2.1 Sendai Reprogramming Kit). In some embodiments, the resulting differentiated cells are then administered to an individual that is not the same individual from which the differentiated cells were derived (e.g., allogeneic cell therapy or allogeneic cell transplantation).

[0329] In any of the embodiments provided, the PSCs described herein can be genetically engineered to be less immunogenic. Methods for reducing immunogenicity are known and include eradicating expression of polymorphic HLA-A / -B / -C and HLA class II molecules, and introducing immune regulators PD-L1, HLA-G, and CD47 into the AAVS1 safe harbor locus in differentiated cells (Han et al., PNAS (2019) 116 (21): 10441-46). Thus, in some embodiments, the PSCs described herein are genetically engineered to delete the highly polymorphic HLA-A / -B / -C genes and introduce immune regulators such as PD-L1, HLA-G, and / or CD47 into the AAVS1 safe harbor locus.

[0330] In some embodiments, PSCs (e.g., iPSCs) are cultured in the absence of feeder cells until they reach 80-90% confluence, at which point they are harvested and further cultured for differentiation (day 0). In some aspects, once iPSCs reach 80-90% confluence, they are washed with phosphate buffered saline (PBS) and subjected to enzymatic dissociation, such as with Accutase™, until the cells can be easily detached from the surface of the culture vessel. The dissociated iPSCs are then resuspended in media for downstream differentiation into desired cell types, such as committed dopaminergic cells and / or dopaminergic neurons.

[0331] In some embodiments, the PSCs are resuspended in a basal induction medium. In some embodiments, the basal induction medium is formulated to contain Neurobasal™ medium and DMEM / F12 medium in a 1:1 ratio, supplemented with N-2 and B27 supplements, non-essential amino acids (NEAA), GlutaMAX™, L-glutamine, β-mercaptoethanol, and insulin. In some embodiments, the basal induction medium is further supplemented with serum replacement, Rho-associated protein kinase (ROCK) inhibitors, and various small molecules for differentiation. In some embodiments, the PSCs are resuspended in the same medium in which they are cultured for at least a portion of the first incubation.

[0332] 5. Exemplary Features of Classified Cells In some embodiments, the cells of the in vitro population of cells identified as having a desired differentiation state, e.g., a second differentiation state, can survive when administered in vivo, e.g., to an animal model. In some embodiments, the cells of the identified in vitro population survive after transplantation into an animal or human subject. In some embodiments, the cells of the identified in vitro population of cells have a therapeutic effect for treating a disease or condition in an animal model. In some embodiments, the cells of the identified in vitro population of cells have a therapeutic effect for treating a disease or condition in a human patient. In some embodiments, the transplanted cells improve or reverse the symptoms of a disease or condition.

[0333] In some embodiments, the cells of the in vitro population of cells identified as having a desired differentiation state, for example a second differentiation state that may be that of a determined dopaminergic neuronal cell, express a marker of midbrain dopaminergic neurons, for example FOXA2 or tyrosine hydroxylase (TH). In some embodiments, the cells express TH (TH+). In some embodiments, the cells express FOXA2 (FOXA2+). In some embodiments, the cells express TH and FOXA2 (TH+FOXA2+).

[0334] In some embodiments, the cells of the identified in vitro population of cells are determined to be or are capable of becoming dopaminergic neurons, i.e., are determined to be dopaminergic cells, as confirmed based on one or more characteristics indicating that the cells may have functional activity of dopaminergic neurons but may not yet express or express at high levels markers of dopaminergic neurons. For example, the cells may exhibit lower levels of TH than dopaminergic neurons, but still exhibit one or more characteristics of the determined dopaminergic cells indicating that the cells may have functional activity of dopaminergic neurons. In some embodiments, the one or more characteristics include the activity of surviving, engrafting, and / or innervating other cells when administered in vivo, for example, to an animal model. In some embodiments, the cells of the identified in vitro population are capable of innervating host tissue after transplantation into an animal or human subject. In some embodiments, the cells of the identified in vitro population exhibit neurite outgrowth after transplantation into an animal or human subject. In some embodiments, the cells of the identified in vitro population survive following transplantation into an animal or human subject. In some embodiments, the cells of the identified in vitro population engraft following transplantation into an animal or human subject.

[0335] In some embodiments, the cells of the in vitro population of identified cells have a therapeutic effect for treating a neurodegenerative disease in an animal model of the neurodegenerative disease. In some embodiments, the neurodegenerative disease is Parkinson's disease. Any suitable animal model of Parkinson's disease can be used for screening. In some embodiments, the animal model is a lesion model in which 6-hydroxydopamine (6-OHDA) is unilaterally stereotaxically injected into the substantia nigra of the animal. In some embodiments, the animal model is a lesion model in which 6-OHDA is unilaterally stereotaxically injected into the medial forebrain bundle of the animal. In some embodiments, the cells are transplanted into the substantia nigra of the animal model. In some embodiments, a behavioral assay is performed to screen the therapeutic effect of transplantation into the animal model. In some embodiments, the behavioral assay includes monitoring amphetamine-induced rotation behavior. In some embodiments, the cells reduce, decrease or reverse Parkinson's disease model brain lesions in the model.

[0336] In some embodiments, the cells of the in vitro population of identified cells have a therapeutic effect for treating a neurodegenerative disease, including in a human patient. In some embodiments, the transplanted cells improve or reverse symptoms of the neurodegenerative disease. In some embodiments, the neurodegenerative disease is Parkinson's disease. In some embodiments, the cells improve Parkinson's symptoms when transplanted into a subject in need thereof, e.g., the substantia nigra of a patient.

[0337] D. Gene expression levels In some embodiments, the gene expression level of any of the test cells or reference cell populations described herein, for example, is determined based on the level of a gene product synthesized using information encoded by one or more genes. In some embodiments, a gene product is any biomolecule assembled, generated, and / or synthesized using information encoded by a gene, and may include polynucleotides and / or polypeptides. In some embodiments, assessing, measuring, and / or determining gene expression includes determining or measuring the level, amount, or concentration of a gene product. In some embodiments, the level, amount, or concentration of a gene product may be transformed (e.g., normalized) or analyzed directly (e.g., raw).

[0338] In some embodiments, the gene product comprises a protein, i.e., a polypeptide encoded and / or expressed by a gene. In certain embodiments, the gene product encodes a protein that is localized and / or exposed on the surface of a cell. In some embodiments, the protein is a soluble protein. In certain embodiments, the protein is secreted by the cell. In certain embodiments, gene expression is the amount, level, and / or concentration of a protein encoded by a gene. In certain embodiments, one or more protein gene products are measured by any suitable means. Suitable methods for assessing, measuring, determining, and / or quantifying the level, amount, or concentration of one or more protein gene products include detection by immunoassays, nucleic acid-based or protein-based aptamer technology, HPLC (high performance liquid chromatography), peptide sequencing (e.g., Edman degradation sequencing or mass spectrometry (e.g., MS / MS), optionally coupled to HPLC), and microarray adaptations of any of the foregoing, including nucleic acid, antibody, or protein-protein (i.e., non-antibody) arrays. In some embodiments, an immunoassay is or includes a method or assay that detects a protein based on an immunological reaction, for example, by detecting binding of an antibody or an antigen-binding antibody fragment to a gene product. Immunoassays include quantitative immunocytochemistry or immunohistochemistry, ELISA (including direct, indirect, sandwich, competitive, multiple, and portable ELISA (see, e.g., U.S. Pat. No. 7,510,687), Western blotting (including one-dimensional, two-dimensional, or higher order blotting or other chromatographic means, optionally including peptide sequencing), enzyme immunoassay (EIA), RIA (radioimmunoassay), and SPR (surface plasmon resonance).

[0339] In certain embodiments, the gene product is a polynucleotide, such as an mRNA or a protein, encoded by a gene. In some embodiments, the gene product is a polynucleotide expressed and / or encoded by a gene. In certain embodiments, the polynucleotide is an RNA. In some embodiments, the gene product is a messenger RNA (mRNA), a transfer RNA (tRNA), a ribosomal RNA, a small nuclear RNA, a small nucleolar RNA, an antisense RNA, a long non-coding RNA, a microRNA, a Piwi-interacting RNA, a small interfering RNA, and / or a short hairpin RNA. In certain embodiments, the gene product is an mRNA.

[0340] In certain embodiments, assessing, measuring, determining, and / or quantifying the amount or level of an RNA gene product comprises generating, polymerizing, and / or obtaining a cDNA polynucleotide and / or a cDNA oligonucleotide from the RNA gene product. In certain embodiments, the RNA gene product is assessed, measured, determined, and / or quantified by directly assessing, measuring, determining, and / or quantifying the cDNA polynucleotide and / or cDNA oligonucleotide derived from the RNA gene product.

[0341] In certain embodiments, the amount or level of a polynucleotide in a sample may be assessed, measured, determined, and / or quantified by any suitable means. For example, in some embodiments, the amount or level of a polynucleotide gene product may be assessed, measured, determined, and / or quantified by: polymerase chain reaction (PCR), including reverse transcriptase (rt) PCR, droplet digital PCR, and real-time and quantitative PCR (qPCR) methods (including, for example, TAQMAN®, Molecular Beacons, LIGHTUP™, SCORPION™, SIMPLEPROBES®; see, for example, U.S. Pat. No. 5,538,848 ... Nos. 5,925,517, 6,174,670, 6,329,144, 6,326,145, and 6,635,427); Northern blotting; Southern blotting of, e.g., reverse transcription products and derivatives; array-based methods, including blot arrays, microarrays, or in situ synthesized arrays; and sequencing, e.g., sequencing by synthesis, pyrosequencing, dideoxysequencing, or ligation, or shendure sequencing. et al., Nat. Rev. Genet. 5:335-44 (2004) or other methods such as those discussed in Nowrousian, Euk. Cell 9(9):1300-1310 (2010) (including specific platforms such as HELICOS®, ROCHE® 454, ILLUMINA® / SOLEXA®, ABI SOLiD®, and POLONATOR® sequencing). In certain embodiments, the levels of nucleic acid gene products are measured by quantitative PCR (qPCR) methods such as qRT-PCR. In some embodiments, qRT-PCR uses a set of three nucleic acids for each gene, which contain primer pairs with probes that bind between the regions of the target nucleic acid to which the primers bind, which are commercially known as TAQMAN® assays.

[0342] In certain embodiments, the expression of two or more of the genes is measured or evaluated simultaneously. In certain embodiments, multiplex PCR, such as multiplex rt-PCR evaluation or multiplex quantitative PCR (qPCR), is used to measure, determine, and / or quantify the level, amount, or concentration of two or more gene products. In some embodiments, microarrays (e.g., AFFYMETRIX®, AGILENT®, and ILLUMINA®-style arrays) are used to evaluate, measure, determine, and / or quantify the level, amount, or concentration of two or more gene products. In some embodiments, microarrays are used to evaluate, measure, determine, and / or quantify the level, amount, or concentration of cDNA polynucleotides derived from RNA gene products. In some embodiments, the expression of one or more gene products, such as polynucleotide gene products, is determined by sequencing the gene products and / or by sequencing cDNA polynucleotides derived from the gene products. In some embodiments, sequencing is performed by non-Sanger sequencing methods and / or next generation sequencing (NGS) technologies. Examples of next generation sequencing technologies include Massively Parallel Signature Sequencing (MPSS), polony sequencing, pyrosequencing, reversible dye terminator sequencing, SOLiD sequencing, ion semiconductor sequencing, DNA nanoball sequencing, heliscope single molecule sequencing, single molecule real time (SMRT) sequencing, single molecule real time (RNAP) sequencing, and nanopore DNA sequencing.

[0343] In some embodiments, the NGS technology is RNA sequencing (RNA-Seq). In certain embodiments, the expression of one or more polynucleotide gene products is measured, determined, and / or quantified by RNA-Seq. RNA-Seq, also called whole transcriptome shotgun sequencing, determines the presence and amount of RNA in a sample. RNA sequencing methods are compatible with the most common DNA sequencing platforms [such as HiSeq systems (Illumina), 454 Genome Sequencer FLX System (Roche), Applied Biosystems SOLiD (Life Technologies) and IonTorrent (Life Technologies)]. These platforms require RNA to be reverse transcribed into cDNA first. Conversely, the single molecule sequencer HeliScope (Helicos BioSciences) can use RNA as a template for sequencing. Proof of principle has also been shown for direct RNA sequencing on the PacBio RS platform (Pacific Bioscience). In some embodiments, one or more RNA gene products are evaluated, measured, determined, and / or quantified by RNA-seq. In some embodiments, the RNA-seq is tag-based RNA-seq. In tag-based methods, each transcript is represented by a unique tag. Initially, tag-based approaches were developed as sequence-based methods to measure transcript abundance and identify differentially expressed genes, assuming that the number of tags (count) directly corresponds to the abundance of mRNA molecules. Reducing the complexity of the sample obtained by sequencing defined regions was essential to make Sanger-based methods affordable. When NGS technology became available, it became possible to generate a large number of reads, which facilitated differential gene expression analysis. Tag-based methods do not encounter transcript length bias in quantification of gene expression levels, as observed for shotgun methods. All tag-based methods are, by definition, strand-specific.In certain embodiments, one or more RNA gene products are assessed, measured, determined, and / or quantified by tag-based RNA-seq.

[0344] In some embodiments, RNA-seq is shotgun RNA-seq. Numerous protocols have been described for shotgun RNA-seq, which have many steps in common: fragmentation (which may be performed at the RNA or cDNA level), conversion of RNA to cDNA (performed by oligo-dT or random primers), second strand synthesis, ligation of adapter sequences at the 3' and 5' ends (at the RNA or DNA level), and final amplification. In some embodiments, RNA-seq is performed by analysing polyadenylated RNA molecules (mainly mRNA, but also some lncRNA, snRNA, etc.) if poly(A)+RNA is selected before fragmentation. Fragmentation may focus only on RNAs (oRNAs, pseudogenes, and even histones) or, if no selection is performed, may also include non-polyadenylated RNAs. In the latter case, ribosomal RNA (>80% of the total RNA pool) must be depleted before fragmentation. Thus, differences in the capture of the mRNA part of the transcriptome are evident with partial overlap in the types of transcripts detected. Furthermore, different protocols may affect the abundance and distribution of sequenced reads. This makes it difficult to compare the results of experiments with different library preparation protocols.

[0345] In some embodiments, RNA is obtained from each sample, fragmented, and used to generate a complementary DNA (cDNA) sample, such as a cDNA library, for sequencing. The reads may be processed and aligned to the human genome, and used to estimate the expected number of mappings per gene / isoform and determine the read count. In some embodiments, the read count is normalized by the length of the gene / isoform and the number of reads in the library to obtain a normalized FPKM, for example, by the length of the gene / isoform and the number of reads in the library to obtain fragments per kilobase of exon per million mapped reads (FPKM) according to the gene length and the total reads mapped. In some aspects, normalization between samples is achieved by normalization, such as, for example, 75th quantile normalization, where each sample is scaled by the median of the 75th quantile from all samples to obtain a quantile-normalized FPKM (FPKQ) value. The FPKQ value may be log-transformed (log2).

[0346] In some embodiments, RNA from each sample is obtained, fragmented, and used to generate a complementary DNA (cDNA) sample, such as a cDNA library, for sequencing. The reads may be processed and aligned to the human genome, and used to estimate the expected number of mappings per gene / isoform and determine read counts. In some embodiments, the read counts are normalized by the length of the gene / isoform and the number of reads in the library. In some embodiments, the read counts are provided as counts per million (CPM). In some embodiments, the CPM read counts are log-transformed (e.g., log2).

[0347] In some embodiments, the relative gene expression is measured by comparing the CPM of target gene with the CPM of housekeeping gene. In some embodiments, the housekeeping gene is GAPDH. In some embodiments, the relative gene expression of target gene is determined as the CPM / CPM ratio of target gene to housekeeping gene (e.g., GAPDH).

[0348] In some embodiments, the gene expression levels are obtained using microarray analysis. In some embodiments, the gene expression levels are obtained using RNA sequencing. In some embodiments, the gene expression levels are obtained using both microarray analysis and RNA sequencing. In some embodiments, RNA sequencing is performed on bulk RNA from a plurality of cells. In some embodiments, bulk RNA sequencing data is obtained from pooled RNA from a plurality of cells. In some embodiments, RNA sequencing is performed on a single cell. In some embodiments, RNA sequencing is performed on bulk RNA from a plurality of cells and a single cell.

[0349] Any suitable method for obtaining bulk RNA sequencing data can be used (see, e.g., Chao et al., 2019, BMC Genomics 20:571, incorporated herein by reference in its entirety). For example, total RNA from a sample, e.g., multiple cells from a population of cells, can be isolated using TRIZOL, treated with DNase I, and purified. The concentration and quality of the isolated RNA can be measured and confirmed prior to library preparation of total RNA or mRNA. For library preparation, total RNA or mRNA can be fragmented and converted to cDNA using reverse transcription. After construction of double-stranded cDNA, amplification, and optional barcoding, the library can be processed for next-generation sequencing using any suitable library preparation technique, sequencing platform, and genome alignment tool.

[0350] In some embodiments, gene expression levels are obtained using single-cell RNA sequencing. In some embodiments, the use of single-cell RNA sequencing data provides certain advantages. In some embodiments, the use of single-cell RNA sequencing data allows characterization of subpopulations of cells, such as determined dopaminergic cells within a larger population of cells. In some embodiments, the use of single-cell RNA sequencing data reduces the number of cells required for use in the methods provided herein, for example, reducing the number of cells required to obtain data for training machine learning models. In some embodiments, the use of single-cell RNA sequencing data improves characterization of biological variability across cells. In some embodiments, the use of single-cell RNA sequencing data allows easier validation and interpretation of gene expression levels.

[0351] Any suitable method for single cell RNA sequencing can be used (see, e.g., Zheng et al., 2017 (Nature Communications 8:14049), and Haque et al., 2017 (Genome Medicine 9:75, which are incorporated herein by reference in their entirety). For single RNA sequencing, single cells from a sample, e.g., an in vitro population of cells, can be isolated using flow cytometric cell sorting, microfluidic platforms, or droplet-based methods. The isolated cells are lysed to allow capture of the RNA molecules. Poly[T] primers can be used specifically for the analysis of polyadenylated mRNA molecules, and the primed mRNA molecules are converted to cDNA using reverse transcription. In some examples, unique molecular identifiers can be used to identify single mRNA molecules based on their cellular origin. The cDNA pool can then be amplified, optionally barcoded, and sequenced, for example using next-generation sequencing (NGS), with library preparation techniques, sequencing platforms, and genome alignment tools similar to those used for bulk RNA samples. In some instances, unbiased cell type classification of a mixed population of different cell types can be achieved with as few as 10,000-50,000 reads per cell, and single-cell libraries from various common protocols can approach saturation when sequenced to a depth of 1,000,000 reads.

[0352] In some embodiments, the gene expression level comprises bulk RNA sequencing data and single-cell RNA sequencing data. In some embodiments, the bulk RNA sequencing data and the single-cell RNA sequencing data are obtained from the same population of cells. In some embodiments, the single-cell RNA sequencing data can be used to approximate the bulk RNA sequencing data obtained from the same population of cells. In some embodiments, the approximated bulk RNA sequencing data is obtained by averaging the single-cell RNA sequencing data from cells in the same population of cells. In some embodiments, the gene expression level comprises the approximated bulk RNA sequencing data.

[0353] III. Computing Devices In some embodiments, also provided herein is a computing device for classifying the differentiation state of an in vitro population of cells, hi some embodiments, a computing device is provided for identifying an in vitro population of cells having a desired differentiation state.

[0354] In some embodiments, the computing device includes a memory including a first reference data set and a second reference data set. Exemplary first and second reference data sets are described in Section II-A. In some embodiments, the first and second reference data sets are optional, as described in Section II-A.

[0355] In some embodiments, the memory further comprises one or more additional reference datasets, hi some embodiments, the one or more additional reference datasets comprise any of the first and second reference datasets described in Section II-A.

[0356] In some embodiments, the memory further comprises a control dataset. Exemplary control datasets are described in Section II-A. In some embodiments, the control dataset is any of those described in Section II-A.

[0357] In some embodiments, the computing device includes instructions stored in the memory for performing any of the methods provided. In some embodiments, the computing device further includes a processor that implements the instructions stored in the memory. In some embodiments, the processor includes one or more processing elements in communication with a system data store (SDS) that comprises one or more storage elements. In some embodiments, the processor includes one or more processing elements such as CELERON, PENTIUM, XEON, CORE 2 DUO, or CORE 2 QUAD class microprocessors (Intel Corp., Santa Clara, Calif.), or SEMPRON, PHENOM, OPTERON, ATHLON X2, or ATHLON 64 X2 (AMD Corp., Sunnyvale, Calif.), although other general-purpose processors may be used. In some embodiments, functionality may be distributed across multiple processing elements. The term processing element may refer to (1) a process running on or across specific hardware, (2) specific hardware, or either (1) or (2), as the context permits. Some implementations may include one or more limited special-purpose processors, such as digital signal processors (DSPs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs), etc. Additionally, some implementations may use a combination of general-purpose and special-purpose processors.

[0358] In some embodiments, the computing device includes one or more input devices for receiving input from a user and / or a software application. In some embodiments, the input includes a test data set. Exemplary test data sets are described in Section II-A. In some embodiments, the test data set is any of those described in Section II-A.

[0359] In some embodiments, the computing device includes one or more output devices for presenting output to a user and / or a software application. In some embodiments, the output device presents the output of any of the methods provided. In some embodiments, the output device includes a monitor capable of displaying a graphical representation of the output to a user.

[0360] In some embodiments, the computing device further includes an SDS, which may include various primary and secondary storage elements. In one implementation, the SDS includes registers and RAM as part of the primary storage. The primary storage may include other forms of memory, such as cache memory or non-volatile memory (e.g., FLASH, ROM, or EPROM), in some implementations. The SDS may also include secondary storage, including single, multiple, and / or various server and storage elements. For example, the SDS may use an internal storage device connected to the system processor. In implementations where a single processing element supports all functions, a local hard disk drive may function as the secondary storage of the SDS, and a disk operating system running on such a single processing element may function as a data server that receives and services data requests.

[0361] Those skilled in the art will appreciate that the different information used in the systems and methods disclosed herein may be logically or physically separated within a single device that serves as secondary storage for an SDS; multiple related data stores accessible through an integrated management system that functions together as an SDS; or multiple independent data stores individually accessible through entirely different management systems that may in some implementations be collectively considered an SDS. The various storage elements that make up the physical architecture of an SDS may be centrally located or distributed across a variety of diverse locations.

[0362] Additionally or alternatively, the above-described functions and approaches, or portions thereof, may be embodied in computer-executable instructions, such instructions being stored in and / or on one or more computer-readable storage media. Such media may include primary and / or secondary storage devices integrated with and / or within the computer, such as RAM and / or magnetic disks, and / or detachable from the computer, such as solid-state devices or removable magnetic or optical disks. The media may use any technology, including ROM, RAM, magnetic, optical, paper, and / or solid media technologies. In some embodiments, the computing device may be a general-purpose machine with modules and / or components dedicated to performing the disclosed methods.

[0363] IV. Compositions and Formulations In some embodiments, provided herein are pharmaceutical compositions containing a population of cells, including a population of cells, e.g., stem cell-derived cells, identified as having a desired differentiation state by any of the provided methods, such as any of the methods described in Section II.

[0364] In some embodiments, the cells in the therapeutic compositions provided comprise stem cell-derived neural cells. In some embodiments, the stem cell-derived neural cells are suitable for treating neurodegenerative diseases when transplanted into the brain of a subject in need of such treatment. In some embodiments, the cells in the therapeutic compositions provided comprise determined dopaminergic (DA) neural cells. In some embodiments, the cells in the therapeutic compositions provided are stem cell-derived neural cells that can engraft into brain regions after transplantation.

[0365] In some embodiments, the cells in the composition are an in vitro stem cell derived neural cell population. In some embodiments, the in vitro stem cell derived neural cell population is characterized by cells expressing one or more genes selected from the group consisting of CCNB2, AURKB, PTTG1, TOP2A, NEUROG2, HES1, REST, E2F4, FOXM1, SIN3A, NFYA, LIN28A, FLRT3, ITGA5, NES, SOX2, SOX9 and RFX4. In some embodiments, the cells in the population are characterized by expressing only one of the above genes. In some embodiments, the cells in the population are characterized by expression 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 or 18 of the above genes. In some embodiments, at least one of the one or more genes is REST.

[0366] In some embodiments, at least 50% of the cells in the in vitro stem cell derived neural cell population express one or more genes. In some embodiments, at least 60% of the cells in the in vitro stem cell derived neural cell population express one or more genes. In some embodiments, at least 70% of the cells in the in vitro stem cell derived neural cell population express one or more genes. In some embodiments, at least 80% of the cells in the in vitro stem cell derived neural cell population express one or more genes. In some embodiments, at least 90% of the cells in the in vitro stem cell derived neural cell population express one or more genes.

[0367] In some aspects, the expression of the one or more genes is RNA expression. In some embodiments, the RNA expression is measured by RNA sequencing. In other aspects, the expression of the one or more genes is protein expression.

[0368] In some embodiments, the population of cells in the compositions provided are differentiated in vitro from pluripotent stem cells (PSCs). Differentiation can be by any of the methods as described in Section C. In certain embodiments, the method comprises differentiating iPSCs into neural progenitor cells to produce committed dopaminergic neurons.

[0369] In some of the optional embodiments, the one or more genes are overexpressed in the cells of the population compared to iPSCs. In some embodiments, the one or more genes are overexpressed in the cells of the population compared to the cells of a precursor population differentiated from iPSCs. For example, in some embodiments, the one or more genes are overexpressed in the cells of the precursor population of cells at a differentiation stage before the cells are determined to be or may be dopaminergic neurons. For example, the precursor population of cells can be cells at day 13 of a dopaminergic differentiation protocol as described herein. In some embodiments, the one or more genes are overexpressed in the cells of the population compared to the cells of a mature committed dopaminergic neural cell population differentiated from iPSCs. For example, in some embodiments, the one or more genes are overexpressed in the cells of the population compared to the mature or committed population of cells at a differentiation stage before the cells are determined to be or may be dopaminergic neurons. For example, the mature committed cells can be cells at day 25 of a dopaminergic differentiation protocol as described herein. In some embodiments, mature committed dopaminergic neurons express LMX1A and / or NR4A2 (NURR1). In some embodiments, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of cells in the committed dopaminergic neuronal population express LMX1A and / or NR4A2. In some embodiments, overexpression is a positive log2 fold change of more than or about 1.5-fold, more than or about 2.0-fold, more than or about 3.0-fold, more than or about 4.0-fold, or more than or about 5-fold.

[0370] In some of the embodiments of any of the in vitro stem cell derived neural cell populations, the one or more genes are genes that are reduced in expression in the cells of the population compared to iPSCs. In some embodiments, the one or more genes are genes that are reduced in expression in the cells of the population compared to cells of a precursor population differentiated from iPSCs. For example, in some embodiments, the one or more genes are genes that are reduced in expression compared to cells of a precursor population of cells at a differentiation stage before the cells are determined to be or may be dopaminergic neurons. For example, the precursor population of cells can be cells at day 13 of a dopaminergic differentiation protocol as described herein. In some embodiments, the one or more genes are genes that are reduced in expression in the cells of the population compared to cells of a mature committed dopaminergic neural cell population differentiated from iPSCs. For example, in some embodiments, the one or more genes are genes that are reduced in expression compared to a mature or committed population of cells at a differentiation stage before the cells are determined to be or may be dopaminergic neurons. For example, the mature committed cells can be cells at day 25 of a dopaminergic differentiation protocol as described herein. In some embodiments, mature committed dopaminergic neurons express LMX1A and / or NR4A2 (NURR1). In some embodiments, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of cells in the committed dopaminergic neuronal population express LMX1A and / or NR4A2. In some embodiments, the decreased expression is a negative log2 fold change of more than or about 1.5-fold, more than or about 2.0-fold, more than or about 3.0-fold, more than or about 4.0-fold, or more than or about 5-fold.

[0371] In some of any of the embodiments of the in vitro stem cell derived neural cell populations, less than 30%, less than 20%, or less than 10% of the cells in the population express LMX1A and / or NR4A2.

[0372] In some of the embodiments of any of the in vitro stem cell derived neural cell populations, the cells in the population can engraft and innervate other cells in vivo. In some embodiments, the cells in the population can exhibit neurite outgrowth when administered to the brain of a subject. In some embodiments, the cells in the population can produce dopamine. In some embodiments, the cells in the population do not produce or substantially do not produce norepinephrine.

[0373] In some embodiments, the cells in the provided therapeutic compositions can produce dopamine (DA). In some embodiments, the cells in the provided therapeutic compositions do not produce or substantially do not produce norepinephrine (NE). Thus, in some embodiments, the cells in the provided therapeutic compositions can produce DA, but do not produce or substantially do not produce NE.

[0374] In some embodiments, the determined DA neuronal cells express EN1. In some embodiments, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, or at least about 80% of the total cells in the therapeutic composition express EN1. In some embodiments, at least about 20% of the cells in the therapeutic composition express EN1. In some embodiments, at least about 25% of the cells in the therapeutic composition express EN1. In some embodiments, at least about 30% of the cells in the therapeutic composition express EN1. In some embodiments, at least about 35% of the cells in the therapeutic composition express EN1. In some embodiments, at least about 40% of the cells in the therapeutic composition express EN1. In some embodiments, at least about 45% of the cells in the therapeutic composition express EN1. In some embodiments, at least about 50% of the cells of the therapeutic composition express EN1. In some embodiments, at least about 55% of the cells of the therapeutic composition express EN1. In some embodiments, at least about 60% of the cells of the therapeutic composition express EN1. In some embodiments, at least about 65% of the cells of the therapeutic composition express EN1. In some embodiments, at least about 70% of the cells of the therapeutic composition express EN1. In some embodiments, at least about 75% of the cells of the therapeutic composition express EN1. In some embodiments, at least about 80% of the cells of the therapeutic composition express EN1.

[0375] In some embodiments, the therapeutic composition comprises about 1×10 -4 In some embodiments, the ratio of CPM EN1 to GAPDH is greater than about 1.5×10 -3 ~1×10 -2 It is.

[0376] In some embodiments, the determined DA neurons express CORIN. In some embodiments, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, or at least about 80% of the total cells in the therapeutic composition express CORIN. In some embodiments, at least about 20% of the cells in the therapeutic composition express CORIN. In some embodiments, at least about 25% of the cells in the therapeutic composition express CORIN. In some embodiments, at least about 30% of the cells in the therapeutic composition express CORIN. In some embodiments, at least about 35% of the cells in the therapeutic composition express CORIN. In some embodiments, at least about 40% of the cells in the therapeutic composition express CORIN. In some embodiments, at least about 45% of the cells in the therapeutic composition express CORIN. In some embodiments, at least about 50% of the cells of the therapeutic composition express CORIN. In some embodiments, at least about 55% of the cells of the therapeutic composition express CORIN. In some embodiments, at least about 60% of the cells of the therapeutic composition express CORIN. In some embodiments, at least about 65% of the cells of the therapeutic composition express CORIN. In some embodiments, at least about 70% of the cells of the therapeutic composition express CORIN. In some embodiments, at least about 75% of the cells of the therapeutic composition express CORIN. In some embodiments, at least about 80% of the cells of the therapeutic composition express CORIN.

[0377] In some embodiments, the therapeutic composition comprises about 1×10 -4 In some embodiments, the ratio of CPM CORIN to GAPDH is greater than about 5×10. -2 ~5×10 -1 It is.

[0378] In some embodiments, the determined DA neuronal cells express EN1 and CORIN. In some embodiments, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, or at least about 80% of the total cells in the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 20% of the cells in the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 25% of the cells in the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 30% of the cells in the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 35% of the cells in the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 40% of the cells in the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 45% of the cells of the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 50% of the cells of the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 55% of the cells of the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 60% of the cells of the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 65% of the cells of the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 70% of the cells of the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 75% of the cells of the therapeutic composition express EN1 and CORIN. In some embodiments, at least about 80% of the cells of the therapeutic composition express EN1 and CORIN.

[0379] In some embodiments, the therapeutic composition comprises: (a) about 1×10 -4(b) counts per million (CPM) / CPM ratio of EN1 to GAPDH of approximately 2 × 10 -2 In some embodiments, the CPM / CPM ratio of CORIN to GAPDH is greater than about 1.5×10. -3 ~1×10 -2 and the CPM / CPM ratio of CORIN to GAPDH is approximately 5 × 10 -2 ~5×10 -1 It is.

[0380] In some embodiments, less than 10% of the determined DA neurons express TH. In some embodiments, the determined DA neurons express low levels of TH. In some embodiments, the determined DA neurons do not express TH. In some embodiments, the determined DA neurons express TH at a lower level than cells harvested or collected on other days. In some embodiments, a portion of the determined DA neurons express EN1 and CORIN, and less than 10% of the determined DA neurons express TH. In some embodiments, less than 8% of the determined DA neurons express TH. In some embodiments, less than 5% of the determined DA neurons express TH.

[0381] In some embodiments, about 2%-10%, about 2%-8%, about 2%-6%, about 2%-4%, about 4%-10%, about 4%-8%, about 4%-6%, about 6%-10%, about 6%-8%, or about 8%-10% of the total cells in the therapeutic composition express TH.

[0382] In some embodiments, the therapeutic composition comprises about 3×10 -2 In some embodiments, the ratio of CPM TH to CPM GAPDH is less than about 1×10 -3 ~2.5×10 -2 It is.

[0383] In some embodiments, less than 10% of the total cells in the therapeutic composition express TH and at least about 20% of the cells in the therapeutic composition express EN1. In some embodiments, less than 10% of the total cells in the therapeutic composition express TH and at least about 25% of the cells in the therapeutic composition express EN1. In some embodiments, less than 10% of the total cells in the therapeutic composition express TH and at least about 30% of the cells in the therapeutic composition express EN1. In some embodiments, less than 10% of the total cells in the therapeutic composition express TH and at least about 35% of the cells in the therapeutic composition express EN1. In some embodiments, less than 10% of the total cells in the therapeutic composition express TH and at least about 40% of the cells in the therapeutic composition express EN1. In some embodiments, less than 10% of the total cells in the therapeutic composition exp...

Claims

1. A computing device for classifying the differentiation state of a population of cells in vitro, wherein the device comprises: A first reference dataset containing representations of gene expression levels for one or more genes that are differentially expressed between cells in a first differentiated state and cells in a second differentiated state, A second reference dataset containing representations of gene expression levels for one or more genes differentially expressed between cells in the second differentiated state and cells in the third differentiated state, A computing device equipped with memory including [this].

2. The method further comprises a processor that implements instructions stored in the memory in order to perform the method, the method (a) Receiving as input a test dataset containing expression levels for genes expressed in one or more test cells included in an in vitro cell population, wherein the expression levels in the test dataset include (i) expression levels for one or more of the genes included in the first reference dataset, and (ii) expression levels for one or more of the genes included in the second reference dataset, (b) Using the test dataset and the first reference dataset, calculate a first similarity score indicating whether the differentiation state of the test cells is more similar to the first differentiation state or more similar to the second differentiation state, (c) Using the test dataset and the second reference dataset, calculate a second similarity score indicating whether the differentiation state of the test cells is more similar to the second differentiation state or more similar to the third differentiation state, (d) Classifying the differentiation state of one or more test cells based on either or both of the first similarity score and the second similarity score, A computing device according to claim 1, including the following:

3. The computing device according to claim 2, wherein the memory further includes a control dataset containing representations of gene expression levels for one or more genes expressed in cells in one or more control differentiation states, the control differentiation state may be the same as or different from one of the first, second, or third differentiation states.

4. The aforementioned test dataset includes gene expression levels for one or more of the genes included in the control dataset, where the expression level representation is: The instruction includes calculating the degree of correlation between the representation of the gene expression level for one or more genes in the control dataset and the gene expression level for one or more genes in the test dataset, in order to calculate a correlation score. Classifying the differentiation state of one or more test cells is based on the correlation score and one or both of the first similarity score and the second similarity score. The computing device according to claim 3.

5. The computing device according to claim 4, wherein the correlation score is calculated before calculating the first similarity score and the second similarity score, and the method terminates if the correlation score of the test cells does not meet a predetermined cutoff value.

6. The computing device according to claim 3, wherein the control dataset includes gene expression levels that are normalized by counts per million mapped reads (CPM) and filtered to include only gene expression levels that exceed a threshold CPM value.

7. The computing device according to claim 4, wherein the control dataset includes the centroid of the gene expression levels of one or more genes in the control dataset.

8. The computing device according to claim 7, wherein the correlation score is calculated by normalizing the gene expression levels of one or more genes in the test dataset and calculating the correlation between the gene expression levels of one or more genes in the test dataset and the centroid.

9. The computing device according to claim 8, wherein the control dataset includes coefficient of variation (CV) values ​​of the gene expression levels of one or more genes in the control dataset, and the correlation with respect to the centroid is weighted by the reciprocal of the CV value.

10. The computing device according to any one of claims 1 to 9, wherein the population of cells is derived from a culture of cells differentiated from pluripotent cells subjected to appropriate differentiation conditions.

11. The computing device according to any one of claims 1 to 9, wherein the first differentiation state is earlier in the stem cell differentiation pathway than the second differentiation state.

12. The computing device according to any one of claims 1 to 9, wherein the second differentiation state is earlier in the stem cell differentiation pathway than the third differentiation state.

13. The computing device according to any one of claims 1 to 9, wherein the first differentiation state is located in a cell differentiation pathway parallel to the cell differentiation pathway of the second differentiation state.

14. The computing device according to any one of claims 1 to 9, wherein the group of cells is selected from the group consisting of stem cell-derived cardiomyocytes, stem cell-derived skeletal muscle cells, stem cell-derived tubular cells, stem cell-derived erythrocytes, stem cell-derived smooth muscle cells, stem cell-derived lung cells, stem cell-derived thyroid cells, stem cell-derived pancreatic cells, stem cell-derived epidermal cells, stem cell-derived pigment cells, and stem cell-derived nerve cells.

15. The computing device according to any one of claims 1 to 9, wherein the population of cells is nerve cells derived from stem cells.

16. The computing device according to any one of claims 1 to 9, wherein the second differentiation state is the determined differentiation state of dopaminergic neurons or hematopoietic progenitor cells.

17. The computing device according to any one of claims 1 to 9, wherein the first reference dataset includes representations of gene expression levels for one or more genes or at least 20 genes selected from Table E1.

18. The computing device according to any one of claims 1 to 9, wherein the second reference dataset includes representations of gene expression levels for one or more genes or at least 20 genes selected from Table E2.

19. The computing device according to any one of claims 1 to 9, wherein at least one of the first, second, and third differentiation states is characterized using an in vitro assay or an in vivo assay.

20. The computing device according to claim 19, comprising determining whether the in vivo assay, when administered to an animal or human subject, is capable of reference cells surviving, engrafting, and / or innervating tissue, or improving or reversing symptoms of neurodegenerative disease.

21. The memory further includes one or more additional reference datasets, each of which includes a representation of gene expression levels for one or more genes that are differentially expressed between cells in the second differentiated state and cells in the additional differentiated state. The processor executes an instruction to calculate one or more additional similarity scores indicating whether the differentiation state of the test cells is more similar to the second differentiation state or one of the one or more additional differentiation states, using the additional reference dataset. Classifying the differentiation state of one or more test cells is based on the first similarity score, the second similarity score, and the one or more additional similarity scores. The computing device according to claim 2.

22. The computing device according to any one of claims 1 to 9, wherein the representation of gene expression levels in the first reference dataset and / or the second reference dataset is obtained using machine learning.

23. The computing device according to claim 22, wherein the machine learning includes principal component analysis.

24. The computing device according to any one of claims 1 to 9, wherein the representation of gene expression levels in the first reference dataset and / or the second reference dataset includes normalized gene expression levels.

25. The computing device according to claim 2, wherein if one or both of the first and second similarity scores indicate that the differentiation state of one or more test cells is more similar to the second differentiation state, the differentiation state of one or more test cells is classified as the second differentiation state.

26. A method for selecting a population of cells having a desired differentiation state, (a) Calculating a first similarity score using the test dataset and the first reference dataset, The first reference dataset includes a representation of gene expression levels for one or more genes that are differentially expressed between cells in a first differentiated state and cells in a second differentiated state. The test dataset includes expression levels for genes expressed in one or more test cells included in an in vitro cell population, and the expression levels in the test dataset include expression levels for one or more of the genes included in the first reference dataset. The first similarity score indicates whether the differentiation state of the test cells is more similar to the first differentiation state or more similar to the second differentiation state. (b) Calculating a second similarity score using the test dataset and a second reference dataset, The second reference dataset includes a representation of gene expression levels for one or more genes that are differentially expressed between cells in the second differentiated state and cells in the third differentiated state, The expression levels in the aforementioned test dataset include the expression levels of one or more of the genes included in the second reference dataset, The second similarity score indicates whether the differentiation state of the test cells is more similar to the second differentiation state or to the third differentiation state. (c) Classifying the differentiation state of one or more test cells based on either or both of the first similarity score and the second similarity score, Methods that include...

27. The test dataset includes the gene expression levels of one or more genes expressed in a control dataset which includes the gene expression levels of one or more genes expressed in cells in a control differentiated state, and the control differentiated state may be the same as or different from one of the first, second, or third differentiated states. The method further includes calculating the degree of correlation between the expression of the gene expression level for one or more genes in the control dataset and the gene expression level for one or more genes in the test dataset, in order to calculate a correlation score. Classifying the differentiation state of one or more test cells is based on the correlation score and one or both of the first similarity score and the second similarity score. The method according to claim 26.

28. The method according to claim 27, wherein the correlation score is calculated before calculating the first similarity score and the second similarity score, and the method terminates if the correlation score of the test cells does not meet a predetermined cutoff value.

29. The method according to claim 27, wherein the control dataset includes gene expression levels that are normalized by counts per million mapped reads (CPM) and filtered to include only gene expression levels that exceed a threshold CPM value.

30. The method according to claim 27, wherein the control dataset includes the centroid of the gene expression levels of one or more genes in the control dataset.

31. The method according to claim 30, wherein the correlation score is calculated by normalizing the gene expression levels of one or more genes in the test dataset and calculating the correlation between the gene expression levels of one or more genes in the test dataset and the centroid.

32. The method according to claim 31, wherein the control dataset includes coefficients of variation (CV) values ​​of the gene expression levels of one or more genes in the control dataset, and the correlation with respect to the centroid is weighted by the reciprocal of the CV value.

33. The method according to any one of claims 26 to 32, wherein the first differentiation state is earlier in the stem cell differentiation pathway than the second differentiation state.

34. The method according to any one of claims 26 to 32, wherein the second differentiation state is earlier in the stem cell differentiation pathway than the third differentiation state.

35. The method according to any one of claims 26 to 32, wherein the first differentiation state is located in a cell differentiation pathway parallel to the cell differentiation pathway of the second differentiation state.

36. The method according to any one of claims 26 to 32, wherein the group of cells is selected from the group consisting of stem cell-derived cardiomyocytes, stem cell-derived skeletal muscle cells, stem cell-derived tubular cells, stem cell-derived erythrocytes, stem cell-derived smooth muscle cells, stem cell-derived lung cells, stem cell-derived thyroid cells, stem cell-derived pancreatic cells, stem cell-derived epidermal cells, stem cell-derived pigment cells, and stem cell-derived nerve cells.

37. The method according to any one of claims 26 to 32, wherein the population of cells is stem cell-derived nerve cells.

38. The method according to any one of claims 26 to 32, wherein the second differentiation state is the determined differentiation state of dopaminergic neurons or hematopoietic progenitor cells.

39. The method according to any one of claims 26 to 32, wherein the first reference dataset includes representations of gene expression levels for one or more genes or at least 20 genes selected from Table E1.

40. The method according to any one of claims 26 to 32, wherein the second reference dataset includes representations of gene expression levels for one or more genes or at least 20 genes selected from Table E2.

41. The method according to any one of claims 26 to 32, wherein at least one of the first, second, and third differentiation states is characterized using an in vitro assay or an in vivo assay.

42. The method according to claim 41, further comprising determining whether the in vivo assay, when administered to an animal or human subject, is capable of reference cells surviving, engrafting, and / or innervating tissue, or improving or reversing symptoms of neurodegenerative disease.

43. This further includes calculating one or more additional similarity scores using one or more additional reference datasets, Each of the aforementioned additional reference datasets includes a representation of gene expression levels for one or more genes that are differentially expressed between cells in the second differentiated state and cells in the additional differentiated state, The one or more additional similarity scores indicate whether the differentiation state of the test cells is more similar to the second differentiation state or one of the one or more additional differentiation states. Classifying the differentiation state of one or more test cells is based on the first similarity score, the second similarity score, and the one or more additional similarity scores. The method according to any one of claims 26 to 32.

44. The method according to any one of claims 26 to 32, wherein the representation of gene expression levels in the first reference dataset and / or the second reference dataset is obtained using machine learning.

45. The method according to claim 44, wherein the machine learning includes principal component analysis.

46. The method according to any one of claims 26 to 32, wherein the representation of gene expression levels in the first reference dataset and / or the second reference dataset includes normalized gene expression levels.

47. The method according to any one of claims 26 to 32, further comprising classifying the differentiation state of one or more test cells as the second differentiation state if one or both of the first and second similarity scores indicate that the differentiation state of one or more test cells is more similar to the second differentiation state.

48. The method according to any one of claims 26 to 32, further comprising selecting the in vitro population of cells, which includes one or more test cells classified as having the second differentiation state, as having the desired differentiation state.

49. A method for training a machine learning model for classifying the differentiation state of a population of cells in vitro, (a) Obtain gene expression levels for one or more genes differentially expressed between cells in a first differentiated state and cells in a second differentiated state in a reference population of multiple cells, and train a first machine learning model to predict whether an in vitro population of cells contains one or more test cells having a differentiated state more similar to the first differentiated state or the second differentiated state, (b) Obtain gene expression levels for one or more genes differentially expressed between cells in the second differentiated state and cells in the third differentiated state for a reference population of multiple cells, and train a second machine learning model to predict whether an in vitro population of cells contains one or more test cells having a differentiated state more similar to the second differentiated state or the third differentiated state, Methods that include...

50. A method for training a machine learning model for classifying the differentiation state of a population of cells in vitro, (a) Selecting one or more genes that are differentially expressed between cells in a first differentiated state and cells in a second differentiated state, and training a first machine learning model to predict whether an in vitro population of cells includes one or more test cells having a differentiated state more similar to the first differentiated state or the second differentiated state, (b) Selecting one or more genes that are differentially expressed between cells in the second differentiated state and cells in the third differentiated state, and training a second machine learning model to predict whether an in vitro population of cells contains one or more test cells having a differentiated state more similar to the second differentiated state or the third differentiated state, Methods that include...