Single-cell screening methods

By screening transcription factor combinations through single-cell analysis and split-and-merge culture methods, the problem of low cell reprogramming efficiency in existing technologies has been solved, achieving efficient and scalable cell type conversion.

CN122641692APending Publication Date: 2026-08-25PLASTISEL GMBH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202480086131.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-11-28
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and scalably determine the combination of transcription factor cocktails required for cell reprogramming, resulting in inefficient and poorly reproducible cell reprogramming, especially in the process of converting from one cell type to another.

Method used

By employing single-cell analysis combined with split and merge culture methods, multiple reagent combinations were added and evaluated one by one at the single-cell level. Combined with genetic barcode markers, the optimal combination and sequence of transcription factors were systematically screened to achieve cell type transformation.

Benefits of technology

It improves the efficiency and reproducibility of cell reprogramming, enables the determination of the optimal cell transformation pathway under high-throughput screening conditions, and reduces experimental costs and time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention relates to a method for determining conditions required for reprogramming a first cell type into a second cell type by analyzing one or more markers in single cells that have been exposed to a combination of reagents.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Cell therapy, which uses cells expanded and potentially modified in the laboratory to treat patients with serious diseases, is a promising technology that is yielding excellent results in oncology (for leukemia and solid tumors) (Ottaviano et al., 2022) and regenerative medicine (e.g., in the rehabilitation of burn patients) (Cossu et al., 2018). Cells can be safely derived from patients or donors, manipulated to expand, differentiate, and potentially correct genetic defects or enhance their therapeutic capabilities, and then administered back to the patient. However, the production of these advanced therapeutic medicine products (ATMPs) remains challenging. One of the main limitations is the lack of pathways to obtain therapeutically useful amounts of cells.

[0002] The promise of stem cells in regenerative medicine lies in the unprecedented possibility of deriving any type of cell from human pluripotent stem cells, such as induced pluripotent stem cells (iPSCs) or human embryonic stem cells (hESCs). These pluripotent stem cells have the potential to generate any human tissue in vitro because they are part of or highly similar to embryonic structures that produce all the primitive structures in our bodies. However, the main challenge lies in guiding these stem cells to develop into specific cell types that can be used for therapeutic purposes in a robust and scalable manner. Expanding and / or differentiating donor-derived stem cells into clinically useful cell products requires cell culture conditions tailored to specific applications and target cell types, and testing and optimizing these conditions is both expensive and time-consuming. Cell culture conditions must be formulated by carefully considering developmental cues from the human embryo, many of which are not yet fully understood. In any case, these conditions may be ineffective in the absence of a physiological microenvironment that provides additional stimuli that are difficult to simulate in cell culture, such as specific cell-cell interactions, extracellular matrix composition, and tissue stiffness. There is no prior method to determine which of these conditions will produce a highly reproducible population of highly homogeneous, correctly differentiated cells, and which is among the scalable methods with the potential to meet regulatory requirements.

[0003] Cell reprogramming, the process of transforming one cell type into another by manipulating gene regulatory networks that define cell identity, holds great promise in regenerative medicine but has not yet been widely applied. While it is now known that the phenotype of one cell type can be converted to that of another by manipulating gene circuits, the components required for cell fate conversion are difficult to identify and, in most cases, unknown. The identification of factors that directly reprogram cell type identity is currently limited (among other things) by the cost of exhaustive experimental testing on seemingly plausible sets of factors, making this approach inefficient and unscalable.

[0004] One of the most important examples of cell reprogramming is the use of exogenous transcription factors (TFs) to convert somatic cells into a pluripotent state. For example, Yamanaka (Yamanaka S. Cell 2006; 126:663-76; PMID:16904174) demonstrated that a small group of TFs, namely OCT4, KLF4, SOX2, and MYC (OKSM), can reprogram human fibroblasts into induced pluripotent stem cells (iPSCs) when introduced into cells. Subsequently, other research groups demonstrated that it is possible to transform fibroblasts into hepatocytes, cardiomyocytes, and various other cell types (Huang P. et al., Nature 2011; 475:386-9; Sekiya S, Suzuki A., Nature 2011; 475:390-3; Kogiso T, et al., Hepatol Int 2013; 7:937-44; Ieda M, et al., Cell 2010; 142:375-86; Song K, et al., Nature 2012; 485:599-604; Qian L, et al., Nature 2012; 485:593-8; Takahashi K, Yamanaka S., Nat Rev Mol Cell Biol 2016; 17:183-93; Sadahiro T, et al., Circ. Res. 2015;). 116:1378-91. Tsunemoto RK, et al., EMBO J 2015; 34:1445-55; Heinaniemi M, et al., Nat Methods 2013; 10:577-83; Lang AH, et al., PLoS Comput Biol 2014; 10:e1003734; Del Sol ICA, StemCells 2013; 31:2127-35; Davis FP, Eddy SR. PLoS One 2013; 8:1-8; D'Alessio et al., Stem Cell Reports 2015; 5:763–75).

[0005] In addition to reprogramming cells to obtain pluripotent cells (iPSCs), many related reprogramming strategies have been demonstrated, such as: (i) "forward programming," in which less differentiated cells, such as iPSCs, are reprogrammed into a specific cell type (Dalby et al., Stem Cell Reports 11:1462-1478); (ii) "direct conversion," in which one differentiated cell type is converted into another differentiated cell type without undergoing a pluripotent or progenitor state (Zhou et al., Nature 455:627-632); and combinations thereof, such as (iii) "indirect (or induced) reprogramming," in which ectopic OKSM is transiently expressed to achieve partial reprogramming into a less differentiated (but not fully pluripotent) intermediate, followed by forward programming to obtain a more differentiated cell type different from the starting cell type (Efe et al., Nat Cell Biol 13:215-222).

[0006] It should be noted that "cell reprogramming" differs qualitatively from traditional "cell differentiation" in both its biological mechanisms and the actual methods by which it is achieved. Cell differentiation occurs by progressively limiting developmental potential until a terminal state is reached; Waddington's analogy of a ball rolling down a mountain to a lower, less energetic resting place in a valley perfectly illustrates this point. The Strategy of the Genes (pub. Allen & Unwin, London, 1957), but this is not necessarily true in cell reprogramming (the opposite is true for iPSCs). In cell differentiation methods, cultured (stem) cells are exposed to developmental cues such as morphogenetics, growth factors, hormones, and small molecules, which are extracellular effector molecules added to the cell culture medium. These molecules typically function by binding to protein elements of cell signaling mechanisms (e.g., cell surface receptors) and cytoplasmic signaling mechanisms (e.g., mechanisms mediated by kinase cascades). In the case of cell reprogramming, cell fate switching is guided by the ectopic expression of transcription factors (TFs), which translocate to the nucleus and directly control gene expression, typically by binding to gene regulatory elements in enhancer and promoter DNA and stimulating (or inhibiting) mRNA transcription. These two different strategies for manipulating cell fate (the first is mediated by the control of the cellular environment (i.e., directed differentiation), and the second is mediated by the reconnection of gene circuits that determine cell fate (cell reprogramming)) are also referred to as the “outside-in” and “inside-out” approaches, respectively (Shakiba et al., Cell Syst. 12:561-592).

[0007] A carefully orchestrated gene flow network (GRN) for cell differentiation or reprogramming can regulate one or more GRNs, among other genes, thereby inducing and / or stabilizing gene regulatory networks that specify cell fate. One of the best-understood GRNs is the pluripotency gene regulatory network (PGRN), which is induced and stabilized by adding the Yamanaka factor OKSM to (differentiated) cells. At the heart of the PGRN are the three musketeers of pluripotency-inducing GRNs: OCT4, SOX2, and NANOG. These GRNs regulate a larger network of interconnected secondary genes. The core GRNs are regulated in a synergistic manner, forming self-regulating and feedforward gene loops that contribute to network stability. It is known that NANOG and LIN28 can substitute for KLF4 and MYC in the Yamanaka cocktail, indicating that specific GRNs can be induced by substitution. Indeed, reprogramming is also known to be influenced by non-coding RNAs that regulate gene expression, such as LincRoR and Let7, which affect key nodes in the PGRN; therefore, in this application, non-coding RNAs that regulate gene expression are included in the definition of a GRN.

[0008] GRNs are essentially composed of genomic components (such as genes at nodes and their cis-regulatory modules) and regulatory state components (i.e., TFs that provide regulatory input to these modules, including non-coding RNAs). The connections between different regulatory gene products (i.e., TFs) and homologous cis-regulatory elements that control other network nodes define network circuits that can explain developmental processes (such as the maintenance of pluripotency in blastocysts or the differentiation of pluripotent stem cells into pancreatic β cells).

[0009] Few GRNs are as well-defined as PGRNs, and the TFs that influence specific cell reprogramming transitions are largely unknown. Recently, numerous computational methods have been developed to predict TFs that can be used to facilitate reprogramming from one cell type to another (reviewed by Kamaraj et al. (Cell Cycle 15: 3343–3354)). These methods typically consider differential gene expression between the starting and target cell types and leverage knowledge of GRN architecture to rank TFs according to their influence within the core GRN that establishes the target cell. The availability of such methods to a priori assess the ability of TFs to influence cell fate is a significant advantage of “inside-out” approaches, which, in contrast, currently lack widely applicable strategies for predicting guiding components of cell culture media that drive cell differentiation.

[0010] While the computational methods cited above have been used to predict whether OKSM can induce PGRN, as well as a relatively small number of other cell fate transitions (TFs), these algorithms (and their required input data) are far from perfect and often result in lists of TF candidates with little or no overlap. Therefore, the literature suggests that cell reprogramming is difficult to achieve and, even when successful, is typically inefficient and / or has low reproducibility. Experimentally identifying useful transcription factor cocktails using methods used in the prior art is laborious and subject to two related major limitations.

[0011] First, the conventional approach to identifying the reprogramming cocktail is to screen candidate TFs in a large pool (or a set of smaller pools) and then identify the minimum core TF by sequentially eliminating each TF, as illustrated in the pioneering experiments of Takahashi and Yamanaka (Cell 126:663-676), where a pool containing 24 TFs was screened in a cell-based assay to identify the OKSM in reprogramming fibroblasts into iPSCs. Yamanaka used mouse fibroblasts in cell culture to observe the population effect of large cell samples.

[0012] In further demonstration of this strategy, Zhou et al. (Nature 455:627-633) introduced a pool of nine TFs into the mouse pancreas in vivo and observed that pancreatic exocrine cells were reprogrammed into insulin-expressing β cells. In this case, Zhou et al. were also able to infer, through the elimination process, that the effective reprogramming cocktail contained PDX1, NGN3, and MAFA. In both instances, these pooled screening experiments failed to identify functional combinations of TFs beyond the minimum core cocktail because the presence of the core TF in the pool would mask the effect of eliminating other effective TF cocktails from the pool. Furthermore, it was impossible to track the effect of individual TFs at the cellular level in either experiment.

[0013] This limitation could be overcome if a set of TFs could be systematically screened to evaluate all combinations in individual experiments. However, the required number of experiments is relatively large, as there are 10,626 different combinations of four TFs among a set of 24 candidate TFs in Yamanaka.

[0014] Furthermore, the number of cells required for such large-scale screening would be enormous, as experiments need to be replicated, and the efficiency of reprogramming human cells with OKSM, for example, is only 0.01% to 0.02% (Takahashi et al., Cell 131:861–872). Therefore, scalable screening methods that can systematically evaluate individual TF combinations are needed, as “large-scale screening of combined TF cocktails remains challenging” (Li and Hon, Frontiers in Bioengineering and Biotechnology (2021) vol 9 article 748942).

[0015] Secondly, pooled screening methods cannot uncover optimal approaches where the order of TF addition is crucial for certain types of reprogramming. Recent advances in single-cell omics technologies (such as RNA-Seq and ATAC-Seq) have enabled us to study cells during reprogramming at unprecedented resolution. These technologies reveal that cell reprogramming is a dynamic process in which cells transition through a continuum of cellular states characterized by different TF expression patterns accompanying GRN remodeling. For example, the optimal reprogramming of pancreatic progenitor cells into insulin-expressing β cells using TFs PDX1, NGN3, and MAFA can only be achieved through differential expression dynamics (Saxena et al., Nature Communications 7 Article number 11247). Therefore, screening methods capable of exploring TF combinations expressed in sequence are needed to discover optimal cell reprogramming approaches.

[0016] In our 2004 patent application WO2004 / 031369, we described a method for identifying factors required to differentiate pluripotent cells into more targeted cell types by using split and merge culture methods and labeling multicellular units in the form of beads, wherein the labeling identifies the conditions to which the beads are exposed.

[0017] In this invention, we provide an improvement on our 2004 process, in which cells are optimally converted from one cell type to another. We combine single-cell analysis of the effects of reagents (including TF) on inter-cell type conversion with the flexibility of adding reagents to cells to better understand the genetic pathways leading to cell type conversion. Summary of the Invention

[0018] In a first aspect, the present invention provides a method for determining the conditions required to convert a first cell type into a second cell type by analyzing one or more markers in a single cell exposed to a combination of reagents.

[0019] Gene expression markers associated with cell differentiation can be used to track cell differentiation pathways. Many gene expression markers are transcription factors (TFs), but many others are not, such as CD antigens, glycoproteins like synaptophysin (Smith et al., Clin Neuropathol. 1993 Nov-Dec;12(6):335-42), and β-tubulin (Korzhevskii et al., Neurosci Behav Physi 42, 215–222 (2012)), desmin (Capetanaki et al., Cell Struct Funct. 1997 Feb;22(1):103-16), etc. Therefore, the markers analyzed in cells can be TF, or they can be markers other than TF. The expression patterns of differentiation markers can be resolved to identify GRNs involved in cell type transformation.

[0020] In one embodiment, the present invention provides a method comprising the following steps: a. Expose a single cell of the first cell type to the first reagent or a combination of reagents; b. Optionally remove the first reagent and expose the cells to different reagents or combinations of reagents; c. Optionally, repeat step (b) above; and d. Identify one or more cells displaying one or more markers of a second cell type, and determine the identity and order of the reagents to which the cells were exposed.

[0021] In some implementations, cells are labeled to identify their exposure to each reagent or combination of reagents. Cell labeling is advantageously performed via genetic barcoding methods, such as retroviral infection and / or genome editing.

[0022] The splitting / merging technique described in WO2004 / 031369 is applicable to single-cell screening regarding the effects of cell exposure to multiple individual reagents and combinations of reagents. The cell unit in WO2004 / 031369 is envisioned as a bead containing multiple cells that can be sorted together; in this invention, single cells are used and sorted individually.

[0023] In some embodiments, the present invention provides a method comprising the following steps: a. Subdivide multiple single cells into a first group of single cell populations, expose each single cell population to a different reagent or combination of reagents, and label the single cells to indicate the exposure to the reagent or combination of reagents; b. Merge the single-cell populations and subdivide the cell pools to form a second group of single-cell populations; c. Expose the second group of single-cell populations to different reagents or combinations of reagents; d. Optionally iterate and repeat steps (a) to (c) as needed; e. Identify cell type markers in cells and unconvolve the markers to identify the identity, sequence, and timing of the reagents that cells were exposed to; f. Optionally, use the data from step (e) to guide the selection of reagents used in step (a), and repeat steps (a) to (e) as needed; g. Identify cells of the second cell type based on cell type markers, and deconvolve the markers to identify the identity and order of the reagents required to reprogram cells of the first cell type into cells of the second cell type.

[0024] This invention employs the method described in WO2004031369 (see...) Figure 1 In this method, known as CombiCult®, sequential cell splitting and merging are used to expose multiple cells to a number of different reagent combinations. The successive method steps in the procedure according to the foregoing aspects of the invention include adding a single reagent or adding a combination of reagents.

[0025] In some implementations, the successive method steps in the procedure involve adding only a single reagent.

[0026] The reagents or combinations of reagents used in the first iteration of step (a) of the above method can be selected from those known in the art for promoting cell reprogramming. This is the only available method in the prior art, and the literature contains numerous publications that will guide those skilled in the art in selecting obviously suitable reagents.

[0027] In an alternative embodiment of the invention, the reagent or combination of reagents used in the first iteration of step (c) is selected from reagents known in the art that promote cell type conversion.

[0028] Referring to the method described above for analyzing data from the selection procedure, the reagents or combinations of reagents used in the second and further iterations of steps (a) and (c) are selected by analyzing the data generated in step (e).

[0029] Therefore, in one embodiment, the reagents used in steps (a) and / or (c) can be selected based on computational analysis of known cell transformation events. The data used for analysis can be based on previous executions of the method of the present invention or derived from existing technology. The computational selection, combined with experimental selection provided by a split / merge selection methodology, enables extremely precise analysis in determining the identity of factors required for cell transformation and the ideal timing of their application.

[0030] In one implementation, the data generated by step (e) is processed by an algorithm that analyzes the impact of the identity and application time of each reagent on the reprogramming of the first cell type.

[0031] For example, data generated by step (e) in a previous execution of the method can be used to select the reagent or combination of reagents to be used in the first iteration of step (a).

[0032] Cellular labeling can be performed in various ways, such as by modifying the nucleic acids of each cell, which allows for precise tracking of their exposure to reagents. Nucleic acids can be modified using techniques such as retroviral integration or genome editing.

[0033] In some implementations, genome editing includes CRISPR.

[0034] Repeating reagents at different times during the split-merge sequence allows for the assessment of the reagent's effect at different points in cell development. For example, a pool may be exposed to a combination of four reagents during the first culture, followed by splitting. Different split portions of the pool may be exposed to different reagent combinations, which may or may not include any of the reagents used in the first culture. After merging and splitting again, further different split portions may be exposed to further reagent combinations, which may include reagents from the first or second culture as well as new reagents. Thus, cells in different split portions are exposed to different reagents, as well as to the same reagents at different times.

[0035] In some implementations, reagents are added to the cells one at a time in successive method steps. For example, successive method steps in a procedure may include adding a single reagent or adding a combination of reagents; alternatively, successive method steps in a procedure may include adding only a single reagent. Thus, a procedure may combine method steps for adding multiple groups of reagents with steps for adding a single reagent, or it may add only a single reagent.

[0036] In this context, “procedure” refers to the execution of the entire method as represented by steps (a) through (f) above.

[0037] While repeated cycles of splitting and merging can be used extremely efficiently, similar to combinatorial chemistry schemes, schemes involving at least two successive splitting steps (without re-merging) can be used if the necessary processing capacity is available. The disadvantage of such schemes is that they can rapidly produce a large number of individually treated samples. However, their advantage is that each sample does not require laborious deconvolution, as the cells within are exposed to only one set of conditions. Therefore, with suitable sample processing equipment, splitting methods can produce rapid results.

[0038] As explained above, another advantage of this invention is that reagents can be added sequentially and individually, or in bulk combinations in predetermined groups. This prevents the effect of one reagent from being masked by a dominant reagent that may be present in a mixed reagent combination used in conventional experiments.

[0039] In a preferred embodiment, the reagents are added to the cells sequentially and individually so that the effect of each reagent can be assessed separately.

[0040] The method of the present invention enables the testing of thousands or millions of reagents and reagent combinations in multiple high-throughput assays to determine the conditions necessary to achieve desired results in any cellular process.

[0041] In contrast to methods that combine multiple reagents and then remove them individually to identify the effective reagent, the method of this invention can test the sequential addition of reagents to determine the effect of each reagent when added in multiple sequential combinations, as well as to test multiple combinations of reagents. This can elucidate the individual steps in the pathway from one cell type to a second cell type, rather than just the overall conversion.

[0042] Advantageously, the reagents used induce alterations in cellular processes; these reagents are selected for generating cells in which gene expression is analyzed. Gene expression can be conveniently analyzed using any comparative expression monitoring technique, including PCR-based techniques such as RT-PCR, gene expression serial analysis (SAGE), RNA sequencing (RNAseq) including single-cell RNA seq, in situ hybridization including single-cell hybridization, or microarray techniques, such as those widely available from vendors (e.g., Affymetrix). In another aspect, the invention provides a method for generating nucleic acids encoding gene products that influence cellular processes, comprising identifying the gene as described above and generating at least the coding region of said gene by nucleic acid synthesis or biological replication.

[0043] In a further aspect, a method for inducing cellular processes in cells is provided, comprising the steps of: a) identifying one or more genes differentially expressed in relation to cellular processes according to the invention; and b) regulating the expression of said one or more genes in the cells. The expression of a gene in a cell can be regulated, for example, by transfecting or otherwise transferring the gene into the cell to cause it to be overexpressed transiently or permanently. Alternatively, the expression of an endogenous gene can be altered, for example by targeting enhancer insertion or by administering an exogenous agent that causes an increase (e.g., a placebo inducer) or a decrease (e.g., an antisense, RNAi, a transcription factor) in gene expression. Furthermore, the gene product itself can be administered to or introduced into the cell to achieve an increase in its activity. Additionally, agents that increase or decrease the activity of gene products (e.g., competitive and non-competitive inhibitors, drugs, pharmaceuticals) can be administered to the cell. In a further aspect, the invention provides a method for identifying the state of cellular processes in cells, comprising the steps of: a) identifying one or more genes differentially expressed in relation to cellular processes as described above; and b) detecting the regulation of the expression of said one or more genes in the cell, thereby determining the state of the cellular processes in the cell. Advantageously, the gene used in this analysis encodes a cellular marker that can be detected, for example, by immunoassay. Alternatively, the gene product can be an enzyme whose activity can be determined using fluorescence, colorimetry, radiometry, or other methods.

[0044] The present invention further provides a method for regulating cellular processes, comprising the steps of: a) determining the effect of one or more reagents on cells according to the foregoing aspects of the present invention; b) exposing cells to reagents that cause changes in cellular processes; and c) isolating the desired cells.

[0045] Therefore, the present invention provides a method for generating differentiated cells from a second different differentiated cell type according to the foregoing aspects of the present invention.

[0046] A method for identifying reagents capable of inducing cellular processes is also provided, comprising the steps of: a) determining the effect of one or more reagents on cells according to the foregoing aspects of the invention; and b) identifying those reagents that induce desired cellular processes in cells. Reagents identified according to the invention can be expressed or synthesized using conventional chemical, biochemical, or other techniques and used in methods for regulating specific cellular processes in cells, such as those described herein.

[0047] In a further aspect, the present invention relates to a method for developing and using a database of cellular gene expression data related to reagent exposure, which can provide improved reagent selection in step (a) of the method of the first aspect of the invention. By using machine learning and inputting experimental data from the method of the present invention, the database can be improved, reagent selection can be further optimized, and thus the performance of the method of the present invention can be optimized. Attached Figure Description

[0048] Figure 1 This is a diagram illustrating the operation of the CombiCult® split / merge selection method.

[0049] Figure 2 This is a schematic diagram of the genetic barcode marking of cells in the method of the present invention.

[0050] Figure 3 This is a schematic diagram of a 27-well assay design. This assay uses a fluorescent protein tag sequence to track cell differentiation without splitting / merging (one well out of 27). This protocol is derived from Feng et al., Stem Cell Reports 2014.

[0051] Figure 4 UMAP plots of cells isolated from control, GFP, and luciferase wells are shown. GFP wells were exposed to the final reagent containing the GFP marker and therefore were exposed to all conditions. Luciferase wells were not exposed to any conditions. BFP, RFP, and GFP markers were screened from these wells.

[0052] Figure 5 Showing with Figure 4 A similar UMAP plot, but the difference is that the presence or absence of labels is scored as either present or absent, rather than using a threshold level.

[0053] Figure 6 Showing with Figure 4 A similar UMAP plot, where RFP and BFP are detected together, allows scoring to be performed only when both markers are present.

[0054] Figure 7 , Figure 8 and Figure 9 Gene expression analysis and gene expression clustering from UMAP plots of 27-well experiments were elucidated.

[0055] Figure 10 The experimental design used in Examples 2 and 3 is explained.

[0056] Figure 11 This is a UMAP graph from the experimental results of Example 2.

[0057] Figure 12 It is a cluster analysis of the genes identified as expressed in the cells identified in Example 2.

[0058] Figure 13 This is a UMAP plot of the experimental results from Example 3.

[0059] Figure 14This is a schematic diagram of the experimental design in Example 4.

[0060] Figure 15 and Figure 16 This is a UMAP map of the cells isolated from Example 4, and the presence of markers in these cells. Detailed Implementation

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art, such as in peptide chemistry, cell culture and phage display, nucleic acid chemistry, and biochemistry. Molecular biological, genetic, and biochemical methods were employed using standard techniques (see Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., 2001, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Ausubel et al., Short Protocols in Molecular Biology (1999) 4th ed., John Wiley & Sons, Inc.; Patil and Sivaram, A Complete Guide to GeneCloning: From Basic to Advanced, Springer: ISBN 978-3-030-96853-3, 28 April 2023; Jaroszewicz et al., Phage display and other peptide display technologies, FEMS Microbiology Reviews, fuab052, 46, 2022, 1-25, and resources at Addgene.org), which are incorporated herein by reference.

[0062] As used herein, the term "culture conditions" refers to the environment in which cells are placed or exposed to promote the growth or differentiation of said cells. Therefore, the term refers to culture media, temperature, atmospheric conditions, substrate, agitation conditions, etc., that may affect cell growth and / or differentiation. More specifically, the term refers to specific reagents that can be incorporated into the culture medium and may affect cell growth and / or differentiation. In a preferred embodiment, culture conditions are the reagents or combinations of reagents to which cells are exposed.

[0063] The term "cell" as used herein refers to the smallest structural unit in an organism capable of independent function, or a single-celled organism composed of one or more nuclei, cytoplasm, and various organelles, all surrounded by a semi-permeable cell membrane or cell wall. Cells can be prokaryotic, eukaryotic, or archaea. For example, a cell can be a eukaryotic cell. Mammalian cells are preferred, especially human cells. Cells can be natural or modified, for example, through genetic manipulation or passage in culture to achieve desired characteristics. Stem cells, defined in more detail below, are totipotent, pluripotent, or multipotent cells capable of producing more than one differentiated cell type. Stem cells can differentiate in vitro to produce differentiated cells, which themselves can be multipotent or terminally differentiated. Cells differentiated in vitro are artificially produced cells by exposing stem cells to one or more agents that promote cell differentiation.

[0064] Cellular processes refer to any characteristic, function, process, event, cause, or result that occurs or is observed or can be attributed to a cell, whether intracellular or extracellular. Examples of cellular processes include, but are not limited to, vitality, aging, death, pluripotency, morphology, signal transduction, binding, recognition, molecular production or destruction (degradation), mutation, protein folding, transcription, translation, catalysis, synaptic transmission, vesicle transport, organelle function, cell cycle, metabolism, proliferation, division, differentiation, phenotype, genotype, gene expression, or the control of these processes.

[0065] Single cells can exist in cultures containing multiple single cells, existing as individual cell pools. These pools can be sorted, subdivided, and processed using any available technique to ultimately separate them into single cells. Examples include manual picking under a microscope, where individual cells are manually picked using a micropipette; flow cytometry and cell sorting (FACS), where cells are suspended in a fluid and excited with fluorescence, for example, using a laser beam. Detectors analyze the cells by size, shape, and fluorescence, then sort and collect droplets containing single cells. In magnetically activated cell sorting (MACS), cells are labeled with magnetic beads that bind specific cell surface markers. The cell suspension is then passed through a magnetic column that captures the labeled cells, which can then be eluted as single cells. Cells can also be sorted on microfluidic devices, by limiting dilution, and by automated robotic systems that physically separate cells.

[0066] A totipotent cell is a cell that has the potential to differentiate into any type of somatic cell or germ cell found in an organism. Therefore, any desired cell can be derived from a totipotent cell in some way.

[0067] Pluripotent cells are cells that can differentiate into more than one, but not all, cell types.

[0068] As used herein, a marker or tag is a means of identifying cell units and / or determining the culture conditions or culture condition sequences to which the cell units have been exposed. Preferably, the marker is a genetic barcode as further described herein.

[0069] Cells are exposed to culture conditions or reagents when they come into contact with a culture medium or grow under conditions that affect one or more cellular processes, such as cell growth, differentiation, or metabolic state. Therefore, if the culture conditions involve culturing cells in a medium in the presence of a reagent, the cells are placed in the medium containing the reagent for a sufficient duration to allow it to exert its effect. Similarly, if the conditions are temperature conditions, the cells are cultured at the desired temperature.

[0070] The merging of one or more cells or cell populations involves mixing cells or cell populations to produce a single group or pool containing cells from more than one background, i.e., cells that have been exposed to more than one different group of culture conditions. The pool may be further subdivided into multiple groups, randomly or non-randomly; for the purposes of this invention, such groups are not themselves “pools,” but they can be merged by combination, for example, after exposure to different group of culture conditions.

[0071] The terms "proliferating cell growth" and "cell proliferation" are used interchangeably herein to refer to a multiple increase in the number of cells without differentiation into different cell types or lineages. In other words, these terms refer to an increase in the number of living cells. Preferably, proliferation is not accompanied by significant changes in phenotype or genotype.

[0072] Cell transformation, as referred to in this article, encompasses all means of cell type transformation, including differentiation, transdifferentiation, reprogramming, and other definitions of cell type plasticity and modification.

[0073] Cell differentiation is the process by which a cell type develops into different cell types. For example, bispotent, pluripotent, or totipotent cells can differentiate into nerve cells. Differentiation can occur alongside proliferation or independently of it. The term "differentiation" generally refers to the phenotype of acquiring a mature cell type (e.g., neuron or lymphocyte) from a cell type with a less defined developmental stage, but it does not exclude transdifferentiation, i.e., the transformation of a mature cell type into another mature cell type, such as a neuron transforming into a lymphocyte.

[0074] The differentiation state of a cell is the level at which a cell differentiates along a specific pathway or lineage.

[0075] Cellular process state refers to whether a cellular process is occurring, and in complex cellular processes, it can represent a specific step or stage in that process. For example, a cell differentiation pathway may be inactive or may have been induced, and may involve multiple discrete steps or components, such as signal transduction events characterized by the presence of a particular set of enzymes or intermediates.

[0076] A gene is a nucleic acid that encodes a gene product (whether a polypeptide or an RNA gene product). As used herein, a gene includes at least a coding sequence that encodes the gene product; optionally, it may include one or more regulatory regions required for transcription and / or translation of that coding sequence.

[0077] Gene products are typically proteins encoded by genes in a conventional manner. However, the term also encompasses non-peptide gene products encoded by genes, such as ribonucleic acid (RNA).

[0078] Nucleic acid synthesis can be performed using any available technology. Preferably, nucleic acid synthesis is automated. Furthermore, nucleic acids can be produced through biological replication (e.g., by cloning and replicating in bacteria or eukaryotic cells) according to procedures known in the art.

[0079] Differentially expressed genes, which are expressed at different levels in response to cell culture conditions, can be identified by gene expression analysis (e.g., on a gene chip), sequencing (e.g., RNA-seq), or any method known in the art. Differentially expressed genes, relative to overall gene expression levels, show more or fewer mRNA or gene products in cells under the test conditions than under alternative conditions.

[0080] Transfected genes can be transferred into cells by any appropriate means. The term is used in this article to refer to routine transfection (e.g., using calcium phosphate), but also includes other techniques used to transfer nucleic acids into cells, including transformation, viral transduction, electroporation, etc.

[0081] The term regulation is used to describe the increase and / or decrease of a regulated parameter. Therefore, the regulation of gene expression includes both increasing and decreasing gene expression.

[0082] stem cells in Stem Cells: Scientific Progress and Future Research Directions. The Department of Health and Human Services, June 2001. http: / / www.nih.gov / news / stemcell / scireport.htm provides a detailed description of stem cells. The contents of that report are incorporated herein by reference.

[0083] Stem cells are cells capable of differentiating to form at least one, and sometimes many, specialized or differentiated cell types. The entire set of different cells that can be formed from stem cells is considered exhaustive; that is, it includes all the different cell types that make up an organism. Stem cells are present throughout the entire life cycle of an organism, from the early embryo where stem cells are relatively abundant to the adult where stem cells are relatively scarce. Stem cells, present in many tissues of adult animals, play an important role in normal tissue repair and homeostasis. The existence of these cells raises the possibility that they could provide a means of generating specialized functional cells that could be transplanted into the human body to replace dead or nonfunctional cells in diseased tissues. The list of diseases for which this could provide therapy includes Parkinson's disease, diabetes, spinal cord injury, stroke, chronic heart disease, end-stage renal disease, liver failure, and cancer.

[0084] Different stem cells possess varying potentials to form different cell types: spermatogonial stem cells are unipotent because they naturally produce only sperm, while hematopoietic stem cells are pluripotent, and embryonic stem cells are considered capable of generating all cell types and are referred to as totipotent or pluripotent. To date, three types of mammalian pluripotent stem cells have been isolated. These cells can generate cell types that are typically derived from all three germ layers of the embryo (endoderm, mesoderm, and ectoderm). These three types of stem cells are: embryonic carcinoma (EC) cells derived from testicular tumors; embryonic stem (ES) cells derived from preimplantation embryos (usually blastocysts); and embryonic germ (EG) cells derived from postimplantation embryos (usually fetal cells destined to become part of the gonads).

[0085] Differentiated cell types can alter their phenotype. This phenomenon, known as transdifferentiation, is the transformation of one differentiated cell type into another, with or without intervening cell division. Differentiation can sometimes be reversed or altered. In vitro protocols that can induce transdifferentiation in cell lines are now available. Furthermore, specialized cell types can be dedifferentiated to generate stem cell-like cells with the potential to differentiate into other cell types.

[0086] Many systems designed for the in vitro differentiation of stem cells are complex, multi-stage procedures in which the exact nature of each step and the timing of each step are important. For example, Lee et al. (2000, Nature Biotechnology, vol. 18, p 675-679) used a five-stage protocol to derive dopaminergic neurons from mouse ES cells: 1) Undifferentiated ES cells were expanded in ES cell culture medium on a gelatin-coated tissue culture surface in the presence of LIF; 2) Embryoids were generated in ES cell culture medium in suspension culture for 4 days; 3) Nestin-positive cells were screened from the embryoids in ITSFn medium for 8 days after being plated on a tissue culture surface; 4) Nestin-positive cells were expanded in N2 medium containing bFGF / laminoids for 6 days; 5) Finally, the expanded neuronal precursor cells were induced to differentiate by removing bFGF from the N2 medium containing laminoids. In a second example of continuous cell culture, Bonner-Weir et al. [Proc.Natl. Acad. Sci. (2000) 97: 7999-8004] derived insulin-producing cells from human pancreatic ductal cells by: 1) selecting ductal cells instead of islet cells for 2 to 4 days by selective adhesion to a solid surface in the presence of serum; 2) subsequently removing the serum and adding keratinocyte growth factor to select ductal epithelial cells instead of fibroblasts for 5 to 10 days; and 3) covering the cells with the extracellular matrix preparation “Matrigel” for 3 to 6 weeks. In a further example of continuous cell culture, Lumelsky et al. [Science (2001) 292: 1389-1394] derived insulin-secreting cells through directed differentiation of mouse embryonic stem (ES) cells by: 1) expanding ES cells in the presence of LIF for 2 to 3 days; 2) generating embryoid bodies in the absence of LIF for 4 days; 3) selecting nestin-positive cells using ITSFn medium for 6 to 7 days; 4) expanding pancreatic endocrine precursors in N2 medium containing 827 medium supplement and bFGF for 6 days; and 5) inducing differentiation into insulin-secreting cells by removing bFGF and adding nicotinamide.

[0087] However, in determining cell differentiation, it is important not only for the sequence and duration of individual steps, or the sequential addition of different factors. Since embryonic development is regulated by the action of a gradient of signaling factors that impart positional information, it is predictable that the concentration of a single signaling factor, as well as the relative concentrations of two or more factors, will be crucial for determining the fate of cell populations in vitro and in vivo. Factor concentrations change during development, and stem cells respond differently to different concentrations of the same molecule. For example, stem cells isolated from the CNS of late-stage embryos respond differently to different concentrations of EGF: low concentrations of EGF induce proliferative signals, while higher concentrations induce proliferation and differentiation into astrocytes. Many factors that influence stem cell self-renewal and differentiation in vitro have been found to be naturally occurring molecules. This is predictable because differentiation is induced and controlled by signaling molecules and receptors acting along signal transduction pathways. However, similarly, many synthetic compounds are likely to influence stem cell differentiation. Such synthetic compounds, which are likely to interact with cellular targets (so-called druggable targets) within signal transduction and signaling pathways, are routinely synthesized, for example, for drug screening by pharmaceutical companies. Once known, these compounds can be used to direct the differentiation of stem cells under in vitro conditions, or they can be administered in vivo, in which case they will act on resident stem cells in the patient's target organs.

[0088] Common variables in tissue culture: When developing conditions for successful culture of a specific cell type, or to achieve or regulate cellular processes, it is often important to consider various factors. One important factor is deciding whether to proliferate cells in suspension or as a monolayer attached to the matrix. Most cells tend to adhere to the matrix, although some cells (including transformed cells, hematopoietic cells, and cells derived from ascites) can proliferate in suspension.

[0089] Assuming adherent cells are being cultured, a crucial factor is the choice of adhesion matrix. Most laboratories use single-use plastics as tissue culture matrices. These plastics include polystyrene (the most common type), polyethylene, polycarbonate, Perspex, PVC, Teflon, cellophane, and cellulose acetate. While virtually any plastic may be used, many require processing to make them wettable and suitable for cell attachment. Furthermore, it is quite possible that any properly prepared solid matrix can be used to support cells; matrices used to date include glass (e.g., aluminoborosilicate glass and soda-lime glass), rubber, synthetic fibers, polymeric dextran, and metals (e.g., stainless steel and titanium). Certain cell types, such as bronchial epithelial cells, vascular endothelial cells, skeletal muscle cells, and neurons, require growth matrices coated with biological products (typically extracellular matrix materials such as fibronectin, collagen, laminin, polylysine, etc.). The growth matrix and application method (wet or dry coating, or gelation) can influence cellular processes such as cell growth and differentiation characteristics, which must be determined empirically as described above. Among the variables in cell culture, the choice of cell culture medium and supplements (such as serum) is perhaps the most obviously important. They provide the aqueous compartment for cell growth and contain nutrients and various factors, some of which are listed above, while others are less well-defined. Some of these factors are essential for adhesion, others are used for signal transduction (e.g., hormones, mitogens, cytokines), and still others act as detoxifiers. Commonly used media include RPMI 1640, MEM / Hank salt, MEM / Earle salt, F12, DMEM / F12, L15, MCDB 153, etc. The composition of various media can vary considerably; some common differences include sodium bicarbonate concentration, concentration of divalent ions (e.g., Ca and Mg), buffer composition, antibiotics, trace elements, nucleosides, peptides, synthetic compounds, drugs, etc. It is well known that different media are selective, meaning they promote the growth of only certain cell types. Culture medium supplements (such as serum, pituitary extracts, and other extracts) are often essential for cell growth in cultures. Furthermore, they frequently determine the phenotype of cells in the culture; that is, they can determine cell survival or guide differentiation. The role of supplements in cellular processes (such as differentiation) is complex and depends on their concentration, the timing of their addition to the culture, the cell type used, and the culture medium. The uncertainty surrounding the properties of these supplements and their potential to influence cell phenotype has driven the development of serum-free media. As with all media, their development has been largely achieved through trial and error, as discussed above. The gas phase in tissue culture is also important, and the composition and volume of gas used may depend on the type of culture medium used, the required buffer volume, whether the culture vessel is open or sealed, and whether specific cellular processes need to be modulated.Common variables include the concentrations of carbon dioxide and oxygen. Other conditions important for tissue culture include the choice of culture vessel, headspace size, seeding density, temperature, frequency of medium changes, enzyme treatment, and the rate and mode of agitation or stirring. Therefore, altering cell culture conditions is one way to achieve desired cellular processes. One aspect of the invention recognizes that altering cell culture conditions in a continuous manner can be a highly efficient method for achieving cellular effects. In various applications, such as in studies of cell differentiation, a specific series of different tissue culture conditions is often required to achieve cellular processes. Different conditions may include adding or removing substances from the culture medium at specific time points, or changing the culture medium. As mentioned above, such sets of conditions (examples of which are given below) are typically developed through trial and error.

[0090] Combined continuous cell culture (split-merge cell culture) A particularly efficient method for sampling large numbers of cell culture conditions is called combinatorial cell culture or split-and-combination cell culture. Figure 1 In one embodiment, this involves sequentially subdividing and combining multiple groups of cells to sample various combinations of cell culture conditions. In one aspect of the invention, the method operates by taking an initial starting culture of cells (or different starting cultures), dividing it into X aliquots, each aliquot containing multiple beads (groups / colonies / carriers / cells) grown individually under different culture conditions. After a given time of cell culture, cells can be merged by combining and mixing the beads from the different aliquots. This pool can be further subdivided into X² aliquots, each cultured under different conditions for a period of time, and subsequently merged. This iterative procedure of splitting, culturing, and merging cells (or merging, splitting, and culturing; depending on where the cycle begins) enables systematic sampling of many different combinations of cell culture conditions.

[0091] The complexity of the experiment, or in other words, the number of different combinations of cell culture conditions tested, is equal to the product of the number of different conditions sampled in each round (X1 x X2 x ... Xn). It should be noted that the step of merging all cells before subsequent splitting can be optional, and a step where a finite number of cells are merged can have the same effect.

[0092] Therefore, this invention embodies many related methods for systematically sampling multiple combinations of cell culture conditions, wherein multiple test units are processed in batches. Regardless of the exact manner in which multiple cell culture conditions are sampled, the procedure is efficient because multiple cells can share a single container in which they are cultured under the same conditions, and it can be performed at any one time using only a few culture containers (the number of culture containers in use equals the number of samples split).

[0093] In many respects, the procedure is analogous to the split-synthesis of large chemical libraries (known as combinatorial chemistry), which samples all possible combinations of bonds between chemical structural unit groups (see, for example: Combinatorial Chemistry, Oxford University Press (2000), Hicham Fenniri (Editor)). Split-and-merge cell cultures can be repeated an arbitrary number of rounds, and an arbitrary number of conditions can be sampled in each round. As long as the number of test units (cells or colonization beads in this example) is greater than or equal to the number of different conditions sampled across all rounds, and assuming that cell splitting occurs completely randomly, it can be expected that at least one cell unit has been cultured according to every possible combination of the experimentally sampled culture conditions. The procedure can be used to sample growth or differentiation conditions for any cell type, or to sample the efficiency of any cell type in producing biomolecules (e.g., erythropoietin or interferon). Because the procedure is iterative, it is well-suited for testing multi-step tissue culture protocols, such as those described above in conjunction with stem cell differentiation. Variables that can be sampled using this technique include cell type, duration of cell culture cycles, temperature, different culture media (including different concentrations of components), growth factors, conditioned media, co-culture with various cell types (e.g., feeder cells), animal or plant extracts, drugs, other synthetic chemicals, viral infections (including transgenic viruses), addition of transgenes, addition of antisense or antigene molecules (e.g., RNAi, triple helix), sensory input (for organisms), electrical stimulation, light stimulation, or redox stimulation.

[0094] The purpose of performing a split-and-merge process on cells in a split-and-merge cell culture is to systematically expose them to a predefined combination of conditions. Those skilled in the art will devise many different means to achieve this result. In addition to the split-and-merge process and its variations, it is worthwhile to briefly discuss the split-and-merge process. A split-and-merge process involves subdividing a group of cells at least twice without cell merging in between. If the split-and-merge process is used in a large number of rounds, the number of individual samples produced increases exponentially. In this case, a certain level of automation is essential, such as using robotic platforms and sophisticated sample tracking systems. The advantage of the split-and-merge step is that (because the cells are not merged) various cell lineages can be separated based on their cell culture history. Therefore, the split-and-merge step can be used to infer whether a particular cell culture condition is responsible for any given cell process, and thus to infer the cell's culture history.

[0095] The splitting and / or merging of cells according to a predetermined protocol can be performed completely randomly or can follow a predetermined protocol. When cells are randomly split and / or merged, the separation of a given cell unit into any group is not predetermined or biased in any way. To induce exposure of at least one cell unit to every possible combination of cell culture conditions with a high probability, it is advantageous to use a larger number of cells than the total number of cell culture condition combinations being tested. Thus, in some cases, it is advantageous to split and / or merge cells according to a predetermined protocol, with the overall effect of preventing accidental duplication or omission of combinations. The predetermined treatment of cells can optionally be planned in advance and recorded in a spreadsheet or computer program, and the splitting and / or merging operations can be performed using automated methods such as robotics. Cell labeling (see below) can be done by any of a number of means, such as RFID tagging, optical tagging, or spatial coding. Robotic devices capable of identifying samples and thus distributing samples according to a predetermined protocol have been described (see 'Combinatorial Chemistry, A practical Approach', Oxford University Press (2000), Ed H. Fenniri). Alternatively, standard laboratory liquid handling and / or tissue culture robots (such as those manufactured by Beckman Coulter Inc, Fullerton, CA; The Automation Partnership, Royston, UK) are capable of spatially encoding the identities of multiple samples and adding, removing, or transferring these samples according to pre-programmed protocols.

[0096] Cell analysis and / or isolation can be performed after each round of cell culture, or after a specified number of rounds, to observe any given cellular processes that may be affected by tissue culture conditions. The following examples are illustrative and are not intended to limit the scope of the invention.

[0097] After each round of cell culture, or after a specified number of rounds, the cells can be assayed to determine the presence of members exhibiting increased cell proliferation. This can be achieved through various techniques, such as visual examination of the cells under a microscope, or by quantifying cell-specific marker products. These can be endogenous markers, such as specific DNA sequences, or cellular proteins that can be detected by ligands or antibodies. Alternatively, exogenous markers, such as green fluorescent protein (GFP), can be introduced into the cells being assayed to provide cell-specific readings. Live cells can be visualized using a variety of live dyes, or conversely, dead cells can be labeled using various methods, such as propidium iodide. Furthermore, labeled cells can be separated from unlabeled cells using various techniques, both manual and automated, including affinity purification (“panning”), or by fluorescence-activated cell sorting (FACS) or substantially similar techniques (Figure 17). Depending on the application, standard laboratory equipment may be used, or the use of specialized instruments may be advantageous. For example, some analytical and sorting instruments (see, for example, Union Biometrica Inc., Somerville MA, USA) have flow cell diameters up to one millimeter, enabling flow sorting of beads up to 500 micrometers in diameter. These instruments provide readings of bead size and optical density, as well as two fluorescence emission wavelengths from tags such as GFP, YFP, or OS-red. Sorting speeds of up to 180,000 beads per hour, and dispensing into multiwell plates or batch receivers, can be achieved. After each round of cell culture, or after a specified number of rounds, cells can be assayed to determine the presence of members displaying a specific genotype or phenotype. Genotyping can be performed using well-known techniques such as polymerase chain reaction (PCR), fluorescence in situ hybridization (FISH), DNA sequencing, etc. Phenotyping can be performed using a variety of techniques, such as visual examination of cells under a microscope or by detecting cell-specific biomarker products. This could be endogenous markers such as specific DNA or RNA sequences, or cellular proteins that can be detected through ligands, enzyme substrate conversion, or antibodies that recognize specific phenotypic markers (see Appendix Eof). Stem Cells: Scientific Progress and Future Research Directions.Department of Health and Human Services. June 2001; the appendix mentioned is incorporated herein by reference. Genetic markers can also be exogenous, i.e., markers introduced into a cell population, for example, through transfection or viral transduction. Examples of exogenous markers are fluorescent proteins (e.g., GFP) or cell surface antigens that are not typically expressed in a particular cell lineage, are epitope-modified, or are from a different species. Transgenic or exogenous marker genes with associated transcriptional control elements can be expressed in a manner that reflects a pattern representing an endogenous gene. This can be achieved by associating the gene with a minimal cell type-specific promoter or by integrating the transgene into a specific locus (see, for example, European Patent No. EP0695351). Labeled cells can be separated from unlabeled cells using a variety of techniques, whether manual or automated, including affinity purification (“panning”) or by fluorescence-activated cell sorting (FACS). Nishikawa et al. (1998, Development vol 125, pp. 1747-1757) used cell surface markers recognized by antibodies to track the differentiation of pluripotent mouse ES cells. Using FACS, they were able to identify and purify hematopoietic lineage cells at various stages of differentiation.

[0098] Alternative or complementary techniques for enriching cells with specific genotypes or phenotypes are essential components of genetic selection. This can be achieved, for example, by introducing a selective marker into cells and measuring viability under selective conditions; see, for example, Soria et al. (2000, Diabetes vol 49, p1-6), who used such a system to select insulin-secreting cells from differentiated ES cells. Li et al. (1998, Curr Biol vol 8, p 971-974) identified neural progenitor cells by integrating a bifunctional selection marker / reporter gene f3geo (which provides f3-galactosidase activity and G418 resistance) into the Sox2 locus in mouse ES cells via homologous recombination. Since one of the characteristics of neural progenitor cells is the expression of Sox2, and therefore the expression of the integrated marker gene, these cells can be selected from non-neuronal lineages by adding G418 after differentiation induced with retinoic acid. Cell viability can be determined by microscopic examination or by monitoring f3-gal activity. Unlike phenotype-based selection methods, which may be limited by the availability of suitable ligands or antibodies, genetic selection can be applied to any differentially expressed gene.

[0099] Determining cell identity or cell culture history can become confusing when dealing with large numbers of cells, whose identity and / or cell culture history (e.g., the temporal sequence and exact nature of a series of culture conditions any given group or cell may have been exposed to) can be complex. For example, split-and-merge schemes for cell culture necessarily involve mixing cells in each round, making it difficult to track individual cells. Determining the cell culture history of a particular cell within a mixture of cells that has undergone multiple culture conditions is sometimes referred to as “deconvolution” of cell culture history. One way to perform this is by labeling cells, and therefore labeling cells is advantageous. Labeling can be done at the beginning of the experiment or in each round of the experiment, and can involve unique labels (which may be modified or unmodified during the experiment) or a set of labels containing a unique set. Similarly, the reading of labels can be done during each round or only at the end of the experiment.

[0100] Cell labeling methods involve sequentially associating unique tags with cells each time they are cultured under different conditions, so that subsequent detection and identification of the tags provides a clear record of the temporal sequence and identity of the cell unit's exposure to the cell culture conditions. Tags can be taken up by the cells or attached to the cell surface via adsorption, suitable ligands, or antibodies. For example, a simple tag that can be introduced into a cell is an oligonucleotide of a specified length and / or sequence. Oligonucleotides can contain any class of nucleic acids (e.g., RNA, DNA, PNA, linear, circular, or viral) and can contain specific sequences for amplification (e.g., primer sequences for PCR) or markers for detection (e.g., fluorophores or quenchers, or isotopic tags). Detection of these tags can be direct, such as by sequencing the oligonucleotides or hybridizing them with complementary sequences (e.g., on microarrays or chips), or indirect, such as by monitoring the gene product encoded by the oligonucleotide or the interference of the nucleotide on cell viability (e.g., antisense repression of a specific gene). A favorable method for amplifying nucleic acids is rolling circle amplification (RCA; 2002, V. Demidov, Expert Rev. Mo / . Diagn. 2(6), p.89-95), where the nucleic acid tag may contain a circularized strut of an RCA template, an extension primer, or an auxiliary microloop template.

[0101] Preferred labeling techniques advantageously used when the test unit is a single cell include nucleic acid modification of the cell by inserting a nucleic acid “barcode” through viral infection (e.g., via lentivirus) or gene editing (e.g., via CRISPR).

[0102] Reading of genetic modifications is advantageously performed via single-cell sequencing. Single-cell sequencing can be used to examine the genome, transcriptome, and epigenome landscape at the resolution of individual cells. Exemplary techniques include: Single-cell RNA sequencing (scRNA-seq), including full-length transcript methods such as SMART-seq and SMART-seq2, captures full-length transcripts and provides information about transcript subtypes; and 3' and 5' end methods, including Drop-seq, 10x Genomics Chromium, and CEL-seq2. These methods capture the 3' or 5' ends of mRNA, enabling the counting of gene expression levels. See Picelli, S., Faridani, OR, Björklund, AK, Winberg, G., Sagasser, S., & Sandberg, R. (2014). Full-length RNA-seq from single cells using Smart-seq2. Nature Protocols, 9(1), 171-181. DOI: 10.1038 / nprot.2014.006; Highly Parallel Genome-wide Expression Profiling ofIndividual Cells Using Nanoliter Droplets. Cell, 161(5), 1202-1214. DOI:10.1016 / j.cell.2015.05.002; Zheng, GXY, Terry, JM, Belgrader, P., Ryvkin, P., Bent, ZW, Wilson, R., Ziraldo, SB, Wheeler, TD, McDermott, GP, Zhu, J., Gregory, MT, Shuga, J., Montesclaros, L.,Underwood, JG, Masquelier, DA, Nishimura, S.Y., Schnall-Levin, M.,Wyatt, PW, Hindson, CM, ... & Livak, KJ (2017). Massively paralleldigital transcriptional profiling of single cells. Nature Communications, 8,14049. DOI: 10.1038 / ncomms14049. .

[0103] Another approach is single-cell DNA sequencing (scDNA-seq). Examples include whole-genome amplification (WGA). Techniques such as MDA (multiple substitution amplification), DOP-PCR (degenerate oligonucleotide primer PCR), and MALBAC (multiple annealing circular amplification) are used to amplify DNA from single cells.

[0104] See Zong, C., Lu, S., Chapman, AR, & Xie, XS (2012). Genome-wide Detection of Single-Nucleotide and Copy-Number Variations of a SingleHuman Cell. Science, 338(6114), 1622-1626. DOI: 10.1126 / science.1229164;Navin, N., Kendall, Tumour evolutioninferred by single-cell sequencing. Nature, 472(7341), 90-94. DOI: 10.1038 / nature09807.

[0105] If the tag is an epigenetic tag, it can be detected using single-cell epigenomics, including using techniques such as single-cell bisulfite sequencing to study DNA methylation at the single-cell level.

[0106] See Smallwood, SA, Lee, HJ, Angermueller, C., Krueger, F., Saadeh, H., Peat, J., Andrews, SR, Stegle, O., Reik, W., & Kelsey, G. (2014). Single-cell genome-wide bisulfite sequencing for assessing epigeneticheterogeneity. Nature Methods, 11(8), 817-820. DOI: 10.1038 / nmeth.3035.

[0107] Some technologies involve multiple approaches and can be referred to as single-cell multi-omics. Examples include G&T-seq (genome and transcriptome sequencing), which can simultaneously sequence the genome and transcriptome of a single cell.

[0108] CITE-seq (Cell Indexing of Transcriptome and Epitopes via Sequencing) combines scRNA-seq with antibody-derived tags to analyze proteins and RNA from the same cell.

[0109] See Macaulay, I. C., Haerty, W., Kumar, P., Li, Y. I., Hu, T. X., Teng, M. J., Goolam, M., Saurat, N., Coupland, P., Shirley, L. M., Smith, M., Van der Aa, N., Banerjee, R., Ellis, P. D., Quail, M. A., Swerdlow, H. P., Zernicka-Goetz, M., Livesey, F. J., & Ponting, C. P. (2015). G&T-seq: parallel sequencing of single-cell genomes and transcriptomes. Nature Methods, 12(6), 519-522. DOI: 10.1038 / nmeth.3370; - Stoeckius, M., Hafemeister, C., Stephenson, W., Houck-Loomis, B., Chattopadhyay, P. K., Swerdlow, H., Satija, R., & Smibert, P. (2017). Simultaneous epitope and transcriptome measurement in single cells. Nature Methods, 14(9), 865-868. DOI: 10.1038 / nmeth.4380。

[0110] Protein expression can also be studied at the single-cell level, for example using flow cytometry (CyTOF), which uses heavy metal-labeled antibodies to study protein expression at the single-cell level. See Bendall, SC, Simonds, EF, Qiu, P., Amir, ED, Krutzik, PO, Finck, R., Bruggner, RV, Melamed, R., Trejo, A., Ornatsky, OI, Balderas, RS, Plevritis, SK, Sachs, K., Pe'er, D., Tanner, SD, & Nolan, GP (2011). Single-cell masscytometry of differential immune and drug responses across a humanhematopoietic continuum. Science, 332(6030), 687-696. DOI: 10.1126 / science.1198704.

[0111] Cell reprogramming involves altering the identity or function of a cell. This may involve transforming a cell from one type to another, or restoring a mature cell to a more pluripotent or stem cell-like state before it redifferentiates into a more directed state.

[0112] In the context of this invention, reprogramming refers to the transformation of a differentiated, directed cell type into another differentiated, directed cell type. Reprogramming does not include the simple conversion of pluripotent cells into differentiated cell types.

[0113] Many techniques for reprogramming are known to those skilled in the art; in one implementation, reprogramming as referred to herein may encompass any one or more of these techniques.

[0114] Induced pluripotent stem cells (iPSCs): This method, discovered by Shinya Yamanaka and colleagues in 2006, involves introducing a specific set of transcription factors (often referred to as Yamanaka factors: Oct4, Sox2, Klf4, and c-Myc) into mature cells.

[0115] These factors induce mature cells to revert to a pluripotent state, similar to embryonic stem cells. These pluripotent cells can then be further differentiated into the desired cell type.

[0116] Yamanaka uses retroviral vectors to deliver TF, but there are many other available methods, including non-integrative viral vectors, mRNA delivery, plasmids and direct addition of TF protein, as well as the use of non-coding RNA.

[0117] Various types of ncRNAs, including microRNAs (miRNAs) and long non-coding RNAs (lncRNAs), can participate in the regulation of transcription factors. These ncRNAs can be used to replace TF in the assay of this invention.

[0118] MicroRNAs (miRNAs) are short (approximately 22 nucleotides) ncRNAs that typically bind to the 3' untranslated region (UTR) of target mRNAs, leading to their degradation or translational repression. Through this mechanism, miRNAs can negatively regulate the levels of specific transcription factors (Bartel DP. (2009) MicroRNAs: target recognition and regulatory functions. *Cell*. 136(2):215-33).

[0119] Long noncoding RNAs (lncRNAs) are longer than 200 nucleotides and have multiple functions. Some lncRNAs can directly interact with transcription factors, regulating their activity, and therefore can be used to regulate TF activity in the assays according to the present invention. See Rinn JL, Chang HY. (2012) Genome regulation by long noncoding RNAs. Annual Review of Biochemistry. 81:145-66.

[0120] Non-coding RNAs have been shown to bind to or act as alternatives to classical Yamanaka factors: 1. miR-302 / 367 cluster: This microRNA cluster has been shown to improve reprogramming efficiency when used in combination with Yamanaka factors (Oct4, Sox2, Klf4, and c-Myc). See Anokye-Danso F, et al. (2011) Highly efficient miRNA-mediated reprogramming of mouse and human somatic cells topluripotency. Cell Stem Cell. 8(4):376-88.

[0121] 2. **lncRNA-RoR**: This lncRNA is involved in the regulation of reprogramming and the maintenance of pluripotency. lncRNA-RoR has been shown to regulate the core transcriptional network of pluripotency, including influencing the level of Yamanaka factor. See Loewer S, et al. (2013) Large intergenic non-coding RNA-RoR modulates reprogramming of human induced pluripotent stem cells. *Nature Genetics*. 45(12):1504-9.

[0122] Direct lineage transformation (transdifferentiation): This method bypasses the pluripotent state. Instead, it uses a combination of lineage-specific transcription factors to allow cells to directly transform from one mature cell type to another.

[0123] For example, fibroblasts can directly differentiate into neurons, cardiomyocytes, or other cell types, depending on the factors used. Currently, there are many techniques available for transdifferentiation: 1. Viral vector-mediated introduction of transcription factors: Viral vectors are used to deliver lineage-specific transcription factors into cells. For example: Fibroblasts are converted into neurons using factors such as Ascl1, Brn2 and Myt1l (Vierbuchen et al., (2010). Nature, 463(7284), 1035-1041).

[0124] Fibroblasts were converted into cardiomyocytes using factors such as Gata4, Mef2c and Tbx5 (Ieda et al., (2010) Cell, 142(3), 375-386).

[0125] 2. RNA-based transcription factor delivery: This method uses synthetic modified mRNA to transiently express essential transcription factors. RNA-based approaches can alleviate concerns associated with potential DNA integration into the host genome. Warren et al., (2010) Cell Stem Cell, 7(5), 618-630.

[0126] 3. Small molecules and chemicals: Although transcription factors are most commonly used for direct lineage conversion, some studies have successfully used combinations of small molecules to improve the efficiency of the process, or in rare cases, transdifferentiation can be achieved without genetic factors.

[0127] For example: Hou et al., (2013). Science, 341(6146), 651-654, successfully reprogrammed mouse somatic cells into pluripotent stem cells using only small molecules, highlighting the potential for generating iPSCs without introducing exogenous genes.

[0128] Shi et al. (2008). Cell stem cell, 2(6), 525-528 showed how combinations of small molecules can significantly improve the efficiency of iPSC generation from mouse fibroblasts.

[0129] Fu et al., (2015). Cell Research, 25(9), 1013-1024 used a set of small molecules to directly convert mouse fibroblasts into cardiomyocytes.

[0130] 4. CRISPR / Cas9-based endogenous gene activation: The CRISPR / Cas9 system, initially known for its gene-editing capabilities, can be modified to achieve gene activation by using inactivated Cas9 (dCas9) fused with a transcription activator. This can be used to activate endogenous lineage-specific genes, thereby promoting transdifferentiation. Gilbert et al., (2013) Cell, 154(2), 442-451.

[0131] 5. microRNA (miRNA): Some miRNAs have been found to play a role in cell fate determination and can be used in conjunction with transcription factors or alone to drive direct lineage conversion. Jayawardena, et al., (2012) Circulation Research, 110(11), 1465-1473.

[0132] 6. Epigenetic regulators: Direct lineage transformation involves significant changes in the cellular epigenetic landscape. Incorporation of molecules that regulate epigenetic markers can improve the efficiency of transdifferentiation. Polo, J et al., (2012) Cell, 151(7), 1617-1632.

[0133] 7. Extracellular clues and microenvironment: Although transcription factor-mediated transformation is the main driving factor, the efficiency and success rate of the transformation process may be affected by the microenvironment, including extracellular matrix composition, neighboring cells, and culture conditions. Engler, AJ, et al., (2006). Cell, 126(4), 677-689.

[0134] In the context of this invention, reprogramming is preferably transdifferentiation, through which cells directly transform from one differentiation state to another without going through a pluripotent state.

[0135] The term "cell type" as used herein refers to a differentiated cell belonging to a specific lineage. A cell type is not a totipotent cell. In the implementation scheme, the cell type is not a pluripotent cell capable of differentiating into more than one lineage. A cell lineage can be any of the various cell types capable of developing during normal embryogenesis. Examples include: Hematopoietic lineage: This lineage produces all the different blood cells. Hematopoietic stem cells (HSCs) in the bone marrow differentiate into various types of blood cells, such as red blood cells, various white blood cells (neutrophils, lymphocytes, and monocytes), and platelets.

[0136] Neuronal lineages: Neural stem cells differentiate into various types of neurons and glial cells. These include motor neurons, sensory neurons, astrocytes, oligodendrocytes, and microglia.

[0137] Epidermal lineage: Epidermal stem cells in the skin differentiate into various skin cells, including keratinocytes, melanocytes, and hair follicle cells.

[0138] Muscle lineage: Myogenic stem cells (or satellite cells) can differentiate into myofibrils or myocytes.

[0139] Mesenchymal stem cell (MSC) lineages: MSCs are pluripotent and can generate various cell types, including osteoblasts, chondrocytes, fibroblasts, and adipocytes.

[0140] Intestinal lineage: Stem cells in the intestinal crypts differentiate into various types of intestinal cells, such as absorptive intestinal epithelial cells, goblet cells, Paneth cells, and intestinal endocrine cells.

[0141] Liver lineage: Liver stem cells can differentiate into hepatocytes (the main functional cells of the liver) and bile duct cells (arranged in the bile ducts).

[0142] Cardiac lineage: Cardiac progenitor cells can differentiate into various cell types of the heart, such as cardiomyocytes, endothelial cells, and smooth muscle cells.

[0143] Germ cell lineage: This lineage causes the formation of eggs in females and sperm in males.

[0144] Retinal lineage: Retinal progenitor cells can differentiate into various types of retinal cells, such as photoreceptors, bipolar cells, and ganglion cells.

[0145] Reprogramming can involve the transformation of cells from one lineage to another, or the transformation of one cell type within a lineage to another cell type within the same lineage, such as the transformation from fibroblasts to osteoblasts.

[0146] The term "reagent" as used herein can refer to any chemical, small molecule, transcription factor, nucleic acid, protein, compound, or ion capable of interacting with cells. In a preferred embodiment, the reagent is a transcription factor, small molecule, or protein compound, or nucleic acid. In the most preferred embodiment, the reagent is a transcription factor (TF) or an agent that affects the activity of transcription factors, such as ncRNA.

[0147] As cited in the references cited herein and described in more detail below, certain reagents are known in the art to promote reprogramming. Transcription factors are commonly used to influence reprogramming.

[0148] Label deconvolution refers to the identification of labels on cells and, optionally, the order in which cells acquired these labels, thereby determining which reagents or combinations of reagents the cells have been exposed to. Since cells are labeled whenever they are exposed to a given reagent, the reagent leaves an "imprint" on the cell that can be used to map the history of the cell's exposure to the reagent and the timing of that exposure.

[0149] Predicting the most effective transformation-inducing reagents requires a combination of experimental data, computational modeling, and optimization techniques. Algorithms can be used to investigate the outcomes of cell exposure to a given reagent, thereby predicting which reagents can be used to induce the desired transformation in a given cell type. The results of reagent exposure should be presented to the algorithm as a consistent metric; for example, the results could be in the form of gene expression data, cell morphology scores, specific gene activation events, etc. The output from this algorithm can be used to select the initial reagent for executing the method of this invention. In the case of multiple executions of the method, the output data can be used to inform further execution instances of the method.

[0150] Machine learning (ML, AI) can be used to better predict the input reagents used for the desired transformation results. A database of gene expression data related to the reagents can be provided, and a gene regulatory network can be constructed that further includes machine learning models configured to output gene regulatory maps.

[0151] The merging, splitting, or subdivision of cell populations are performed as described in WO2004013369. Iterative repetition of the steps of the method of this invention can also be performed in a manner similar to WO2004013369 and as described above.

[0152] Cell type identification can be performed by any suitable means, including genetic, morphological, immunochemical, or other characteristics.

[0153] Gene expression analysis includes analyzing one or more reporter genes, as well as analysis based on multiple cellular genes, including the use of arrays that analyze expression across the entire genome. In one embodiment, gene expression analysis focuses on one or more genes as markers of a desired cell type. The method of the present invention can be used to identify these genes by analyzing changes in gene expression in cells exposed to reagents that, according to analysis of alternative cell differentiation markers (e.g., immunochemical markers or cell morphology), lead to a desired transformation outcome.

[0154] Genes that influence cell transformation can then be purposefully modulated to induce transformation; for example, information about how genes are modulated in response to exposure to reagents can be used to select gene regulation techniques (such as genome editing) to further modulate the gene in place of reagent application.

[0155] CombiCult® Procedures for splitting and merging cell populations to determine the effect of reagents on cell fate are described in WO2004031369. This combinatorial screening platform is capable of testing thousands of time-resolved combinations of culture conditions simultaneously. It is based on combinatorial science, which has been successfully used in the chemical synthesis of new compounds. In short, stem cells are grown in or on beads that can be labeled with condition-specific fluorescent markers. At each step of the differentiation process, the beads are merged and split, and subjected to different conditions under which they are simultaneously labeled in a condition-specific manner.

[0156] This process repeats multiple steps to ensure that each bead is exposed to a specific combination of conditions, and that multiple beads undergo the same combination. Cells within the beads are then screened using antibodies against markers and selected using a large-particle sorter. Positive beads are isolated and dissociated, and their tags are read using FACS, thus detecting all tag combinations enriched in positive beads. A computational deconvolution strategy identifies and quantifies the most frequent tag combinations, which correspond to molecules added to a specific set and sequence in the cell culture (also known as a “scheme”). Highly representative schemes are then independently validated in vitro and compared to gold-standard differentiation or expansion procedures. CombiCult® has been successfully deployed in multiple cases, enabling the discovery of novel schemes, such as those for hematopoietic stem cell expansion, macrophage, neutrophil, natural killer cell, megakaryocyte, smooth muscle cell, and oligodendrocyte progenitor differentiation.

[0157] The decision of which molecules to include in the CombiCult® screening comes from a careful reading of the literature, meaning its ability to discover new combinations is somewhat limited by the hypotheses already put forward. By leveraging the power of modern machine learning models combined with a large number of high-resolution single-cell transcriptomics datasets, this invention is able to generate new data-driven hypotheses with an unprecedented breadth.

[0158] Furthermore, CombiCult®'s multiplex analysis potential is limited by the use of beads and the fact that each screening selects one positive cell fate while discarding the remaining material.

[0159] This invention provides: 1. Define a custom computational analysis of the molecular inputs to be tested, including a machine learning algorithm to prioritize molecular drivers of cell fate acquisition and prevent contradictory or conflicting pathways from being tested together; 2. Single-cell labeling methods completely bypass the need for sorting beads and large particles, while obtaining single-cell clarity and monitoring of differentiation results based on gene expression through low-cost single-cell sequencing (e.g., Oxford Nanopore benchtop sequencing); 3. The second computational pipeline analyzes the results of combinatorial screening by identifying viral barcodes and their representativeness in obtaining target cell fates.

[0160] These three elements create a virtuous cycle, where information from one screening can be used to better optimize the initial predictions, as each round generates increasing experimental evidence linking molecules to their transcriptional functions. Ultimately, once a critical amount of evidence is generated from multiple experiments, a novel AI-based approach can be implemented where, for a given spectrum, an ordered set of molecular drivers can be efficiently predicted, effectively reducing the need for combinatorial testing.

[0161] Computational analysis The results of a single CombiCult screening can be used to select input reagents for further screening, or alternatively, to determine the optimal reagents for the desired transdifferentiation protocol.

[0162] A large amount of publicly available single-cell RNA sequencing (scRNA-seq) and single-cell chromatin accessibility data (scATAC-seq) from actively differentiating cells reveal a continuum of gene expression changes.

[0163] Typically, the algorithm used to select reagents can process the screening results as follows: enter: 1. A list of cells with their respective identities: cell_list 2. Reagent list: reagent_list 3. Data on the effect of each reagent on each cell: effect_data (this can be gene expression profiles, changes in cell phenotype, or any measurable measure after exposure) 4. Required cell type profile: desired_profile Output: Selected reagent list: selected_reagents step: 1. Initialize an empty list: selected_reagents = [] 2. For each cell in cell_list: a. For each reagent in reagent_list: i. Obtain the effect of the reagent on the cell from effect_data.

[0164] ii. Calculate the similarity score between this influence and the desired_profile. (This can be a relevance score, a frequency measure, or any relevant similarity metric.) b. Rank the reagents used to process the cells based on similarity scores.

[0165] c. Store the reagent that ranks highest for this cell.

[0166] 3. Compile a list of the reagents that rank highest across all cells.

[0167] 4. For each reagent in the summary list: a. Calculate the frequency at which it is selected in all cells.

[0168] b. If the frequency is higher than a certain threshold (indicating consistent efficacy): i. Add the reagent to selected_reagents.

[0169] 5. Returns selected_reagents.

[0170] By further sorting cells into cell types, the algorithm is able to return selected_reagents associated with transdifferentiation events of a specific cell type.

[0171] Machine learning can be applied to this algorithm to further optimize reagent selection. For example, machine learning algorithms can be applied in the following ways: enter: 1. Training dataset: contains cell identity, time and identity of reagent exposure, and the observed effects after exposure.

[0172] 2. Required cell type profile: desired_profile Output: The optimal reagents for transdifferentiation: optimal_reagents step: 1. Data preprocessing: a. Normalize and standardize the data (e.g., z-score normalization for gene expression levels).

[0173] b. Split the dataset into a training set and a validation set.

[0174] 2. Feature selection: a. Select relevant features (e.g., specific gene expression levels, cell markers).

[0175] b. If the dataset is high-dimensional, consider dimensionality reduction methods such as PCA.

[0176] 3. Model selection: a. Choose an appropriate ML model. Regression models (such as random forests or gradient boosting machines (GBM)) are examples.

[0177] b. Train the model on the training set.

[0178] c. Validate the model on the validation set.

[0179] 4. Hyperparameter tuning: a. Use techniques such as grid search or random search to find the optimal hyperparameters for the selected model.

[0180] b. Retrain the model using the optimal hyperparameters.

[0181] 5. Prediction: a. Use trained models to predict the effect of each reagent on transdifferentiation toward the desired cell type spectrum.

[0182] b. Rank the reagents based on the effectiveness of the prediction.

[0183] 6. Model Explanation (Optional but Recommended): a. Use techniques such as SHAP (SHapley Additive exPlanations) or feature importance scoring to explain which features (e.g., time-specific, cellular markers) are most influential in prediction.

[0184] 7. Return: a. Return the highest-ranked reagent as optimal_reagents.

[0185] The algorithm of this invention applies gold-standard data preprocessing steps and machine learning methods to identify temporal (or pseudo-temporal) trajectories and their associated gene expression changes. These trajectories connect stem cells to terminally differentiated cells through a continuum of gene expression changes. This continuum can be viewed as a high-resolution sequence of molecular events specific to cell fate acquisition. Within this continuum, a set of agents driving gene expression changes can be detected. The algorithm of this invention prioritizes and ranks these agents based on four main characteristics: 1. Their trends along the differentiation trajectory (differential expression) 2. Their gene regulatory networks were inferred through multiple regression and externally validated reagent-target networks and protein-protein interaction networks (PPIN). 3. Their specificity, that is, their preferential activity within a particular trajectory compared to all other trajectories in similar physiological processes (e.g., selection of T cell-specific driver groups compared to other lymphoid lineages during hematopoiesis). 4. Their dynamics in terms of switching time and / or instantaneous expression.

[0186] In the case of TF reagents, if matching public scATAC data are available, they can be supplemented by TF activity analysis in open chromatin regions, thereby providing orthogonal evidence for gene regulatory networks and supplementing other datasets.

[0187] Once the TFs are identified, the target networks from each of them are screened for annotated interactions. If the interactors of two TFs show a significant degree of overlap but involve opposite interactions (e.g., one TF is reported to inhibit these interactors, while the other is reported to activate them), they will be marked as incompatible in the same screening, thus providing a more reasonable way to arrange the combinations.

[0188] The algorithm of this invention can utilize regression-based gene-gene association networks and existing PPINs to identify molecules that act as transducers and receptors in signal transduction pathways upstream of the TF and also change along the trajectory. Once candidate receptors are identified, their ligands can be inferred and included in combinatorial screening.

[0189] Finally, the small molecule-protein interaction network overlaid on PPIN, and the drug-transcriptome correlation database whose features can be explored along the trajectory, can also be included in the software for prioritization and inclusion in the screening.

[0190] With the emergence of new hypotheses, improvements can be made by significantly expanding the multi-analysis and resolution capabilities of the CombiCult platform. The implementation in WO2004031369 is limited to identifying one cell differentiation at a time, meaning that many other possibilities that may arise during culture (e.g., alternative cell fates) remain unexplored. In this invention, the use of viral transduction markers (“barcoding”) of single cells avoids this problem. To avoid using beads (whose encapsulation is a time-consuming and resource-intensive step) and to track all single cells during screening, barcoding is achieved through a unique, concise, expressed DNA sequence that acts as a unique conditional tag (UCT) delivered via a lentiviral vector, which is stably integrated into the cultured cells. Figure 2 When cells undergo different splits and are treated with different molecules / TFs / growth factors (“driver factors”), they are tagged with lentiviral vectors carrying barcodes specific to the particular condition. By using an integrated viral tag at low doses (so that cells do not become saturated in each transduction cycle), cells can be uniquely barcoded for each condition, and UCT is retained within each cell, tracking the condition under which they were inserted. In the first split, candidate driver factors and their corresponding vectors are added to the culture medium, which can be replaced after transduction to remove the viral vectors, depending on the protocol requirements.

[0191] The cells were then separated, resuspended, merged together, and split in different plates for a second split.

[0192] The merged and split cells underwent a second round of transduction and driver factor treatment; the culture medium was changed and the driver factor was re-added until the third merge and split process.

[0193] Repeat this process for as many splits as needed until the final split is achieved and the cells are isolated, optionally sorted via FACS and ready for single-cell RNA sequencing. The fundamental difference between this method and the method in WO2004031369 is that, in this invention, barcoding occurs at the single-cell level, and the data consists of the transcriptome of each cell and its corresponding UCT combination.

[0194] Viral tagging and reagent delivery to cells can be combined. The viral delivery and overexpression of TFs are described in Joung et al., Cell 186, 209–229, January 5, 2023; the same method can be used to deliver tags to label cells, or the reagent transgene itself can be considered a tag and detected by sequencing.

[0195] Develop new analytical pipelines for selecting options. Single-cell RNA analysis combined with UCT detection resolved the impact of different cell fate choices (including successful and unsuccessful ones) in a single snapshot. Analysis of transcriptome data annotated cells through cell similarity and the expression of cell type-specific traits. By reconstructing the UCT “lineage” tree and matching it with traits, it was possible to map the sequences of variables that cause cell fate-specific transcriptional profiles. Using single-cell data combined with UCT has three fundamental advantages: first, it can resolve even minute differences in cell fate acquisition; second, if the entire experiment is sequenced and analyzed, it can identify multiple cell fates at once; and third, it produces reusable and integrable data.

[0196] Regarding the first two aspects, scRNA-seq has been deployed in a similar manner to track guide RNA constructs in CRISPR screening (Dixit et al., 2016. Cell 167, 1853–1866) or lentivirus-mediated TF transduction (Joung et al., 2023. Cell 186, 209–229). Furthermore, complex studies using sequential genome editing (“trace”) at the single-cell level have been successfully conducted (Spanjaard et al., Nat Biotechnol. 2018 June; 36(5):469–473), providing the first analytical framework for connecting different events in a single experiment.

[0197] These pioneering methods have been used to study transcriptional responses to perturbations at the system level, including cell fate acquisition, demonstrating in principle the feasibility of tracking experimental and / or physiological conditions based on genetic tags at the single-cell level. The innovative and groundbreaking aspect of this proposal lies in the application of the CombiCult® principle, which can track combinations of thousands of molecules or TFs added sequentially in a culture, thereby resolving a much larger experimental space in a single assay. A third advantage greatly enhances predictive power and links the three elements of the platform together: the generation of datasets matching transcriptional changes at the single-cell level with specific cell culture conditions provides an excellent opportunity to re-input experimentally validated driver-target interaction data into the algorithm. This virtuous cycle allows the algorithm of this invention to become increasingly accurate, learning from its own experimental results and paving the way for supervised learning frameworks that, in the long term, can predict the molecules required for any type of differentiation with high precision. In fact, given the large amount of ground truth data (i.e., data generated by known interventions), a language model (LM) can, in principle, be created capable of predicting the steps required to generate a specific cell type. Different types of LMs have been successfully deployed in studies of viral and antibody evolution (Hie et al., 2023. Nat biotechnol doi:10.1038 / s41587-023-01763-2) and chemical modification (Kosonocky et al., arXiv 2023 doi:10.48550 / arXiv.2305.16330).

[0198] The successful application of the sequencing-based platform of this invention will rapidly generate many such data points, thereby greatly accelerating the engineering and production of ATMP.

[0199] transcription factors The most widely accepted agents for promoting transdifferentiation are transcription factors. Transdifferentiation can be driven by overexpressing specific transcription factors, as observed in various transdifferentiation experiments: 1. Neuron: - Ascl1, Brn2, and Myt1l can reprogram fibroblasts into induced neurons (iNs).

[0200] NeuroD1 alone, or in combination with other factors, can also convert fibroblasts into neurons.

[0201] 2. Cardiac cardiomyocytes: Gata4, Mef2c, and Tbx5 can transform fibroblasts into induced cardiomyocytes (iCM).

[0202] 3. Endoderm cells: FoxA2 and Hnf1α can convert fibroblasts into hepatocyte types.

[0203] 4. Pancreatic β cells: Pdx1, Ngn3, and MafA, sometimes along with other factors, can transform various cell types into insulin-producing β-like cells.

[0204] 5. Endothelial cells: - Fli1, FoxC2, and Etv2 can transform fibroblasts into endothelial-like cells.

[0205] 6. Muscle cells: MyoD is a classic example of the ability to convert fibroblasts into skeletal muscle cells.

[0206] Other transcription factors that can be used for transdifferentiation can be obtained from the literature.

[0207] The foregoing are examples of agents known in the art that promote cell reprogramming.

[0208] Improved reagent database The method of the present invention can generate a database that can improve the starting parameters of experiments designed to identify reagents suitable for transforming cell types. Therefore, the present invention provides a method for selecting starting reagents in a method according to the foregoing aspects of the present invention, comprising the following steps: A database of reagents and cell transformation results is compiled; and a set of reagents or combinations of reagents is selected from this database for use in determining the conditions required for cell transformation.

[0209] The present invention also provides a method for improving a database of reagents and cell transdifferentiation results, comprising using the database to select reagents as in the foregoing embodiments, and inputting the results of a method for determining the conditions required for reprogramming cells into the database, thereby providing further data points.

[0210] The algorithms and machine learning methods described above can also be used to further improve the database.

[0211] In another aspect, the method according to the invention provides: a database of useful reagents (e.g., transcription factors) for reprogramming cells; at least one computer processor; and a memory operatively communicating with the processor, the memory containing instructions for configuring the processor to: (a) Gene regulatory networks (GRNs) are generated from a database of transcription factors. (b) Identify candidate transcription factors for the desired reprogramming event; (c) Analyze the effects of candidate transcription factors on cell reprogramming; (d) Optionally, iterate through steps (a) to (c); and (e) Output a set of optimal transcription factors.

[0212] The database preferably contains gene expression data, such as RNA-seq data.

[0213] Identifying candidate transcription factors involves querying a gene regulatory map to identify an optimal set of transcription factors for differentiating any given cell type from any given starting cell. Machine learning models can be used to generate metric calculations that function as a function of the gene regulatory map.

[0214] Examples of metric computation include critical algorithms, which can be configured for time-series RNA-seq data; see Oh et al., Biomed Res Int. 2013;2013:203681. doi: 10.1155 / 2013 / 203681. Epub 2013Mar 24.

[0215] The present invention further provides a method for determining the optimal transcription factors for cell reprogramming, comprising: using a computing device to organize a gene expression database related to the use of reagents in cell culture; generating a gene regulatory network from multiple gene expression datasets; determining candidate optimal reagents; analyzing the effect of the optimal reagents on cell differentiation; and outputting a set of optimal reagents.

[0216] As described in Kamaraj, Cell Cycle 2016, vol. 15, no. 24, 3343–3354, gene regulatory networks can be predicted computationally.

[0217] Methods included CellNet ( et al., Cell 2014; 158:903-15), which used microarray gene expression data and correlation-based network scoring; D'Alessio et al.'s method (D'Alessio et al., Stem Cell Reports 2015; 5:763–75), which used microarray gene expression data and information theory methods (Jensen-Shannon divergence); and Mogrify (Rackham et al., Nat Genet 2016; 48(3):331–5; PMID:26711105), which used cap analysis (CAGE) data of gene expression and network-based scoring.

[0218] Regardless of the computational method chosen, the iterative experimental procedures and computational analysis offer novel advantages over either the computational or experimental methods used individually.

[0219] Bioinformatics methods To detect viral tags, Parse Bio split-pipe (v 1.2.0) was used. mkref The pattern creates a new reference genome and adds the following sequences to the Ensembl GRCh38 Release 108 FASTA file: These sequences are referenced in a custom GTF file, in which they are included as genes, transcripts, and exons.

[0220] Use split-pipe tubing all The pattern aligns the original FASTQ file to this reference and uses... combine The pattern combines 8 sub-libraries, all using default parameters. The downloaded normalized h5ad file is then used... schard The package (https: / / github.com / cellgeni / schard) contains h5ad2sce() The function is imported into R v. 4.3.2 as a SingleCellExperiment object.

[0221] For 27-well and cancer-iPSC experiments, cells in which the T2A_PURO signature was detected with at least a 1 / 2 log-normalized count were isolated and partitioned according to the experiment, as these were the cells in which we were able to definitively detect that the lentiviral vector had been integrated and was being actively transcribed. This was not possible for 3TF experiments, and all cells were used.

[0222] Log-normalized counts were used as input for variable feature selection, the selection using data from... scran The package (https: / / bioconductor.org / packages / release / bioc / html / scran.html) modelGeneVar() The function selects the top 2000 features with the largest variance as input for principal component analysis (PCA) for dimensionality reduction. It uses features from... scaterThe package (https: / / bioconductor.org / packages / release / bioc / html / scater.html) runPCA() The function performs PCA and retains 20 dimensions for all subsequent analyses.

[0223] use uwot The package (https: / / cran.r-project.org / web / packages / uwot / index.html) creates a Universal Manifold Approximation and Projection (UMAP) 2D embedding from the first 20 dimensions of PCA, with the following parameters: n_neighbors = 30, min_dist = 0.4, seed = 1 .

[0224] A tag is considered to have been detected if it has a log-normalized count greater than 0.

[0225] Use from bluster The package (https: / / bioconductor.org / packages / release / bioc / html / bluster.html) makeSNNGraph() The function uses k = 20 and a Jaccard weighted scheme to construct a shared nearest neighbor (SNN) graph to identify clusters. Then, it uses data from... igraph The package (https: / / cran.r-project.org / web / packages / igraph / index.html) cluster_leiden() The function uses a resolution of 0.3 to identify clusters in the graph.

[0226] Use from presto The package (https: / / github.com / immunogenomics / presto) wilcoxauc () The function identifies cluster markers and prioritizes them based on a score calculated by multiplying the area under the receiver operating characteristic curve (auROC) by the difference in the percentage of cells expressing the marker in the cluster compared to the rest of the data.

[0227] UMAP and dot plots use from ggplot2 Visualization is achieved using the plotting functions of the package (https: / / cran.r-project.org / web / packages / ggplot2 / index.html), with heatmaps using... ComplexHeatmapCreate the package (https: / / bioconductor.org / packages / release / bioc / html / ComplexHeatmap.html).

[0228] Example 1 Genetic markers for differentiating single cells This experiment proves that: Genetic markers in single cells during CombiCult experiments • Tracking single cells exposed to various reagent conditions • Deconvolution of genetic tags using single-cell RNA sequencing (scRNAseq) • Correlating cell culture conditions history with differentiation outcomes The experiment consisted of three splits, each with three conditions, resulting in a total of 3 × 3 × 3 = 27 combinations. Induced pluripotent stem cells (iPSCs) served as the starting cell type. In each split, only one condition involved the differentiation medium used for iPSC-megakaryocyte (MK) differentiation. This condition was referred to as the "specified" condition. The other two conditions, containing only basal medium and fetal bovine serum (FBS), did not specify differentiation into MK and were referred to as the "unspecified" conditions. Therefore, out of the 27 combinations, only one combination allowed for iPSC-MK differentiation.

[0229] Cells were cultured in the following media: Specified conditions (using fluorescent proteins as genetic tags): Day 1 and Day 4: APEL2 1% anti-anti, 50 ng / ml BMP4, 50 ng / ml VEGF, 50 ng / ml bFGF; Day 7 (amplification): APEL2 1% anti-anti, TPO 25 ng / ml, SCF 25 ng / ml, Flt3L 25 ng / ml, IL-3 10 ng / ml, IL-6 10 ng / ml, heparin 5 U / ml.

[0230] Unspecified conditions (using luciferase as a genetic tag): DMEM 1% anti-anti 5% FBS.

[0231] Each split is labeled with a unique lentivirus carrying a different fluorescent protein transgene for the specified conditions, while unspecified conditions in all splits are labeled with a lentivirus carrying a luciferase transgene. Theoretically, only cells that meet the specified conditions in all three splits will fully differentiate, and these cells will be labeled with three different lentiviruses carrying three different fluorescent proteins. A schematic diagram of the described procedure is shown below. Figure 3At the end of the experiment, all cells were collected for scRNA-seq. Gene expression of all three fluorescent proteins was expected to be observed only in differentiated cells. Although some differentiated cells may show gene expression of fewer than three fluorescent proteins due to low labeling efficiency or detection thresholds, no differentiated cells were expected to show luciferase gene expression, as cells passed through unspecified conditions were not expected to differentiate.

[0232] Detailed procedures • On day 0, iPSCs were seeded as single cells into 27 wells of a multi-well culture plate. Cells from 9 of the 27 wells were transduced with a lentivirus carrying the blue fluorescent protein gene (LV-BFP). Figure 3 (a). The cells in the remaining 18 wells were transduced using a lentivirus (LV-LUC) carrying the luciferase gene. Figure 4 (a) The MOI used for all lentiviral transductions is 5.

[0233] • On day 1, remove the virus and wash the cells twice to remove residual virus before switching to fresh medium. Switch cells transduced with LV-BFP to MK differentiation medium (specified conditions) and switch cells transduced with LV-LUC to basal medium with FBS (unspecified conditions).

[0234] • On day 4, wash the cells and switch to the second split conditions. Transfer three wells from the specified conditions and six wells from the unspecified conditions from the first split to MK differentiation medium (specified conditions). In this specified condition, transduce the cells using a lentivirus carrying the red fluorescent protein gene (LV-RFP). Figure 3 (b) Switch the cells in the remaining wells to basal medium with FBS (unspecified conditions) and transduce them with LV-LUC ( Figure 4 (b) On day 5, the virus was removed, and the cells were rinsed twice to remove any residual virus before adding fresh culture medium.

[0235] • On day 7, wash the cells and switch to the third split condition following the same procedure. Cells under the specified condition are transduced with a lentivirus carrying the green fluorescent protein gene (LV-GFP). Figure 4 c). Cells under unspecified conditions were transduced using LV-LUC ( Figure 3 (c). On day 8, the virus was removed, and the cells were rinsed twice before adding fresh culture medium.

[0236] • On day 18, cells were collected for scRNAseq.

[0237] Cells are collected and fixed so that they can all be processed simultaneously. The final split / labeling is identified using ParseBio's combined barcode system, where the first of three barcodes indicates the original sample.

[0238] The RNA reads were aligned into a custom-designed reference human genome, in which the following cDNA sequences were added: • EGFP (from GATA1 lentivirus) mRFP BFP · copGFP dTomato (from TAL1 lentivirus) • LSSmOrange (from FLI1 lentivirus) · LUC • Puro_T2A (LV skeleton, excluding TF carrier) Only cells with sufficient barcode reads and sufficient RNA reads can be analyzed (approximately 78K good-quality cells were obtained from a total of 100K cells).

[0239] The results are shown in Figures 4 to 7 middle.

[0240] Figure 4 The image shows the detection of BFP tags in various cell clusters on a UMAP plot. The plot shown is for control (CTRL) cells, including adherent (ADH) and suspension (SUS); "GFP" cells are cells taken from the final GFP-positive wells in the final step of the procedure. Therefore, in Figure 4 In the study, the "GFP" profile was used to detect BFP expression in cells exposed to all three tags. LUC cells were cells exposed only to luciferase, i.e., cells that had not been tagged with any viral tags and had not been exposed to any specified conditions.

[0241] Figure 4 The same UMAP plots for cells screened for RFP and GFP expression are also shown. Although GFP expression is low, this is presumably due to increased viral resistance during cell differentiation. All “GFP” cells were derived from the final GFP-positive wells and therefore inevitably exposed to the specified conditions of the third stage.

[0242] exist Figure 5 The image shows a UMAP plot where markers are scored for presence or absence, rather than based on intensity. Triple-positive markers are shown in... Figure 6 The 27W GFP ADH is visible in the image.

[0243] Figure 7 and8 Analysis of the mid-cluster revealed a strong presence of the three-marker cell (27W ADH GFP) in cluster 7. Gene expression associated with cluster 7 was observed in... Figure 9 The results were elucidated and further characterized, indicating that these cells are stromal cells. Further analysis is needed. However, the results clearly demonstrate that single cells can be labeled and their differentiation tracked in a variety of culture media.

[0244] This experiment successfully demonstrated that single cells can be genetically labeled and their exposure to a variety of reagent conditions can be tracked, thus providing insights into how specific exposure sequences affect cell differentiation.

[0245] Example 2 Markers of cells exposed to transcription factors (cells on microcarrier beads) This experiment proves that: • Genetic markers of cells on microcarrier beads during CombiCult experiments • Deconvolution of genetic tags using single-cell RNA sequencing (scRNAseq) • By deconvolving genetic tags, the TF assemblies responsible for reprogramming cells into specific lineages are traced. This experiment involved using transcription factors (TFs) to drive the reprogramming of induced pluripotent stem cells (iPSCs) into megakaryocytes (MKs). The experiment consisted of three splits, each with three conditions, resulting in a total of 27 combinations. In each split, cells on beads were transduced with different lentiviruses carrying different transgenes. Only cells in one condition were transduced with a lentivirus carrying the TF involved in driving iPSC reprogramming into MKs. This condition was referred to as the "TF condition." Each lentivirus in the TF condition also carried a transgene for a fluorescent protein, which served as a genetic tag. Cells in the other two conditions were transduced with lentiviruses carrying only the luciferase gene, which served as the genetic tag for those conditions. Before each split, beads were washed, merged, mixed, and split into three different conditions for the next round of lentivirus transduction. All conditions were cultured in the same medium, which does not promote differentiation into MKs in the absence of TFs. Successful differentiation of iPSCs into MKs was driven solely by the presence of TFs. At the end of the experiment, cells were collected for scRNA-seq. Cell identity and corresponding genetic tags were characterized. Theoretically, only cells that pass all three TF conditions in all three splits will be transduced with lentiviruses carrying all the necessary TFs for MK reprogramming. Since these lentiviruses also carry different fluorescent protein tags, it is expected that only MKs will express all three fluorescent proteins. Although some MKs may show fewer than three fluorescent proteins due to detection thresholds, it is not expected that any MK will express the luciferase gene.

[0246] Cell passage culture Human free iPSC line (Thermo Fisher Scientific, A18945) was maintained in Essential 8 medium (Thermo Fisher Scientific, A1517001) in containers coated with telonein (Thermo Fisher Scientific, A14700) and passaged using Accutase (Thermo Fisher Scientific, A1110501). After resuscitation or passage, Y-27632 (ROCK inhibitor) (STEMCELL Technologies, 72304) was added to the medium for maintenance for 1 day.

[0247] Differentiation culture medium The cells were cultured in the following culture medium: Days 1-2: E8 medium Days 2 to 4: E6, FGF2 20 ng / ml, BMP4 10 ng / ml, VEGF 50 ng / ml Days 4-11: Cellgenix SCGM, SCF 25 ng / ml, TPO 20 ng / ml Days 11-17: SFEM, TPO 50 ng / ml, SCF 20 ng / ml, IL-6 10 ng / ml, IL-9 10 ng / ml, ETP 500 ng / ml, arachidonic acid 5 μM, nicotinamide 2.5 mM Cells were seeded onto microcarrier beads. On day -1, hydrated and sterilized plasticell microcarrier beads were distributed at 4,000 beads per well into 100 mm square culture dishes (Thermo Fisher Scientific, 103). Subfusion (approximately 80%) hiPSCs were dissociated using Accutase and seeded onto the beads at 300 cells per bead in Essential 8 medium containing Y-27632.

[0248] Combined lentiviral transduction on beads • On day 0, prior to lentiviral transduction, the number of cells on the beads was counted using a hemocytometer to calculate the volume of each lentivirus at an MOI of 2. After replacing the medium with fresh Essential 8 medium, one condition was transduced using a lentivirus carrying the TF1 and enhanced green fluorescent protein (EGFP) genes (LV-TF1-EGFP). Other conditions were transduced using a lentivirus carrying the luciferase gene (LV-LUC).

[0249] • On day 1, the virus was removed, and the cells were rinsed twice to eliminate residual virus. Cells on the microcarrier beads were pooled, mixed, and split into three conditions in fresh Essential 8 medium. One condition was transduced with a lentivirus carrying the TF2 and dTomato genes (LV-TF2-dTomato), and the other conditions were transduced with LV-LUC.

[0250] • On day 2, the virus was removed, and the cells were rinsed twice to eliminate residual virus. Cells on the microcarrier beads were pooled and mixed in fresh reprogramming medium 1 and split into three conditions. One condition was transduced with a lentivirus carrying the TF3 and LSSmOrange genes (LV-TF3-LSSmOrange), and the other conditions were transduced with LV-LUC.

[0251] • On day 3, remove the virus, wash the cells twice, and then replace the culture medium with fresh reprogrammed medium 1. From day 3 to day 17, change the culture medium as specified in Tables 1 and 2. A schematic diagram of the experimental procedure is shown below. Figure 10 As shown.

[0252] · Single-cell RNA sequencing and deconvolution On day 17, cells were collected for scRNAseq.

[0253] result scRNAseq results revealed the expression of all three genetic tags (EGFP, dTomato, and LSSmOrange); Figure 11 and Figure 12 Three clusters of cells: 1. A cluster of CD41a + CD235a - Cells, which are MK 2. A cluster of CD41a + CD42b + CD235a - Cells, representing more mature MK 3. A cluster of CD41a + CD235a+ Cells, which are bipotent progenitor cells in conclusion This experiment demonstrates MK reprogramming via combined TF transduction incorporating genetic tags. It showcases a method for labeling cells on microcarrier beads in a CombiCult experiment and using genetic tag deconvolution via scRNAseq to trace the TFs responsible for lineage-specific reprogramming. By screening a broad range of TF pathways and employing an efficient labeling strategy, this method facilitates the discovery of novel reprogramming schemes.

[0254] Example 3 Markers of cells exposed to transcription factors (single cells on a monolayer) This experiment proves that: Genetic markers in single cells during CombiCult experiments • Single-cell splitting-merging in CombiCult experiments • Deconvolution of genetic tags using single-cell RNA sequencing (scRNAseq) • By deconvolving genetic tags, trace the combination of transcription factors (TFs) responsible for reprogramming cells into a specific lineage. This experiment used transcription factors (TFs) to drive the reprogramming of induced pluripotent stem cells (iPSCs) into megakaryocytes (MKs). The experiment consisted of three splits, each with three conditions, resulting in a total of 27 combinations. In each split, cells were transduced with different lentiviruses carrying different transgenes. Only cells in one condition were transduced with a lentivirus carrying the TF involved in driving iPSC reprogramming into MKs. This condition was referred to as the "TF condition." Each lentivirus in the TF condition also carried a transgene for a fluorescent protein, which served as a genetic tag. Cells in the other two conditions were transduced with lentiviruses carrying only the luciferase gene, which served as the genetic tag for those conditions. Before each split, single cells were pooled, mixed, and split into three different conditions for the next round of lentivirus transduction. All cells were cultured in the same medium, which does not promote differentiation into MKs in the absence of TFs. Successful differentiation of iPSCs into MKs was driven solely by the presence of TFs. At the end of the experiment, cells were collected for scRNA-seq. Cell identity and corresponding genetic tags were characterized. Theoretically, only cells that pass all three TF conditions in all three splits will be transduced with lentiviruses carrying all the necessary TFs for MK reprogramming. Since these lentiviruses also carry different fluorescent protein tags, it is expected that only MKs will express all three fluorescent proteins. Although some MKs may show fewer than three fluorescent proteins due to detection thresholds, it is not expected that any MK will express the luciferase gene.

[0255] Detailed procedures • On day -1, iPSCs were seeded into three wells of a multi-well plate. On day 0, cells in one well were transduced with a lentivirus carrying the TF1 and enhanced green fluorescent protein (EGFP) genes (LV-TF1-EGFP). Cells in the other two wells were transduced with a lentivirus carrying the luciferase gene (LV-LUC). Figure 13 ).

[0256] On day 1, the virus was removed, and the cells were washed twice to remove residual virus. All cells were dissociated into single cells, merged, mixed, and evenly distributed into three wells of a multi-well plate. Cells in one well were transduced with a lentivirus carrying the TF2 and dTomato genes (LV-TF2-dTomato), and cells in the other two wells were transduced with LV-LUC (LV-LUC). Figure 13 ).

[0257] On day 2, the virus was removed, and the cells were washed twice to remove residual virus. All cells were dissociated into single cells, merged, mixed, and evenly distributed into three wells of a multi-well plate. Cells in one well were transduced with a lentivirus carrying the TF3 and LSSmOrange genes (LV-TF3-LSSmOrange), and cells in the other two wells were transduced with LV-LUC. Figure 13 ).

[0258] • On day 3, the virus was removed by rinsing the cells twice to remove residual virus, and then fresh culture medium was added. The MOI used for all lentiviral transductions was 2.

[0259] • On day 17, cells were collected for scRNAseq.

[0260] result The results were similar to those of Example 2.

[0261] Example 4 iPSCs and markers of cancer cells This experiment proves that: Genetic markers in single cells during CombiCult experiments • Deconvolution of genetic tags using single-cell RNA sequencing (scRNAseq) This experiment involved three different cell lines. In the first split, only one cell line was labeled with a lentivirus carrying a fluorescent protein gene. The other two cell lines were labeled with a lentivirus carrying a luciferase gene. The cells were then merged and split again for a second split, where only one condition was labeled with fluorescent protein, and the other two conditions were labeled with luciferase. This process was repeated for the third split. All cells were cultured in the same medium (DMEM, anti-anti 1%, FBS 10%) to maintain viability and prevent differentiation, as the cells already possessed distinct identities. At the end of the experiment, cells were collected for scRNA-seq. Theoretically, only one cell line or type should carry all three fluorescent protein tags. All cells expressing all three fluorescent proteins are expected to belong to the cell line or type labeled with the fluorescent protein gene in the first split.

[0262] Detailed procedures • On day -1, A549 lung cancer cell line, T47D breast cancer cell line, and iPSC line were seeded onto culture plates. On day 1, A549 cells were transduced with lentivirus carrying the green fluorescent protein (GFP) gene (LV-GFP). Figure 13 T47D cells and iPSCs were transduced using a lentivirus carrying the luciferase gene (LV-LUC). Figure 14 ).

[0263] On day 2, the virus was removed, and the cells were washed twice to remove residual virus. All cells were dissociated into single cells, merged, mixed, and evenly distributed into three wells of a multi-well plate. Cells in one well were transduced with a lentivirus carrying the blue fluorescent protein (BFP) gene (LV-BFP), and cells in the other two wells were transduced with LV-LUC. Figure 14 ).

[0264] On day 3, the virus was removed, and the cells were washed twice to remove residual virus. All cells were dissociated into single cells, merged, mixed, and evenly distributed into three wells of a multi-well plate. Cells in one well were transduced with a lentivirus carrying the dTomato gene (LV-dTomato), and cells in the other two wells were transduced with LV-LUC (LV-LUC). Figure 14 ).

[0265] • On day 4, the virus was removed by rinsing the cells twice to remove residual virus, and then fresh culture medium was added. The MOI used for all transductions was 10.

[0266] • On day 6, cells were collected for scRNAseq.

[0267] result Figure 15Clustering of single cells using iPSCs and cancer cell markers as detection probes is shown. The distribution of the tested cell populations is clear.

[0268] exist Figure 16 In this process, genetic tags were screened, and cells were clustered as shown in Figure 17. As expected, the tags co-segregated with cells clustered according to cell type markers, indicating that individual cells could be successfully genetically labeled and tracked during the split-merge procedure. GFP was observed only in A549 cells, while after split-merge, BFP and RFP tags were distributed across all cell types.

[0269] This experiment successfully demonstrated that single cells from different cell lines can be labeled and tracked under various reagent conditions, thus providing a method for distinguishing and tracking individual cells in a mixed population.

[0270] In all embodiments, the experiments effectively demonstrated the following capabilities: • Labeling single cells: Using lentiviruses carrying fluorescent protein genes or luciferase, individual cells are uniquely labeled under each reagent condition.

[0271] • Tracking multiple conditions experienced: The sequential introduction of tags can reconstruct the history of each cell's exposure to different reagent conditions.

[0272] • Deconvolution of genetic tags: scRNAseq enables the identification of genetic tags within single cells, thereby facilitating the analysis of how specific combinations of conditions affect cell fate.

[0273] • Understanding differentiation pathways: By linking genetic tags to differentiation outcomes, we gain insights into the necessary conditions and factors that drive specific lineage orientations.

[0274] These methodologies provide powerful tools for studying complex cellular processes and can be applied to a wide range of research areas in cell biology and regenerative medicine.

Claims

1. A method for determining the conditions required to promote the conversion of a first cell type to a second cell type by analyzing one or more markers in a single cell exposed to a combination of reagents.

2. The method according to claim 1, comprising the following steps: a. Expose a single cell of the first cell type to the first reagent or a combination of reagents; b. Optionally, remove the first reagent and expose the cells to different reagents or combinations of reagents; c. Optionally, repeat step (b) above; and d. Identify one or more cells displaying one or more markers of a second cell type, and determine the identity and order of the reagents to which the cells were exposed.

3. The method of claim 1 or claim 2, wherein cells are labeled to identify exposure to each TF or reagent combination.

4. The method according to claim 3, comprising the following steps: a. Subdivide multiple single cells into a first group of single cell populations, expose each single cell population to a different reagent or combination of reagents, and label the single cells to indicate the exposure to the reagent or combination of reagents; b. Merge single-cell populations and subdivide the cell pools to form a second group of single-cell populations; c. Expose the second group of single-cell populations to different reagents or combinations of reagents; d. Optionally iterate and repeat steps (a) to (c) as needed; e. Identify cell type markers in cells and unconvolve the markers to identify the identity, sequence, and timing of the reagents that cells were exposed to; f. Optionally, use the data from step (e) to guide the selection of reagents used in step (a), and repeat steps (a) to (e) as needed; g. Identify cells of the second cell type based on cell type markers, and deconvolve the markers to identify the identity and order of the reagents required to reprogram cells of the first cell type into cells of the second cell type.

5. The method of claim 4, wherein successive method steps in the procedure include adding a single reagent or adding a combination of reagents.

6. The method of claim 4, wherein the successive method steps in the procedure comprise only the addition of a single reagent.

7. The method according to claim 2 or claims 4 to 6, wherein the TF or combination of TF used in the first iteration of step (a) is selected from agents known in the art for promoting cell type transformation.

8. The method of claim 7, wherein the reagent or combination of reagents used in the first iteration of step (c) is selected from reagents known in the art for promoting cell type conversion.

9. The method according to claim 2 or claims 4 to 8, wherein the reagents or combinations of reagents used in the second and further iterations of steps (a) and (c) are selected by analyzing the data generated by step (e).

10. The method of claim 10, wherein the data generated in step (e) is processed by an algorithm that analyzes the effect of the identity and application time of each reagent on the transformation of the first cell type.

11. The method according to claims 4 to 10, wherein the reagent or combination of reagents used in the first iteration of step (a) is selected using data generated by step (e) in a previous execution of the method.

12. The method according to any one of claims 3 to 11, wherein the cells are labeled by modifying the nucleic acids of the cells.

13. The method of claim 12, wherein the nucleic acids of the cell are modified by genome editing or viral integration.

14. The method of claim 13, wherein genome editing comprises CRISPR.

15. The method of claim 12, wherein data derived from the fate of cells exposed to different reagents are combined to correlate transcriptional changes at the single-cell level with specific cell culture conditions, thereby generating a model that predicts the culture conditions required to generate a specific cell type.

16. The method of claim 15, wherein the model is used to generate the initial reagent selections for steps (a) to (c), and the iterative execution of the steps results in an improved model.

17. The method according to any of the preceding claims, wherein the transformation of cells from a first cell type to a second cell type is a procedure selected from cell differentiation, cell reprogramming, cell transdifferentiation, and cell dedifferentiation.

18. A method for identifying genes that influence cellular processes, comprising the steps of: a) Determine the effect of one or more reagents or combinations of reagents on single cells according to any one of the preceding claims; b) When exposed to the reagent or combination of reagents, analyze gene expression in the cells; and c) Identify genes that are differentially expressed upon exposure to the reagent or combination of reagents.

19. A method for producing a nucleic acid encoding a gene product that affects cellular processes, comprising identifying a gene as described in claim 18 and producing at least the coding region of the gene by nucleic acid synthesis or biological replication.

20. A method for inducing a cellular process, comprising the following steps: a) According to claim 18, identifying one or more genes differentially expressed in relation to cellular processes; and b) Regulate the expression of one or more of the genes mentioned above in the cell.

21. The method of claim 20, wherein regulation of gene expression in the cell comprises transfecting the one or more genes into the cell.

22. A method for selecting a starting reagent in the method according to claim 4, comprising the steps of: A database compiling reagents and cell transformation results; and Select a set of reagents or a combination of reagents from the database for use in the method according to claim 4.

23. A method for improving a database of reagents and cell transformation results, comprising using the database to select reagents according to claim 22 and inputting the results of the method according to claim 4 into the database to provide further data points.

24. The method according to any of the preceding claims, wherein the reagent is a transcription factor (TF).

25. A system comprising a database for converting cells from a first cell type to a second cell type TF, at least one computer processor, and a memory operatively communicative to the processor, the memory containing instructions for configuring the processor to perform the following operations: (a) Gene regulatory networks (GRNs) are generated from a database of transcription factors. (b) Identify candidate transcription factors for the desired transformation event; (c) Analyze the effects of candidate transcription factors on cell reprogramming; (d) Optionally, iterate through steps (a) to (c); and Output a set of optimal transcription factors.

26. The system of claim 25, wherein the database contains gene expression data.

27. A method for determining the optimal transcription factor for cell reprogramming, comprising: Using computing devices, compile a database of gene expression related to the use of TF in cell culture; Gene regulatory networks were generated from multiple gene expression datasets. Identify candidate optimal TFs; analyze the impact of optimal TFs on cell differentiation; and output a set of optimal TFs.

Citation Information

Patent Citations

  • Isolation, selection and propagation of animal transgenic stem cells

    EP0695351A1

  • Porous media and method of manufacturing same

    WO2004013369A1

  • Cell culture

    WO2004031369A1