Integrated pipeline for cell culture design
An integrated pipeline using single-cell RNA sequencing and computational analysis generates targeted cell differentiation protocols, enhancing the efficiency and purity of stem cell culture production for therapeutic and drug development applications.
Patent Information
- Application Number
- PCT/US2025/012053
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-24
AI Technical Summary
Current methods for controlling stem cell differentiation into specific cell types are inefficient and lack a deep understanding of the underlying mechanisms, limiting their therapeutic potential.
An integrated pipeline using single-cell RNA sequencing and computational analysis to generate a low-dimensional graph from genomics data, analyze trajectories, and assign scores based on gene expression, enabling the development of targeted cell differentiation protocols.
This approach allows for more effective and efficient production of high-purity cell cultures, such as skeletal muscle fibers and satellite cells, facilitating cell therapy and drug development by understanding molecular mechanisms of differentiation.
Smart Images

Figure US2025012053_24072025_PF_FP_ABST
Abstract
Description
INTEGRATED PIPELINE FOR CELL CULTURE DESIGNGOVERNMENT FUNDING
[0001] The present invention was made with government support under Grant No. 5RO1AR07452604 (agreement 2018A016392), awarded by the National Institutes of Health. The US government has certain rights in this invention.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 621,711, filed January 17, 2024, which is incorporated herein by reference.BACKGROUND
[0003] Stem cells are promising candidates for use in cell therapy, as they have the potential to differentiate into a wide range of cell types, making them ideal for repairing or replacing damaged tissues. However, the therapeutic potential of stem cells is dependent on their ability to differentiate into specific cell types, which is difficult to achieve without a deeper understanding of the mechanisms controlling their differentiation.SUMMARY OF THE INVENTION
[0004] In one example, a system includes a processor and a non-transitory computer readable medium storing instructions executable by the processor. The machine-executable instructions include a graph generator that receives a genomics dataset representing a target cell type and generates a low-dimensional graph from the genomics dataset and a graph analyzer that analyzes the low-dimensional graph to determine a set of trajectories within the generated graph and assigns a score to each of the set of trajectories based on gene expression associated with the target gene within the trajectory.
[0005] In another example, a method includes generating a low-dimensional graph from a genomics dataset representing a target cell type. The low-dimensional graph is analyzed to determine a set of trajectories within the low-dimensional graph. A score is assigned to each of the set of trajectories based on gene expression associated with the target gene within the trajectory. A trajectory of the set of trajectories having a best score is selected.
[0006] In another example, a method includes generating a low-dimensional graph from a genomics dataset representing a target cell type. The low-dimensional graph is analyzed to determine a set of trajectories within the low-dimensional graph. A score is assigned to each of the set of trajectories based on gene expression associated with the target gene within the trajectory. A trajectory of the set of trajectories having a best score is selected. A cell differentiation protocol is developed from the trajectory of the set of trajectories having the best score, and a cell culture is produced from the cell differentiation protocol.BRIEF DESCRIPTION OF THE FIGURES
[0007] The present invention may be more readily understood by reference to the following figures, wherein:
[0008] FIG. 1 illustrates one example of a system for determining a cell differentiation protocol.
[0009] FIG. 2 illustrates another example of a system for determining a cell differentiation protocol.
[0010] FIG. 3 illustrates one method for determining a cell differentiation protocol.
[0011] FIG. 4 illustrates another method for determining a cell differentiation protocol.
[0012] FIGS. 5 A and 5B provide graphs showing the effect of SHH or SAG on amplification of human-induced pluripotent stem cells (hiPSC), with A) showing the SAG percentage PAX7 by condition, and B) showing the SAG percentage PAX_K167 by condition.
[0013] FIG. 6 is a schematic block diagram illustrating an exemplary system of hardware components capable of implementing examples of the systems and methods disclosed herein.DETAILED DESCRIPTION OF THE INVENTION
[0014] The systems and methods described herein combine single-cell RNA sequencing and advanced computational analysis to predict the fate of cells, allowing researchers to gain a deeper understanding of the molecular mechanisms underlying cell differentiation, which in turn can lead to the development of more effective therapeutic interventions. This technologyhas the potential to enable researchers to identify the specific genetic factors that drive differentiation, allowing for more targeted approaches to cell therapy. By analyzing the gene expression profiles of cells, researchers can identify the key genetic pathways involved in differentiation, and develop strategies to control these pathways, ultimately leading to more effective cell therapies. In addition to cell therapy, this technology can have many other potential applications in the biotech industry, including drug development and personalized medicine. By understanding the molecular mechanisms of cell differentiation, researchers can develop more effective drugs that target specific cell types, and also develop personalized therapies based on an individual's genetic expression profile.
[0015] In one embodiment, the disclosed invention comprises an integrated pipeline that combines bioinformatic cell fate prediction with ex vivo cell culture. This technology enables the generation of a new protocol for the creation of a given cell type ex vivo. It can understand pathways involved in the development of a given cell type based on in vivo genomics data and create a differentiation culture protocol for this cell type that can be used to re-create this cell type from differentiated embryonic stem cells or induced pluripotent stem cells in culture.
[0016] In one example, referred to as the human iPSC-derived Skeletal Muscle Fibers and Satellite Cells (hiPSCSkMFSC) differentiation method, a highly pure culture of hiPSCSkMFSC is developed. The systems and methods described herein allow for discovery of the main pathways controlling the differentiation of PSM derivatives, allowing for manipulation of the fate of cells in culture to create one of the most advanced hiPSC-SkMFSC differentiation method. This protocol is considerably more efficient than other patented protocols found in the scientific literature to this day and yields functional human skeletal muscle fibers and muscle stem cells. The hiPSC-SkMFSC differentiation method can provide an unlimited source of muscle fibers and stem cells for cell therapy to treat muscle injuries and defects. In addition, it can be used for drug discovery to identify molecules that can cure genetic disorders, repair muscle injury or enhance the growth of muscle tissue function in vivo.Definitions
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which these exemplary embodiments belong. The terminology used in the description herein is for describingparticular exemplary embodiments only and is not intended to be limiting of the exemplary embodiments. As used in the specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety.
[0018] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0019] A "gene," or a "sequence which encodes" a particular protein, is a nucleic acid molecule which is transcribed (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vitro or in vivo when placed under the control of one or more appropriate regulatory sequences. A gene of interest can include, but is no way limited to, cDNA from eukaryotic mRNA, genomics DNA sequences from eukaryotic DNA, and even synthetic DNA sequences. A transcription termination sequence will usually be located 3' to the gene sequence. Typically, a polyadenylation signal is provided to terminate transcription of genes inserted into a recombinant virus.
[0020] An “effective amount,” as used herein, refers to an amount sufficient to achieve an intended result. An “effective amount” includes an amount that is 100% effective in achieving that result but also includes amounts that are less effective but still exhibit a significant effect. For example, an effective amount of an antibiotic compound used for antibiotic selection is an amount sufficient to kill a sufficient number of microorganisms not having an antibiotic selection marker for microorganisms having that selection marker to be readily separated and identified.
[0021] An “average,” as used herein, is a measure of central tendency, and should be read to include any of the arithmetic mean, the median, the mode, or the geometric mean of a set of values.
[0022] All scientific and technical terms used in the present application have meanings commonly used in the art unless otherwise specified. The definitions provided herein are to facilitate understanding of certain terms used frequently herein and are not meant to limit the scope of the present application.
[0023] FIG. 1 illustrates one example of a system 100 for determining a cell differentiation protocol. The system 100 includes a processor 102 and a non-transitory computer readable medium 110 storing executable instructions, executed by the processor 102. It will be appreciated that the executable instructions can be spread across multiple non-transitory computer readable media that are operatively connected via an appropriate data connection, such that the executable instructions can be executed by multiple processors. In particular, one or more machine learning models associated with the system 100 can be trained across multiple graphics processing units, with a number of graphics processing units and storage at the non- transitory computer readable medium 110 used for a given application be scalable with the application and the amount of available training data.
[0024] The executable instructions stored on the non-transitory computer readable medium 1 10 include a graph generator 102 that receives a genomics dataset and generates a lowdimensional graph from the genomics dataset. In one implementation, the genomics dataset can be selected to include evolution of a cell culture over time, and thus includes, for each cell culture, a time series of at least two observations. In practice, the genomics dataset can be selected to include samples representing a specific cell type of interest. In one example, the graph generator 112 conditions the genomics data for analysis, applies a dimension reduction technique to the data, and the resulting data can be embedded into graph based on the reduced- dimensionality time series. For example, a clustering algorithm can be applied to identify distinct stages of cell differentiation or cell types, and a simplified representation of the clusters as a graph can be generated, for example, via a principal tree algorithm.
[0025] A graph analyzer 114 analyzes the generated graph to determine the main trajectories within the generated graph and assign a score to each node of the tree. For example, the scoremay represent the expression of a particular gene or set of genes relative to a reference set of genes. The score can then be distributed among the various trajectories within the graph, such that the genes activated at each step of the graph can be determined, particularly for trajectories leading to the desired cell type. As a result, a candidate protocol for generating the desired cell type can be determined and tested in vitro from the trajectory with a best score (e.g. , highest, lowest, or closest to a target number).
[0026] The system 100 can be applied to any kind of cell type to reach high level of purity in the populations and high efficiency in the cell fate decisions guidance using specific molecules. This can decrease the cost and increase the speed of cell therapy production and permits generation of new hypotheses and protocols with limited costs. Further, while data can be generated specifically for the genomics dataset, the system 100 can operate using only publicly available data generated for other purposes.
[0027] FIG. 2 illustrates another example of a system 200 for determining a cell differentiation protocol. The system 200 includes a processor 202 and a non-transitory computer readable medium 210 storing executable instructions, executed by the processor 202. It will be appreciated that the executable instructions can be spread across multiple non-transitory computer readable media that are operatively connected via an appropriate data connection, such that the executable instructions can be executed by multiple processors. The executable instructions stored on the non-transitory computer readable medium 210 include a data interface 212 that receives a genomics dataset from a local or remote source and conditions the dataset for analysis. In one implementation, the data interface 212 is configured to operate in concert with a network interface 204 to retrieve the genomics dataset from a remote non- transitory storage location.
[0028] In one example, the genomics dataset is an embryonic genomics dataset (e.g. , human, mouse, etc.) that contains multiple timepoints. For example, the datasets could include single cell RNA sequencing datasets of mouse embryos at stages ell, el2, and el3. The data associated with a given target cell type can be separated into a subset containing the timepoints in which the target cell type can be found. It will be appreciated that the subset of the genomics dataset can be a proper subset. A matrix formed from this subset of the data (e.g., gene x cells or nuclei) and preprocessed to prepare the data for analysis. For example, one or more of logarithmization of the data matrix, normalization by total counts over all genes, scaling, andbatch correction can be applied to the matrix. The data interface 212 can also apply a dimension reduction technique to the preprocessed data.
[0029] A graph generator 214 embeds the data into a low-dimensional graph, for example, via a two-dimensional Uniform Manifold Approximation and Projection (UMAP) algorithm or a two-dimensional force atlas algorithm, In this process, the different cell types are spread over different branches leading to target cell type as well as one or more other non-desired cell types present in the dataset. In the illustrated example, a clustering algorithm, for example, leiden clustering, can be applied to refine the different stages of differentiation or cell identities. A principal tree algorithm, such as simplePPT, can then be applied to the data to learn a simplified representation on a space composed of nodes. A trajectory algorithm, such as pseudotime, is then applied to analyze the main trajectories inside the dataset along the tree.
[0030] A graph analyzer 216 calculates a score for each genomics entity (e.g., cells or nuclei) using a set of genes associated with biological pathways. In one example, the score can be calculated as the average expression of the gene set subtracted with an average expression of a reference set of genes randomly sampled from all the genes in the subset of the geonomics dataset. The score for each pathway is fitted to the trajectory tree, for example, via a generalized additive model, and then distributed along trajectories. This fitting technique can also be done for genes (e.g., transcription factors) or other genomics information that can be directly or indirectly relevant to cell differentiation. Accordingly, the pathways or genes activated in the different branches leading to each cell type can be determined and used for experiments in vitro.
[0031] A protocol testing platform 220 can be used to test one or more protocols developed at the graph analyzer. It will be appreciated that the protocol testing platform can be automated or semi-automated to generate a cell culture based on a provided cell differentiation protocol. Molecules known to modulate pathways upregulated in the desired cell type and downregulated in other cell types can be tested in vitro using differentiation protocols recapitulating the desired differentiation leading to the first timepoint. The dosage can be found using ECso for upregulation and IC50 for downregulation and can be adjusted based on, for example, cell toxicity, literature knowledge, and similar factors. The selected molecules can be added to the culture media and the order in which they are added can be extrapolated from the timepoints in the graph where the fitted score is higher for the desired cell type and lower for the non-desired cell type, resulting in a new protocol. The new protocol can then be applied to cellular cultures, and the resulting culture can be sequenced at any time after the use of at least one of the molecules using genomics techniques, for example single cell RNA sequencing, to obtain a measure of the quality of the culture, such as a quantity of the desired cell type, a quality of the desired cell types, or a purity of the culture. If the culture quality metric meets a threshold value, the protocol can be retained for use in other applications, such as cell therapy, drug screening, protocol licensing, and other industrial applications.
[0032] If the culture quality metric does not meet the threshold, the data generated at the protocol testing platform 220 can be provided back to the data interface 212 to generate a second protocol at the graph generator 214 and graph analyzer 216. The second protocol can then be analyzed at the protocol testing platform 220. Additionally, a similarity between the initial dataset and the newly generated culture data can be evaluated using a classifier 218. The classifier 218 can be trained using a machine learning algorithm, for example, using hierarchically organized SVMs on the initial subset of the genomics dataset with the pathway scores or the cell types as a feature. Where pathway scores are used, a threshold can be applied to convert continuous data scores to categorical parameters. To enhance the speed and accuracy of the classification, the classifier 218 can be trained only on highly variable genes. The classifier can then be used to classify the pathways or cell types based on a genetic similarity of the cells between the newly generated culture data and the initial subset of the genomics dataset before testing the new protocol at the protocol testing platform.
[0033] In view of the foregoing structural and functional features described above, methods in accordance with various aspects of the present invention will be better appreciated with reference to FIGS. 3 and 4. While, for purposes of simplicity of explanation, the methods of FIGS. 3 and 4 are shown and described as executing serially, it is to be understood and appreciated that the present invention is not limited by the illustrated order, as some aspects could, in accordance with the present invention, occur in different orders and / or concurrently with other aspects from that shown and described herein. Moreover, not all illustrated features may be required to implement a method in accordance with an aspect of the present invention.
[0034] FIG. 3 illustrates one method 300 for determining a cell differentiation protocol. At 302, a low-dimensional graph is generated from a genomics dataset representing a target cell type. It will be appreciated that the genomics dataset can be a proper subset of a receivedgenomic dataset, with the subset comprising a plurality of time series in the received genomics dataset containing time points associated with the target cell type. In one implementation, one or more of normalization, scaling, and batch correction can be applied to the genomics dataset. In one example, the low-dimensional graph is generated by embedding the genomics dataset into a graph, applying a clustering algorithm to identify different stages of cell differentiation within the graph, and generating the low-dimensional graph as a simplified representation of the graph.
[0035] At 304, the low-dimensional graph is analyzed to determine a set of trajectories within the low-dimensional graph. In one example, a trajectory algorithm is applied to the lowdimensional graph to identify the set of trajectories. At 306, a score is assigned to each of the set of trajectories based on gene expression associated with the target gene within the trajectory. For example, a score can be calculated for each genomics entity within the low-dimensional graph, and the score for each trajectory can be fitted in the trajectory tree to assign the score to each of the set of trajectories. At 308, a trajectory of the set of trajectories having a best score is selected. Once a trajectory is selected, a cell differentiation protocol can be developed from the selected trajectory and a cell culture can be generated from the cell differentiation protocol.
[0036] FIG. 4 illustrates another method 400 for determining a cell differentiation protocol. At 402, a low-dimensional graph is generated from a genomics dataset representing a target cell type. It will be appreciated that the genomics dataset can be a proper subset of a received genomic dataset, with the subset comprising a plurality of time series in the received genomics dataset containing time points associated with the target cell type. In one implementation, one or more of normalization, scaling, and batch correction can be applied to the genomics dataset. In one example, the low-dimensional graph is generated by embedding the genomics dataset into a graph, applying a clustering algorithm to identify different stages of cell differentiation within the graph, and generating the low-dimensional graph as a simplified representation of the graph.
[0037] At 404, the low-dimensional graph is analyzed to determine a set of trajectories within the low-dimensional graph. In one example, a trajectory algorithm is applied to the lowdimensional graph to identify the set of trajectories. At 406, a score is assigned to each of the set of trajectories based on gene expression associated with the target gene within the trajectory. For example, a score can be calculated for each genomics entity within the low-dimensionalgraph, and the score for each trajectory can be fitted in the trajectory tree to assign the score to each of the set of trajectories. At 408, a trajectory of the set of trajectories having a best score is selected. At 410, a cell differentiation protocol can be developed from the selected trajectory and a cell culture can be generated from the cell differentiation protocol at 412.
[0038] At 414, a quality metric is generated from the second genomics dataset. In one example, a second genomics dataset is generated from the cell culture, for example, via single cell RNA sequencing, and the quality metric is determined from the second genomics dataset. At 416, it is determined if the quality metric meets a threshold value. If so (Y), the protocol is accepted. Otherwise (N), the method returns to 402 to analyze a set of genomics data associated with the cell culture to generate another cell differentiation protocol.
[0039] In one example, a multi-time points single cell RNA sequencing (scRNAseq) dataset of mouse embryos was analyzed from stages E8.5, E9.5, E10.5 and El 1.5 (whole embryos staged in one-somite increments as they match the time frame of Presomitic Mesoderm (PSM) differentiation into myogenic precursors. The paraxial mesoderm was extracted from this dataset to study the bifurcation from PSM to sclerotome and dermomyotome characterized by the expression of PAX 1 / 9 and PAX3 respectively. Palantir diffusion maps and force atlas were used to create a 2-D representation of a final subset of interest, and the annotation of the different cell types were confirmed using differentially expressed genes. Then, the PSM node was assigned as the trajectory root, and a simple principal tree algorithm was used to leam a tree structure on the data. A pseudotime algorithm (scfates) was used to assign temporality within the tree branches. The two main identified branches coming from the anterior PSM (characterized by ME0X1) are the myogenic derivatives on one branch (characterized by PAX7 and TTN) and the cartilage / connective tissue progenitors (SOX5, PAX9 and PAX1) on the other branch. Then, for each cell, we scored the average expression of sets of genes related to known signaling pathways described in BioPlanet. The score is the average expression of the set of genes subtracted with the average expression of the reference set of genes (being randomly sampled from the genes contained within the dataset). This score is calculated for every signaling pathway, then fitted to the trajectory using a machine learning algorithm (mgcv : a generalized additive model) and can be visualized in the force atlas space over pseudotime. As a result, the Tyrosine receptor kinase B (TRKB), Bone Morphogenetic Protein (BMP), Platelet-derived growth factor (PDGF) and Transforming growth factor beta (TGFB) pathwayswere identified as being activated in the sclerotome branch and inhibited in the dermomyotome branch.
[0040] The resulting protocol can be used for altering cell fate, such as creating new muscle fibers and progenitors derived from induced pluripotent stem cells culture protocol using this pipeline. Publicly available mouse embryo genomics datasets were used for generating the low-dimensional graph, and during generation of the graph, the sclerotome / dermomyotome bifurcation (characterized by Sox9 and Pax respectively) was highlighted by two main branches coming from the Presomitic Mesoderm (characterized by Meoxl) and going to the myogenic derivatives (characterized by Pax7 and Myodl) on one branch and going to the chondrocytes / connective tissue progenitors (Sox5 and Col3al respectively) on the other branch. Analysis of the low-dimensional graph found some of the main pathways involved in the branch that we want to downregulate in order to promote the myogenic content of the culture: Tyrosine receptor kinase (TRKB), Bone Morphogenetic Protein (BMP), Platelet- derived growth factor receptor (PDGFR) and Transforming growth factor beta (TGFB) pathways. The hypothesis generated by the pipeline was tested to successfully generate new muscle fibers and satellite cells derived from induced pluripotent stem cells culture protocol using this pipeline.
[0041] Human skeletal muscle fibers and satellite cells constitute integral components of the musculoskeletal system, playing pivotal roles in essential biological functions like movement and force generation. Satellite cells, specifically, serve as muscle stem cells that can regenerate damaged skeletal muscle fibers. Their remarkable regenerative capacity holds tremendous potential for disease modeling and cell therapy applications, which is particularly useful in treating a wide array of conditions, from genetic muscle diseases to age-related muscle wasting. The systems and methods herein not only offer a renewable source of human muscle cells for biological research and therapeutic development but also circumvent the limitations associated with biopsies, including invasiveness and limited cell numbers. As a revolution in the field of regenerative medicine, it opens an unprecedented window of opportunity for both scientific exploration and commercial exploitation.EXAMPLESDEVENIR cells; an integrated genomics pipeline for cell culture design
[0042] The directed differentiation method of FIGS. 3 and 4 was used to obtain functional Skeletal Muscle Fibers and Satellite cells (muscle progenitors) derived from human induced Pluripotent Stem Cells (hiPSC-SkMFSC). The method improves upon a previously published patent, specifically U.S. Patent No. 10,240,123 which is hereby incorporated by reference, that describes a generation of presomitic mesoderm (PSM) from induced pluripotent stem cells by downregulating Bone Morphogenetic Protein (BMP) and upregulating Wnt pathways. The systems and methods described herein were employed to identify the main pathways controlling the differentiation of PSM derivatives. Based on these new findings, starting from the human induced pluripotent stem cells and using the PSM protocol of U.S. Patent No. 10,240,123 to obtain PSM at day 6, BMP and Transforming Growth Factor beta (TGFB) pathways were downregulated from day 6 to day 9, BMP, TGFb, Platelet-Derived Growth Factor (PDGF), Tyrosine receptor kinase B (TrkB) pathways were downregulated from day 9 to day 12, and BMP and PDGFR were downregulated starting from day 12.
[0043] These modulations resulted in a highly enriched culture of Skeletal Muscle Fibers and Satellite cells (muscle progenitors) derived from human induced Pluripotent Stem Cells with higher amount of PAX7+ cells (52.3% on average at day 31) compared to other published protocols where the PAX7+ SC population represents a fraction of up to 25-35% of the mononucleated population after 30-80 days in vitro. BMP regulation was performed with LDN-193189, TGFbeta regulation was performed with SB 431542, PDGFR regulation was performed with AC 710, and TrkB regulation was performed with ANA 12. Briefly, human induced pluripotent stem cells (iPSCs) were dissociated to single cells using Accutase and seeded at a density of 28, 000-33, 000 / cm2 on Matrigel-coated dishes in mTESR-1 supplemented with 10 pM Y-27632. The next day, the cells made small pluripotent colonies, and they were treated with DMEM / F12 GlutaMAX supplemented with 1 % ITS, 3 pM CH1R99021 and 0.5 pM LDN- 193189 (PSM medium) for 3 days. For the next 3 days cells were cultured in DMEM / F12 GlutaMAX supplemented with 1 % ITS, 3 pM CHIR99021, 0.5 pM LDN-193189 and 20 ng / ml FGF-2.
[0044] The information given by the integrated bioinformatic platform of FIGS. 1 and 2 was used to establish a new myogenic differentiation protocol in which we aimed to enrich in SCs by reducing the contribution of sclerotome-derived lineages. An existing differentiation protocol, referred to as the Chai protocol, is based on myogenic cues described in the literature and used BMP inhibition (LDN), Hepatocyte growth factor (HGF) activation, Fibroblastgrowth factor (FGF-2) activation and Insulin-Like Growth Factor (IGF-1) activation from day 6 to day 8, then IGF-1 activation from day 8 to day 12, finally IGF-1 activation and HGF activation starting from day 12. Based on the findings of our bioinformatics pipeline of FIGS. 1 and 2, the anterior PSM cells were treated with BMP (LDN) and TGFB (SB 431542) inhibitors from day 6 to day 9 to block sclerotome differentiation. Next, from day 9 to day 12, Platelet- Derived Growth Factor (PDGF) (AC 710) and Tyrosine receptor kinase B (TrkB) (ANA 12) inhibitors were added. Finally, BMP and PDGF pathways were inhibited starting from day 12. This new protocol is serum free as inhibitors were added to Dulbecco's Modified Eagle Medium (DMEM) with 15% Knock-Out Serum Replacement (KSR) and no factors used in the previous muscle differentiation protocol were used. Using flow cytometry, it can be shown that the optimized conditions increase the number of PAX3+ cells (+13.4%) and decrease the SOX9+ cells (-2.4%) at day 9. To trace the fate of myogenic cells, an mCherry fluorescent protein was introduced in the MYODI locus in our PAX7-Venus reporter line. At day 31 , with the optimized protocol, a significant increase in PAX7+ cells (+23.1 % on average), MY0D+ cells (+9.8% on average), and PAX7+MY0D+ cells (+21.1% on average) was observed together with a decrease of PAX7-MY0D- cells (-15.5% on average) (Fig 2F). Thus, the protocol optimized with the bioinformatics pipeline of FIGS. 1 and 2 generates a cell population significantly enriched in SCs.
[0045] ScRNAseq was used to perform a comparison of both protocols. Such an analysis allows capturing myofibers nuclei that are otherwise not quantified in FACS or scRNASeq analyses. Nuclei were extracted from day 31 cultures followed by lOx encapsulation and sequencing. In total, 1881 cells from the Chai protocol and 6763 cells from the new protocol. Both datasets were similarly processed to generate UMAP projections and Leiden clustering. The identity of clusters in both datasets was established using differential gene expression. Both datasets contain the same cell types corresponding to PSM derivatives: mesenchymal progenitors (PDGFRA+), satellite cells (PAX7+), muscle fibers (ACTN2+) and neural cells (MAP2+). However, the two protocols yield different proportions of these cell types. The new protocol resulted in a higher proportion of SC nuclei (25. 1%) than the Chai protocol (10.8%). This was accompanied by a lower proportion of all other populations (mesenchymal, muscle fibers and neural cells). Next, a ‘satellite cell score’ based on expression of the SC markers PAX7, NOTCH3, FGFR4, CD82, CDH15 was created and displayed in a matrix plot. This revealed that the SC cluster from the new protocol exhibits the highest score than any other cluster from both protocols.
[0046] A normalized enrichment score was also calculated for the signaling pathways that we are modulating in the two myogenic protocols. This score was computed using the pre-rank function of GSEA and Bio Planet 2019 signaling pathway gene sets. The pathways inhibited in the new protocol were all decreased compared to the Chai protocol. To identify the developmental stage of the myogenic cells generated in vitro, their gene expression profiles were compared to a scRNAseq dataset of human embryonic limbs ranging from PCW 5 to 9. Cells of the myogenic lineage were extracted (based on the original annotations) and performed clustering of this myogenic dataset using Leiden. This analysis identified three myogenic clusters, including SCs (PAX7+), myoblasts (MYMK+) and myocytes (MYH3+). STREAM pseudotime analysis identified the developmental trajectory linking these three clusters. Comparison of the three clusters identified in the in vitro datasets using a classifier trained against the human embryo myogenic dataset indicated that the three clusters detected in vitro are similar to their in vivo counterparts. Further, the classifier indicated that the SCs generated in vitro are closest to the 8 weeks PAX7+SCs in the human embryo, thus, staging the in vitro SC at the embryonic / fetal transition. Overall, this new protocol allows the generation of up to 52.3% of PAX7+ cells on average in the mononucleated fraction, resulting in a significant enrichment in the PAX7+ population compared to our previous myogenic induction protocol.Activation of the Hedgehog Pathway for Amplification of Human Muscle Stem Cells (Satellite Cells) in vitro
[0047] In some embodiments, the method further includes activating the Hedgehog pathway in the pluripotent stem cells to amplify muscle cell production. The Hedgehog signaling pathway is a signaling pathway that transmits information to embryonic cells required for proper cell differentiation. Different parts of the embryo have different concentrations of hedgehog signaling proteins. The treatment with Hedgehog pathway activators such as the small molecule small molecule smoothened agonist (SAG) could be successfully employed for the expansion of hiPSC-derived muscle stem cells. The process of expansion is currently one of the limiting steps in the development of cell therapy treatments for muscle diseases such as Duchenne Muscular Dystrophy.
[0048] Human-induced pluripotent stem cells (hiPSC) are reprogrammed cells capable of differentiating into any type of embryonic tissue such as nerve cells, lungs or skeletal muscle. Since its initial development, hiPSC-derived tissues have emerged as a promising tool to studyhuman biology, providing a virtually unlimited source of tissues for disease modelling applications. Importantly, as hiPSC-derived tissues can be engineered with patient-specific cells, they constitute an emerging tool in the field of cellular therapy. Muscle diseases, known as myopathies, are a significant cause of functional disability. Our laboratory, being one of the pioneers in the field, developed a protocol to produce hiPSC-derived skeletal muscle, which results in a dense culture of myofibers interspersed with muscle stem cells.
[0049] However, to be able to establish cell therapies there is a need to scale-up the production of hiPSC-derived muscle stems cells. The inventors identified that the activation of Hedgehog pathway via treatment with human Sonic Hedgehog recombinant protein (SHH N-terminus, Bio-Techne Corporation #1845-SH-025) or with small molecules such as the Smoothened agonist (SAG, Sigma-Aldrich #566661) induces an increase in the proliferation of hiPSC- derived muscle stem cells, thus promoting their amplification allowing to double the proportion of PAX7+ muscle stem cells in myogenic cultures derived from human IPS cells. Accordingly, in some embodiments, the Hedgehog pathway is activating using a Sonic Hedgehog recombinant protein or a Smoothened agonist. The addition of SAG to culture media used for muscle stem cells is considerably more efficient in amplifying the muscle stem cell population than the original culture media and represents an interesting addition for scaling up production of these cells for cell therapy applications.
[0050] To test whether the activation of the Hedgehog pathway could promote the amplification of hiPSC-derived muscle stem cells, we have tested the treatment of both SHH recombinant protein (SHH N-terminus, Bio-Techne Corporation #1845-SH-025) and the small molecule Smoothened agonist (SAG, Sigma-Aldrich #566661), two well-known activators of the Hedgehog pathway, in hiPSC-derived muscle stem cells isolated by Fluorescence- activated cell sorting (FACS). The parental hiPSC line used for these experiments is a genetically engineered double reporter PAX7-YFP and ACTN2-RFP which allows us to monitor the presence of both muscle stem cells (PAX7-YFP+) and the presence of the differentiated myofibers (ACTN2-RFP+). Treatment of the isolated hiPSC-derived muscle stem cells with SHH (doses tested: 0.5-1 pg / ml) or SAG (doses tested: 125-1000 nM) resulted in higher proliferation (characterized by the co-staining of PAX7-YFP and KI67, an established proliferative marker) of the hiPSC-derived muscle stem cells. Differentiation of the muscle stem cells was not affected by the treatments, as no decrease in the expression of ACTN2- mKATE2 was identified upon imaging with fluorescence microscopy. This led to a doublingof the population of PAX7+ muscle stem cells in the treated cultures compared to controls. These experiments have been conducted in three independent batches of hiPSC-derived muscle stem cells. The results are shown in Figures 5A and 5B.
[0051] FIG. 6 is a schematic block diagram illustrating an exemplary system 600 of hardware components capable of implementing examples of the systems and methods disclosed herein. The system 600 can include various systems and subsystems. The system 600 can be a personal computer, a laptop computer, a workstation, a computer system, an appliance, an applicationspecific integrated circuit (ASIC), a server, a server BladeCenter, a server farm, etc.
[0052] The system 600 can include a system bus 602, a processing unit 604, a system memory 606, memory devices 608 and 610, a communication interface 612 (e.g., a network interface), a communication link 614, a display 616 (e.g., a video screen), and an input device 618 (e.g., a keyboard, touch screen, and / or a mouse). The system bus 602 can be in communication with the processing unit 604 and the system memory 606. The additional memory devices 608 and 610, such as a hard disk drive, server, standalone database, or other non-volatile memory, can also be in communication with the system bus 602. The system bus 602 interconnects the processing unit 604, the memory devices 606-610, the communication interface 612, the display 616, and the input device 618. In some examples, the system bus 602 also interconnects an additional port (not shown), such as a universal serial bus (USB) port.
[0053] The processing unit 604 can be a computing device and can include an applicationspecific integrated circuit (ASIC). The processing unit 604 executes a set of instructions to implement the operations of examples disclosed herein. The processing unit can include one or multiple processing cores. The additional memory devices 606, 608, and 610 can store data, programs, instructions, database queries in text or compiled form, and any other information that may be needed to operate a computer. The memories 606, 608 and 610 can be implemented as computer-readable media (integrated or removable), such as a memory card, disk drive, compact disk (CD), or server accessible over a network. In certain examples, the memories 606, 608 and 610 can include text, images, video, and / or audio, portions of which can be available in formats comprehensible to human beings. Additionally or alternatively, the system 600 can access an external data source or query source through the communication interface 612, which can communicate with the system bus 602 and the communication link 614.
[0054] In operation, the system 600 can be used to implement one or more parts of a system for generating protocols for in vitro generation of cells. Computer executable logic for implementing the diagnostic system resides on one or more of the system memory 606, and the memory devices 608 and 610 in accordance with certain examples. The processing unit 604 executes one or more computer executable instructions originating from the system memory 606 and the memory devices 608 and 610. The term "computer readable medium" as used herein refers to a medium that participates in providing instructions to the processing unit 604 for execution. This medium may be distributed across multiple discrete assemblies all operatively connected to a common processor or set of related processors.
[0055] Implementation of the techniques, blocks, steps, and means described above can be done in various ways. For example, these techniques, blocks, steps, and means can be implemented in hardware, software, or a combination thereof. For a hardware implementation, the processing units can be implemented within one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, other electronic units designed to perform the functions described above, and / or a combination thereof.
[0056] Also, it is noted that the embodiments can be described as a process which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart can describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations can be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in the figure. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination corresponds to a return of the function to the calling function or the main function.
[0057] Furthermore, embodiments can be implemented by hardware, software, scripting languages, firmware, middleware, microcode, hardware description languages, and / or any combination thereof. When implemented in software, firmware, middleware, scripting language, and / or microcode, the program code or code segments to perform the necessary tasks can be stored in a machine-readable medium such as a storage medium. A code segment or machine-executable instruction can represent a procedure, a function, a subprogram, aprogram, a routine, a subroutine, a module, a software package, a script, a class, or any combination of instructions, data structures, and / or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, and / or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, ticket passing, network transmission, etc.
[0058] For a firmware and / or software implementation, the methodologies can be implemented with modules (e.g., procedures, functions, and so on) that perform the functions described herein. Any machine-readable medium tangibly embodying instructions can be used in implementing the methodologies described herein. For example, software codes can be stored in a memory. Memory can be implemented within the processor or external to the processor. As used herein the term "memory" refers to any type of long term, short term, volatile, nonvolatile, or other storage medium and is not to be limited to any particular type of memory or number of memories, or type of media upon which memory is stored.
[0059] Moreover, as disclosed herein, the term "storage medium" can represent one or more memories for storing data, including read only memory (ROM), random access memory (RAM), magnetic RAM, core memory, magnetic disk storage mediums, optical storage mediums, flash memory devices and / or other machine-readable mediums for storing information. The term "machine-readable medium" includes, but is not limited to, portable or fixed storage devices, optical storage devices, wireless channels, and / or various other storage mediums capable of storing that contain or carry instruction(s) and / or data.
[0060] In the preceding description, specific details have been set forth in order to provide a thorough understanding of example implementations of the invention described in the disclosure. However, it will be apparent that various implementations may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the example implementations in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the examples. The description of the example implementations will provide those skilled in the art with an enabling description for implementing an example of the invention, but it should be understood that various changes may be made in the functionand arrangement of elements without departing from the spirit and scope of the invention. Accordingly, the present invention is intended to embrace all such alterations, modifications, and variations that fall within the scope of the appended claims.
Claims
CLAIMSWhat is claimed is:
1. A system comprising: a processor; and a non-transitory computer readable medium storing instructions executable by the processor, the machine-executable instructions comprising: a graph generator that receives a genomics dataset representing a target cell type and generates a low-dimensional graph from the genomics dataset; and a graph analyzer that analyzes the low-dimensional graph to determine a set of trajectories within the generated graph and assigns a score to each of the set of trajectories based on gene expression associated with the target gene within the trajectory.
2. The system of claim 1, further comprising a protocol testing platform configured to produce a cell culture based on a trajectory of the set of trajectories having a best score.
5. The system of claim 1, wherein the genomics dataset is a first genomics dataset, the system further comprising a data interface that receives a second genomics dataset and selects the first genomics dataset as a proper subset of the second genomics dataset.
6. The system of claim 1, wherein the graph generator embeds the genomics dataset into a graph, applies a clustering algorithm to identify different stages of cell differentiation within the graph, and learns a simplified representation of the graph to provide the low-dimensional graph7. The system of claim 1, wherein the graph analyzer applies a trajectory algorithm to the low-dimensional graph to identify the set of trajectories, calculates a score for each genomics entity within the low-dimensional graph, and fits the score for each trajectory in trajectory tree to assign the score to each of the set of trajectories.
8. A method comprising: generating a low-dimensional graph from a genomics dataset representing a target cell type;analyzing the low-dimensional graph to determine a set of trajectories within the lowdimensional graph; assigning a score to each of the set of trajectories based on gene expression associated with the target gene within the trajectory; and selecting a trajectory of the set of trajectories having a best score.
9. The method of claim 8, further comprising: developing a cell differentiation protocol from the trajectory of the set of trajectories having the best score; and producing a cell culture from the cell differentiation protocol.
10. The method of claim 9, wherein the genomics dataset is a first genomics dataset and the method further comprises: generating a second genomics dataset from the cell culture; generating a quality metric from the second genomics dataset; and accepting the cell differentiation protocol if the quality metric meets a threshold value.
11. The method of claim 10, wherein the low-dimensional graph is a first low-dimensional graph, the set of trajectories is a first set of trajectories, the method further comprising: generating a second low-dimensional graph from a genomics dataset representing a target cell type if the quality metric does not meet the threshold value; analyzing the second low-dimensional graph to determine a second set of trajectories within the low-dimensional graph; assigning a score to each of the second set of trajectories based on gene expression associated with the target cell within the trajectory; and selecting a trajectory of the second set of trajectories having a best score.
12. The method of claim 8, wherein the genomics dataset is a first genomics dataset, the method further comprising: receiving a second genomics dataset; and selecting the first genomics dataset as a proper subset of the second genomics dataset, the first genomics dataset comprising a plurality of time series in the second genomics dataset containing time points associated with the target cell type.
13. The method of claim 8, wherein generating a low-dimensional graph from a genomics dataset representing a target cell type comprises applying at least one of normalization, scaling, and batch correction to the genomics dataset.
14. The method of claim 8, wherein generating the low-dimensional graph from the genomics dataset comprises: embedding the genomics dataset into a graph; applying a clustering algorithm to identify different stages of cell differentiation within the graph; and generating the low-dimensional graph as a simplified representation of the graph.
15. The method of claim 8, wherein analyzing the low-dimensional graph to determine the set of trajectories and assigning the score to each of the set of trajectories comprises: applying a trajectory algorithm to the low-dimensional graph to identify the set of trajectories; calculating a score for each genomics entity within the low-dimensional graph; and fitting the score for each trajectory in the trajectory tree to assign the score to each of the set of trajectories.
16. A method comprising: generating a low-dimensional graph from a genomics dataset representing a target cell type; analyzing the generated graph to determine a set of trajectories within the generated graph; assigning a score to each of the set of trajectories based on gene expression associated with the target gene within the trajectory; selecting a trajectory of the set of trajectories having a best score; developing a cell differentiation protocol from the trajectory of the set of trajectories having the best score; and producing a cell culture from the cell differentiation protocol.
17. The method of claim 16, wherein the genomics dataset is a first genomics dataset and the method further comprises: generating a second genomics dataset from the cell culture; generating a quality metric from the second genomics dataset; and accepting the cell differentiation protocol if the quality metric meets a threshold value.
18. The method of claim 17, wherein the low-dimensional graph is a first low-dimensional graph, the set of trajectories is a first set of trajectories, the method further comprising: generating a second low-dimensional graph from a genomics dataset representing a target cell type if the quality metric does not meet the threshold value; analyzing the second low-dimensional graph to determine a second set of trajectories within the low-dimensional graph; assigning a score to each of the second set of trajectories based on gene expression associated with the target cell within the trajectory; and selecting a trajectory of the second set of trajectories having a best score.
19. The method of claim 16, wherein generating the low-dimensional graph from the genomics dataset comprises: embedding the genomics dataset into a low-dimensional graph; applying a clustering algorithm to identify different stages of cell differentiation within the low-dimensional graph; and learning a simplified representation of the low-dimensional graph.
20. The method of claim 16, wherein analyzing the low-dimensional graph to determine the set of trajectories and assigning the score to each of the set of trajectories comprises: applying a trajectory algorithm to the low-dimensional graph to identify the set of trajectories; calculating a score for each genomics entity within the low-dimensional graph; and fitting the score for each trajectory in the trajectory tree to assign a score to each of the set of trajectories.
Citation Information
Patent Citations
Methods utilizing single cell genetic data for cell population analysis and applications thereof
US20200370112A1
System for image-driven cell manufacturing
US20210301247A1
Systems and methods for joint low-coverage whole genome sequencing and whole exome sequencing inference of copy number variation for clinical diagnostics
US20220215900A1