Pedigree tracking method
By introducing viral barcodes into the human ventral midbrain-hindbrain differentiation model through the SISBAR method and combining potential perspective and origin perspective analysis, the problem that the existing technology cannot effectively define the lineage relationship between cells and their progeny is solved, and the drawing of multi-level lineage maps and the revelation of developmental trajectories are achieved.
Patent Information
- Application Number
- CN202380094586.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-21
- Filing Date
- 2023-12-21
- Publication Date
- 2025-09-16
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Background Art
[0001] A fundamental interest in developmental and stem cell biology is mapping the developmental trajectories and lineage relationships of functional cells within a tissue or organ. Understanding these lineage relationships provides insights into the molecular mechanisms underlying normal development as well as pathological states, including cancer and developmental disorders. It also provides insights into manipulating cell fate in vivo, improving cell differentiation processes in vitro, and enhancing cell-based therapeutic approaches for regenerative medicine.
[0002] Recent advances in single-cell RNA sequencing (scRNA-seq) have enabled transcriptome analysis of thousands of single cells, providing a powerful method for identifying single-cell identities and resolving the diversity of cell types in tissues or cell cultures that emerge during in vitro development or cell differentiation. Computational methods based on single-cell transcriptomics have been developed to describe lineage trajectories during differentiation by sampling cells at multiple stages. Most of these algorithms rely on the assumption that cells with similar transcriptome profiles are more likely to originate from the same lineage. Therefore, lineage trajectories based on single-cell transcriptomics do not necessarily reflect the true lineage relationships between cells.
[0003] The classic and gold-standard method for defining the kinship between different cell types is lineage tracing. Historically, there have been two main paradigms for studying cell lineages: prospective lineage tracing and retrospective lineage tracing. Methods developed for prospective lineage tracing rely on experimentally introducing lineage tracers into progenitor cells at an early time point and tracing their descendants at a later time point. In contrast, retrospective lineage tracing methods use naturally occurring genetic mutations (e.g., copy number variations, single nucleotide variants, or microsatellites) to infer the past developmental relationships of differentiated cells. Recently, some studies have combined scRNA-seq with prospective and retrospective lineage tracing, enabling the simultaneous measurement of the whole transcriptome of single cells as well as their lineage information, opening up new avenues for studying cell lineages. Although prospective and retrospective lineage tracing methods have provided valuable insights into the mechanisms and principles of embryonic development, cell differentiation, and pathology, these methods are mainly used to elucidate the lineage relationships between cells at a given time point, but cannot provide rich cell state information about their ancestors or directly resolve the lineage relationships between ancestors and their descendants. Therefore, to fully understand the developmental trajectories of tissues and functional cells, there is a great need to develop methods to define the lineage relationships between cells and their progeny across time and developmental stages, ideally combined with rich single-cell transcriptome data. Summary of the Invention
[0004] The applicants developed single-cell division barcodes (SISBARs) to enable cross-stage clonal tracking of single-cell transcriptomes in an in vitro model of human ventral midbrain-hindbrain differentiation. The applicants further developed "potential perspective" and "origin perspective" analyses to explore cross-stage lineage relationships and created multi-level clonal lineage maps depicting the entire differentiation process. Using SISBARs, the applicants discovered a large number of previously uncharacterized convergent and divergent trajectories.
[0005] In one aspect, the present application provides a lineage tracing method, which includes harvesting cells and their sister cells at different stages of differentiation, dividing the cells and their sister cells into at least two parts, one part for genetic analysis, and the other part for genetic analysis after further differentiation, and finally analyzing the cross-stage lineage.
[0006] On the other hand, the present application provides a method for predicting cell differentiation fate, which includes harvesting cells and their sister cells at different differentiation stages, dividing the cells and their sister cells into at least two parts, one part for genetic analysis, and the other part for genetic analysis after continuing to differentiate, and finally analyzing the cross-stage lineage.
[0007] On the other hand, the present application provides a method for identifying the origin of cells, which includes harvesting cells and their sister cells at different differentiation stages, dividing the cells and their sister cells into at least two parts, one part for genetic analysis, and the other part for genetic analysis after further differentiation, and finally analyzing the cross-stage lineage.
[0008] In some embodiments, cells at different stages of differentiation undergo cell division and give rise to sister cells.
[0009] In some embodiments, cells at different stages of differentiation are derived from the same kind of progenitor cell.
[0010] In some embodiments, cells at different stages of differentiation are derived from different types of progenitor cells.
[0011] In some embodiments, the cells are derived from pluripotent stem cells.
[0012] In some embodiments, the cells are derived from human pluripotent stem cells.
[0013] In some embodiments, the pluripotent stem cells include embryonic stem cells (ESCs) and / or induced pluripotent stem cells (iPSCs).
[0014] In some embodiments, the progenitor cells comprise neural progenitor cells.
[0015] In some embodiments, the differentiation is in vitro differentiation.
[0016] In some embodiments, the method comprises introducing a marker into the differentiation-initiating cells.
[0017] In some embodiments, a marker is used to label differentiation-initiating cells.
[0018] In some embodiments, the label comprises a barcode.
[0019] In some embodiments, the barcode comprises a viral barcode.
[0020] In some embodiments, the viral barcode comprises a lentiviral vector.
[0021] In some embodiments, the viral barcode comprises a retroviral vector.
[0022] In some embodiments, viral barcoding comprises inserting a 32 bp semi-random barcode into a retroviral vector.
[0023] In some embodiments, the viral barcode contains 16 repeats of an alternating S (G or C) / W (A or T) nucleic acid sequence.
[0024] In some embodiments, the viral barcode contains a random nucleic acid sequence.
[0025] In some embodiments, genetic analysis includes recovery of viral barcodes.
[0026] In some embodiments, cells with the same viral barcode are defined as a single clone.
[0027] In some embodiments, the marker comprises a Cre-LoxP system.
[0028] In some embodiments, markers include insertions and / or deletions of traceable elements introduced by the CRISPR-Cas9 system.
[0029] In some embodiments, the method comprises one or more stages of differentiation.
[0030] In some embodiments, the method comprises two or more stages of differentiation.
[0031] In some embodiments, the method comprises three or more stages of differentiation.
[0032] In some embodiments, the genetic analysis comprises sequencing the cell or its sister cells.
[0033] In some embodiments, sequencing comprises transcriptome sequencing.
[0034] In some embodiments, the sequencing comprises single-cell RNA sequencing (scRNA-seq).
[0035] In some embodiments, the method comprises performing data analysis after sequencing the cell or its sister cells.
[0036] In some embodiments, the method includes potential viewing angle analysis.
[0037] In some embodiments, potential perspective analysis involves examining the potential of a given progenitor cell type at an early stage by calculating the distribution of clonal cells at a later stage and comparing it to a random distribution of permutations.
[0038] In some embodiments, the method comprises a provenance perspective analysis.
[0039] In some embodiments, origin-perspective analysis involves examining the origin of a given cell type at a later stage by calculating the distribution of clonal cells at an earlier stage and comparing it to a random distribution of permutations.
[0040] In some embodiments, the distribution of clonal cell types can be visualized by heatmap clustering based on the Pearson correlation coefficient of the distribution.
[0041] In some embodiments, the lineage relationships of different cell clusters within a developmental stage are determined by calculating the Jaccard similarity coefficient score (the number of shared clones divided by all clones of a pair of clusters), P value (given the null hypothesis that cells in a clone are randomly distributed among all clusters), lineage coupling correlation analysis and / or clonal coupling (observed / expected) between each pair of cell clusters.
[0042] In some embodiments, the method includes the following steps: (a) introducing a viral barcode into differentiation-initiating cells; (b) culturing the differentiation-initiating cells and obtaining cells at the next differentiation stage and their sister cells; (c) dividing the cells at the next differentiation stage and their sister cells into two parts, one part for viral barcode recovery and scRNA-seq, and the other part for continued differentiation; (d) if differentiation includes two or more stages, repeating (c) until the differentiation process is completed; (e) after sequencing, constructing a lineage tree using potential analysis and origin analysis.
[0043] In another aspect, the present application provides use of the method of the present application in lineage tracing of pluripotent stem cell differentiation products.
[0044] In another aspect, the present application provides use of the method of the present application in lineage tracing of a heterogeneous cell population.
[0045] Other aspects and advantages of the present disclosure will become apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be appreciated, the present disclosure is capable of other and different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not restrictive.
[0046] Incorporated by reference
[0047] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The novel features of the present invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description and accompanying drawings (also referred to herein as "figures"), which set forth illustrative embodiments embodying the principles of the present invention, wherein:
[0049] Figures 1A-1F Demonstrates the use of SISBAR for clonal tracking of single-cell transcriptomes at different stages
[0050] (1A) Schematic diagram of SISBAR for cross-stage lineage tracing. Each dot represents a cell. Cells with unique viral barcodes are marked with different colors.
[0051] (1B) hESC-based ventral midbrain-hindbrain neural differentiation protocol.
[0052] (1C) Paired experimental design for studying lineage relationships at different stages in an hPSC-based in vitro model of ventral midbrain-hindbrain differentiation.
[0053] (1D) Annotated UMAP plots of scRNA-seq data from the indicated differentiation stages. scRNA-seq data from the same stage from different experiments and three independent biological replicates were integrated.
[0054] (1E) Number of cells and clones captured in the experiment.
[0055] (1F) Reconstruction and visualization of cells across stage lineages from Experiment #1 (768 clones, 2,762 cells), Experiment #2 (1,213 clones, 5,856 cells), and Experiment #3 (559 clones, 2,816 cells) using a force-directed network. Each node represents a single cell linked by clonal relationships. Purple nodes: Stage I cells; yellow nodes: Stage II cells; green nodes: Stage III cells; blue nodes: Stage IV cells.
[0056] Figures 2A-2N Describes mapping lineage relationships between cell types across stages through potential and origin perspectives
[0057] (2A and 2B) Examples of stage III sister cells (each color represents a clone) (A) and quantitative analysis of stage III sister cells based on gene expression correlation and graphical distance on UMAP plots (B). ***P < 0.001, one-sided t-test.
[0058] (2C) Heatmaps showing the distribution of clonal fates across stages III and IV. Each row on both heatmaps represents a clone across stages; the color depth indicates the proportion of clonal cells in each cluster across stages. Proportional.
[0059] (2D) Schematic diagram of potential- and origin-based analyses. Each dot represents a clone. Different clones are marked with different colors.
[0060] (2E) Dot plot showing lineage coupling scores (P values and Jaccard scores) between stage III and stage IV cell type pairs derived from potential- and origin-perspective analyses. The color depth of the dots reflects the significance of the P value (dots with P ≤ 0.05 are shown). The size of the dots indicates the Jaccard score (Jac.), which is the proportion of shared clones between cell type pairs. The number within each dot indicates the number of clones that overlap between each cluster pair.
[0061] (2F and 2K) depict III-NP vMB (2F, clone number = 72) or III-NP pHB (2K, clones = 204) Lineage trajectory of differentiation from stage III to stage IV. Clones with 100% progenitor cells in a given progenitor cluster were selected as progenitor cluster-associated clones, and all cluster-associated clones were plotted on UMAP plots for both stages. Expression of marker genes for the corresponding cell types is shown. P and O, P values for potential- or origin-based analyses, respectively.
[0062] (2H) Experimental design for tracking the fate of ventral midbrain progenitor cells.
[0063] (2I) Representative immunofluorescence images of TH, vGlut2 (RNA probe), and COL1A1 in LMX1A+ / EN1+-derived cultures. The boxed area on the right is magnified. Scale bars: 100 μm in full-length micrographs and 50 μm in magnified micrographs.
[0064] (2J) Percentage of TH+, vGlut2+, and COL1A1+ cells in the tdtomato+ population in LMX1A+ / EN1+-derived cultures.
[0065] (2G and 2L) depict III-NP vMB (2G) and III-NP pHB (2L) Schematic diagram of lineage divergence.
[0066] (2M and 2N) Lineage hierarchy between stage III and stage IV cell types inferred based on shared barcodes (P-value significance) (2M) or transcriptome similarity (2N).
[0067] Figures 3A-3N Illustrates the construction of multi-level lineage trees and clonal lineage analysis of convergent trajectories
[0068] (3A and 3B) Heat maps showing the distribution of clonal fates between stages I and II (A) and II and III (B).
[0069] (3C and 3D) Dot plots showing lineage coupling scores (P values and Jaccard scores) between cell type pairs from stage I and stage II (3C) or stage II and stage III (3D) using potential- and origin-perspective analyses. The numbers within each dot are the number of clones that overlap between each cluster pair.
[0070] (3E) Multilevel lineage tree showing statistically significant lineage relationships of cell types in four differentiation stages.
[0071] (3F and 3G) Schematic diagrams of two types of convergent lineage trajectories.
[0072] (3H) Highlighted cells in stage III are colored according to the initial state of the clonally related progenitors observed in stage II.
[0073] (3I) II-NP aHB Derivatization (left) and II-NP MHB Derivatized (right) III-NP aHB Differentially expressed genes in . FDR-corrected P values are shown. Selected genes are labeled.
[0074] (3J) II-NP aHB (Left) and II-NP MHB (Right) Differentially expressed genes. FDR-corrected P values are shown. Selected genes are labeled.
[0075] (3K) II-NP aHB and II-NP MHB Marker gene expression was projected onto the Phase II UMAP plot.
[0076] (3L) Schematic diagram of lentiviral labeling of FGF8-lineage cells three days before stage II and identification of their lineage fate using scRNA-seq at stage III.
[0077] (3M) scRNA-seq analysis of the fate of stage III FGF8 lineage cells (4,101 cells).
[0078] (3N)Description and Figure 1D UMAP embedding of stage III Fgf8 lineage cells during stage III cell integration.
[0079] Figures 4A-4I Clonal lineage analysis illustrating divergent trajectories
[0080] (4A) Model that can explain the observed population-level divergence trajectories.
[0081] (4B) III-NPs with multipotent or unipotent fates vMB A typical example of a single cross-stage clone.
[0082] (4C) Left, III-NP vMB Cells are colored according to their fate at stage IV. Right panel shows III-NP vMB Statistical pie chart of clone fate distribution.
[0083] (4D) II-NPs with multipotent or unipotent fates pHB A typical example of a single cross-stage clone.
[0084] (4E) Left, II-NP pHB Cells are colored according to their fate at stage III. Right panel shows II-NP pHB Statistical pie chart of clone fate distribution.
[0085] (4F and 4G) Compared with other II-NPs with different fates pHB , in III-NP pHB Fate II-NP pHBEnriched genes (4F) and top biological processes (4G) are shown. FDR-corrected P values are shown. Selected genes are labeled.
[0086] (4H and 4I) relative to other III-NPs with different fates pHB , in IV-Motor fate III-NP pHB Enriched genes (4H) and top biological processes (4I) in the . FDR-corrected P values are shown. Selected genes are labeled.
[0087] Figures 5A-5H Identifying the time points of lineage divergence in ventral midbrain neural progenitor cells
[0088] (5A) Exploring NP vMB Experimental design for lineage divergence time points: Progenitor cells were labeled with a barcoded retroviral library two days after stage II or III, followed by 7 days of clonal expansion and 14 days of maturation, and then cells were harvested for scRNA-seq and viral barcode recovery.
[0089] (5B) UMAP plot of integrated scRNA-seq data from the experiment performed in (5A), and feature plots showing expression of marker genes for mDA (PITX3+ / SLC17A6+), mGlut (PITX3– / SLC17A6+), and VLMCs (COL1A1+).
[0090] (5C) Representative examples of individual clones with mDA neuronal fate and statistical pie charts showing the clonal fate distribution of mDA-associated clones two days after Phase II labeling and Phase III labeling experiments.
[0091] (5D) Density contour plots of mDA, mGlut, and cells in OTX2+VLMC-associated clones.
[0092] (5E) Histogram of lineage correlations between mDA and other terminal cell types. Blue bars show the distribution of lineage correlation scores for the isopotency null model, where clonal cells are randomly distributed in clusters. Red circles represent observed lineage correlations.
[0093] (5F) Tracking individual NPs vMB Cell fate experimental design.
[0094] (5G) Representative immunofluorescence images showing the three clonal fates of single EN1+ / LMX1A+ progenitor cells. Scale bar, 100 μm.
[0095] (5H) Statistical pie chart showing the clonal fate distribution of single EN1+ / LMX1A+ cells sorted at stage II (127 clones) or two days after stage III (26 clones).
[0096] Figures 6A-6J Describes barcoded retroviral library construction and viral barcode recovery
[0097] (6A) Top, semirandom DNA barcode structure (left and right anchor sequences are considered as homology arms for Gibson cloning and recognition sites for barcode sequence positioning). Bottom, probability of bases at each position of the 32 bp semirandom barcode sequence in the barcoded retroviral library.
[0098] (6B) Left, image of neurospheres labeled with a barcoded retroviral library at stage III. Scale bar, 500 μm. Right, diagram of the FACS sorting scheme used to isolate EGFP-positive cells at stage III.
[0099] (6C) UMAP and histogram analysis of scRNA-seq data for stage III virus-labeled cells (n = 5000 cells, labeled at stage II) and unlabeled cells (n = 5000 cells). Virus-labeled and control cells are evenly distributed among these populations, although some variation is expected between these independent biological replicates.
[0100] (6D) Comparison of EGFP expression levels with the expression level of the housekeeping gene GAPDH.
[0101] (6E) EGFP expression levels are higher than 95% of other endogenous genes in the single-cell transcriptome.
[0102] (6F) Diversity of the barcoded retroviral library, estimated by high-throughput sequencing of approximately 48 million reads from the plasmid library.
[0103] (6G) Barcode library diversity estimation. Left: Observed number of overlapping viral barcodes between three separate scRNA-seq experiments transduced with the same retroviral library. Right: Expected number of overlapping viral barcodes between these three experiments simulated using three million unique barcodes. We repeated the simulation 500 times and took the median number of overlapping barcodes between each sample pair as the expected number of overlapping barcodes between samples. Barcodes in the blacklist were removed before calculating overlap.
[0104] (6H) Distribution of minimum barcode distances (the distance between a given barcode and the most similar barcode in the entire library pool) within the barcoded retroviral library.
[0105] (6I) Scatter plot of clone size and barcode frequency in deep sequencing pools.
[0106] (6J) Schematic diagram of viral barcode recovery from single-cell transcriptomes: During scRNA-seq, EGFP barcoded mRNA transcripts are captured and reverse-transcribed into cDNA on gel beads in parallel with other transcripts. EGFP barcode sequences are then enriched by PCR amplification using specific primers and sequenced using next-generation sequencing. UMI, unique molecular identifier.
[0107] Figures 7A-7E SISBAR bioinformatics processing is illustrated.
[0108] (7A) Schematic diagram of the SISBAR data processing and filtering pipeline.
[0109] (7B) Read count distribution of cellular barcodes, UMIs, and viral barcodes combined before (left) and after (right) filtering.
[0110] (7C) Read count distribution of cellular barcodes, UMIs, and viral barcodes combined before (left) and after (right) string distance filtering.
[0111] (7D and 7E) Whitelist (7D, cell barcodes extracted by Cellranger) and blacklist (7E, high-copy viral barcodes in the genomic barcode pool of 293T cells infected with the barcoded virus library).
[0112] Figures 8A-8G Transcriptional analysis of an hPSC-based midbrain-hindbrain differentiation dataset is illustrated.
[0113] (8A) Proportions of major cell types at four differentiation stages in three biological replicates.
[0114] (8B) In situ hybridization (ISH) images of selected genes in E11.5 mouse brain from the Allen Brain Atlas.
[0115] (8C) Representative marker gene expression for each cell type at stages I, II, III, and IV.
[0116] (8D) Immunofluorescence images of stage IV TH+ and COL1A1+ cells. Scale bar, 100 μm.
[0117] (8E and 8G) Heatmaps depicting pairwise transcriptional correlations between hPSC-derived ventral midbrain-hindbrain cells and cells in the Developing Mouse Brain Atlas (Stage IV) (8E) and the Human Ventral Midbrain dataset (La Manno et al., 2016) (8G).
[0118] (8F) Integrated UMAP plot of scRNA-seq data from hPSC-derived ventral midbrain-hindbrain cells (stage IV) and the mouse nervous system, highlighting clusters of COL1A1+ cells (IV-OTX2+VLMC, IV-LUM+VLMC), vascular and leptomeningeal cells (VLMC), and enteric glial cells.
[0119] Figures 9A-9H Clonal and lineage relationship analysis for phase III-IV is illustrated.
[0120] (9A) Heatmaps showing the distribution of clonal fates during stages III and IV. Each row in both heatmaps represents a clone; the color depth indicates the number of cells in each cluster at each stage.
[0121] (9B and 9C) Comparison of transformed p-values (9B) and Jaccard scores (9C) for lineage relationships across stages (stages III and IV) between three biological replicates. Transformed p-value = log2(-log2(p-value) + 0.001). Lineage relationships for cell type pairs with fewer than five overlapping clones are not shown.
[0122] (9D) EN1 and LMX1A gene expression on UMAP plots during stage III.
[0123] (9E and 9G) Heatmaps of stage IV lineage relatedness scores (P values) between pairs of clusters based on shared barcodes (9E) or transcriptome similarity (9G).
[0124] (9F and 9H) Lineage hierarchies among stage IV clusters inferred based on shared barcodes (9F) or transcriptome similarity (9H).
[0125] Figures 10A-10J Analysis of clones and lineages spanning stages I and II or stages II and III is illustrated.
[0126] (10A and 10B) Heatmaps showing the distribution of clonal fates at stages I-II (A) and II-III (B). Each row in both heatmaps represents a clone; the color depth indicates the number of cells in each cluster at each stage.
[0127] (10C to 10F) Examples of sister cells at early stages (C and E, each color represents a clone) and quantitative analysis of sister cells based on gene expression correlation and graphical distance on UMAP plots (D and F).
[0128] (10G and 10I) Comparison of transformed p-values for lineage relationships between three biological replicates for stages I-II (10G) and II-III (10I). Transformed p-value = log²(-log²(p-value) + 0.001). Lineage relationships for cell type pairs with fewer than five overlapping clones are not shown.
[0129] (10H and 10J) Comparison of Jaccard scores for lineage relationships at stages I-II (10H) and II-III (10J) between three biological replicates. Lineage relationships for cell type pairs with fewer than 5 overlapping clones are not shown.
[0130] Figures 11A-11D Describes cross-stage lineage trajectory analysis and the construction of FGF8-iCre hESC lines.
[0131] (11A) Describing II-NP pHB The lineage trajectory of differentiation from stage II to stage III is shown. The expression of marker genes for the corresponding cell types is also shown next to them. Clones with 100% progenitor cells in a given progenitor cluster were selected as progenitor cluster-associated clones, and all II-NP pHB Cluster-associated clones are plotted on UMAP plots from two phases. Number of clones = 251. P, P value from potential-perspective analysis; O, P value from origin-perspective analysis.
[0132] (11B and 11C) Schematic diagrams illustrating the difference between latent-view analysis and origin-view analysis. Each dot represents a clone. Different clones are marked with different colors.
[0133] (11D) Left panel, strategy for generating FGF8-iCre knock-in hESC lines. Exons are shown as gray boxes. Right panel, PCR genotyping of FGF8-iCre hESC clones. Clones with red and blue horizontal arrows indicate the expected size of PCR products used to assess target locus insertion and homozygosity, respectively. Heterozygous clones (red asterisks) were selected for downstream experiments.
[0134] Figures 12A-12H Shows NP vMB and NP pHB Clonal analysis of relevant divergent trajectories and associated molecular features.
[0135] (12A) Characteristic graph showing the expression of marker genes for stage IV mDA (PITX3+ / SLC17A6+), mGlut (PITX3– / SLC17A6+), and VLMC (COL1A1+).
[0136] (12B) Cell density contour plots in stage IV IV-mDA-associated clones (left) and a representative example of a single tripotent clone with an mDA neuronal fate (right).
[0137] (12C) Pie chart showing the clonal fate distribution of stage IV IV-mDA-associated clones.
[0138] (12D) III-NPs with multipotent or unipotent fates pHB A typical example of a single cross-stage clone.
[0139] (12E) Left, III-NP pHB Cells are colored according to their fate at stage IV. Right panel, showing III-NP pHB Statistical pie chart of clone fate distribution.
[0140] (12F) Histogram showing the proportion of symmetric and asymmetric self-renewing clones for all progenitor cell types in stages I and II. To quantify the self-renewal potential of each progenitor cell type, we defined clones as self-renewing clones that had ≥1 clonal cell (late stage) with the same fate as the corresponding early stage progenitor cell. The numbers above the columns indicate the number of clones associated with the progenitor cell type (clones with 100% progenitor cells in the specified progenitor cell cluster were selected as progenitor cell type-associated clones).
[0141] (12G) Density distribution of p-values for all single-fate progenitors grouped by clone size.
[0142] (12H)IV-hGluts fate of III-NP pHB (Left) or IV-Seros fate of III-NP pHB (Right) Enriched genes compared to other III-NPs with different fates pHB FDR-corrected P values are shown. Selected genes are labeled.
[0143] Figures 13A-13D Illustrated are the hGlut-associated clonal fate distributions in prospective lineage-tracing experiments of cells labeled at stage II or two days after stage III.
[0144] (13A) Annotated UMAP plot of integrated scRNA-seq data from two prospective lineage tracing experiments.
[0145] (13B) Density contour plots of cells in hGlut-associated clones during the Phase II labeling experiment (left) and two days after the Phase III labeling experiment (right).
[0146] (13C and 13D) Representative examples of individual clones with hGlut fate. Pie charts showing the clonal fate distribution of hGlut-associated clones at Phase II labeling (13C) and two days after Phase III labeling experiments (13D).
[0147] Figure 14 illustrates the ventral midbrain lineage trajectory and associated biological processes.
[0148] (14A) Top: UMAP plots showing clonal trajectories of the ventral midbrain lineage from stage I to stage II (clone number = 112) and from stage II to stage III (clone number = 212). Bottom: Expression of marker genes (LMX1A, EN1, and OTX2) of the ventral midbrain lineage. Clones with 100% progenitor cells in a given progenitor cluster were selected as progenitor cluster-associated clones, and all cluster-associated clones were plotted on the UMAP plots of the two stages.
[0149] (14B to 14D) In Phase I, II-NP vMB Top biological processes enriched in fate-committed progenitors compared to other fate-committed progenitors (14B); in phase II, III-NP vMB Top biological processes enriched in stage III IV-mDA-destined progenitors compared to other stage III progenitors (14C); and top biological processes enriched in stage III IV-mDA-destined progenitors compared to other stage III progenitors (14D). DETAILED DESCRIPTION
[0150] Although various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided as examples only. Many variations, changes, and substitutions may be envisioned by those skilled in the art without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed.
[0151] As used herein, the term "barcode" generally refers to a label or identifier that conveys or is capable of conveying information about an analyte. A barcode can be part of an analyte. A barcode can be independent of the analyte. A barcode can be a label attached to an analyte (e.g., a nucleic acid molecule) or a combination of a label and an endogenous characteristic of the analyte (e.g., the size of the analyte or a portion of its sequence). A barcode may be unique. A barcode can have a variety of different formats. For example, a barcode can include: a polynucleotide barcode; random sequence nucleic acids and / or amino acids; and synthetic nucleic acids and / or amino acids. A barcode can be attached to an analyte in a reversible or irreversible manner. For example, a barcode can be added to a fragment of a deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) sample before, during, and / or after sequencing the sample. A barcode can identify and / or quantify individual sequencing reads.
[0152] As used herein, the term "lineage tracing" generally refers to a method for determining cell fate and / or defining the relationship between different cell types. The term used herein may include all techniques known to those skilled in the art that can be used for lineage tracing or the techniques used in this application.
[0153] As used herein, the term "differentiation" generally refers to the process by which a non-specific or less specific cell acquires specific cell characteristics. A differentiated cell or a differentiation-induced cell is a cell that occupies a more specific position in a cell lineage.
[0154] As used herein, the term "genetic analysis" may include all techniques for analyzing genetic information known to those skilled in the art. For example, use of the Cre-LoxP system as a marker, use of the CRISPR / Cas9 system as a marker, use of barcodes (e.g., viral barcodes) as markers, sequencing, prospective lineage tracing, retrospective lineage tracing, potential perspective analysis, origin perspective analysis, lineage coupling correlation (Bandler, RC, Vitali, I., Delgado, RN, Ho, MC, Dvoretskova, E., Ibarra Molinas, JS, Frazel, PW, Mohammadkhani, M., Machold, R., Maedler, S. et al. (2021). Single-cell delineation of lineage and genetic identity in the mouse brain. Nature.) and / or clonal coupling (observed / expected) (Weinreb, C., Rodriguez-Fraticelli, A., Camargo, F., and Klein, AM (2020). Lineage tracing on transcriptional landscapes links state to fate during differentiation. Science 367(6479): eaaw3381) for genetic analysis.
[0155] As used herein, the term "potential perspective analysis" generally refers to examining the potential of a given progenitor cell type at an early stage by calculating the distribution of clonal cells at a later stage.
[0156] As used herein, the term "origin-perspective analysis" generally refers to examining the origin of a given cell type at a later stage by calculating the distribution of clonal cells at an earlier stage.
[0157] As used herein, the term "sequencing" generally refers to methods and techniques for determining the nucleotide base sequence in one or more polynucleotides. Polynucleotides can be, for example, nucleic acid molecules, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including variants or derivatives thereof (such as single-stranded DNA). Sequencing can be performed by currently available various systems, such as, but not limited to, Illumina®, Pacific Biosciences (PacBio®), Oxford Nanopore®, or Life Technologies (Ion Torrent®) sequencing systems. Alternatively or in addition, sequencing can be performed using nucleic acid amplification, polymerase chain reaction (PCR) (for example, digital PCR, quantitative PCR, or real-time PCR) or isothermal amplification. Such systems can provide a plurality of original genetic data corresponding to the genetic information of a subject (for example, the human race), such as generated by the sample provided by the system according to the subject. In some instances, such systems provide sequencing reads (also referred to herein as "reads"). Reads can include a string of nucleic acid bases corresponding to the sequence of sequenced nucleic acid molecules. In some cases, provided herein are systems and methods that can be used together with proteomics information.
[0158] As used herein, the term "pluripotent stem cell" generally refers to a class of cells that have the potential to differentiate into any cell type in the human body. Pluripotent stem cells can be derived from fertilized eggs or somatic cells, including blood cells, urine cells, skin cells and / or umbilical cord blood cells. Depending on the source, pluripotent stem cells include human embryonic stem cells (derived from fertilized eggs) and human induced pluripotent stem cells (somatic cells). Pluripotent stem cells have the ability to continuously proliferate and differentiate into various cell types.
[0159] As used herein, "progenitor cell," "precursor," and "precursor cell" are used interchangeably.
[0160] As used herein, the term "neural progenitor cell" generally refers to an undifferentiated progenitor cell that has not yet expressed terminal differentiation characteristics and is capable of proliferating and / or differentiating into a mature neuronal cell.
[0161] As used herein, the term "cell population" may include human stem cells; progenitor cells or their precursors; and differentiated cells. For example, the cell population may include neural progenitor cells and / or neural precursor cells or mature neurons derived from human stem cells or precursors, and neuronal derivatives derived therefrom, but is not limited thereto. Specifically, examples of human stem cells or precursor cells may include, but are not limited to, embryonic stem cells, embryonic germ cells, embryonic carcinoma cells, induced pluripotent stem cells (iPSCs), adult stem cells, and fetal cells.
[0162] Methods and uses
[0163] Constructing lineage trees depicting the individual development of functional cells and tissues has been a long-standing goal in developmental and stem cell biology. The applicants developed SISBAR to map lineage trajectories at multiple developmental stages in an hPSC-based ventral midbrain-hindbrain differentiation in vitro model. The applicants used "latent perspective analysis" and "origin perspective analysis" to statistically identify the fate outcomes or origins of individual transcriptome-defined cell types and discovered a large number of previously unknown divergent and convergent trajectories. Importantly, through clonal analysis, the applicants demonstrated that lineage trajectories at the population level are the consensus outcomes of highly variable individual lineage trajectories. Further sister cell analysis revealed the underlying molecules and signaling pathways behind the different fates of individual progenitor cells within the same cell type. These results demonstrate the practicality of SISBAR in revealing lineage dynamics during human neural differentiation and highlight its application in regenerative medicine.
[0164] Compared with methods that use a serial clonal sampling strategy across multiple or all stages (i.e., I-II-III-IV), which tends to oversample large clones, the SISBAR method uses a pairwise approach (i.e., I-II, II-III, and III-IV) to assess lineage relationships across stages to avoid interference caused by the large size differences typically seen in developmental systems. This allows for a more balanced detection of lineage relationships between consecutive stages and, therefore, better resolution of lineage trajectories, particularly those at the end of differentiation stages.
[0165] The SISBAR approach provides valuable insights into defining lineage trajectories and can reveal divergent or convergent lineage trajectories that cannot be inferred from traditional prospective or retrospective lineage tracing methods. In prospective lineage tracing combined with genetic barcoding, the identification of divergent or convergent lineage trajectories is based on the assumption that, at a single terminal time point, the terminating cell type must share clonality with at least one other cell type to form a divergent trajectory, or with at least two different starting cell types to form a convergent trajectory. However, this approach can be problematic in two situations: 1) when two subpopulations of the same type of unipotent or fate-committed progenitor cells with similar transcriptional states (within the same UMAP cluster) give rise to different progeny cell types; and 2) when two different unipotent progenitor cell types converge to a single progeny cell type. In both cases, shared clonality or lineage relationships would not be identified by prospective lineage analysis. In contrast, in our SISBAR approach and lineage analysis, scenarios 1 and 2 would be identified as divergent or convergent trajectories. Therefore, the SISBAR approach provides a more comprehensive perspective on the complex lineage trajectories underlying cell differentiation.
[0166] The cellular basis of divergent trajectories is that specific types of stem or progenitor cells are multipotent. In this study, we show that the clonal output of individual multipotent progenitors with similar states is variable, and each clonal output consists of a variety of combinations of one cell type or different cell types. With the support of the SISBAR strategy, the applicants demonstrated that individual progenitors of the same type (in the same cluster of UMAP) with similar transcription can have different fate outcomes. Importantly, the fates of individual clones collectively represent the cell type composition at the population level, which is found in almost all transcriptionally defined multipotent progenitor cell types in the study. These results suggest that the clear divergent trajectories or tree-like lineage structures pursued in the field of development may only occur at the population level, which is actually composed of variable individual trajectories.
[0167] In one aspect, the present application provides a lineage tracing method, which includes harvesting cells and their sister cells at different stages of differentiation, dividing the cells and their sister cells into at least two parts, one part for genetic analysis, and the other part for genetic analysis after further differentiation, and finally analyzing the cross-stage lineage.
[0168] On the other hand, the present application provides a method for predicting cell differentiation fate, which includes harvesting cells and their sister cells at different differentiation stages, dividing the cells and their sister cells into at least two parts, one part for genetic analysis, and the other part for genetic analysis after continuing to differentiate, and finally analyzing the cross-stage lineage.
[0169] On the other hand, the present application provides a method for identifying the origin of cells, which includes harvesting cells and their sister cells at different stages of differentiation, dividing the cells and their sister cells into at least two parts, one part is used for genetic analysis, and the other part continues to differentiate and then undergoes genetic analysis, and finally analyzes the cross-stage lineage.
[0170] In this application, cells at different stages of differentiation can undergo cell division and produce sister cells.
[0171] In the present application, cells at different stages of differentiation are derived from the same type of progenitor cells. In the present application, cells at different stages of differentiation are derived from different types of progenitor cells. In the present application, cells at different stages can be a heterogeneous cell population.
[0172] In the present application, the cells may be derived from pluripotent stem cells. For example, the pluripotent stem cells are human pluripotent stem cells.
[0173] In the present application, pluripotent stem cells include embryonic stem cells (ESC) and / or induced pluripotent stem cells (iPSC). As used herein, the term "embryonic stem cell (ESC)" generally refers to cells with unlimited proliferation, self-renewal and multidirectional differentiation. Embryonic stem cells are stem cells obtained from the undifferentiated inner cell mass of the blastocyst (early embryonic stage), and their source and preparation method are not restricted, and no matter in vitro or in vivo, they can be induced to differentiate into almost all cell types of the human body. For example, the cell type mentioned can be a nerve cell etc. As used herein, the term "induced pluripotent stem cell (iPSC)" generally refers to a class of pluripotent stem cells artificially prepared by non-pluripotent somatic cells. Induced pluripotent stem cells can be obtained by introducing specific transcription factors that reprogram terminally differentiated somatic cells. For example, the terminally differentiated somatic cells can be fibroblasts, hematopoietic stem cells, myoblasts, neurons, epidermal cells etc.
[0174] In the present application, progenitor cells may include neural progenitor cells.
[0175] In the present application, the lineage tracing method may include introducing a marker into a differentiation initiating cell. All markers that can be used to mark cells are included in the present application. For example, the marker can be a barcode. For example, the barcode is a viral barcode. For example, the viral barcode is a lentiviral vector. For example, the viral barcode is a retroviral vector. In certain embodiments, the viral barcode includes inserting a 32 bp semi-random barcode into a retroviral vector. In certain embodiments, the viral barcode contains 16 repeats of an S (G or C) / W (A or T) alternating nucleic acid sequence.
[0176] In this application, genetic analysis can include the recovery of viral barcodes. Cells with the same viral barcode are defined as a single clone.
[0177] In the present application, the marker may include the Cre-LoxP system.
[0178] In this application, markers may include insertions and / or deletions of traceable elements introduced by the CRISPR-Cas9 system.
[0179] In the present application, the lineage tracing method may include analyzing a lineage that has undergone one or more stages of differentiation. In the present application, the lineage tracing method may include analyzing a lineage that has undergone two or more stages of differentiation. In the present application, the lineage tracing method may include analyzing a lineage that has undergone three or more stages of differentiation.
[0180] In the present application, genetic analysis may include sequencing a cell or its sister cells. For example, sequencing may include transcriptome sequencing. For example, sequencing may include single-cell RNA sequencing (scRNA-seq). After sequencing, a data analysis process may be performed.
[0181] In the present application, all data analysis methods known to a person skilled in the art can be used.
[0182] For example, gene expression-related cell clusters can be estimated by calculating the Pearson correlation coefficient of the average gene expression values between clusters; the gene expression correlation between two sister cells A (Xa, Ya) and B (Xb, Yb) can be estimated by the Pearson correlation coefficient of their gene expression and their Euclidean distance on the UMAP graph.
[0183] For example, single-cell differential gene expression analysis can be performed using the FindMarkers function in Seurat, which performs fold-change Wilcox test p-value calculations and FDR (false discovery rate) corrected t-tests.
[0184] For example, genetic analysis can include a potential perspective analysis, which involves examining the potential of a given progenitor cell type at an early stage by calculating the distribution of clonal cells at a later stage and comparing it to a random distribution of permutations.
[0185] For example, genetic analysis can include origin-perspective analysis, which involves examining the origin of a given cell type at a later stage by calculating the distribution of clonal cells at an earlier stage and comparing it to a random distribution of permutations.
[0186] For example, the distribution of clonal cell types can be visualized by heatmap clustering based on the Pearson correlation coefficient of the distribution.
[0187] For example, the lineage relationships of different cell clusters within a developmental stage can be determined by calculating the Jaccard similarity coefficient score between each pair of cell clusters (the number of shared clones divided by all clones of a pair of clusters), P value (given the null hypothesis that cells in a clone are randomly distributed among all clusters), lineage coupling correlation analysis, and / or clonal coupling (observed / expected).
[0188] In the present application, the lineage tracing method may include the following steps: (a) introducing viral barcodes into differentiation-initiating cells; (b) culturing the differentiation-initiating cells and harvesting cells at the next differentiation stage and their sister cells; (c) dividing the cells at the next differentiation stage and their sister cells into two parts, one part for viral barcode recovery and scRNA-seq, and the other part for continued differentiation; (d) if differentiation includes two or more stages, repeating (c) until the differentiation process is completed; (e) after sequencing, constructing a lineage tree using potential analysis and origin analysis.
[0189] In another aspect, the present application provides use of the method of the present application in lineage tracing of pluripotent stem cell differentiation products.
[0190] In another aspect, the present application provides use of the method of the present application in lineage tracing of a heterogeneous cell population.
[0191] On the other hand, the applicants developed single-cell division barcodes (SISBARs) by combining viral-mediated cell barcoding, scRNA-seq, and clonal division strategies to longitudinally track the clonal fate of individual progenitor cells while capturing the transcriptome of the founding progenitor cells in an in vitro differentiation model. Furthermore, the applicants developed "potential perspective" and "origin perspective" analyses to statistically identify the fate potential or origin of transcriptome-defined cell types, respectively.
[0192] Here, we develop the SISBAR method to characterize the comprehensive cross-stage clonal lineage landscape in human neural differentiation, which may provide new insights into the molecular programs of neuronal specification and improve cell-based replacement therapies in regenerative medicine.
[0193] Example
[0194] The following examples are set forth in order to provide one of ordinary skill in the art with a complete disclosure and description of how to make and use the invention, and are not intended to limit the scope of what the inventors regard as their invention, nor are they intended to represent that the following experiments are all or the only experiments performed. Efforts have been made to ensure correctness with respect to the numbers used (e.g., amounts, temperatures, etc.), but some experimental errors and deviations should be considered. Unless otherwise indicated, parts are parts by weight, molecular weight is average molecular weight, temperature is in degrees Celsius, and pressure is at or near atmospheric pressure. Standard abbreviations may be used, such as bp, base pairs; kb, kilobase; pl, picoliter; s or sec, seconds; min, minutes; h or hr, hours; aa, amino acid; nt, nucleotide; im, intramuscular; ip, intraperitoneal; sc, subcutaneous; etc.
[0195] Materials and methods
[0196] experimental animals
[0197] SCID Beige mice were purchased from Vital River Laboratory. All animals used in this study were group-housed on a 12:12-hour light / dark cycle with free access to water and food. All animal experiments were performed in accordance with protocols approved by the Animal Care and Use Committee of the Institute of Neuroscience, Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences.
[0198] human embryonic stem cells
[0199] hESCs (H9 hESCs and reporter H9 hESCs, less than 50 passages) were grown on irradiated mouse embryonic fibroblast (MEF) feeder layers in ESC medium consisting of Dulbecco's Modified Eagle's Medium / F-12 (DMEM / F12, Gibco), Knockout Serum Replacement (KOSR, Gibco), 0.5X Glutamax (Gibco), 1X Nonessential Amino Acids (NEAA, Life Technologies), 0.1 mM β-mercaptoethanol (Sigma), and 8 ng / ml recombinant human FGF-basic (bFGF, R&D Systems). The medium was changed daily, and hESCs were passaged weekly using Dispase II (Gibco, 1 mg / ml) onto new 6-well plates coated with fresh MEFs (prepared at least 24 hours before passage).
[0200] hESC-based human ventral midbrain-hindbrain differentiation
[0201] The hESC neural differentiation protocol to mimic human ventral midbrain-hindbrain embryonic development was performed according to previously described methods (Chen, Y., Xiong, M., Dong, Y., Haberman, A., Cao, J., Liu, H., Zhou, W., and Zhang, SC (2016). Chemical Control of Grafted Human PSC-Derived Neuronsin a Mouse Model of Parkinson's Disease. Cell Stem Cell 18, 817–826.; Xiong, M., Tao, Y., Gao, Q., Feng, B., Yan, W., Zhou, Y., Kotsonis, TA, Yuan, T., You, Z., Wu, Z. et al. (2021). Human Stem Cell-Derived Neurons Repair Circuitsand Restore Neural Function. Cell Stem Cell 28, 112-126.e6; Xu, P., Xiong, M., and Chen, Y. (2022a). Human Midbrain dopaminergic neuronal differentiation markers predict cell therapy outcome in a Parkinson's disease model.) was modified. One day after passaging, the hESC culture medium was changed to neural induction medium (NIM) containing DMEM / F12, 1× N2 supplement (Gibco), and 1× NEAA, and supplemented with SHH (C25II, R&D Systems, 500 ng / ml), CHIR99021 (Tocris, 0.4 μM), DMH-1 (Tocris, 2 μM), and SB431542 (Stemgent, 2 μM) and cultured for 8 days (days 1 to 9). On day 9, single colonies were gently blown off and transferred to 6-well culture dishes coated with NIM containing fresh MEFs supplemented with SHH (100 ng / ml), SAG (Millipore, 1 μM), and CHIR99021 (0.4 μM) and maintained for 4 days until stage I (day 13).In stage I, individual colonies were gently dislodged and transferred to non-adherent 25 ml culture dishes and cultured in suspension medium containing NIM supplemented with SHH (20 ng / ml), SAG (Millipore, 0.5 μM), and FGF8b (PeproTech, 100 ng / ml) for 8 days until stage II (day 21). From stage II to stage III (day 30), neurospheres were allowed to continue differentiation and proliferation in suspension medium containing NIM supplemented with SHH (20 ng / ml) and FGF8b (20 ng / ml). In stage III, neurospheres were dissociated using Accutase (Innovative Cell Technologies) at 37°C for 6 minutes and then replated onto 24-well culture dishes coated with Matrigel (BD Biosciences). From stage III to stage IV (day 45), differentiated cells were fed with neural differentiation medium (NDM), which contained neurobasal medium, 1× N2 supplement (Gibco), and 1× B27 (Life Technologies), supplemented with brain-derived neurotrophic factor (BDNF, Peprotech, 10 ng / ml), glial-derived neurotrophic factor (GDNF, Peprotech, 10 ng / ml), transforming growth factor β3 (TGFβ3, R&D Systems, 1 ng / ml), ascorbic acid (AA, Sigma-Aldrich, 200 μM), cAMP (Sigma-Aldrich, 1 μM), and compound E (Calbiochem, 1 μM). In addition, to improve cell survival during passaging, vitamin A-free B-27 supplement (Life Technologies) and a Rho-kinase (ROCK) inhibitor (Tocris, 0.5 mM) were added to the culture medium.
[0202] Donor plasmid construction
[0203] Human codon-optimized Streptococcus pyogenes wild-type Cas9 (Cas9-2A-GFP) and Cas9 nickase (Cas9D10A-2A-GFP) were obtained from Addgene (plasmid #44719, plasmid #44720). By replacing the loxP sequence of PL552 (#68407) with the vloxP sequence, a PL752 donor plasmid vector containing a vloxP-flanked PGK-puromycin expression cassette was constructed. To generate the FGF8-iCre donor plasmid, a DNA fragment with a left or right homology arm was amplified from H9 hESC genomic DNA just upstream or downstream of the FGF8 gene stop codon (certain mutations were present on the right homology arm to avoid self-targeting of the sgRNA on the donor plasmid) by PCR. A DNA fragment of the P2A-iCre gene was amplified from pDIRE (plasmid #80945) by PCR. These three fragments were then cloned into the multiple cloning site of the plasmid PL752.
[0204] Generation of hESC reporter lines
[0205] H9 hESCs were pretreated with ROCK inhibitor (0.5 mM) for 24 hours and then digested with TrypLE™ Express enzyme at 37°C for 6 minutes to dissociate into single cells. Single cells were then resuspended in 500 μl of electroporation buffer (5 mM MgCl2, 5 mM KCl, 102.94 mM Na2HPO4, 47.06 mM NaH2PO4, and 15 mM HEPES, pH = 7.2), supplemented with Cas9 plasmid, donor plasmid, and sgRNA, and electroporated using the Gene Pulser Xcell System (Bio-Rad) at 250 V and 500 mF in 4 mm cuvettes (Phenix Research Products). Cells were then immediately plated in 6-well culture dishes coated with fresh MEFs and cultured in MEF-conditioned ESC medium containing ROCK inhibitor (0.5 mM for 24 hours). After three days, cells were selected for two weeks by treatment with G418 (50-100 μg / ml) or puromycin (0.5 μg / ml). Subsequently, cells were treated with a ROCK inhibitor for 24 hours and then individually removed for further proliferation. Clones were then genotyped to check for transgene integration.
[0206] To generate FGF8-iCre, a cassette containing a codon-optimized Cre recombinase sequence linked to a P2A peptide and vloxp-flanked PGK-Pur (PGK promoter-driven puromycin followed by a polyA signal) was inserted into the endogenous FGF8 gene of H9 hESCs just upstream of the stop codon by Cas9 nickase.
[0207] sgRNA targeting the first 100 bp of the right homology arm was designed on the website https: / / benchling.com / . Genotyping primers for checking transgene integration were designed on the website https: / / www.bioinformatics.nl / cgi-bin / primer3plus / primer3plus.cgi.
[0208] Cell sorting and transplantation
[0209] Neurospheres were gently dissociated into single cells using Accutase (Innovative Cell Technologies) at 37°C for 8 minutes and filtered through a 40 μm cell strainer. The isolated cells were resuspended in neural induction medium (NIM) supplemented with 1× penicillin-streptomycin. Fluorescence-activated cell sorting (FACS) was then performed using a MA900 multi-application cell sorter (Sony, Japan) and analyzed using cell sorting software. The sorting strategy is shown in Supplementary Figure 10. tdTomato was determined based on the fluorescence intensity in the tdTomato channel using 568 nm laser excitation. +And tdTomato-fraction. The sorted and unsorted (control group) cells were seeded into 96-well conical (V) bottom plates (Thermo Scientific) coated with Lipidure-CM5206 (NOFCORPORATION) at a density of 5,000 cells / well. B-27 supplement (Life Technologies) without vitamin A and Rho-kinase (ROCK) inhibitor (Tocris, 0.5 mM) were added to the culture medium to improve the survival rate of cells during passage. Two days later, a small portion of the re-aggregated neurospheres were separated into single cells and the number of cells per neurosphere was estimated. Approximately 100,000 cells were implanted in each transplant. The neural stem cell transplantation procedure was based on a previous study (Chen, Y., Xiong, M., Dong, Y., Haberman, A., Cao, J., Liu, H., Zhou, W., and Zhang, SC (2016). Chemical Control of Grafted Human PSC-Derived Neurons in a Mouse Model of Parkinson's Disease. Cell Stem Cell 18, 817–826.). Briefly, PD model mice were randomly divided into groups and transplanted with marker-sorted or unsorted progenitor cells (resuspended in 1 mL of aCSF containing 0.5 mM Rock inhibitor, B27, and 20 ng / ml BDNF) at the following coordinates relative to the bregma: AP +0.6 mm, ML -1.8 mm, DV -3.2 mm.
[0210] Immunohistochemistry
[0211] For in vitro cultured cells, neurospheres were collected at stage IV (day 45) and fixed with 4% paraformaldehyde (PFA) for 15 minutes at room temperature. Neurospheres were then washed three times with Dulbecco's phosphate-buffered saline (DPBS) for 10 minutes each wash and incubated in 30% sucrose at 4°C until they sank to the bottom of the centrifuge tube. Dehydrated neurospheres were then sliced into thin sections (16 μm) using a cryostat (Leica) and mounted on SuperFrost Plus slides for immunohistochemistry. For mouse brains, animals were sacrificed and perfused with saline dissolved in diethylpyrocarbonate (DEPC)-treated water, followed by treatment with 4% ice-cold PFA. Brain tissue was removed and immersed in 4% PFA for 4 hours, followed by transfer to 20% and then 30% sucrose until it sank. Mouse brains (coronal sections) were sliced into thin sections (30 μm) using a cryostat (Leica) and stored in cryopreservation solution at -20°C. For immunohistochemistry, neurosphere sections were slide-mounted for immunocytochemistry, while brain sections were processed for free-floating immunohistochemistry. Sections were blocked with 0.3% Triton X-100 and 10% donkey serum at 37°C for 60 minutes. Sections were then incubated with primary antibodies diluted in 0.2% Triton X-100 and 5% donkey serum for 2 hours at room temperature and then transferred to 4°C overnight. Sections were then incubated with corresponding secondary antibodies conjugated to Alexa fluorescent dyes for 60 minutes at room temperature and mounted using Fluoromount-G (Southern Biotech).
[0212] RNA in situ hybridization
[0213] RNA in situ hybridization was performed using the RNAscope Multiplex Fluorescence Detection Kit v2 (Advanced Cell Diagnostics) essentially according to the manufacturer's instructions. Animals were sacrificed and perfused with saline dissolved in DEPC-treated water, followed by treatment with 4% ice-cold paraformaldehyde (PFA). Brain tissue was removed and immersed in 4% PFA for 4 hours, then transferred to 20% and then 30% sucrose until submerged. Brains were sectioned coronally into thin sections (20 μm) using a freezing microtome (Leica) and mounted on SuperFrost Plus slides. After dehydration, slides were either stained with RNAscope or stored at −80°C. For RNAscope staining, slides were equilibrated to room temperature and held for approximately 10 minutes. Slides were then pretreated with hydrogen peroxide for 10–12 minutes and subjected to target retrieval. Slides were then incubated in proteinase III (1:15 dilution in DEPC-PBS) at 40°C for 30 minutes. The probes used were all human-specific (COL1A1-C1, SLC13A6-C1) and hybridization was performed at 40°C for 2 hours in strict accordance with the manufacturer's instructions. Images were captured using an OLYMPUS FV3000 microscope at 20x magnification.
[0214] Imaging and cell quantification
[0215] To quantify the number of TH, COL1A1, and vGlut2 expressing cells in the total cells (Hoechst labeled or human nuclear marker), at least five images randomly selected from the coverslips were counted using ImageJ software. The data were repeated at least three times and expressed as the mean ± SEM. + To determine the percentage of cells, we first counted the number of Hoechst-labeled nuclei in neurospheres or human nuclei-positive (hN) cells in transplants using ImageJ from fluorescent images acquired at 20× magnification (Olympus FV3000 microscope). + ) cells, we then counted TH surrounded by cytoplasmic TH expression in cells with nuclei labeled with Hoechst or human nuclear antibodies. + The number of cells.
[0216] Construction of a high-complexity barcoded retroviral library
[0217] The barcoded viral plasmid library construction process was primarily adapted and modified from a previous study (Bhang et al., 2015). For the retroviral barcoded library, an 84-bp single-stranded template was synthesized containing a semi-random 32-bp barcode sequence (AGCGGGTTTAAACGGGCCCTGGATCCSWSWSWSWSWSWSWSWSWSWSWSWSWSWSWSWSWSWSWGAATTCAGTCAGTCACGCGTTTATTT, SEQ ID NO: 1) and a flanking primer pair for barcode amplification (forward, GGCTGGCAACTAGAAGGCACAGTCGAGGCTGATCAGCGGGTTTAAACGGGCCCTGGATCC, SEQ ID NO: 2; reverse, CCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAATAAACGCGTGACTGACTGAATT, SEQ ID NO: 3). S represents a degenerate base sequence of G and C, and W represents a degenerate base sequence of A and T. Using the above template and primer pair, double-stranded DNA fragments were generated by PCR reaction (Q5 High-Fidelity DNA Polymerase (NEB, M0492L)) with the following program: 1, 98°C, 30 seconds; 2, 98°C, 10 seconds; 3, 60°C, 10 seconds; 4, 72°C, 15 seconds; 5, repeat steps 2-4 15 times; 6, 72°C, 3 minutes; 7, 4°C, hold. After PCR amplification, the PCR fragments were recovered using a gel extraction kit (QIAGEN). The viral vector was digested with EcoRI / BamHI and then recovered using a gel recovery kit. Before Gibson ligation, the recovered PCR fragments and digested viral vector were purified by isopropanol precipitation to remove impurities. Next, 300 ng of the purified digested viral vector and 50 ng of the purified PCR fragment were mixed with 0.5× NEbuilder (New England BioLabs) in a 20 μl reaction volume and incubated at 50°C for 1 hour. The Gibson ligation product was then further purified by isopropanol precipitation and transformed into Escherichia coli DH5α electrocompetent cells (Takara). Finally, the transformed bacterial pool was inoculated into 2000 ml of LB medium containing 50 μg / ml ampicillin (Sigma-Aldrich) and cultured overnight at 20°C with an OD600 of no more than 0.3 to ensure balanced bacterial amplification, followed by mega-prep extraction of plasmid DNA (Invitrogen).
[0218] To measure the diversity of the barcoded viral library, we designed primer pairs with Illumina adapters to amplify random barcode sequences in the retroviral barcode library plasmid pool (forward, AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTTTCCGATCTTATTCGGAGGACGACCCTATTTG, SEQ ID NO: 4; reverse, CAAGCAGAAGACGGCATACGAGATAGGATTCGGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGGATCACTCTCGGCATGGACG, SEQ ID NO: 5) using the following PCR program: 1, 98°C, 3 minutes; 2, 98°C, 10 seconds; 3, 60°C, 10 seconds; 4, 72°C, 25 seconds; 5, repeat steps 2-4 15 times; 6, 72°C, 3 minutes; 7, 4°C, hold. PCR products were purified by 0.85x left size selection using Ampure XP beads (Beckman Coulter, catalog number A63880) and then sequenced on an Illumina HiSeq (Novogene).
[0219] Barcoded retrovirus production and blacklist generation
[0220] Barcoded retroviruses were generated by transfecting HEK293T cells with a pooled DNA plasmid library and packaging plasmids gag / pol (Addgene plasmid 14887) and VSV.G (Addgene plasmid 14888). Viruses were collected 48 hours after transfection and concentrated by ultracentrifugation at 27,000 rpm (HITACHI P28S rotor). The titer of the barcoded virus was determined by serial dilution on HEK293T cells. To generate a blacklist containing high-copy viral barcodes, HEK293T cells were infected with the aliquoted virus and genomic DNA was then extracted using the Blood Mini Kit (QIAGEN). Using genomic DNA as a template, the integrated viral barcode sequence was amplified using the same primer pair as used to measure the diversity of the barcoded viral library. PCR fragments were then purified using Ampure XP beads by 0.85× left size selection and sequenced using Illumina HiSeq (Novogene). The top 5% of viral barcodes as measured by read counts were added to the blacklist of the barcode virus library.
[0221] Cellular barcoding, scRNA-seq, and viral barcode recovery
[0222] For each cross-stage lineage tracing experiment, the number of neurospheres used in each experiment ranged from 50 to 100. After labeling with a low-efficiency (approximately 10%) viral barcode library, the neurospheres were expanded in culture for four days, dissociated into single cells, and then split into two parts. To estimate the number of cell divisions between viral barcoding and cell division, we first dissociated 50 neurospheres into single cells and evenly divided them in half on day 14. Half of the cells were counted after 1 day (n cells), and the other half of the cells were cultured for 5 days before cell counting (m cells). Therefore, we estimated that the average cell division rate during viral barcoding and cell division was 3-4 times (m / n). The cell capture rate depends on the following factors: (1) the cell loss rate during FACS sorting (EFACS loss), which is approximately 20%; (2) the cell loss rate during centrifugation and resuspension of sorted cells (Ecentrifugation loss), which is approximately 10%; (3) the 10X genomic capture efficiency (E10X capture) (usually approximately 65%); and (4) the barcode detection efficiency (E detection bar), which is approximately 65%. The final efficiency of cell capture for each clone is proportional to the following factors:
[0223] (1-EFACS loss) * (1-E centrifugation loss) * E10X capture * E strip = ~30%
[0224] For scRNA-seq, neurospheres were digested with TrypLE enzyme at 37°C for 10 minutes and dissociated into single cells. Digested cells were filtered through a 35 μm cell strainer and sorted by FACS to collect EGFP-positive cells, which were then concentrated by centrifugation at 400 g for 5 minutes. All sorted EGFP-positive cells were used to generate single cells using the Chromium Single Cell 3' Kit (v2 and v3) following the manufacturer's instructions (10X Genomics). Single-cell libraries were sequenced using an Illumina HiSeq. After mRNA reverse transcription and first-strand cDNA amplification, 50 ng of unfragmented cDNA library was used as template with the primer pair: AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTC (SEQ ID NO: 6) and CAAGCAGAAGACGGCATACGAGATAGGATTCGGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGGCATGGACGAGCTGTACAAATAA (SEQ ID NO: 7) to specifically amplify cellular and viral barcodes using the following program: 1, 98°C, 3 minutes; 2, 98°C, 10 seconds; 3, 60°C, 15 seconds; 4, 72°C, 1 minute; 5, repeat steps 2-4 20 times; 6, 72°C, 3 minutes; 7, 4°C, hold. PCR products were then purified using Ampure XP beads with 0.8-0.6× duplex size selection and sequenced using an Illumina HiSeq.
[0225] scRNA-seq data processing
[0226] For data processing, scRNA-seq data were preprocessed using the Cellranger v.3.1 pipeline (10X Genomics) with GRCh38 as the reference genome to generate digital gene expression (DGE) matrices for downstream analysis. The DGE matrices were then analyzed using Seurat (v3) (Satija, R., Farrell, JA, Gennert, D., Schier, AF, & Regev, A. (2015). Spatial reconstruction of single-cell gene expression data. Nat. Biotechnol. 33, 495–502.). First, cells with fewer than 1,000 detected genes and cells with more than 30,000 total unique molecular identifiers (UMIs) were considered low-quality or doublets and filtered out. Second, dying or stressed cells were filtered out if >15% of the UMI counts were attributed to mitochondrial genes. DEGs were then normalized using the 'NormalizeData' function. Highly variable genes were calculated using 'FindVariableFeatures', and genes associated with TOP2A expression were excluded (Pearson correlation coefficient greater than 0.15). Afterwards, we used the 'CellCycleScoring' function to generate cell cycle scores and assign a cell cycle difference score to each cell. The 'ScaleData' function was then used to regress the Seurat object using UMI counts and cell cycle difference scores as parameters. Next, the first 100 principal components (PCs) were calculated and inspected using the 'RunPCA' function to determine the number of input components used to cluster the cells. Finally, the cells were clustered using the 'FindClusters' function of the Seurat (v3) package and visualized using UMAP. Clusters that were not reproducible between biological replicates were excluded from further analysis.
[0227] Single cell virus barcoding identification
[0228] First, we extracted cellular barcodes and UMI sequences using UMItools (V 0.5.5) to construct a data frame containing cellular barcodes, UMIs, and read sequences. Next, to identify viral barcodes, we used the R package Biostrings (V2.58) to locate upstream and downstream anchor sequences (sequences flanking the 32-bp viral barcode) and extracted the sequence between them as the viral barcode sequence. We then combined the cellular barcode, UMI, and viral barcode as the label for each read. We generated a distribution of label counts and observed a bimodal distribution (Figure S2), suggesting that labels with fewer reads may be due to PCR or sequencing errors. Therefore, we used labels with a read count of at least 5 as a cutoff based on the distribution to filter out those problematic labels. In addition, we retained viral barcodes with at least 2 unique UMIs in a cell for downstream analysis. To correct for errors caused by PCR and NGS (next-generation sequencing), we used Starcode (Zorita, E., Cuscó, P., and Filion, GJ (2015). Starcode: Sequence clusteringbased on all-pairs search. Bioinformatics 31, 1913–1919.) to calculate the Levenshtein distance between each pair of tags and merge similar tags (with a string distance of less than 6 nucleotides) into the same cluster. These similar tags were then re-examined and re-separated into distinct tags whose cellular or viral barcodes had a string distance of less than 4 nucleotides. To avoid situations where two or more starting cells were labeled with the same viral barcode, we performed blacklist filtering to remove tags with high-copy viral barcodes. After these steps, the viral barcode of each cell was identified, and cells with the same viral barcode were grouped into the same clone for downstream lineage analysis.
[0229] Clonal relationships between cell clusters within a stage
[0230] To determine the lineage relationships of different cell clusters within a developmental stage, we calculated the Jaccard similarity coefficient score (the number of shared clones divided by the total number of clones in a pair of clusters (e.g., clusters A and B)) (1) and the P value (given the assumption that cells in a clone are randomly distributed across all clusters) between each pair of cell clusters. To calculate the P value, we performed a permutation test by shuffling the cells and calculating the Jaccard similarity coefficient 500 times to obtain the theoretical distribution of the Jaccard scores for each pair of clusters. Next, we calculated the Z score based on the actual Jaccard similarity coefficient of each pair and its permutation distribution (2), and further converted it into a P value using the following formula (3).
[0231]
[0232] Clonal relationships between cell clusters across two stages
[0233] To estimate the clonal relationships between different pairs of cell clusters in two developmental stages, we considered clones that coexisted in both stages and calculated the clonal relationships from two perspectives. In the potential perspective analysis, we estimated the differentiation potential of a given progenitor cluster (e.g., cluster A). We first calculated the Jaccard score (J) between progenitor cluster A and a later stage cluster (e.g., cluster B) by dividing the number of intersecting clones by the number of cross-stage clones associated with progenitor cluster A (4). 潜力 Next, we randomly permuted the cell identities of clonally associated cells of late-stage progenitor cluster A 500 times and compared the random jaccard score distributions with the actual jaccard scores between progenitor cluster A and late-stage clusters to calculate Z scores (2), which were further converted into P values (3).
[0234] J 潜力 (A, b) = |CloneA ∩ Cloneb| / CloneA (4)
[0235] In the origin-perspective analysis, we estimated the differentiation origin of a given differentiation cluster (e.g., cluster A) by first calculating the Jaccard score (J) between differentiation cluster A and an earlier stage cluster (e.g., cluster B) by dividing the number of intersecting clones by the number of cross-stage clones associated with differentiation cluster A (5). 起源 Next, we randomly permuted the cell identities of clone-associated cells in early-stage differentiation cluster a 500 times and compared the random Jaccard score distributions with the actual Jaccard scores between differentiation cluster a and the early-stage clusters to calculate Z scores (2), which were further converted into P values (3).
[0236] J 起源 (a, B) = |clone a ∩ clone B| / clone a (5)
[0237] Whether from the perspective of potential or origin, the pedigree relationship of cross-stage cluster pairs with a P value <= 0.05 was considered significant.
[0238] Gene expression correlation analysis of sister cells and cell clusters
[0239] The gene expression correlation between cell clusters was estimated by calculating the Pearson correlation coefficient of the average gene expression values between clusters (6). The gene expression correlation between two cells A (Xa, Ya) and B (Xb, Yb) was estimated by the Pearson correlation coefficient of their gene expression (7) and their Euclidean distance on the UMAP plot (8).
[0240]
[0241] Single-cell differentially expressed gene analysis
[0242] Single-cell differential gene expression analysis was performed using the FindMarkers function in Seurat (V3), which performs fold-change Wilcox test p-value calculations and FDR (false discovery rate) corrected t-tests. To enhance the statistical significance of molecular features associated with differential fate biases of individual progenitors within the same transcriptome-defined progenitor cell type, we applied a filter of clonal cell number ≥ 3 (at later stages) as the minimum cell number.
[0243] The number of viral barcodes expected to overlap between experiments
[0244] Before calculating the number of overlapping barcodes between samples, frequently occurring (blacklisted) barcodes were removed. Next, we calculated the expected number of overlapping barcodes between samples as follows: We randomly sampled three barcode sets with replacement from the barcode library pool (the number of barcodes in each set was equal to the number of barcodes in Pools #1, #2, and #3) and calculated the number of overlapping barcodes between these three sample sets. After repeating this process 500 times, we took the median number of overlapping barcodes between each sample pair as the expected number of overlapping barcodes between samples.
[0245] Statistical test of monoclonal fate clones
[0246] For individual clones with a single clonal fate, the null hypothesis is that each clone has at least two different fates at later stages. To calculate the p-value, we performed a permutation test by shuffling the clonal cells of the late-stage progenitors and counting the clonal fate number 500 times to obtain a random distribution of clonal fate numbers. Next, we calculated the Z score of each single-fate progenitor based on the actual clonal fate number (which is 1) and its permutation distribution (9), and further converted it into a p-value using formula (10).
[0247]
[0248] Example 1 SISBAR method
[0249] The SISBAR method, which integrates virus-mediated cell barcoding, clonal division, and scRNA-seq, was developed to study the lineage relationship between cells at different developmental stages in hPSC-based neural differentiation. The working principle of SISBAR is to genetically label each progenitor cell with a unique viral barcode in the early differentiation stage, followed by limited cell division, and finally randomly divide the cells into two parts. Half of the cells are immediately subjected to scRNA-seq and viral barcode recovery, and the remaining cells continue to differentiate to the later stage, and then scRNA-seq and viral barcode recovery ( Figure 1A ). SISBAR thus provides both the transcriptome and viral barcode sequences of individual cells to infer their identity and lineage information. Barcoded sister cells analyzed at early stages can serve as surrogates for the ancestors of clonally related descendants analyzed at later stages, thereby approximating individual cross-stage lineage trajectories based on single cross-stage clones. Each clone is defined as a group of cells sharing the same viral barcode. When scaled to many cross-stage clones, statistically significant lineage relationships between various early and late cell types are revealed ( Figure 1A , right side).
[0250] A classic retroviral vector was modified by inserting a 32 bp semirandom barcode into the 3' untranslated region (3' UTR) of the EGFP coding sequence, driven by the universal EF1α promoter ( Figure 1A and 6A ).
[0251] This allows for high-level expression of EGFP and viral barcodes without affecting cell differentiation behavior when labeled at low multiplicity of infection (MOI approximately 0.1). Figures 6B-6E The viral barcode was designed to contain 16 repeats of alternating S (G or C) / W (A or T) nucleic acid sequences with balanced GC content to ensure uniform efficiency of PCR amplification and error correction in downstream analysis ( Figure 6A Deep sequencing results showed that the viral barcode plasmid library consisted of more than three million unique barcodes with relatively uniform count distribution, sufficient to label more than 30,000 cells, with a barcode collision rate of less than 1% ( Figure 6F We further validated the high complexity of the barcoded viral library by three independent labeling experiments using the same viral pool, showing that the observed number of overlapping viral barcodes between samples was very close to the expected number simulated using three million unique barcodes ( Figure 6GFurther analysis of the barcode composition revealed that the barcodes were distinguishable from each other and that a given barcode had a substantial minimum distance from other barcodes in the library pool ( Figure 6H The high complexity and high diversity of the library reduce the possibility that different starting cells will absorb the same barcode or highly similar barcodes. scRNA-seq was used to capture barcoded EGFP transcripts in parallel with endogenous transcripts ( Figure 6J and Figure 7).
[0252] The developed SISBAR method can track the clonal history of different developmental stages at the single-cell transcriptome level.
[0253] Example 2 SISBAR Elucidates Human Ventral Midbrain-Hindbrain Differentiation
[0254] SISBAR was applied to an hPSC-based in vitro neural differentiation model that recapitulated human embryonic ventral midbrain-hindbrain development, including four differentiation stages from early patterned neural progenitors to terminal mature cells ( Figure 1B Three consecutive but independent labeling experiments were designed to reveal the lineage relationships between stages I-II (experiment #1), II-III (experiment #2), and III-IV (experiment #3) ( Figure 1C A total of 173,564 high-quality single-cell transcriptomes were captured from three independent populations, each of which contained three consecutive but independent cross-stage scRNA-seq datasets and visualized using Uniform Manifold Approximation and Projection (UMAP). Figure 1D The results showed that the proportions of the main cell types in the four differentiation stages were reproducible in three biological replicates ( Figure 8A ). First, cell clusters were annotated based on their expression patterns. Based on the expression of brain region markers, progenitor cells were divided into four major cell types ( Figure 8B ), including ventral midbrain neural progenitor cells (NP vMB , OTX2+ / EN1+ / FOXA2+), midbrain-hindbrain boundary (MHB) neural progenitor cells (NP MHB , FGF8+ / OTX2- / EN1+), anterior hindbrain neural progenitor cells (NP aHB , FGF8- / OTX2- / EN1+) and posterior hindbrain neural progenitor cells (NP pHB , EN1- / HOXA2+) ( Figure 1D and Figure 8C ). NP pHB The high expression of HOXA2, HOXB2, and HOXB1 indicates that the rhombomere (r) has 2 to 4 segments, which are the same as those in the embryonic hindbrain. MHBSpecifically expressed fibroblast growth factor 8 (FGF8), a marker gene for the MHB (also known as the isthmus organizer), which plays a key role in the development of the midbrain and hindbrain in vivo (Crossley et al., 1996). We observed these NPs in both stage I and II. MHB , but not observed in stages III and IV. Neurons were divided into midbrain dopaminergic neurons (mDA, NR4A2+ / PITX3+ / EN1+), midbrain glutamatergic neurons (mGlut, PITX3- / SLC17A6+ / EN1+), GABAergic neurons (GABA, GAD2+ / SLC32A1+), hindbrain glutamatergic neurons (hGlut, SLC17A6+ / HOXA2+ / EN1-), serotonergic neurons (Sero, FEV+ / GATA2+) and motor neurons (Motor, ISL1+ / PHOX2A+) ( Figure 1D and Figure 8C Among terminally differentiated cells (stage IV), a non-neuronal cell type, COL1A1+ cells, was also identified that was transcriptionally and morphologically distinct from neuronal cells ( Figure 1D and Figure 8D COL1A1+ cells were also found in transplants from Parkinson's disease (PD) mice that received hESC-derived mDA progenitors. By matching a phase IV scRNA-seq dataset with a dataset from the mouse nervous system, we found that COL1A1+ cells are transcriptionally similar to vascular and leptomeningeal cells (VLMCs) and enteric glial cells ( Figure 8E and Figure 8F ). Therefore, we termed these cells VLMC, which includes three distinct subclusters: IV-OTX2+VLMC, IV-LUM+VLMC, and IV-APOD+VLMC. We further compared our scRNA-seq data with a public scRNA-seq dataset from human fetal midbrain using MetaNeighbor. The similarity scores and clustering results demonstrated that the cell types generated in our protocol are similar to their in vivo counterparts ( Figure 8G ).
[0255] In these experiments, 26% of cells were assigned to 10,702 clones containing two or more cells, of which 2,540 spanned two differentiation stages ( Figure 1E No significant correlation was observed between clone size in cross-stage lineage tracing experiments and barcode frequency in overall sequencing ( Figure 6I), indicating that large clones are unlikely to be derived from multiple initial cells labeled by viruses with the same high-copy barcode. Using a force-directed graph, thousands of lineage networks were constructed for these clones, where nodes are cells colored by their corresponding stages and edges are clonal relationships ( Figure 1F ).
[0256] Thus, the described protocol generates multiple cell types that recapitulate human ventral midbrain-hindbrain development. Furthermore, SISBAR is a high-throughput method for single-cell transcriptome profiling of clonally associated cells at various stages during hPSC neural differentiation.
[0257] Example 3: Cross-stage pedigree relationships determined through “potential perspective” and “origin perspective” analysis
[0258] For sister cells of putative founder cells, which are clonally related descendants, to have similar transcriptional states, we first examined the transcriptome similarity of clonally related sister cells at stage III. We found that sister cell pairs were closely positioned on the UMAP map, with higher expression correlation scores (P < 0.001) and closer UMAP distances (P < 0.001) compared to random cell pairs. Figure 2A and Figure 2B ), indicating that sister cells at stage III in our system can replace each other transcriptionally. 559 clones totaling 2,816 cells were recovered, spanning stages III-IV, and clustered using heatmap visualization of the Pearson correlation coefficient of clonal cell type distribution ( Figure 2C and Figure 9A We quantitatively analyzed the lineage relationships between different cell types at stages III and IV based on the null hypothesis that cells in each clone are randomly distributed. We examined the potential of a given progenitor type at early stages by calculating the distribution of clonal cells at later stages and comparing it to a permuted random distribution, which we termed a “potential perspective” analysis ( Figure 2D Alternatively, we can examine the origin of a given cell type at later stages by calculating the distribution of clonal cells at earlier stages and comparing it to a permuted random distribution, which is called an “origin-perspective” analysis ( Figure 2D , left). Therefore, the combination of potential perspective analysis and origin perspective analysis provides a more comprehensive approach for cross-stage lineage analysis ( Figure 2E Importantly, we found that lineage relationships (p-values and Jaccard scores) across pairs of stage clusters were significantly correlated across biological replicates ( Figure 9B 、 9C and 9G-9J), demonstrating the reproducibility of the SISBAR method and our results. vMBThey specifically expressed typical ventral midbrain markers, such as EN1, OTX2, and LMX1A, and had significant lineage (clonal) relationships with IV-mDA, IV-mGlut, and IV-OTX2+VLMC (all P values < 0.05), but had no significant relationship with other terminal cell types of stage IV ( Figure 2E These results indicate that III-NP vMB At least tripotential and can generate mDA neurons, mGlut neurons and OTX2+VLMC cells. Origin perspective analysis also revealed that IV-mDA, IV-mGlut and IV-OTX2+VLMC originated from III-NP vMB ( Figure 2E , right), confirming the cross-stage lineage relationships between these cell types ( Figure 2F In the developing mouse brain, endogenous En1 is expressed in the caudal midbrain and anterior hindbrain, and endogenous Lmx1a is expressed in the midbrain floor plate and diencephalon. The brain region expressing both markers can be designated as the ventral midbrain floor plate, where mDA neurons originate (Ang, 2006; Arenas et al., 2015). In our dataset, III-NP vMB Specifically express EN1 and LMX1A, and the combination of EN1 and LMX1A is more characteristic for VM progenitor cells ( Figure 9D ). Generate LMX1A / EN1 dual reporter hESC line (LMX1A-tdtomato / EN1-NeoGreen-hESC) to verify the multipotency of ventral midbrain progenitor cells ( Figure 2G LMX1A+ / EN1+ progenitor cells derived from this dual-reporter hESC line were sorted by flow cytometry at stage III and further differentiated and matured until stage IV ( Figure 2H Immunohistochemistry revealed that LMX1A+ / EN1+ progenitor cells can generate mDA neurons (TH+), glutamatergic neurons (vGlut2+) and VLMCs (COL1A1+), demonstrating the tripotentiality of VM progenitor cells ( Figure 2I and 2J These results are consistent with those of stage IV prospective lineage tracing based on clonal relationship analysis, indicating that mDA, mGlut, and OTX2+ VLMCs have significant lineage relationships with each other and are close neighbors on the reconstructed lineage tree, although this approach does not provide information about which cell clusters represent their stage III ancestors ( Figure 9E and 9F ). Similarly, through the potential perspective and origin perspective analysis, we found that the posterior hindbrain progenitor cells (III-NP pHB) are multipotent and can differentiate into GABAergic neurons, glutamatergic neurons, serotonin neurons, and motor neurons at stage IV ( Figure 2E and 2K ). Therefore, we identified III-NP vMB and III-NP pHB Different lineage trajectories from stage III to stage IV during differentiation ( Figure 2G and 2L ). In contrast, we found that III-NP aHB It is unipotent and mainly differentiates into LUM+VLMC cells at stage IV ( Figure 2E ).
[0259] It is worth noting that the lineage hierarchies constructed based on clonal analysis are very different from those constructed based on transcriptome similarity. For example, the hierarchy derived from transcriptome similarity cannot elucidate the relationship between III-NP vMB Cross-stage relationships with IV-mDA, IV-mGlut, and IV-OTX2+VLMC ( Figure 2M and 2N ), nor can the lineage relationship between IV-mDA, IV-mGlut and IV-OTX2+VLMC in stage IV be elucidated ( Figures 9E-9H One reason for the observed differences is that mDA and mGlut are transcriptionally distinct from OTX2+ VLMC and NP vMB These results further demonstrate the advantage of barcode-based clonal tracing in inferring lineage relationships more accurately, especially in the absence of intermediate cell states.
[0260] In summary, through the potential perspective and origin perspective analysis, the lineage relationship between progenitor cell types and terminal cells across differentiation stages was elucidated, and various unipotent and multipotent transcriptome-defined progenitor cell types were identified. Specifically, the III-NP vMB The three lineage fates of OTX2+VLMCs are mDA neurons, mGlut neurons and non-neuronal cell types.
[0261] Example 4 Construction of a multi-level pedigree tree
[0262] We further examined the lineage relationships of different cell types between stages I and II or between stages II and III ( Figure 1B ) to map the complete lineage trajectory of human ventral midbrain-hindbrain differentiation. 768 clones spanning stages I and II and 1,213 clones spanning stages II and III were captured ( Figure 3A 、 3B, 10A and 10B). After demonstrating the transcriptome similarity between stage I and stage II sister cells ( Figures 10C-10F ), quantitatively analyzed the lineage relationships between different cell type pairs across different stages through potential perspective or origin perspective analysis, and identified many previously uncharacterized convergent or divergent lineage trajectories ( Figure 3C 、 3D and 10G-10J). For the divergent trajectory, it is found that NP pHB and NP MHB NPs in stage I have multiple lineage fates, ranging from stage I to stage II or from stage II to stage III. pHB (I-NP pHB ) can produce II-NP in phase II pHB , Motor and Sero, while II-NP pHB The resulting offspring have more diverse fates, including stage III-NP pHB , Glut, Motor, Sero and GABA ( Figure 3C 、 3D and 11A). NP in stage I MHB (I-NP MHB ) can differentiate into II-NP in stage II MHB and Glut, while II-NP MHB Can produce III-NP in phase III aHB and III-NP vMB ( Figure 3C and 3D ). For the convergence trajectory, we find that II-NP MHB and II-NP aHB Can be convergently differentiated into III-NP aHB , and II-NP MHB and II-NP vMB Both can produce III-NP vMB ( Figure 3C and 3D ). We also identified non-branching lineage trajectories. For example, NP vMB Are unipotent and differentiate specifically into late stage NPs from differentiation stages I to III vMB ( Figure 3C and 3D ). It is noteworthy that non-overlapping lineage relationships of some cell type pairs across stages were observed from the potential perspective or origin perspective analysis ( Figure 3C and 3D), probably because clones randomly distribute to different objects in the null hypothesis of these two analyses, indicating that potential perspective or origin perspective analysis serves different purposes ( Figure 11B and 11C By aggregating all clonal data across stages, we generated a multi-level lineage tree containing all cell types and significant lineage associations (P ≤ 0.05 in potential- or origin-perspective analyses) across stages I to IV, depicting the complete lineage landscape of human ventral midbrain-hindbrain differentiation in vitro ( Figure 3E ).
[0263] Example 5 Progenitor cells of different lineages can generate the same cell type and leave molecular imprints on their progeny
[0264] We identified two forms of convergent lineage trajectories: one that converges to molecularly distinct subtypes of the same type, and another that converges to cells of the same type with similar transcriptomes (within the same transcriptome-defined cell clusters). aHB and III-NP vMB Both can differentiate into COL1A1+ VLMCs, but VLMCs derived from these two progenitor cells have different transcriptomes, showing different VLMC subtypes ( Figure 1D 、 3E and 3F). In contrast, the aHB or II-NP MHB III-NP aHB In the UMAP map, the cell clusters defined by the same transcriptome partially overlap ( Figure 3E 、 3G and 3H). It is worth noting that we classified cell clusters based on transcriptome similarity, however, a transcriptome-centric view does not provide a comprehensive view of cell state. Other factors such as DNA methylation, histone modifications, and chromatin accessibility are also important parameters that determine cell state (Wagner and Klein, 2020). Therefore, although the II-NP aHB or II-NP MHB The derived III-NPaHBs have similar transcriptomic states, but may differ in other aspects of their cellular states. We next investigated the relationship between the III-NPs derived from different ancestors. aHB Whether it has specific molecular characteristics. We found that the II-NP aHB or II-NP MHB III-NP aHB Differentially expressed genes (DEGs) can be isolated in the samples derived from II-NP MHB III-NP aHBCNPY1, PCDH15, and PAX5 are specifically expressed in the mitochondria or in the mitochondria derived from II-NP aHB III-NP aHB CCDC80, DOCK10 and LPL are specifically expressed in Figure 3I ). Notably, most of these DEGs were closely related to the differentiation of II-NP aHB and II-NP MHB The DEGs are similar ( Figure 3J and 3K These results suggest that progenitor cells from different lineages can differentiate into the same transcriptome-defined cell types and leave molecular imprints on their progeny.
[0265] Midbrain-hindbrain boundary progenitor cells (NPs) MHB ) during embryonic development is not yet fully understood. Recent studies have shown that Fgf8 lineage cells are distributed in the midbrain and forebrain (rhombomere 1, R1) of the mouse brain. Using the SISBAR method and potential and origin perspectives, it was shown that NP MHB Specialized to produce late-stage NPs MHB , Glut, NP aHB and NP vMB ( Figure 3E ). In order to experimentally verify the MHB To investigate the fate of these lineages, we applied classical recombinase-based prospective lineage tracing (fate mapping) combined with scRNA-seq to analyze NPs that specifically express the MHB marker gene FGF8. MHB The pedigree results ( Figure 8C ). An FGF8-iCre knock-in hESC line was constructed ( Figure 11D After neural differentiation, cells were infected with a lentiviral vector encoding Cre-dependent (DIO, or double-slash inverted orientation) EGFP three days before stage II to mark FGF8+ cells ( Figure 3L Ten days after differentiation (stage III), EGFP+ cells were sorted and scRNA-seq was performed. The results showed that the progeny of FGF8+ cells marked in stage II were composed of NP aHB NP vMB NP MHB , Glut and a small amount of NP pHB and Motor Figure 3M and 3N ), and the NP deduced by analysis MHB The fate of the lineage is consistent.
[0266] Example 6 The multi-lineage fate of a progenitor cell type represents the collective outcome of the fates of many distinct but dissimilar individual progenitor cell clones, each with a distinct molecular signature.
[0267] Three models can explain the observed population-level divergence of lineage trajectories inferred from either a potential perspective or an origin perspective. Each lineage trajectory can be composed of many identical or similar clonal trajectories (Model 1). Alternatively, it can be a consensus of many constrained but distinct clonal trajectories (Models 2 and 3). Figure 4A ).
[0268] Individual III-NP was examined vMB Many unipotent III-NPs have been identified. vMB , which specifically differentiate into mDA, mGlut or OTX2+VLMC, as well as many bipotent III-NPs vMB , which can simultaneously produce mGlut and OTX2+VLMC or mDA and OTX2+VLMC ( Figure 4B and 4C ). Although individual tri-energetic III-NP vMB did not appear in our cross-stage data, but clones with triple fates including mDA, mGlut, and OTX2+VLMC were observed in our prospective lineage tracing analysis at stage IV ( Figures 12A-12C ). It is speculated that our data underestimate the true number of multipotent progenitors because some clonal fates may be missed due to inevitable cell loss during the experiment. Similarly, it was observed that III-NP pHB The different terminal clonal fates of individual progenitor cells in the human body are composed of a variety of unipotent and multipotent progenitor cells ( Figure 12D and 12E Intermediate progenitor cell types (e.g., II-NPs) were also examined. pHB , which produces III-NP in phase III pHB , GABA, Glut, Sero and Motor) ( Figure 3E ). Individual II-NP pHB The cloning results were also highly diverse, with only III-NP pHB , single neuron type or III-NP pHB and various combinations of neuron types ( Figure 4D and 4EWe then calculated the proportion of symmetric self-renewing clones with a unique clonal fate identical to that of the corresponding progenitor, and the proportion of asymmetric self-renewing clones with multiple clonal fates (≥2 fates), including one identical to that of the corresponding progenitor, across all progenitor cell types ( Figure 12F ). Specifically, 10.7% of II-NP was found pHB Related clones specifically produced NP in phase III pHB (III-NP pHB ), which indicates that these NPs pHB Related clones undergo symmetrical self-renewal during this period ( Figure 12F ).
[0269] We further explored the molecular signatures associated with the different fates of individual progenitor cells within the same transcriptome-defined progenitor cell type. Statistical testing was performed on all clones with a single clonal fate (unipotency). The results showed that if a clone had at least three cells that all developed into a single fate at a later stage, the clone was significantly unipotent (p value < 0.05) ( Figure 12G Although different fate progenitors are mixed in the same progenitor cluster in the UMAP map ( Figure 4C 、 4E and 12E), but DEG analysis revealed genes and related pathways that were differentially enriched in these different fate progenitors (clonal cell number ≥3 was used as a filter for the minimum cell number at the later stage) ( Figures 4F-4I and 12H). For example, self-renewing II-NPs were found pHB (III-NP pHB Fate) highly express cell cycle and self-renewal-related genes, such as HES5 and PAX6, as well as many other genes that have not yet been described in neural stem cell self-renewal ( Figure 4F and 4G Paired mesoderm homeobox protein 2b (PHOX2B) plays an important role in the generation of ventral hindbrain motor neurons and is closely related to other III-NPs with different fates. pHB Compared to the III-NP in Motor Fate pHB Moderately high expression ( Figure 4H and 4I These results suggest that individual progenitors within the same transcriptome-defined cell type can commit to distinct fates, which are associated with specific molecular signatures.
[0270] These results suggest that individual progenitor cells within the same type have distinct but restricted clonal fates associated with their unique molecular signatures and that the clonal fates of individual cells collectively contribute to the lineage outcome of the progenitor cell type.
[0271] Example 7 Multipotent ventral midbrain progenitor cells differentiate into specialized lineages
[0272] The potential of progenitor cells is often more restricted during neural differentiation. vMB The cells are multipotent and can produce mDA, mGluts and OTX2+VLMC ( Figure 4B 、 4C and 12A-12C). Exploring specialized NPs vMB Whether cells become fate-restricted in more specialized lineages during differentiation. Progenitor cells were labeled using a barcoded viral library at stage II or 2 days after stage III, and then further differentiated / matured and subjected to scRNA-seq ( Figure 5A A large number of bi- or tri-potent clones containing mDA, OTX2+VLMC, and / or mGlut were observed in the early-stage but not the late-stage labeled groups ( Figures 5B-5D and 13A). Testing the clonal relationships between mDA and other clusters, we found that the clonal correlations between mDA and mGlut (P < 10-3) or between mDA and OTX2+VLMC (P < 10-9) in early-stage labeling experiments were significantly higher than those in the isoenergetic null model (in which the fate of each clone is assumed to be randomly distributed among all clusters) ( Figure 5E These results suggest that early multipotent progenitors that produce mDA, mGluts, and OTX2+ VLMCs divide into more fate-restricted lineages at later stages of differentiation (2 days after stage III). In contrast, we observed clonal relationships between hGluts, Seros, and GABA both at stage II and 2 days after stage III ( Figures 13B to 13D ), indicating that different types of multipotent progenitors differentiate into specialized lineages at different time points.
[0273] To further verify the individual NP vMB To investigate the multi-lineage fate of NPs, traditional fluorescent labeling experiments were designed to track the fate of single NPs. vMB Single-cell differentiation experiments were performed by mixing single LMX1A+ / EN1+ progenitor cells derived from LMX1A-tdtomato / EN1-NeoGreen-hESCs with hundreds of neural progenitor cells of the same stage derived from wild-type hESCs, and then further differentiated and matured to stage IV ( Figure 5F The clonal cell type composition of more than one hundred clones was analyzed by immunostaining for TH (tyrosine hydroxylase) and COL1A1 ( Figure 5G), found that 18% of clones in early (stage II) sorted LMX1A+ / EN1+ cells had both TH+ neurons (mDA) and COL1A1+ non-neuronal cells (VLMC) ( Figure 5H , top), indicating that mDA neurons and VLMCs share a common clonal origin. In contrast, we did not observe clones containing both TH+ cells and COL1A1+ cells in the group sorted at a later stage (2 days after stage III) ( Figure 5H , bottom), which supports the late NP vMB It is noteworthy that approximately half of the clones sorted at either early or late stages consisted of cells that did not express TH or COL1A1 ( Figure 5H ), suggesting that other cell types besides mDA neurons and VLMCs can also differentiate from LMX1A+ / EN1+ cells.
[0274] In summary, through prospective lineage tracing and single-cell differentiation experiments of progenitors at different differentiation time points, we experimentally validated the role of individual NPs in vMB multi-lineage fate and demonstrated that specialized NPs vMB Can be divided into specialized lineages as they differentiate.
[0275] Although preferred embodiments of the present invention have been shown and described herein, such embodiments are provided by way of example only, which will be apparent to those skilled in the art. The present invention is not limited to the specific embodiments provided in the specification. Although the present invention has been described with reference to the foregoing description, the description and illustration of the embodiments herein are not intended to be interpreted in a limiting sense. Various modifications, variations and replacements will now be made by those skilled in the art without departing from the present invention. In addition, it should be understood that all aspects of the present invention are not limited to the specific descriptions, configurations or relative proportions described herein that depend on various conditions and variables. It should be understood that, in practicing the present invention, various alternatives to the embodiments of the present invention described herein may be adopted. Therefore, it is contemplated that the present invention will also encompass any such alternatives, modifications, variations or equivalents. This means that the appended claims define the scope of the present invention and that methods and structures within the scope of these claims and their equivalents are therefore encompassed.
Claims
1. A lineage tracing method comprising obtaining cells and their sister cells at different differentiation stages, dividing the cells and their sister cells into at least two parts, one part being used for genetic analysis and the other part continuing to differentiate and undergo genetic analysis, and performing cross-stage lineage analysis.
2. A method for predicting cell differentiation fate, comprising obtaining cells and their sister cells at different differentiation stages, dividing the cells and their sister cells into at least two parts, one part for genetic analysis, and the other part for continued differentiation and genetic analysis, as well as cross-stage lineage analysis.
3. A method for identifying the origin of a cell, which comprises obtaining a cell and its sister cells at different stages of differentiation, dividing the cell and its sister cells into at least two parts, one part being used for genetic analysis, the other part continuing to differentiate and undergo genetic analysis, and performing cross-stage lineage analysis.
4. The method according to any one of claims 1 to 3, wherein the cells at different stages of differentiation undergo cell division and produce sister cells.
5. The method according to any one of claims 1 to 4, wherein the cells at different stages of differentiation are derived from the same kind of progenitor cells.
6. The method according to any one of claims 1 to 4, wherein the cells at different stages of differentiation are derived from different kinds of progenitor cells.
7. The method according to any one of claims 1 to 6, wherein the cells are derived from pluripotent stem cells.
8. The method of any one of claims 1 to 7, wherein the cells are derived from human pluripotent stem cells.
9. The method according to any one of claims 7-8, wherein the pluripotent stem cells comprise embryonic stem cells (ESCs) and / or induced pluripotent stem cells (iPSCs).
10. The method of any one of claims 5-9, wherein the progenitor cells comprise neural progenitor cells.
11. The method according to any one of claims 1 to 10, wherein the differentiation is in vitro differentiation.
12. The method according to any one of claims 1 to 11, comprising introducing a marker into the differentiation-initiating cells.
13. The method according to any one of claims 1 to 12, wherein the marker is used to mark differentiation-initiating cells.
14. The method of any one of claims 12-13, wherein the marker comprises a barcode.
15. The method of claim 14, wherein the barcode comprises a viral barcode.
16. The method of claim 15, wherein the viral barcode comprises a lentiviral vector.
17. The method of any one of claims 15-16, wherein the viral barcode comprises a retroviral vector.
18. The method of any one of claims 15-17, wherein the viral barcode comprises inserting a 32-bp semi-random barcode into a retroviral vector.
19. The method of any one of claims 15-19, wherein the viral barcode comprises a random nucleic acid sequence.
20. The method according to any one of claims 15 to 19, wherein the viral barcode contains 16 repeats of an alternating S (G or C) / W (A or T) nucleic acid sequence.
21. The method of any one of claims 1-20, wherein the genetic analysis comprises recovering viral barcodes.
22. The method of any one of claims 12-21, wherein cells with the same viral barcode are defined as a single clone.
23. The method according to any one of claims 12 to 22, wherein the marker comprises a Cre-LoxP system.
24. The method of any one of claims 12-23, wherein the marker comprises an insertion and / or deletion of a traceable element introduced by the CRISPR-Cas9 system.
25. The method of any one of claims 1-24, comprising one or more differentiation stages.
26. The method of any one of claims 1-25, comprising two or more differentiation stages.
27. The method of any one of claims 1-26, comprising three or more differentiation stages.
28. The method of any one of claims 1-27, wherein the genetic analysis comprises sequencing the cell or its sister cells.
29. The method of claim 28, wherein the sequencing comprises transcriptome sequencing.
30. The method of any one of claims 28-29, wherein the sequencing comprises single-cell RNA sequencing (scRNA-seq).
31. The method of any one of claims 28-30, comprising performing data analysis after sequencing the cell or its sister cells.
32. The method of any one of claims 1-31 comprising potential perspective analysis.
33. The method of any one of claims 32, wherein the potential perspective analysis comprises examining the potential of a given progenitor cell type at an early stage by calculating the distribution of clonal cells at a later stage and comparing it to a random distribution of permutations.
34. The method of any one of claims 1-33, comprising a provenance perspective analysis.
35. The method of claim 34, wherein the origin-perspective analysis comprises examining the origin of a given cell type at a later stage by calculating the distribution of clonal cells at an earlier stage and comparing it to a random distribution of permutations.
36. The method of any one of claims 33-35, wherein the distribution can be visualized by heat map clustering based on the Pearson correlation coefficient of the clonal cell type distribution.
37. The method of any one of claims 1-36, wherein the lineage relationships of different cell clusters within a developmental stage are determined by calculating the jaccard similarity coefficient score (the number of shared clones divided by all clones of a pair of clusters), the p-value (given the null hypothesis that cells in a clone are randomly distributed across all clusters), lineage coupling correlation analysis, and / or clonal coupling (observed / expected) between each pair of cell clusters.
38. The method according to any one of claims 1 to 37, comprising the steps of: (a) introducing a viral barcode into the differentiation-initiating cells, (b) culturing the differentiation-initiating cells and obtaining cells at the next differentiation stage and their sister cells, (c) dividing the cells at the next differentiation stage and their sister cells into two parts, one part for viral barcode recovery and scRNA sequencing, and the other part for continued differentiation, (d) repeating (c) if the differentiation includes two or more stages, (e) after sequencing, constructing a lineage tree using potential perspective analysis and origin perspective analysis.
39. Use of the method of any one of claims 1-38 for lineage tracing of differentiation products of pluripotent stem cells.
40. Use of the method of any one of claims 1-38 for lineage tracing of a heterogeneous cell population.