Method for reconstructing lineage of multiple cells differentiated from common ancestor, computer program and recording carrier thereof, and device

By calculating the similarity matrix of DNA methylation status information in human cells and correcting cell type-specific signals, the problem of low-resolution lineage reconstruction was solved, and high-accuracy human cell lineage reconstruction was achieved.

WO2025245662A1PCT designated stage Publication Date: 2025-12-04WESTLAKE LAB OF LIFE SCI & BIOMEDICINE
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/095497
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve high-resolution lineage reconstruction in human cells, especially due to the low natural mutation rate of nuclear genomic DNA and mitochondrial DNA, resulting in insufficient resolution for lineage reconstruction based on DNA methylation status. Furthermore, the dynamic changes in DNA methylation during development and differentiation lead to sparse signals, making accurate non-invasive lineage reconstruction methods lacking.

Method used

By calculating the DNA methylation status information of the cell nuclear genome sequence, a similarity matrix is ​​constructed using the average methylation fraction of the genome segments. The lineage is then reconstructed by combining similarity matrix correction and cell type-specific signal removal.

Benefits of technology

It enables the high-accuracy reconstruction of human cell lineages with low genome coverage, and can directly reconstruct cell lineages in individuals where gene editing is not possible, thus improving the resolution and accuracy of lineage reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024095497_04122025_PF_FP_ABST
    Figure CN2024095497_04122025_PF_FP_ABST
Patent Text Reader

Abstract

A method for reconstructing a lineage of multiple cells differentiated from a common ancestor. A computer program for reconstructing a lineage of multiple cells differentiated from a common ancestor and a recording carrier thereof. A device for reconstructing a lineage of multiple cells differentiated from a common ancestor.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, computer programs, recording media, and devices for reconstructing lineages from multiple cells differentiated from a common ancestor [Technical Field]

[0001] This invention relates to a method for reconstructing lineages from multiple cells differentiated from a common ancestor. The invention also relates to a computer program and its recording medium for reconstructing lineages from multiple cells differentiated from a common ancestor. Furthermore, the invention relates to an apparatus for reconstructing lineages from multiple cells differentiated from a common ancestor. [Background Technology]

[0002] Phylogenetic reconstruction is a technique for tracing cellular relationships, which can be broadly divided into two types: forward tracing and backward tracing. Forward tracing involves pre-labeling a specific type of cell and then tracking the subsequent evolution and fate of these cells (i.e., what cell type they differentiate into). Backward tracing, on the other hand, does not specifically label any particular cell type; instead, it uses algorithms to trace the closeness of relationships between all cells based on the current state of the detected cells. This invention primarily focuses on backward tracing.

[0003] Regarding reverse tracing, individual cells can be labeled by inducing heritable DNA mutations, which can then be analyzed in detail using single-cell sequencing. Several techniques for inducing heritable DNA mutations to label individual cells are currently known, including, for example, random transposon integration, recombinase-induced gene rearrangement, or CRISPR-Cas9 gene editing at defined sites.

[0004] Non-patent literature 1 (Li et al., 2023) mentions a highly efficient lineage reconstruction mouse model called DARLIN, which can induce the generation of approximately 10 lineages at any time point. 18 Different lineage barcodes. Non-patent literature 2 (Wang et al., 2022) mentions a computational package called CoSpar that learns differentiation dynamics by integrating lineage and cell state information. The application of these latest lineage reconstruction tools has revealed important insights into cell fate selection, cell migration dynamics, cancer evolution, and clonal memory.

[0005] Due to the strict prohibition of gene manipulation in humans, research progress on lineage reconstruction of human cells has been limited. Currently, lineage reconstruction of human cells mainly relies on the natural accumulation of mutations in the genomic DNA of the cell nucleus. However, normal human cells accumulate only about 10 mutations per cell division. -9Nucleus-level mutations in nuclear genomic DNA result in extremely low resolution for lineage reconstruction, even far lower than that of DARLIN mentioned in non-patent literature 1 (Li et al., 2023). Compared to natural mutations in human cell nuclear genomic DNA, the natural mutation rate in human cell mitochondrial DNA is approximately 10–100 times higher. Furthermore, due to its smaller size (approximately 16.7 kb), mitochondrial DNA is easier to study and has been proposed as an alternative to lineage reconstruction based on the natural accumulation of mutations in human cell nuclear genomic DNA. However, the results of lineage reconstruction based on the natural accumulation of mutations in human cell mitochondrial DNA often differ significantly from those based on the natural accumulation of mutations in human cell nuclear genomic DNA (as determined by CRISPR-Cas9 editing). This may be due to the significant size difference between human cell nuclear genomic DNA and mitochondrial DNA, and also because the natural mutation rate of human cell mitochondrial DNA remains too low, meaning the resolution of lineage reconstruction based on the natural accumulation of mutations in human cell mitochondrial DNA remains very low.

[0006] To address the aforementioned technical problems, the inventors focused on the natural accumulation of epigenetic variations on the genomic DNA of human cell nuclei and developed a corresponding lineage reconstruction method. Specifically, DNA methylation (primarily occurring on the cytosine residues (C) of CpG dinucleotides), one of the epigenetic variations, accumulates on average at approximately 10 at each CpG site during each cell division. -3 The mutation rate is much higher than that of natural mutations in human cell nuclear genomic DNA and mitochondrial DNA. Furthermore, due to the approximately 29 million CpG sites on human cell nuclear genomic DNA, lineage reconstruction based on the natural accumulation of variations in DNA methylation status on human cell nuclear genomic DNA can achieve sufficiently high resolution. Previous observations have supported the possibility that DNA methylation can distinguish different clones (cell populations originating from a common ancestor). Previous studies have found different methylation patterns in different clones, in somatic cell lines, and in cancer cells. Non-patent literature 3 (Gaiti et al., 2019) used a lineage reconstruction method based on the natural accumulation of variations in DNA methylation status on human cell nuclear genomic DNA to infer the lineage of cancer cells; however, this study was limited to a specific type of cancer cell (B lymphocytes), lacked sufficient validation, and had low accuracy, indicating it was not yet mature.

[0007] How to infer cell lineage from epigenetic mutations in DNA methylation in normal cells and make this tool applicable to a wide variety of cell types is a previously unexplored problem. There are two main challenges: (1) Unlike DNA, DNA methylation undergoes dynamic changes during development and differentiation to produce cell type-specific variations; (2) Single-cell DNA methylation rates typically have very low genome coverage, resulting in sparse signaling.

[0008] Currently, there is no method that can accurately predict cell lineage information from very sparse single-cell DNA methylation data, thus achieving non-invasive, high-precision lineage reconstruction.

[0009] [Existing Technical Documents]

[0010] Non-patent literature 1: Li, L. et al. A mouse model with high clonal barcode diversity for joint lineage, transcriptomic, and epigenomic profiling in single cells. Cell(2023).doi:10.1016 / j.cell.2023.09.019

[0011] Non-patent literature 2: Wang, SW, Herriges, MJ, Hurley, K., Kotton, DN & Klein, AMCoSpar identifies early cell fate biases from single-cell transcriptomic and lineage information.Nat.Biotechnol.1-9(2022).doi:10.1038 / s41587-022-01209-1

[0012] Non-patent literature 3: Gaiti, F. et al. Epigenetic evolution and lineage histories of chronic lymphocytic leukaemia. Nature 569, 576-580 (2019). doi:10.1038 / s41586-019-1198-z

[0013] [Summary of the Invention]

[0014] This invention includes the following:

[0015] 1. A method for reconstructing lineages from multiple cells differentiated from a common ancestor, comprising:

[0016] (a) Nuclear genome sequence data for each cell of multiple cells differentiated from a common ancestor, containing information on DNA methylation status.

[0017] (a1) The cell nuclear genome sequence is divided into discrete or continuous non-overlapping genome segments of a certain unit length. All the divided genome segments of each cell constitute the genome segment set W of that cell, and

[0018] (a2) Calculate the average methylation fraction A within each genomic segment in the set of genomic segments W;

[0019] (b) The set of genomic segments W consists of all genomic segments in cells i and j of the plurality of cells in which the average methylation fraction A is not unknown, forming a total set of genomic segments W. ij Only in the shared genomic segment set W ij The average methylation fraction A within cell i of genomic region x is calculated. ix and the average methylation fraction A of cell j jx The similarity between cells, denoted as S, is the similarity between cells i and j. ij :

[0020] in

[0021] i is the cell number among the plurality of cells, i = 1, 2, ..., N.

[0022] j is the cell number among the plurality of cells, j = 1, 2, ..., N.

[0023] i≠j,

[0024] N is the total number of cells in the plurality of cells.

[0025] x is the set of shared genomic segments W ij The genomic region numbers in the data, x = 1, 2, ..., M, and

[0026] M is the set of shared genomic segments W ij The total number of genomic segments in the genome.

[0027] The similarity S ij The similarity matrix S is formed by N×N; and

[0028] (c) Reconstruct lineages of the plurality of cells using the similarity matrix S prepared in step (b).

[0029] 2. According to the method described in the foregoing embodiments, the cell lineage-specific similarity matrix L is obtained by subtracting the cell type-specific similarity matrix T from the similarity matrix S: L = ST.

[0030] 3. According to the method described in the foregoing embodiments, the cell type-specific similarity matrix T is prepared as follows:

[0031] [i] Define set Ω p and Ω q :

[0032] Ω p It is the set of cells belonging to cell type p among the plurality of cells, and

[0033] Ω q It is the set of cells belonging to cell type q among the multiple cells.

[0034] Where p and q are the same or different;

[0035] [ii] When cell i belongs to cell type p and cell j belongs to cell type q.

[0036] The similarity subset G is formed by extracting the similarity between cell i and all cells k belonging to cell type q from the similarity matrix S. iq :

[0037] G iq ={S ik};k∈Ω q ,k≠i

[0038] The similarity subset G is formed by extracting the similarity between cell j and all cells k belonging to cell type p from the similarity matrix S. jp :

[0039] G jp ={S jk};k∈Ω p ,k≠j

[0040] in,

[0041] k is the set Ω p and Ω q Cell numbering, k = 1, 2, ..., and

[0042] k≠i and k≠j exclude the diagonal elements of the similarity matrix S; and

[0043] [iii] The similarity subset G iq and similarity subset G jp Merge into a merged similarity subset G ij :

[0044] G ij =G iq ∪G jp

[0045] Merging similarity subsets G ij The average similarity between cells belonging to cell type p and cells belonging to cell type q is obtained by averaging all similarities. T represents the cell type-specific similarity between cell i and cell j. ij The cell type specific similarity T ij The cell type-specific similarity matrix T is formed by N×N.

[0046] 4. The method according to the foregoing embodiments, wherein...

[0047] (b1) Regarding the similarity S ij Perform the following iterative correction:

[0048] S ij =Z i S ij Z j

[0049] in,

[0050] = is the assignment operator.

[0051] Z i The average methylation fraction A of cell i ix The noise factor is initialized as the reciprocal of the square root of the maximum value in the i-th row of the similarity matrix S, excluding the values ​​in the diagonal cells:

[0052] And perturb at each loop.

[0053] Z j The average methylation fraction A of cell j jx The noise factor is initialized as the reciprocal of the square root of the maximum value in the j-th column of the similarity matrix S, excluding the values ​​in the diagonal cells:

[0054] And perturb at each loop.

[0055] The Z i and Z j Composition vector

[0056] (b2) For the corrected similarity matrix S obtained in each iteration, calculate the ratio of the standard deviation to the mean of all its elements except the diagonal elements as the cost Cost(S); and

[0057] (b3) The loop ends when Cost(S) is minimized, and the corrected similarity matrix S at this time is used as the final corrected similarity matrix S.

[0058] 5. The method according to the foregoing embodiments, wherein the vector Each element is used average Perform calibration and reassignment

[0059] = is the assignment operator.

[0060] 6. According to the method described in the foregoing embodiments, the cell lineage-specific similarity matrix L is obtained by subtracting the cell type-specific similarity matrix T from the corrected similarity matrix S: L = ST.

[0061] 7. The method according to the foregoing embodiments, wherein the cell type-specific similarity matrix T is prepared as follows:

[0062] [i] Define set Ω p and Ω q :

[0063] Ω p It is the set of cells belonging to cell type p among the plurality of cells, and

[0064] Ω q It is the set of cells belonging to cell type q among the multiple cells.

[0065] Where p and q are the same or different;

[0066] [ii] When cell i belongs to cell type p and cell j belongs to cell type q.

[0067] The similarity subset G is formed by extracting the similarity between cell i and all cells k belonging to cell type q from the similarity matrix S. iq :

[0068] G iq ={S ik};k∈Ω q ,k≠i

[0069] The similarity subset G is formed by extracting the similarity between cell j and all cells k belonging to cell type p from the similarity matrix S. jp :

[0070] G jp ={S jk};k∈Ω p ,k≠j

[0071] in,

[0072] k is the set Ω p and Ω q Cell numbering, k = 1, 2, ..., and

[0073] k≠i and k≠j exclude the diagonal elements of the similarity matrix S; and

[0074] [iii] The similarity subset G iq and similarity subset G jp Merge into a merged similarity subset G ij :

[0075] G ij =G iq ∪G jp

[0076] Merging similarity subsets G ij The average similarity between cells belonging to cell type p and cells belonging to cell type q is obtained by averaging all similarities. T represents the cell type-specific similarity between cell i and cell j. ij The cell type specific similarity T ij The cell type-specific similarity matrix T is formed by N×N.

[0077] 8. The method according to the foregoing embodiments, wherein the DNA methylation status information includes: the location and number of CpG sites on the genome and whether each CpG site is methylated or not.

[0078] 9. The method according to the foregoing embodiments, wherein whether or not the CpG site is methylated includes:

[0079] Methylation of the CpG site was detected.

[0080] The CpG site was detected to be unmethylated, and

[0081] It is unknown whether the CpG site is methylated, i.e., NaN.

[0082] 10. The method according to the foregoing embodiments, wherein the nuclear genome sequence data with DNA methylation status information in step (a) is obtained by single-cell bisulfite sequencing.

[0083] 11. The method according to the foregoing embodiments, wherein in step (a1), the genomic segments are discretely distributed, the size of the genomic segments is 2 bp, and each genomic segment corresponds to a CpG site.

[0084] 12. The method according to the foregoing embodiments, wherein in step (a1), the genomic segments are discrete or continuous and non-overlapping, the size of the genomic segments is 10 to 2000 bp, and each genomic segment contains at least one CpG site.

[0085] 13. The method according to the foregoing embodiments, wherein in step (a1), the genomic segments are continuously distributed without overlap, and the size of the genomic segments is 100-1500 bp.

[0086] 14. The method according to the foregoing embodiments, wherein in step (a2), the methylation fraction m of each CpG site is calculated:

[0087] The methylation fraction m of CpG sites within each genomic segment in the set of genomic segments W is calculated as the average of the methylation fraction m of all CpG sites that are not unknown, and is taken as the average methylation fraction A within that genomic segment.

[0088] 15. The method according to the foregoing embodiments, wherein between steps (a1) and (a2) the method further comprises:

[0089] Calculate the average methylation rate of each genomic region across all said multiple cells.

[0090] For each of the plurality of cells, select a genomic segment that satisfies the following conditions [a] and [b]:

[0091] [a] Average methylation rate in all said multiple cells At the pre-set lower limit and upper limit Between, that is and

[0092] [b]The genomic region selected according to the above conditions [a] is present in more than two of the multiple cells. Not unknown,

[0093] This results in a set W of genomic segments for each cell, consisting of genomic segments that satisfy the above conditions [a] and [b].

[0094] 16. The method according to the foregoing embodiments, wherein...

[0095] Selected from 0 to 0.9,

[0096] Selected from 0.1 to 1, and

[0097] yes:

[0098] 0 to 1, or

[0099] Unknown, i.e., NaN.

[0100] 17. The method according to the foregoing embodiments, wherein condition [b] is: the genomic segment selected in condition [a] is present in more than 5% of the cells of the plurality of cells. It is not unknown.

[0101] 18. The method according to the foregoing embodiments, wherein...

[0102] The genomic segments are continuous and non-overlapping, and

[0103] The set of genomic segments W is replaced by a set of merged genomic segments W′, which consists of merged genomic segments that have merged different adjacent genomic segments, obtained by merging adjacent genomic segments that satisfy the conditions [a] and [b].

[0104] 19. The method according to the foregoing embodiments, wherein the methylation fraction m of the CpG site is:

[0105] 0 to 1, or

[0106] Unknown, i.e., NaN.

[0107] 20. The method according to the foregoing embodiments, wherein the similarity is selected from: Pearson correlation coefficient, cosine similarity and Jaccard similarity.

[0108] 21. The method according to the foregoing embodiments, wherein step (c) is performed using an algorithm selected from the following: unweighted pair-group method with arithmetic means (UPGMA), Neighbor-joining (NJ), and FastME.

[0109] 22. The method according to the foregoing embodiments, wherein a verification step of the reconstructed lineage is further included after step (c).

[0110] 23. The method according to the foregoing embodiments, wherein the cells are selected from prokaryotic cells and eukaryotic cells.

[0111] 24. The method according to the foregoing embodiments, wherein the cells are selected from somatic cells and germ cells.

[0112] 25. The method according to the foregoing embodiments, wherein the cell is an embryonic cell.

[0113] 26. The method according to the foregoing embodiments, wherein the cell is derived from stem cells or progenitor cells.

[0114] 27. A recording medium for a computer program that reconstructs a lineage of multiple cells differentiated from a common ancestor, wherein the computer program causes a computer to perform the method described in the foregoing embodiments.

[0115] 28. An apparatus for reconstructing lineages from multiple cells differentiated from a common ancestor, comprising:

[0116] The processing unit executes the method described in the foregoing embodiments, and

[0117] The display unit shows the lineage of the plurality of cells reconstructed by the processing unit.

[0118] [Technical Effects]

[0119] (1) Lineages can be reconstructed from multiple cells that have differentiated from a common ancestor by using single-cell DNA methylation data with low genome coverage;

[0120] (2) Similarity correction can significantly improve the accuracy of cell lineage reconstruction;

[0121] (3) It can effectively remove the interference of cell type-specific methylation signals in the cell similarity matrix, thereby obtaining a similarity matrix that is only related to cell lineage;

[0122] (4) It can directly reconstruct the lineage between human cells in individuals such as humans who cannot undergo gene editing.

[0123] [Brief Description of the Attached Image]

[0124] [Figure 1A] Criteria for selecting genomic regions [a]: Average methylation rate At the pre-set lower limit and upper limit Between, that is

[0125] [Figure 1B] Genomic segments that are adjacent to each other on the genome and meet the conditions [a] and [b] will be merged to obtain a set of merged genomic segments W′ consisting of merged genomic segments that have merged different adjacent genomic segments.

[0126] [Figure 1C] Similarity matrix between cells derived from HEK293T cells from the same subclone, obtained when each CpG site / 2bp genomic segment is divided into discrete genomic segments.

[0127] [Figure 1D] Similarity matrix between cells derived from HEK293T cells from the same subclone, obtained when the genome is divided into continuous, non-overlapping genomic segments of 500bp / segment.

[0128] [Figure 1E] Lineage tree of cells derived from HEK293T cells from the same subclone.

[0129] [Figure 1F] Simplified phylogenetic tree of cells derived from HEK293T cells from the same subclone.

[0130] [Figure 1G] When selecting genomic regions that have not been merged into adjacent genomic regions for similarity calculation between cells derived from HEK293T cells. and The impact of changes on the accuracy of cell lineage reconstruction.

[0131] [Figure 1H] When selecting merged genomic regions after merging adjacent genomic regions for similarity calculation between cells derived from HEK293T cells. and The impact of changes on the accuracy of cell lineage reconstruction.

[0132] [Figure 1I] The impact of reducing the number of merged genomic regions selected from the merged genomic region set W′ on the accuracy of cell lineage reconstruction.

[0133] [Figure 1J] Correlation between the number of merged genomic regions selected from the merged genomic region set W′ and the accuracy of cell lineage reconstruction.

[0134] [Figure 1K] The effect of pre-correction and post-correction similarity of cells derived from HEK293T cells on the accuracy of cell lineage reconstruction.

[0135] [Figure 2A] Cloning of H9 human embryonic stem cells.

[0136] [Figure 2B] When selecting merged genomic segments that combine adjacent genomic segments for similarity calculation between cells derived from H9 human embryonic stem cells. and The impact of changes on the accuracy of cell lineage reconstruction.

[0137] [Figure 2C] Similarity matrix before and after correction of cells derived from H9 human embryonic stem cells.

[0138] [Figure 3A] Cells derived from mouse hematopoietic progenitor cells include eight different cell types: megakaryocytes (Mk), erythrocytes (Er), basophils (Ba), mast cells (Ma), eosinophils (Eos), neutrophils (Neu), neutrophil-like monocytes (Neu-Mon), and dendritic monocytes (Dc-Mon).

[0139] [Figure 3B] shows a heatmap of cell type composition in each observed clone. Only 52 clones with >1 cell are shown.

[0140] [Figure 3C] Histogram of clone size in a dataset of cells derived from mouse hematopoietic progenitor cells.

[0141] [Figure 3D] Similarity matrix between cells derived from mouse hematopoietic progenitor cells.

[0142] [Figure 3E] When selecting merged genomic segments after merging adjacent genomic segments for similarity calculation between cells derived from mouse hematopoietic progenitor cells. and The impact of changes on the accuracy of cell lineage reconstruction.

[0143] [Figure 3F] Cell type-specific DNA methylation patterns in cells derived from mouse hematopoietic progenitor cells.

[0144] [Figure 4A] Similarity matrix and inferred phylogenetic relationships between cells from 21-week-old human embryos.

[0145] [Figure 4B] Merged genomic segments selected after merging adjacent genomic segments were used for cell similarity calculation from 21-week-old human embryos, after removal of cell type-specific methylation signals. and The impact of changes on the accuracy of cell lineage reconstruction.

[0146] [Figure 4C] Similarity matrix and inferred phylogenetic relationships between cells from 7-week-old human embryos.

[0147] [Figure 4D] The pre-correction similarity matrix and the post-correction similarity matrix of cells from a 21-week-old human embryo.

[0148] [Figure 4E] The effect of pre-correction and post-correction similarity of cells from 21-week-old human embryos and cells from 7-week-old human embryos on the accuracy of cell lineage reconstruction.

Detailed Implementation Methods

[0149] In the context of this invention, "DNA methylation" will also be referred to simply as "methylation".

[0150] In the context of this invention, "=" usually represents equality, but in a special definition it represents an assignment operator.

[0151] In one embodiment, the present invention provides a method for reconstructing lineages from multiple cells differentiated from a common ancestor, which may include:

[0152] (a) Nuclear genome sequence data for each cell of multiple cells differentiated from a common ancestor, containing information on DNA methylation status.

[0153] (a1) The cell nuclear genome sequence is divided into discrete or continuous non-overlapping genome segments of a certain unit length. All the divided genome segments of each cell constitute the genome segment set W of that cell, and

[0154] (a2) Calculate the average methylation fraction A within each genomic segment in the set of genomic segments W;

[0155] (b) The set of genomic segments W consists of all genomic segments in cells i and j of the plurality of cells in which the average methylation fraction A is not unknown, forming a total set of genomic segments W. ij Only in the shared genomic segment set W ij The average methylation fraction A within cell i of genomic region x is calculated. ix and the average methylation fraction A of cell j jx The similarity between cells, denoted as S, is the similarity between cells i and j. ij :

[0156] in

[0157] i is the cell number among the plurality of cells, i = 1, 2, ..., N.

[0158] j is the cell number among the plurality of cells, j = 1, 2, ..., N.

[0159] i≠j,

[0160] N is the total number of cells in the plurality of cells.

[0161] x is the set of shared genomic segments W ij The genomic region numbers in the data, x = 1, 2, ..., M, and

[0162] M is the set of shared genomic segments W ij The total number of genomic segments in the genome.

[0163] The similarity S ij The similarity matrix S is formed by N×N; and

[0164] (c) Reconstruct lineages of the plurality of cells using the similarity matrix S prepared in step (b).

[0165] In one embodiment, the cell lineage-specific similarity matrix L can be obtained by subtracting the cell type-specific similarity matrix T from the similarity matrix S: L = ST. In a further embodiment, the cell type-specific similarity matrix T can be constructed as follows:

[0166] [i] Define set Ω p and Ω q :

[0167] Ω p It is the set of cells belonging to cell type p among the plurality of cells, and

[0168] Ω q It is the set of cells belonging to cell type q among the multiple cells.

[0169] Where p and q are the same or different;

[0170] [ii] When cell i belongs to cell type p and cell j belongs to cell type q.

[0171] The similarity subset G is formed by extracting the similarity between cell i and all cells k belonging to cell type q from the similarity matrix S. iq :

[0172] G iq ={S ik};k∈Ω q ,k≠i

[0173] The similarity subset G is formed by extracting the similarity between cell j and all cells k belonging to cell type p from the similarity matrix S. jp :

[0174] G jp ={S jk};k∈Ω p ,k≠j

[0175] in,

[0176] k is the set Ω p and Ω q Cell numbering, k = 1, 2, ..., and

[0177] k≠i and k≠j exclude the diagonal elements of the similarity matrix S; and

[0178] [iii] The similarity subset G iq and similarity subset G jp Merge into a merged similarity subset G ij:

[0179] G ij =G iq ∪G jp

[0180] Merging similarity subsets G ij The average similarity between cells belonging to cell type p and cells belonging to cell type q is obtained by averaging all similarities. T represents the cell type-specific similarity between cell i and cell j. ij The cell type specific similarity T ij The cell type-specific similarity matrix T is formed by N×N.

[0181] In one implementation, the similarity S obtained in step (b) can be... ij Calibration is performed to eliminate measurement noise. In a further embodiment, measurement noise can be eliminated as follows:

[0182] (b1) Regarding the similarity S ij Perform the following iterative correction:

[0183] S ij =Z i S ij Z j

[0184] in,

[0185] = is the assignment operator.

[0186] Z i The average methylation fraction A of cell i ix The noise factor is initialized as the reciprocal of the square root of the maximum value in the i-th row of the similarity matrix S, excluding the values ​​in the diagonal cells:

[0187] And perturb at each loop.

[0188] Z j The average methylation fraction A of cell j jx The noise factor is initialized as the reciprocal of the square root of the maximum value in the j-th column of the similarity matrix S, excluding the values ​​in the diagonal cells:

[0189] And perturb at each loop.

[0190] The Z i and Z j Composition vector

[0191] (b2) For the corrected similarity matrix S obtained in each iteration, calculate the ratio of the standard deviation to the mean of all its elements except the diagonal elements as the cost Cost(S); and

[0192] (b3) The loop ends when Cost(S) is minimized, and the corrected similarity matrix S at this time is used as the final corrected similarity matrix S.

[0193] In a further implementation, the vector can be... Calibration is performed as long as cell lineage reconstruction can be achieved, and the calibration vector is adjusted accordingly. The method is not particularly limited. In one particular implementation, the vector Each element is available average Perform calibration and reassignment

[0194] = is the assignment operator.

[0195] In a further embodiment, the cell lineage-specific similarity matrix L can be obtained by subtracting the cell type-specific similarity matrix T from the corrected similarity matrix S: L = ST. In a further embodiment, the cell type-specific similarity matrix T can be prepared as follows:

[0196] [i] Define set Ω p and Ω q :

[0197] Ω p It is the set of cells belonging to cell type p among the plurality of cells, and

[0198] Ω q It is the set of cells belonging to cell type q among the multiple cells.

[0199] Where p and q are the same or different;

[0200] [ii] When cell i belongs to cell type p and cell j belongs to cell type q.

[0201] The similarity subset G is formed by extracting the similarity between cell i and all cells k belonging to cell type q from the similarity matrix S. iq :

[0202] G iq ={S ik};k∈Ω q ,k≠i

[0203] The similarity subset G is formed by extracting the similarity between cell j and all cells k belonging to cell type p from the similarity matrix S. jp :

[0204] G jp ={S jk};k∈Ω p ,k≠j

[0205] in,

[0206] k is the set Ω p and Ω q Cell numbering, k = 1, 2, ..., and

[0207] k≠i and k≠j exclude the diagonal elements of the similarity matrix S; and

[0208] [iii] The similarity subset G iq and similarity subset G jp Merge into a merged similarity subset G ij :

[0209] G ij =G iq ∪G jp

[0210] Merging similarity subsets G ij The average similarity between cells belonging to cell type p and cells belonging to cell type q is obtained by averaging all similarities. T represents the cell type-specific similarity between cell i and cell j. ij The cell type specific similarity T ij The cell type-specific similarity matrix T is formed by N×N.

[0211] In a further embodiment of the above-described implementation, the total number of CpG sites detected in multiple cells differentiated from a common ancestor in step (a) is approximately 1 to 10. 1 10 2 10 3 10 4 10 5 10 6 Or 10 7 More than one.

[0212] In a further embodiment of the above embodiments, the DNA methylation status information includes: the location and number of CpG sites on the genome, and whether each CpG site is methylated or not. In a further embodiment, whether a CpG site is methylated or not includes: detecting that the CpG site is methylated, detecting that the CpG site is not methylated, and not knowing whether the CpG site is methylated (NaN).

[0213] In a further embodiment of the above-described embodiments, as long as cell lineage reconstruction can be achieved, the method of obtaining nuclear genome sequence data containing DNA methylation status information is not particularly limited, and any method capable of measuring single-cell DNA methylation information can be used. In a further embodiment, nuclear genome sequence data containing DNA methylation status information can be obtained by single-cell bisulfite sequencing.

[0214] In a further embodiment of the above-described embodiments, the size of the genomic segment in step (a) can be 2bp–2000bp, 10bp–2000bp, 100bp–1500bp, 200bp–1000bp, 300bp–800bp, 400bp–600bp, or 500bp. In a further embodiment, each genomic segment can contain at least one CpG site. In a further embodiment, the genomic segments can be discretely distributed, with a segment size of 2bp, and each segment can correspond to one CpG site. The advantage of this genome partitioning is that it eliminates the need for manual selection of genomic regions and achieves very high accuracy in lineage reconstruction. The disadvantage is that there are approximately 2.9 × 10⁻⁶ CpG sites in human cells. 7 This results in a large computational load. In another further embodiment, the genomic segments can be discrete or continuous without overlap, the size of the genomic segments can be 10–2000 bp, and each genomic segment can contain at least one CpG site. In another further embodiment, the genomic segments can be continuous without overlap, and the size of the genomic segments can be 100–1500 bp.

[0215] In a further embodiment of the above-described embodiments, the calculation of the average methylation fraction A within the genomic segment is not particularly limited. In a further embodiment, the methylation fraction m of each CpG site can be calculated as follows:

[0216] The methylation fraction m of CpG sites within each genomic segment in the set of genomic segments W is calculated as the average of the methylation fraction m of all CpG sites that are not unknown, and is taken as the average methylation fraction A within that genomic segment.

[0217] In a further embodiment of the above implementation, the method may further include the following step between (a1) and (a2):

[0218] Calculate the average methylation rate of each genomic region across all said multiple cells.

[0219] For each of the plurality of cells, select a genomic segment that satisfies the following conditions [a] and [b]:

[0220] [a] Average methylation rate in all said multiple cells At the pre-set lower limit and upper limit Between, that is and

[0221] [b]The genomic region selected according to the above conditions [a] is present in more than two of the multiple cells. Not unknown,

[0222] This results in a set W of genomic segments for each cell, consisting of genomic segments that satisfy the above conditions [a] and [b].

[0223] In a further embodiment of the above-described embodiments, It can be selected from 0 to 0.9, for example, 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, or 0.9; and The value can be selected from 0.1 to 1, for example 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or 1. It can be 0 to 1, or unknown (i.e., NaN).

[0224] In a further embodiment of the above implementation, condition [b] may be: the genomic segment selected in condition [a] accounts for more than 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of all cells in all cells intended to reconstruct the lineage. It is not unknown (not NaN).

[0225] In a further embodiment of the above implementation, a set of merged genomic segments W′, consisting of merged genomic segments that have merged different adjacent genomic segments, can be used for cell lineage reconstruction.

[0226] In a further embodiment of the above embodiments, the methylation fraction m of the CpG site may be 0 to 1, or unknown (i.e., NaN).

[0227] In a further embodiment of the above implementation, the method for calculating similarity is not particularly limited as long as cell lineage reconstruction can be achieved. In a preferred embodiment, the similarity can be selected from: Pearson correlation coefficient, cosine similarity, and Jaccard similarity.

[0228] In a further embodiment of the above implementation, step (c) can be performed using any algorithm for reconstructing the spectral system based on the distance matrix. In a preferred embodiment, the algorithm for reconstructing the spectral system based on the distance matrix can be selected from: unweighted pair-group method with arithmetic means (UPGMA), Neighbor-joining (NJ), and FastME.

[0229] UPGMA assumes a constant rate of change in DNA methylation state, and therefore assumes that outlying branches branching from the same point are of uniform length. The specific process is as follows:

[0230] (a) Find the most similar (D ij If two cells (i≠j) are the smallest, then consider that they are equal in distance to their common branch point and equal to half the distance between them. Then, group the two most similar cells together and treat them as one cell.

[0231] (b) Calculate the distances between the remaining cells and the new node obtained in (a); and

[0232] (c) Repeat steps (a) and (b) above until all cells come together to obtain a complete phylogenetic tree.

[0233] NJ acknowledges that the rate at which DNA methylation states change is not constant, and the specific process is as follows:

[0234] (a) Find the most similar (D ij (i≠j) the smallest two cells, and connect the two points to a new intermediate node (u);

[0235] (b) Calculate the distances between the remaining cells and the new node obtained in (a);

[0236] (c) Repeat steps (a) and (b) above until all cells are connected to the star network.

[0237] The FastME algorithm is similar to the NJ algorithm, but with improvements. For details, see: http: / / www.atgc-montpellier.fr / fastme / .

[0238] In a further embodiment of the above-described embodiments, the cell can be any cell. In a further embodiment, the cell is selected from prokaryotic cells and eukaryotic cells. In a further embodiment, the cell is selected from somatic cells and germ cells. In a further embodiment, the cell is an embryonic cell. In a further embodiment, the cell is a cell derived from stem cells or progenitor cells.

[0239] In another embodiment, the present invention provides a recording medium for a computer program that reconstructs lineages of multiple cells differentiated from a common ancestor, wherein the computer program causes a computer to perform the methods of the present invention as defined above.

[0240] In another embodiment, the present invention provides an apparatus for reconstructing lineages of multiple cells differentiated from a common ancestor, comprising: a processing unit that performs the method of the present invention as defined above, and a display unit that displays the lineages of the multiple cells reconstructed by the processing unit.

[0241]

Example

[0242] [Example 1: Lineage Reconstruction of Cells Derived from HEK293T Cells]

[0243] (1.1) Cloning of HEK293T cells

[0244] Multiple HEK293T cells were taken from single-cell colonies of HEK293T cells, and each HEK293T cell was expanded from a distant lineage into a large clone with approximately 20,000 cells. To increase lineage complexity, two cells were randomly selected from one of the clones and each was expanded into a subclone with approximately 20,000 cells. These single clones were dissociated, and the dissociated cells were sorted at 1 cell / well into each well of a flat-bottomed 96-well plate pre-filled with growth medium (DMEM + 10% FBS) (Thermofisher) and cultured for 10 days for single-clone expansion. For each amplified subclone, cells in each well were dissociated by treating with 0.25% trypsin (Thermofisher, 25200072) at 37°C for 5 minutes. Single cells from each subclone were then sorted at 1 cell / well into each well of a flat-bottomed 96-well plate using FACS for further processing.

[0245] (1.2) Single-cell bisulfite sequencing (scBS-seq)

[0246] The cells obtained in step (1.1) were directly sorted into 96-well plates containing Buffer RLT Plus (Qiagen, 1053393), and their DNA was converted to bisulfite using the EZ-96 DNA Methylation Direct MagPrep Kit (ZYMO, D5045). DNA was extracted from these cells, amplified, and purified. Then, scBS-seq library construction was performed, and 150bp sequences were sequenced from both ends using the Illumina platform to obtain scBS-seq sequence data (nuclear genome sequence data with DNA methylation status information), which was used for the following computer processing.

[0247] (1.3) Preprocessing of scBS-seq sequence data

[0248] Bismark (version 0.24.0) was used to read, align, remove duplicate data, and extract methylation information of CpG sites from the scBS-seq sequence data obtained in step (1.2).

[0249] It is known that there are approximately 2.9 × 10⁻⁶ cells in the human genome. R There are 1 CpG site, and the methylation fraction m for each CpG site is calculated as follows:

[0250] m∈(0,1), where 0 indicates that the CpG site is unmethylated and 1 indicates that the CpG site is methylated. The closer m is to 1, the greater the probability that the CpG site is methylated.

[0251] Because scBS-seq can only cover a portion of the genome (about 5%) and the experimental sequencing depth is relatively shallow, about 95% of the genome region in each cell remains undetected, and the methylation fraction of CpG sites in these undetected genome regions is unknown.

[0252] When the methylation fraction m of a certain CpG site is unknown, it is represented as m = NaN, and its true value can be any value between 0 and 1.

[0253] Selecting a total number of CpG sites ≥ 10 5 The scBS-seq sequence data of the cells were used for further processing.

[0254] (1.4) Genome division and selection of genomic regions

[0255] The genome was divided using two different methods (1.4a) and (1.4b).

[0256] (1.4a) Discretely distributed genomic segments with each CpG site / 2bp genomic segment

[0257] For all cells selected in (1.3) (hereinafter referred to as "all selected cells"), the nuclear genome sequence was divided into discrete genomic segments of 2 bp in length, with each genomic segment corresponding to a CpG site. This yielded a sequence corresponding to all CpG sites (approximately 2.9 × 10⁻⁶). 7 W is a set of discrete genomic regions consisting of genomic regions.

[0258] (1.4b) 500bp / genomic region of continuous non-overlapping genomic region

[0259] For all cells selected in (1.3) (hereinafter referred to as "all selected cells"), the nuclear genome sequence was divided into continuous, non-overlapping genomic segments of 500 bp in unit length, and then the average methylation rate of each genomic segment was calculated across all selected cells.

[0260] Where 0 indicates that there is no methylation in the genome segment, and 1 indicates that all CpG sites in the genome segment are methylated. The closer a value is to 1, the higher the methylation level within that genomic segment. The average methylation rate within a given genomic segment... When unknown, it is represented as Its true value can be any value between 0 and 1.

[0261] For each cell, select genomic regions that meet the following conditions [a] and [b]:

[0262] [a] Average methylation rate At the pre-set lower limit and upper limit Between, that is (Figure 1A), where and

[0263] [b]The genomic regions selected based on the above conditions [a] comprise more than 10% of all selected cells. Not unknown (with observed values), that is

[0264] This results in a set W of genomic segments for each cell, consisting of genomic segments that satisfy the above conditions [a] and [b].

[0265] Alternatively, adjacent genomic segments that satisfy the above conditions [a] and [b] can be merged to obtain a set of merged genomic segments W′ (Figure 1B) consisting of merged genomic segments that have merged different adjacent genomic segments.

[0266] (1.5) Calculation of the average methylation fraction within genomic segments / merged genomic segments

[0267] For each cell i (i = 1, 2, ..., N), within each genomic segment / merged genomic segment x obtained in step (1.4), identify all CpG sites whose methylation fraction m is not NaN. Then, calculate the average methylation fraction m of all CpG sites whose methylation fraction m is not NaN for each genomic segment / merged genomic segment x, and use this average methylation fraction A within that genomic segment / merged genomic segment x (x = 1, 2, ..., M). ix .

[0268] A ix ∈(0,1), where 0 indicates that there is no methylation within the genomic segment / merged genomic segment x, and 1 indicates that all CpG sites within the genomic segment / merged genomic segment x are methylated. A ix The closer a value is to 1, the higher the methylation level of cell i within that genomic segment / merged genomic segment x. When the methylation fraction m of all CpG sites within a certain genomic segment / merged genomic segment x of cell i is unknown, the average methylation fraction A within that genomic segment / merged genomic segment x is... ix Also unknown, represented as A ix =NaN, whose actual value can be any value between 0 and 1.

[0269] (1.6) Similarity Calculation

[0270] To calculate the similarity S between cell i (i = 1, 2, ..., N) and cell j (j = 1, 2, ..., N) ij The average methylation fraction of the two cells in the genomic segment set W or the merged genomic segment set W′ obtained in step (1.4) is not NaN (i.e., A). ix ≠NaN and A jx The set W of all genomic regions / merged genomic regions (≠NaN) constitutes the common genomic region / merged genomic region set. ij That is, x∈W ij .

[0271] Only in the shared genomic region / merged genomic region set W ij The mean methylation fraction A of cell i is calculated internally. ix and the average methylation fraction A of cell j jx The Pearson similarity between cells, as the similarity S between cells i and j. ij :

[0272] Using the similarity S obtained above ij Construct an N×N similarity matrix S.

[0273] (1.7) Removal of measurement noise

[0274] There may be a discrepancy between the measured and actual values ​​of methylation fraction, known as measurement noise. The greater the measurement noise, the smaller the calculated Pearson correlation coefficient will be compared to the actual Pearson correlation coefficient, a phenomenon statistically termed "attenuation." If the standard deviation (also called the "noise factor," denoted as Z) between the measured and actual values ​​of methylation fraction is the same across all cells, the relative similarity will remain constant, ensuring that the cell lineage determined based on similarity is correct. However, if the noise factor differs significantly between different cells (e.g., due to heterogeneity in library preparation and sequencing), the relative similarity will be distorted, leading to errors in cell lineage reconstruction. Therefore, the applicant has developed the following method to correct for similarity.

[0275] (1.7.1) Core Iterative Correction Algorithm

[0276] The average methylation fraction A of cell i ix noise factor Z i Initialize to the reciprocal of the square root of the maximum value in the i-th row of the similarity matrix S created in step (1.5), excluding the values ​​in the diagonal cells:

[0277] The average methylation fraction A of cell j jx noise factor Z j Initialize to the reciprocal of the square root of the maximum value in the j-th column of the similarity matrix S created in step (1.5), excluding the values ​​in the diagonal cells:

[0278] The above Z i and Z j Composition vector

[0279] Will Each element is used average Perform calibration: Corrected vector Still named This is the current noise factor.

[0280] The similarity S between cell i and cell j is corrected as follows. ij The corrected similarity is obtained.

[0281] For the corrected similarity matrix S * The evaluation is performed, and the result is expressed using the cost function Cost(S). * Let S be represented as ). First, let matrix S be represented as ). * After removing the diagonal elements, the remaining off-diagonal elements are flattened into a vector.

[0282] Then calculate the mean and variance of all elements in this vector, and the final cost value is:

[0283] By perturbing the existing noise factor This minimizes the Cost corresponding to the corrected S, and uses S at this point. * This serves as the final corrected similarity matrix. The iterative process is as follows:

[0284] [initialization]

[0285] Iteration boosting factor: Δ = 1

[0286] Current iteration count: n = 0

[0287] Maximum number of iterations: n_max = 1000 # The loop terminates when the maximum number of iterations is reached.

[0288] Current noise factor:

[0289] Current similarity matrix: S 0 =S

[0290] [Loop Iteration]

[0291] while Δ>0:

[0292] n = n + 1

[0293] Record the current noise factor Corresponding corrected descendant value: Cost n

[0294] Initialize a correction factor of length N:

[0295] Execute the following loop to update each element of the correction factor one by one:

[0296] for j in [1,2,…,N]:

[0297] [1] Define a temporary noise factor Initialize to the noise factor obtained in the previous iteration. But add a small perturbation ∈ to the j-th element:

[0298] [2] will Normalization And the normalized vector is still called

[0299] [3] This temporary noise factor Used to correct the original similarity matrix S 0 Calculate the cost corresponding to the corrected similarity matrix. And record it with Cost n The difference:

[0300] After completing the above cycle, the corresponding correction factor for each cell is obtained.

[0301] According to the correction factor Update noise factor

[0302] in It is the absolute value of the element with the largest absolute value in the correction factor vector.

[0303] Will Normalization And the normalized vector is still called Used to correct the similarity matrix S:

[0304] Finally, update S n+7 Cost n+7 And calculate the iteration improvement factor:

[0305] Δ = Cost n -Cost n+1

[0306] The final output similarity matrix S n+7 This is the final corrected similarity matrix, renamed S.

[0307] (1.7.2) Further iterations

[0308] Using the corrected similarity matrix S obtained in (1.7.1) as new input, and applying the iterative correction algorithm in (1.7.1) again, the correction effect can be further improved. This process can continue iterating until matrix S shows no difference before and after correction. The resulting similarity matrix S can then be used for the next step of genealogical reconstruction.

[0309] [initialization]

[0310] Initialize convergence factor

[0311] Current iteration count: n = 0

[0312] Initialize S n =S

[0313] [Loop Iteration]

[0314] while

[0315] [1] with S n As input, the iterative correction algorithm of (1.7.1) is run to obtain the output S. n+1 .

[0316] [2] Calculate convergence:

[0317] Let matrix e ij Expand the difference into a vector and calculate the standard deviation of this difference vector as a measure of convergence.

[0318] (1.8) Cell lineage reconstruction

[0319] Cell lineages were reconstructed using the classic unweighted pair-group method with arithmetic means (UPGMA).

[0320] The maximum value S is selected from the corrected similarity S matrix prepared in step (1.7). max and minimum value S min And calculate the phylogenetic distance D between cell i and cell j as follows. ij =S max -S ij -S min Let D be the diagonal element of the matrix. ii =0.

[0321] The upgma function from the public biotite algorithm package is used to construct the phylogenetic tree Γ from matrix D. o :biotite.sequence.phylo.upgma(D).

[0322] (1.9) Verification of the phylogenetic tree

[0323] (1.9.1) Reliability of the phylogenetic tree

[0324] The following verifies the phylogenetic tree Γ created in step (1.8) above. o Reliability:

[0325] [1.9.1.1] Randomly select 80% of the genomic segment set W or the merged genomic segment set W′ obtained from step (1.4) to form a sub-window set, and perform steps (1.5) to (1.8) to obtain a new phylogenetic tree Γ. s .

[0326] [1.9.1.2] Perform the above step [1.9.1.1] 100 times to obtain 100 sampling trees {Γ s};

[0327] [1.9.1.3] For Γ o Each subtree Traverse the {Γ} obtained in step [1.9.1.2] above. s}, and check if it is possible in Γ s Find similar subtrees The term "similar" means that two subtrees are composed of the same set of cells, regardless of their structure and arrangement within the respective subtrees.

[0328] [1.9.1.4] Subtree Support is defined as the sum of all 100 sampled trees {Γ} s} found with The proportion of similar subtrees (not multiplied by 100, so its value is between 0 and 1), is displayed. Near the roots of the tree. This subtree Support reflects the likelihood that cells belonging to the corresponding subtree cluster together in the actual phylogenetic tree. A support close to 1 indicates a higher likelihood that cells belonging to the corresponding subtree cluster together in the actual phylogenetic tree, and thus indicates a higher likelihood that the phylogenetic tree Γ created in step (1.8) above is clustered together. o The higher the reliability.

[0329] [1.9.1.5] The cells involved in the largest subtree with support above a given threshold α are defined as a clone. The clone is determined as follows:

[0330] [1.9.1.5.1] Start traversing and searching from the root node.

[0331] [1.9.1.5.2] If the current node is not a leaf node (the last node) and its support is less than α, further search all child nodes of the current node.

[0332] [1.9.1.5.3] If the current node is not a leaf node and its support is ≥ α, then identify the cells involved in the subtree rooted at the current node as a clone. Stop searching the current node and all its child nodes, and start searching all other nodes.

[0333] [1.9.1.5.4] If the current node is already a leaf node, then search the remaining nodes.

[0334] [1.9.1.5.5] Repeat steps [1.9.1.5.2] to [1.9.1.5.4] until there are no more nodes left to search.

[0335] (1.9.2) Dimensionality Reduction

[0336] [1.9.2.1] The similarity matrix S obtained in step (1.6) or (1.7) is reduced in dimensionality using the public function sklearn.manifold.spectral_embedding. The parameter involved is n_components.

[0337] [1.9.2.2] Use the scanpy.pp.neighbors function to generate a k-NN graph from the spectral components obtained in [1.9.2.1]. The parameter involved is n_neighbors.

[0338] [1.9.2.3] Use the scanpy.tl.umap function to generate a two-dimensional visualization projection. This involves the parameter min_dist.

[0339] The parameters mentioned in [1.9.2.1] to [1.9.2.3] are represented as the following ternary parameter group: (n_components, n_neighbor, min_dist), and the parameter group used in this data is (7,7,0.7).

[0340] (1.9.3) Accuracy of the phylogenetic tree

[0341] The accuracy Q of the phylogenetic tree is calculated using the following formula:

[0342] In the above formula, g k n is the maximum number of cells in clone k that cluster together according to the inferred phylogenetic order. k is the total number of cells in clone k, and M is the total number of clones.

[0343] In the above formula, the reason for starting from g k and n k Subtracting 1 from the middle ensures that no two cells of clone k are adjacent in the inferred pedigree (i.e., g). k =1), and the corresponding accuracy will become 0. This evaluation only used clones with at least 2 cells, excluding cells where no clonal marker was observed.

[0344] The accuracy Q of the phylogenetic tree calculated above is between 0 and 1. When all cells are grouped according to their clonal origin, Q = 1.

[0345] (1.10) Results

[0346] Figure 1C shows the similarity matrix between cells from the same subclone derived from HEK293T cells when the genome was divided into discrete segments of 2 bp per CpG site (1.4a). Figure 1D shows the similarity matrix between cells from the same subclone derived from HEK293T cells when the genome was divided into continuous, non-overlapping segments of 500 bp per segment (1.4b). Regardless of the genome division method, cells from the same subclone were correctly clustered together (Figure 1E), and the closer phylogenetic relationship between P9_1 and P10_1 was also correctly inferred (Figure 1F).

[0347] In step (1.4b), all genomic segments that have not been merged into adjacent genomic segments are selected. When m1=1), the accuracy of cell lineage reconstruction is 88% (Figure 1G), which is much higher than the accuracy of lineage reconstruction using randomly shuffled cell sorting (17%).

[0348] In step (1.4b), all merged genomic segments that have been merged through adjacent genomic segment merging are selected. At that time, the accuracy of cell lineage reconstruction was 91%, while when other methods were used... and Value, such as At that time, the accuracy of cell lineage reconstruction reached 100% (Figure 1H).

[0349] A subset of merged genomic regions (wherein) are randomly selected from the merged genomic region set W′. For downstream calculations, when more than 20% of the merged genomic segments in the merged genomic segment set W′ are selected, the accuracy of cell lineage reconstruction can approach 100% (Figure 1I).

[0350] The more merged genomic regions selected from the merged genomic region set W′, the higher the accuracy of cell lineage reconstruction (Figure 1J).

[0351] Compared to not merging adjacent genomic regions, merging adjacent genomic regions can not only reduce computational load but also potentially improve the accuracy of cell lineage reconstruction (Figures 1G and 1H).

[0352] The accuracy of cell lineage reconstruction based on corrected similarity (1.00) for cells derived from HEK293T cells was significantly higher than that based on uncorrected similarity (0.84) (Figure 1K).

[0353] These results demonstrate that the method of the present invention can efficiently reconstruct cell lineages from multiple cells differentiated from a common ancestor.

[0354] [Example 2: Lineage Reconstruction of Cells Derived from H9 Human Embryonic Stem Cells]

[0355] (2.1) Monoclonal expansion of H9 human embryonic stem cells

[0356] H9 human embryonic stem cells were dissociated at 37°C using ACCUTASE (STEMCELL, 07920) for 10 minutes. Single cells were then cultured and expanded into clones on the wells of a Methylgel substrate plate (Corning, 354277) in mTeSR1 basal medium (STEMCELL) containing 20% ​​mTeSR1 5× additive and 10% CloneR2 (Figure 2A).

[0357] (2.2) Genealogical Reconstruction

[0358] As described in steps (1.2) to (1.9) of Example 1, the lineage of cells derived from H9 human embryonic stem cells was reconstructed and verified.

[0359] (2.3) Results

[0360] In step (1.4), all merged genomic segments that have been merged through adjacent genomic segment merging are selected. At that time, the accuracy of cell lineage reconstruction reached 62%, while... and Near the same location, the accuracy of cell lineage reconstruction is close to 100% (Figure 2B).

[0361] The accuracy of cell lineage reconstruction derived from H9 human embryonic stem cells based on corrected similarity (1.000) was significantly higher than that based on uncorrected similarity (0.867) (Figure 2C).

[0362] These results demonstrate that the method of the present invention can efficiently reconstruct cell lineages from multiple cells differentiated from a common ancestor.

[0363] [Example 3: Lineage Reconstruction of Cells Derived from Mouse Hematopoietic Progenitor Cells]

[0364] (3.1) Isolation of mouse hematopoietic progenitor cells

[0365] After euthanasia, bone marrow from the femur, tibia, pelvis, and sternum was isolated by pulverizing the cells using a pestle and mortar. The collected bone marrow cells were filtered through a 40 μm filter and washed in cold EasySep buffer (STEMCELL, 20144). Red blood cells and mature lineage cells were magnetically removed using the EasySep Mouse Hematopoietic Progenitor Cell Isolation Kit (STEMCELL, 19856). The resulting Lin... - The components were stained with cKit (CD117-PE, clone 2B8, Biolegend) and Sca-1 (Ly6a-FITC, clone D7, Biolegend), and Lin were isolated by fluorescence-activated cell sorting (FACS) using a 130 μM nozzle on a Sony MA900. - cKit + Sca1 - (LK) cells.

[0366] Sorted LK cells were barcoded using spin infection (800g, 90 min) in LARRY lentiviral concentrate containing polybrene (Sigma), and then seeded in parallel at 1,000–1,500 cells per well in 9-well round-bottom 96-well plates containing medium designed to support panmyeloid differentiation (including StemSpan SFEM medium (STEMCELL, 09650), IL-3 (20 ng / mL), FLT3-L (50 ng / mL), IL-11 (50 ng / mL), IL-5 (10 ng / mL), TPO (50 ng / mL) (Peprotech), and mSCF (50 ng / mL) and IL-6 (10 ng / mL)) (R&D Systems). After 6 days of cell culture, cells in each well were detached by treatment with 0.25% trypsin (Thermofisher) for 5 min. Single GFP cells were then... + Cells were directly sorted into 96-well plates containing mild lysis buffer using FACS. Nuclei were captured using Dynabeads Myone Carboxylic Acid (Invitrogen, 65011), and the supernatant containing the released RNA was transferred to individual 96-well plates.

[0367] (3.2) Improved Camellia-seq

[0368] scBS-seq was performed using Dynabeads containing genomic DNA. The RNA portion was partially reverse transcribed and amplified for 15 cycles. The resulting cDNA was homogenized, with half used for scSTRT-seq to obtain a single-cell transcriptome, and the other half amplified using Larry primer sequences to generate Larry lineage barcodes.

[0369] (3.3) Genealogical Reconstruction

[0370] As described in steps (1.3) to (1.9) of Example 1, the lineage of cells derived from mouse hematopoietic progenitor cells was reconstructed and verified.

[0371] (3.4) Results

[0372] In the transcriptome data, a total of eight different cell types were observed: megakaryocytes (Mk), erythrocytes (Er), basophils (Ba), mast cells (Ma), eosinophils (Eos), neutrophils (Neu), neutrophil-like monocytes (Neu-Mon), and dendritic monocytes (Dc-Mon) (Figure 3A).

[0373] In the LARRY lineage data, a total of 75 LARRY lineage barcodes were detected, with 52 observed in ≥2 cells and 23 found in ≥2 cell types (Figures 3B and 3C). This barcode information reveals the actual lineage of the cells. The similarity matrix calculated using scBS-seq data and the inferred cell lineage based on it reproduced the lineage results observed by the LARRY barcodes, achieving 100% prediction lineage accuracy (Figures 3D and 3E). It is particularly noteworthy that this calculation process did not remove cell type-specific methylation information. Therefore, the cell similarity matrix based on DNA methylation primarily displays cell lineage information, rather than cell type information. This indicates that although cell type-specific DNA methylation patterns were indeed found near key marker genes of hematopoietic cells, including Spi1 which encodes the transcription factor PU.1 that regulates myeloid fate selection (Fig. 3F), the cell type-related DNA methylation changes that occurred during 6 days of culture were relatively small for cells derived from hematopoietic progenitor cells compared to genome-wide epigenetic mutations accumulated during development and adulthood.

[0374] These results demonstrate that the method of the present invention can efficiently reconstruct cell lineages from multiple cells that have differentiated from a common ancestor and have various cell types.

[0375] [Example 4: Lineage Reconstruction of Cells from Human Embryos]

[0376] (4.1) Obtaining scBS-seq sequence data from human embryonic cells

[0377] Multiple scBS-seq data from somatic cells and fetal germ cell (FGC) of human embryos at 7 and 21 weeks of age were obtained from entry HRA000186 in the Human Genome Sequence Archive of the National Genome Science Data Center (https: / / ngdc.cncb.ac.cn / gsa-human) (see also the “Data availability” section of Li et al., “Dissecting the epigenomic dynamics of human fetal germ cell development at single-cell resolution”, Cell Research, volume 31, pages 463-477 (2021) (https: / / www.nature.com / articles / s41422-020-00401-9).

[0378] (4.2) Similarity Calculation

[0379] As described in steps (1.3) to (1.7) of Example 1, preprocess the scBS-seq sequence data, divide the genome and select genome segments, calculate similarity and remove measurement noise.

[0380] (4.3) Removal of cell type-specific methylation signals

[0381] In theory, regardless of whether the two cells (cell i and cell j) whose similarity is being calculated belong to the same cell type or different cell types, the calculated similarity S between the two cells will remain the same. ij The middle part must include cell type specific similarity T ij Similarity to cell lineage specificity L ij S ij =T ij +L ij .

[0382] Although sometimes cell type specific similarity T ij (Whether the two cells (cell i and cell j) belong to the same cell type or different cell types) the similarity is so small as to be negligible (not substantially affecting the accuracy of lineage reconstruction), but sometimes (such as in comparisons between somatic cells and FGCs) it cannot be ignored, requiring consideration of the similarity S between the two cells. ij Subtracting cell type specific similarity T ij And obtain the cell lineage specific similarity L ij Only then can the cell lineage be reconstructed more accurately.

[0383] In this experiment, the cell type-specific similarity T was calculated as follows. ij .

[0384] Define set Ω p and Ω q :

[0385] ●Ω p It is the set of all selected cells that belong to cell type p, and

[0386] ●Ω q It is the set of all selected cells that belong to cell type q.

[0387] p and q may be the same or different.

[0388] When cell i belongs to cell type p and cell j belongs to cell type q

[0389] ● Extract the similarity scores between cell i and all cells k belonging to cell type q from the similarity matrix S to form a similarity subset G. iq :

[0390] G iq ={S ik};k∈Ω q ,k≠i

[0391] ● Extract the similarity scores between cell j and all cells k belonging to cell type p from the similarity matrix S to form a similarity subset G. jp :

[0392] G jp ={S jk};k∈Ω p ,k≠j

[0393] in,

[0394] k is the set Ω p and Ω q Cell numbering, k = 1, 2, ..., and

[0395] k≠i and k≠j exclude the diagonal elements in the similarity matrix S.

[0396] Similarity subset G iq and similarity subset G jp Merge into a merged similarity subset G ij :

[0397] G ij =G iq ∪G jp

[0398] Merging similarity subsets G ij The average similarity between cells belonging to cell type p and cells belonging to cell type q is obtained by averaging all similarities. T represents the cell type-specific similarity between cell i and cell j. ij .

[0399] The cell type specific similarity T obtained above ij An N×N cell type-specific similarity matrix T is constructed, and the cell lineage-specific similarity matrix L is obtained by subtracting the cell type-specific similarity matrix T from the similarity matrix S: L = ST.

[0400] (4.4) Cell lineage reconstruction

[0401] Using the cell lineage-specific similarity matrix L obtained in (4.3), the lineages of somatic cells and FGCs derived from 7-week-old and 21-week-old human embryos were reconstructed and verified as described in steps (1.8) to (1.9) of Example 1.

[0402] (4.5) Results

[0403] The similarity matrix from cells from 21-week-old human embryos without removal of cell type-specific methylation signals as described in step (4.3) clearly shows that somatic cells and FGCs are distinctly distinguished by cell type (Fig. 4A).

[0404] The cell lineage-specific similarity matrix, obtained from cells derived from 21-week-old human embryos after removal of cell type-specific methylation signals as described in step (4.3), clearly shows that cells are clearly grouped according to their embryonic origin, and in order to... When selecting merged genomic segments that have been merged from adjacent genomic segments, the accuracy of cell lineage reconstruction is close to 100% (Figure 4B).

[0405] Approximate results can also be seen from the similarity matrix between cells from 7-week-old human embryos and the cell lineage-specific similarity matrix (Figure 4C).

[0406] When there is no uniform noise factor among cells, the corrected similarity matrix better reflects the lineage of cells (FGCs and somatic cells) (Figs. 4D and 4E).

[0407] These results demonstrate that the method of the present invention can efficiently and directly reconstruct the lineage between human cells in individuals, such as humans, where gene editing is not possible.

Claims

1. A method for reconstructing lineages from multiple cells differentiated from a common ancestor, comprising: (a) Nuclear genome sequence data for each cell of multiple cells differentiated from a common ancestor, containing information on DNA methylation status. (a1) The cell nuclear genome sequence is divided into discrete or continuous non-overlapping genome segments of a certain unit length. All the divided genome segments of each cell constitute the genome segment set W of that cell, and (a2) Calculate the average methylation fraction A within each genomic segment in the set of genomic segments W; (b) The set of genomic segments W consists of all genomic segments in cells i and j of the plurality of cells in which the average methylation fraction A is not unknown, forming a total set of genomic segments W. ij Only in the shared genomic segment set W ij The average methylation fraction A within cell i of genomic region x is calculated. ix and the average methylation fraction A of cell j jx The similarity between cells, denoted as S, is the similarity between cells i and j. ij : in i is the cell number among the plurality of cells, i = 1, 2, ..., N. j is the cell number among the plurality of cells, j = 1, 2, ..., N. i≠j, N is the total number of cells in the plurality of cells. x is the set of shared genomic segments W ij The genomic segment numbers in the data, x = 1, 2, ..., M, and M is the set of shared genomic segments W ij The total number of genomic segments in the genome. The similarity S ij The similarity matrix S is formed by N×N; and (c) Reconstruct lineages of the plurality of cells using the similarity matrix S prepared in step (b).

2. The method according to claim 1, wherein the cell lineage-specific similarity matrix L is obtained by subtracting the cell type-specific similarity matrix T from the similarity matrix S: L = ST.

3. The method according to claim 2, wherein the cell type-specific similarity matrix T is prepared as follows: [i] Define set Ω p and Ω q : Ω p It is the set of cells belonging to cell type p among the plurality of cells, and Ω q It is the set of cells belonging to cell type q among the multiple cells. in, p and q may be the same or different; [ii] When cell i belongs to cell type p and cell j belongs to cell type q. The similarity subset G is formed by extracting the similarity between cell i and all cells k belonging to cell type q from the similarity matrix S. iq : G iq ={S ik };k∈Ω q ,k≠i The similarity subset G is formed by extracting the similarity between cell j and all cells k belonging to cell type p from the similarity matrix S. jp : G jp ={S jk };k∈Ω p ,k≠j in, k is the set Ω p and Ω q Cell numbering, k = 1, 2, ..., and k≠i and k≠j exclude the diagonal elements of the similarity matrix S; and [iii] The similarity subset G iq and similarity subset G jp Merge into a merged similarity subset G ij : G ij =G iq ∪G jp Merging similarity subsets G ij The average of all similarities is used to obtain the cells and genus belonging to cell type p. Average similarity between cells of cell type q T represents the cell type-specific similarity between cell i and cell j. ij The cell type specific similarity T ij The cell type-specific similarity matrix T is formed by N×N.

4. The method according to claim 1, wherein (b1) Regarding the similarity S ij Perform the following iterative correction: S ij =Z i S ij Z j in, = is the assignment operator. Z i The average methylation fraction A of cell i ix The noise factor is initialized as the reciprocal of the square root of the maximum value in the i-th row of the similarity matrix S, excluding the values ​​in the diagonal cells: And perturb at the beginning of each loop. Z j The average methylation fraction A of cell j jx The noise factor is initialized as the reciprocal of the square root of the maximum value in the j-th column of the similarity matrix S, excluding the values ​​in the diagonal cells: And perturb at the beginning of each loop. The Z i and Z j Composition vector (b2) For the corrected similarity matrix S obtained in each iteration, calculate the ratio of the standard deviation to the mean of all its elements except the diagonal elements as the cost Cost(S); and (b3) The loop ends when Cost(S) is minimized, and the corrected similarity matrix S at this time is used as the final corrected similarity matrix S.

5. The method according to claim 4, wherein the vector Each element is used average Perform calibration and reassignment in, = is the assignment operator.

6. The method according to claim 4, wherein the cell type-specific similarity matrix T is subtracted from the corrected similarity matrix S to obtain the cell lineage-specific similarity matrix L: L = ST.

7. The method of claim 6, wherein the cell type-specific similarity matrix T is prepared as follows: [i] Define set Ω p and Ω q : Ω p It is the set of cells belonging to cell type p among the plurality of cells, and Ω q It is the set of cells belonging to cell type q among the multiple cells. in, p and q may be the same or different; [ii] When cell i belongs to cell type p and cell j belongs to cell type q. The similarity subset G is formed by extracting the similarity between cell i and all cells k belonging to cell type q from the similarity matrix S. iq : G iq ={S ik };k∈Ω q ,k≠i The similarity subset G is formed by extracting the similarity between cell j and all cells k belonging to cell type p from the similarity matrix S. jp : G jp ={S jk };k∈Ω p ,k≠j in, k is the set Ω p and Ω q Cell numbering, k = 1, 2, ..., and k≠i and k≠j exclude the diagonal elements of the similarity matrix S; and [iii] The similarity subset G iq and similarity subset G jp Merge into a merged similarity subset G ij : G ij =G iq ∪G jp Merging similarity subsets G ij The average similarity between cells belonging to cell type p and cells belonging to cell type q is obtained by averaging all similarities. T represents the cell type-specific similarity between cell i and cell j. ij The cell type specific similarity T ij The cell type-specific similarity matrix T is formed by N×N.

8. The method according to any one of claims 1 to 7, wherein the DNA methylation status information includes: The location and number of CpG sites on the genome, and whether each CpG site is methylated.

9. The method of claim 8, wherein whether or not the CpG site is methylated comprises: Methylation of the CpG site was detected. The CpG site was detected to be unmethylated, and It is unknown whether the CpG site is methylated, i.e., NaN.

10. The method according to any one of claims 1 to 7, wherein the nuclear genome sequence data with DNA methylation status information in step (a) is obtained by single-cell bisulfite sequencing.

11. The method according to any one of claims 1 to 7, wherein in step (a1), the genomic segments are discretely distributed, the size of the genomic segments is 2 bp, and each genomic segment corresponds to a CpG site.

12. The method according to any one of claims 1 to 7, wherein in step (a1), the genomic segments are discrete or continuous and non-overlapping, the size of the genomic segments is 10 to 2000 bp, and each genomic segment contains at least one CpG site.

13. The method according to any one of claims 1 to 7, wherein in step (a1), the genomic segments are continuously distributed without overlap, and the size of the genomic segments is 100 to 1500 bp.

14. The method according to any one of claims 1 to 7, wherein in step (a2), the methylation fraction m of each CpG site is calculated: The methylation fraction m of CpG sites within each genomic segment in the set of genomic segments W is calculated as the average of the methylation fraction m of all CpG sites that are not unknown, and is taken as the average methylation fraction A within that genomic segment.

15. The method of claim 13, wherein between steps (a1) and (a2) the method further comprises: Calculate the average methylation rate of each genomic region across all said multiple cells. For each of the plurality of cells, select a genomic segment that satisfies the following conditions [a] and [b]: [a] Average methylation rate in all said multiple cells At the pre-set lower limit and upper limit Between, that is and [b]The genomic region selected according to the above conditions [a] is present in more than two of the multiple cells. Not unknown, This results in a set W of genomic segments for each cell, consisting of genomic segments that satisfy the above conditions [a] and [b].

16. The method of claim 15, wherein Selected from 0 to 0.9, Selected from 0.1 to 1, and yes: 0 to 1, or Unknown, i.e., NaN.

17. The method of claim 15, wherein condition [b] is: the genomic region selected in condition [a] is present in more than 5% of the plurality of cells. It is not unknown.

18. The method of claim 15, wherein The genomic segments are continuous and non-overlapping, and The set of genomic segments W is replaced by a set of merged genomic segments W′, which consists of merged genomic segments that have merged different adjacent genomic segments, obtained by merging adjacent genomic segments that satisfy the conditions [a] and [b].

19. The method according to any one of claims 1 to 7, wherein the methylation fraction m of the CpG site is: 0 to 1, or Unknown, i.e., NaN.

20. The method according to any one of claims 1 to 7, wherein the similarity is selected from: Pearson correlation coefficient, cosine similarity and Jaccard similarity.

21. The method according to any one of claims 1 to 7, wherein step (c) is performed using an algorithm selected from the following: unweighted pair group method with arithmetic means (UPGMA), Neighbor-joining (NJ), and FastME.

22. The method according to any one of claims 1 to 7, wherein after step (c) it further comprises a verification step of the reconstructed lineage.

23. The method according to any one of claims 1 to 7, wherein the cell is selected from prokaryotic cells and eukaryotic cells.

24. The method according to any one of claims 1 to 7, wherein the cell is selected from somatic cells and germ cells.

25. The method according to any one of claims 1 to 7, wherein the cell is an embryonic cell.

26. The method according to any one of claims 1 to 7, wherein the cell is derived from stem cells or progenitor cells.

27. A recording medium for a computer program that reconstructs a lineage of multiple cells differentiated from a common ancestor, wherein the computer program causes a computer to perform the method of any one of claims 1 to 26.

28. An apparatus for reconstructing lineages from multiple cells differentiated from a common ancestor, comprising: The processing unit executes the method according to any one of claims 1 to 26, and The display unit shows the lineage of the plurality of cells reconstructed by the processing unit.

Citation Information

Patent Citations

  • Preparation method and application of human hematopoietic progenitor cell with B pedigree differentiation potential

    CN114181968A

  • Cell surface marker for distinguishing and separating pedigree-biased new subpopulation of human pluripotent progenitor cells and application of cell surface marker

    CN115976228A

  • Gene construct and method for producing multilineage hematopoietic stem progenitor cells

    CN116479041A

  • Methods and Systems for Generating Cell Lineage Tree of Multiple Cell Samples

    US20090292482A1

  • Generation of Lineage-Restricted Progenitor Cells from Differentiated Cells

    US20160230144A1